Data Center Production Operations Engineer
Meta
Data Center Production Operations Engineer Responsibilities:
- Monitor and maintain the operational health of large-scale server fleets and production infrastructure across data center environments
- Diagnose and resolve hardware and systems failures, coordinating with engineering teams to drive root cause analysis and implement corrective actions
- Execute and refine server deployment, decommissioning, and lifecycle management processes to support capacity and reliability goals
- Develop and maintain operational runbooks, escalation procedures, and documentation to standardize production operations workflows
- Collaborate with hardware engineering and capacity planning teams to identify systemic issues and propose infrastructure improvements
- Track and analyze operational metrics and failure trends to surface insights that improve fleet reliability and reduce mean time to resolution
- Support the qualification and rollout of new server hardware generations by validating operational readiness and identifying deployment risks
- Partner with cross-functional teams including network engineering, facilities, and software infrastructure to resolve complex production incidents
- Identify opportunities to automate repetitive operational tasks and contribute to tooling improvements that increase operational efficiency
- Provide technical guidance to peers on production operations best practices, hardware troubleshooting methodologies, and process standards
- 2+ years of experience in data center operations, production operations, or systems administration in a large-scale infrastructure environment
- Experience troubleshooting server hardware components including CPUs, memory, storage, and networking hardware in a production setting
- Experience developing or improving operational processes, runbooks, or standard operating procedures for data center or infrastructure teams
- Experience analyzing operational data or failure metrics to identify trends and drive reliability improvements
- Experience collaborating with cross-functional engineering teams to resolve production incidents and implement systemic fixes
- Background in capacity planning, hardware lifecycle management, or server deployment operations for hyperscale data centers
- Experience supporting hardware qualification or new server platform bring-up in a data center production environment
- Experience with fleet management tooling, asset tracking systems, or infrastructure monitoring platforms at scale
- Familiarity with scripting languages such as Python or Bash for automating operational workflows and data analysis tasks
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.