Meta

Data Center Production Operations Engineer

Meta Cluain Aodha, Leinster, Ireland

Software Development · 10,001+ employees

13 h ago
Mid (2-5 yrs) Full-time Ireland
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The engineer will monitor and maintain the operational health of large-scale server fleets and production infrastructure. They will also diagnose hardware failures, execute lifecycle management processes, and collaborate with engineering teams to improve system reliability.

What they look for

Data center operations Server hardware troubleshooting Systems administration Hardware lifecycle management Root cause analysis Operational runbooks Capacity planning Python Bash Fleet management Infrastructure monitoring Production incident resolution Automation Technical documentation Cross-functional collaboration

Requirements

Candidates must have at least 2 years of experience in data center operations or systems administration with strong hardware troubleshooting skills. Proficiency in developing operational processes and analyzing failure metrics is required, along with the ability to work effectively in cross-functional teams.

Full description

Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. In this role, you will be responsible for the day-to-day operational health of server fleets and production systems, working at the intersection of hardware lifecycle management, systems troubleshooting, and operational process improvement. You will partner closely with hardware engineering, capacity planning, and infrastructure teams to ensure Meta's data centers operate at peak performance, directly enabling the products and services that connect billions of people worldwide.

Responsibilities

  • Monitor and maintain the operational health of large-scale server fleets and production infrastructure across data center environments
  • Diagnose and resolve hardware and systems failures, coordinating with engineering teams to drive root cause analysis and implement corrective actions
  • Execute and refine server deployment, decommissioning, and lifecycle management processes to support capacity and reliability goals
  • Develop and maintain operational runbooks, escalation procedures, and documentation to standardize production operations workflows
  • Collaborate with hardware engineering and capacity planning teams to identify systemic issues and propose infrastructure improvements
  • Track and analyze operational metrics and failure trends to surface insights that improve fleet reliability and reduce mean time to resolution
  • Support the qualification and rollout of new server hardware generations by validating operational readiness and identifying deployment risks
  • Partner with cross-functional teams including network engineering, facilities, and software infrastructure to resolve complex production incidents
  • Identify opportunities to automate repetitive operational tasks and contribute to tooling improvements that increase operational efficiency
  • Provide technical guidance to peers on production operations best practices, hardware troubleshooting methodologies, and process standards

Minimum Qualifications

  • 2+ years of experience in data center operations, production operations, or systems administration in a large-scale infrastructure environment
  • Experience troubleshooting server hardware components including CPUs, memory, storage, and networking hardware in a production setting
  • Experience developing or improving operational processes, runbooks, or standard operating procedures for data center or infrastructure teams
  • Experience analyzing operational data or failure metrics to identify trends and drive reliability improvements
  • Experience collaborating with cross-functional engineering teams to resolve production incidents and implement systemic fixes

Preferred Qualifications

  • Experience supporting hardware qualification or new server platform bring-up in a data center production environment
  • Experience with fleet management tooling, asset tracking systems, or infrastructure monitoring platforms at scale
  • Familiarity with scripting languages such as Python or Bash for automating operational workflows and data analysis tasks
  • Background in capacity planning, hardware lifecycle management, or server deployment operations for hyperscale data centers