Halo Media

Site Reliability Engineer (SRE)

Halo Media United States

Technology, Information and Internet · 201-500 employees

11 h ago
Remote sre Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Lead large-scale operating system modernization projects and drive infrastructure automation to improve system reliability. Provide operational support, manage incident response, and partner with engineering teams on cloud migration initiatives.

What they look for

Linux Python Automation CI/CD Infrastructure Engineering RPM Packaging Configuration Management Observability Incident Management Troubleshooting Cloud Migration RHEL Chef CINC System Administration

Requirements

Requires 5+ years of experience in Site Reliability or Infrastructure Engineering with strong proficiency in Python and Linux administration. Candidates must have proven experience in leading infrastructure migrations and managing CI/CD pipelines.

Benefits

100% Remote International and collaborative environment

Full description

About the Role:

We are looking for a Senior Site Reliability Engineer (SRE) to help modernize large-scale infrastructure and improve the reliability, scalability, and operational excellence of critical production systems. In this role, you will lead OS modernization initiatives, drive infrastructure automation, strengthen observability, and partner closely with engineering teams to support cloud migration efforts.

The ideal candidate has a strong background in Linux systems, Python, automation, CI/CD, and production operations, with a passion for building resilient platforms and solving complex infrastructure challenges.

Responsibilities:

  • Lead large-scale operating system modernization projects, including migrations from RHEL7 to EL8/9 across approximately 1,700 systems and virtual machines.
  • Drive infrastructure and packaging migrations, including Chef to CINC and yinst to RPM.
  • Build, maintain, and configure RPM packages to support modern infrastructure deployments.
  • Develop automated operational runbooks and infrastructure automation to improve efficiency and reliability.
  • Harden CI/CD pipelines, rollout/rollback mechanisms, and deployment processes for infrastructure modernization.
  • Strengthen observability by onboarding services to modern monitoring and logging platforms.
  • Triage, investigate, and resolve complex production incidents and critical software bugs.
  • Provide Tier-2 operational support in a follow-the-sun model alongside Site Reliability Engineering and Cloud Infrastructure teams.
  • Support incident response, troubleshooting, and break/fix activities across distributed production environments.
  • Partner with software engineering teams during cloud migration initiatives, providing operational guidance and technical support.
  • Automate repetitive operational tasks and maintain comprehensive technical documentation.
  • Drive reliability, operational excellence, and continuous improvement across production systems.

Required Qualifications:

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Software Engineering with a strong infrastructure focus.
  • Strong hands-on experience with Python.
  • Proven experience leading Linux operating system modernization and infrastructure migration projects.
  • Experience with RHEL, Linux administration, and package management.
  • Hands-on experience building and maintaining RPM packages.
  • Strong experience with infrastructure automation and configuration management.
  • Experience designing, maintaining, and improving CI/CD pipelines.
  • Strong troubleshooting and incident management skills in large-scale production environments.
  • Experience supporting distributed systems with a strong focus on reliability and availability.
  • Excellent scripting, automation, and problem-solving skills.

Benefits:

  • 100% Remote.
  • International and collaborative environment.

Similar roles