Senior Site Reliability Engineer
Elevate Government Solutions Washington, District of Columbia, United States
IT Services and IT Consulting · 11-50 employees
About the role
The Senior Site Reliability Engineer will oversee the reliable and secure operation of critical application infrastructure for a federal financial agency. Responsibilities include maintaining production environments, troubleshooting incidents, managing deployments, and optimizing system performance.
What they look for
Requirements
Candidates must have an active Secret security clearance and professional experience with Python, Linux, Terraform, Ansible, and Docker. A bachelor's degree in Computer Science or a related field is highly preferred.
Benefits
Full description
This position requires an ACTIVE SECRET security clearance. Your application will not be reviewed if you do not have an ACTIVE security clearance at SECRET or above.
Elevate Government Solutions is a high growth Technology Services company focused on driving technological change in the government space. Our teams engage with various government agencies to develop and deploy emerging technology solutions using a tailored Agile methodology.
We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the United States with a Secret security clearance. The Senior Site Reliability Engineer will oversee the efficient, reliable, and secure operation of critical application infrastructure for a federal financial agency. In this operationally focused role, you will be responsible for maintaining production environments, troubleshooting incidents, performing monitoring and diagnostics, supporting deployments, and continually improving system uptime and reliability.
The ideal candidate brings Python experience, deep operational expertise with Linux environments, and hands-on experience managing infrastructure with tools like Terraform, Ansible, and Docker. Experience with CI/CD pipelines, containerization, Git-based workflows, and monitoring/alerting systems are essential. Familiarity with managing AWS services—particularly EKS, S3, and EMR (Elastic Map Reduce)—is highly valued. In addition, familiarity with Spark, JupyterHub, and Hue is preferred. Effective communication and a collaborative mindset are critical to success in this client-facing environment.
Responsibilities and Duties
- Maintain, monitor, and troubleshoot production environments to ensure optimal uptime and performance
- Operate and manage infrastructure using Terraform, Ansible, and Docker
- Oversee and support CI/CD pipelines and automate operational workflows using Git and related tooling
- Ensure ongoing reliability and operational excellence for critical systems running on AWS services, including EKS, S3, and EMR (Elastic Map Reduce)
- Support operational use of Spark, JupyterHub, and Hue
- Diagnose and resolve operational issues, perform root cause analysis, and drive problem resolution
- Implement, refine, and monitor infrastructure and application alerting and diagnostics
- Optimize data flows and storage integrations
- Collaborate with engineering, product, and client stakeholders to communicate issues, coordinate maintenance, and support deployments
- Contribute to continual improvement of platform operational processes, documentation, and best practices
Required Experience, Skills and Qualifications
- Proactive, creative problem-solving mindset with a demonstrated ability to anticipate, diagnose, and resolve complex operational challenges
- Experience with Python, specifically in operational and support environments
- Professional experience with Terraform, Ansible, Docker, and CI/CD pipeline operations
- Strong proficiency with Git and version control workflows
- Comfortable working extensively on the Linux command line and performing operational tasks
- Demonstrated ability to learn new technologies quickly
- Experience monitoring, maintaining, and supporting production infrastructure in a collaborative environment
- Strong incident response and communication skills
- Ability to work East Coast hours
- Ability to work hybrid in Vienna, VA, or remotely
Preferred Qualifications
- Experience operating AWS or other cloud platforms
- Familiarity with Spark, JupyterHub, and Hue in an operational context
- Hands-on experience with EMR (Elastic Map Reduce)
- Experience with Databricks
- Experience supporting federal, regulated, or client-facing environments
Education Requirement
- 4 year Bachelor's Degree in Computer Science/Eng or related (highly preferred)
Benefits
- Health Insurance
- 401k match
- Unlimited Vacation
- 11 Federal Holidays
- Remote work
- Training
- Mentoring/Coaching
Similar roles
-
Summer 2027 Site Reliability Internship
Tradeweb London, England, United Kingdom
-
Digital Site Reliability Engineer
Radisson Hotel Group Madrid, Community of Madrid, Spain
-
Site Reliability Engineer
WorldQuant Montevideo, Montevideo, Uruguay
-
Site Reliability Engineer II
Axon Boston, Massachusetts, United States · $116K–$165K/yr
-
Staff Site Reliability Engineer
Crunchyroll, LLC Los Angeles, California, United States · $210K–$263K/yr
-
Senior Site Reliability Engineer
Ciklum Kyiv, Ukraine