Senior Site Reliability Engineer (Cloud Platform)
Salve.Inno Consulting Denver, Colorado, United States
Human Resources Services · 11-50 employees
About the role
The Senior Site Reliability Engineer will maintain the reliability, availability, and performance of production environments while driving automation and observability. They will also partner with software engineering teams to improve application reliability and participate in incident response and root cause analysis.
What they look for
Requirements
Candidates must hold a Bachelor's or Master's degree in a relevant field and possess strong experience with Kubernetes, AWS, and Linux-based production environments. Proficiency in scripting languages like Python or Go and experience with Infrastructure as Code tools such as Terraform or Ansible are required.
Benefits
Full description
B2B Contract
Role Overview
We're looking for a Senior Site Reliability Engineer to help build, operate, and continuously improve a highly available cloud platform supporting mission-critical production services.
In this role, you'll work at the intersection of cloud infrastructure, software engineering, and operations, helping engineering teams build reliable, scalable systems while driving automation, observability, and operational excellence. You'll play a key role in strengthening platform reliability, improving incident response, and embedding SRE best practices throughout the software development lifecycle.
Key Responsibilities:
- Maintain the reliability, availability, and performance of production and pre-production environments.
- Monitor platform health and improve alerting, automation, and operational processes.
- Respond to production incidents, participate in root cause analysis, and implement long-term improvements.
- Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards.
- Partner with software engineers to improve application reliability throughout the development lifecycle.
- Develop and maintain operational documentation, troubleshooting guides, and runbooks.
- Automate repetitive operational tasks to improve efficiency and reduce manual intervention.
- Participate in on-call rotations while continuously improving incident response processes.
- Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams.
Requirements:
- Bachelor's or Master's degree in Engineering, Computer Science, or a related field.
- Strong experience operating Kubernetes or other container orchestration platforms.
- Experience supporting large-scale production services.
- Hands-on experience with AWS.
- Experience with Prometheus, Grafana, and ELK.
- Strong scripting skills (Bash, Python, or Go).
- Experience administering Linux-based production environments.
- Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible.
- Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing).
- Excellent troubleshooting, communication, and collaboration skills.
- A proactive mindset with a passion for automation and reliability.
Nice to Have:
- Experience with SIP or VoIP technologies.
- Familiarity with MySQL or PostgreSQL.
- Experience with Redis or other NoSQL databases.
What's on Offer:
- Long-term, full-time collaboration.
- Flexible remote working environment.
- Professional development opportunities, including training and technical learning.
- The opportunity to work on innovative cloud technologies used by customers worldwide.
- Collaborative engineering culture focused on knowledge sharing and continuous improvement.
- Modern Apple equipment provided.
Diversity and Inclusion Commitment
We are dedicated to creating and sustaining an inclusive, respectful workplace for all -regardless of gender, ethnicity, or background. We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast-paced, supportive team.
Similar roles
-
Staff Site Reliability Engineer
NinjaTrader Chicago, Illinois, United States · $160K–$210K/yr
-
Senior Site Reliability Engineer – Network Observability
Blueprint Technologies $104K–$114K/yr
-
Amazon Connect SRE
Miratech Surat, Gujarat, India
-
AWS Site Reliability Engineer
Miratech Ahmedabad, Gujarat, India
-
Senior Data SRE / Cloud Platform Engineer
Flywire Valencia, Valencian Community, Spain · €49K–€62K/yr
-
Sr. Site Reliability Engineer III (6804)
MetroStar Washington, District of Columbia, United States · $185K–$230K/yr