About the role
Provide technical and operational support for cloud-based systems while maintaining reliability and performance. Design and implement automation, monitoring, and alerting solutions to proactively manage production environments.
What they look for
Requirements
Requires at least 3 years of experience in SRE, DevOps, or a similar cloud-focused role. Must have hands-on experience with AWS, Kubernetes, Linux, and scripting.
Full description
The Company
Glassbox's mission is to empower enterprises to shape trusted, frictionless digital experiences.
Glassbox is a leading force in shaping digital experiences. It helps organizations uncover digital issues, boost conversion rates, enhance accessibility, prevent fraud, and more. Leveraging AI-driven customer intelligence, Glassbox enables enterprises to deliver secure, proactive, and preventative digital experiences. Its solutions are trusted by highly regulated organizations, including SoFi, Cal, and many others. We are growing and have been recognized by G2 as one of 2024's Top 50 Software Companies in the world.
The Opportunity
Glassbox is looking for an SRE to join our global Cloud team.
What will you do?
- Provide technical and operational support for customers according to defined SLAs.
- Work in cloud-based environments (primarily AWS) and operate/support Kubernetes-based systems (EKS).
- Design, build, and maintain advanced automation systems to enhance reliability, monitoring, and operational efficiency across production environments.
- Develop scalable monitoring and alerting solutions to proactively detect issues before they impact customers.
- Build and maintain runbooks for NOC/SOC teams.
- Serve as Tier-2 escalation for production incidents, including collaboration with DevOps and participation in a 24×7 on-call rotation.
- Implement automation-driven improvements using scripts and configuration management tools to streamline system operations.
- Leverage modern technologies and tooling to optimize system performance, observability, and resilience.
- Work closely with the Cloud DevOps team to transition products from development to the production environment via continuous integration and deployment processes
What will you need?
- At least 3 years of experience as an SRE/DevOps or in a similar cloud/monitoring role.
- Hands-on experience with AWS.
- Strong knowledge and practical experience working with Kubernetes and EKS.
- Scripting experience with Bash and hands-on experience working with Linux systems.
- Experience working with cloud monitoring, management, and alerting tools.
- Strong troubleshooting skills in production environments.
- Willingness to participate in a 24×7 on-call rotation.
- Ability to work effectively as part of a collaborative team, with strong interpersonal skills and a positive, team-oriented mindset.
- Assertive, confident, fast learner, and comfortable working in a fast-paced environment
Advantage:
- Experience with Azure and AKS.
- Knowledge of additional programming languages.
- Experience with Prometheus, Grafana, or similar monitoring/observability tools.
- Bachelor's degree in Computer Information Systems, Management Information Systems, Computer Science, or another related field experience.
- AWS or Azure certifications
Our Commitment
At Glassbox, we value curiosity, ownership, and a constant drive to learn and improve, no matter the gender, age, nationality, religion, or background. We believe diverse perspectives make us better, and that potential matters just as much as experience, and encourage people of all shapes and sizes to apply. If this role excites you and you're motivated to make an impact - we'd love to hear from you, even if you don't meet every listed qualification.
Similar roles
-
Senior Site Reliability Engineer
Precisely US Jobs United States
-
Site Reliability Engineer (SRE) Graduate Programme
mthree United Kingdom
-
Staff Software Engineer, Site Reliability Engineering, Cloud Billing
Google New York, New York, United States · $207K–$300K/yr
-
Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects
Apple Cupertino, California, United States
-
Senior Software Engineer, Site Reliability Engineering
Google Sunnyvale, California, United States · $174K–$252K/yr
-
Senior Site Reliability Engineer
TailorCare United States