Production Site Reliability Engineer
Charles Schwab Inc. Omaha, Nebraska, United States · $120K–$155K/yr
Financial Services · 10,001+ employees
About the role
The Site Reliability Engineer will protect the stability and performance of the Order Management System by resolving complex production incidents. They will also collaborate with cross-functional teams to drive automation, observability, and operational excellence.
What they look for
Requirements
Candidates must have a bachelor's degree and at least 5 years of experience in production support or SRE roles. Strong technical proficiency in Java, SQL, Linux, and monitoring tools is required, along with the ability to work night shifts and weekends.
Benefits
Full description
Your Opportunity
At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site 4-days per week during night shifts (2pm-10pm CST), and weekends as needed, in the specified location(s).
As a Site Reliability Engineer, you will play a critical role in protecting the stability, performance, and resiliency of Schwab’s Order Management System, supporting the technology that enables our clients to navigate their financial futures with confidence. In this highly visible role, you will assess and resolve complex production incidents, drive rapid restoration of critical systems, and collaborate across engineering, infrastructure, databases, and vendor teams to minimize business impact and improve client experiences.
Success in this role requires strong problem-solving capabilities, sound decision-making during high-pressure situations, and the ability to navigate complex distributed environments. You will partner closely with technology teams to strengthen production readiness, improve operational excellence, and identify opportunities to reduce recurring incidents through automation, observability, and continuous improvement initiatives. As a senior member of the team, you will also help shape operational best practices, mentor fellow engineers, and contribute to a culture of collaboration, accountability, and knowledge sharing.
At Schwab, we succeed together as One Schwab. You will have the opportunity to make a meaningful impact while developing your technical expertise, leadership capabilities, and operational excellence skills in a collaborative environment that values innovation, adaptability, and continuous learning.
What you have
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- 5+ years of experience in Production Support, Site Reliability Engineering (SRE), Software Operations, or a related technology support role.
- Advanced experience troubleshooting Java-based distributed applications and SQL-backed enterprise systems.
- Strong experience with application monitoring and observability tools, including AppDynamics, Splunk, Grafana, InfluxDB, and Control-M.
- Strong Oracle Database knowledge with experience in SQL analysis, database troubleshooting, and performance optimization.
- Strong Linux administration experience, preferably supporting RHEL 7/8/9 environments.
- Experience using scripting or automation technologies such as Python or Shell scripting to improve operational efficiency.
- Experience leading incident response activities and managing high-severity production incidents.
- Strong written and verbal communication skills with the ability to communicate effectively with both technical and non-technical stakeholders.
- Availability to support night shifts, weekends, and participation in a rotating on-call support model.
Preferred Qualifications
- Experience driving operational excellence initiatives that improve availability, reliability, and mean time to resolution (MTTR).
- Experience developing or enhancing monitoring strategies, operational runbooks, and escalation procedures.
- Demonstrated ability to identify root causes, implement permanent corrective actions, and reduce incident recurrence.
- Experience partnering with software engineering teams to improve application supportability and production readiness.
- Experience leading change implementation planning and supporting medium-to-high risk production releases.
- Demonstrated mentoring, coaching, or technical leadership experience within production support or SRE teams.
- Knowledge of change management, risk management, security, and compliance practices in highly regulated environments.
In addition to the salary range, this role is eligible for bonus or incentive opportunities.
Similar roles
-
Site Reliability Engineer
Origami Risk LLC Atlanta, Georgia, United States · $100K–$120K/yr
-
Site Reliability Engineer
Anduril Industries Waltham, Massachusetts, United States · $166K–$220K/yr
-
Lead, Site Reliability Engineer
Mastercard Singapore, Singapore
-
Lead Site Reliability Engineer
Mastercard Ciudad de México, Mexico
-
Engineer, Site Reliability Engineering
LSEG Bengaluru, Karnataka, India
-
Engenheiro(a) Sênior de Confiabilidade de Sites (SRE) – Storage
Dell Technologies Eldorado do Sul, Rio Grande do Sul, Brazil