Charles Schwab Inc.

Site Reliability Engineer

Charles Schwab Inc. Chicago, Illinois, United States · $90K–$110K/yr

Financial Services · 10,001+ employees

Yesterday Closes in 5d
sre Mid (2-5 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Site Reliability Engineer will partner with cross-functional teams to improve service resiliency, observability, and automation for critical trading platforms. They will apply engineering principles to solve complex operational challenges and strengthen production readiness across distributed environments.

What they look for

Site Reliability Engineering DevOps Infrastructure as Code Cloud Operations Python Go Java Bash PowerShell Kubernetes Grafana Prometheus OpenTelemetry CI/CD Google Cloud Platform Observability

Requirements

Candidates must have 4+ years of experience in SRE, DevOps, or related technical disciplines with hands-on experience in public cloud environments. Proficiency in Infrastructure as Code, CI/CD automation, and monitoring tools is required.

Benefits

Bonus opportunities Incentive opportunities

Full description

Your Opportunity

Your Opportunity

At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).

The Client Trading Experience Technology team is responsible for ensuring the reliability, scalability, and operational excellence of critical trading platforms that support clients around the clock. As a Site Reliability Engineer, you will partner across application engineering, architecture, platform, cybersecurity, and support teams to improve service resiliency, observability, automation, and cloud adoption for business-critical systems. Your work will directly influence platform stability, incident response effectiveness, deployment reliability, and overall client experience.

In this role, you will apply engineering principles to solve complex operational challenges, build automated and reusable cloud solutions, and strengthen production readiness across distributed environments. You will help advance modern reliability practices through infrastructure such as code, monitoring, CI/CD automation, and AI-assisted operational capabilities while contributing to a collaborative culture focused on continuous improvement, innovation, and engineering excellence.

What you have

Required Qualifications

  • 4+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Cloud Operations, or a related technical discipline.
  • Hands-on experience supporting and troubleshooting production applications in public cloud or large-scale enterprise environments.
  • Experience with Infrastructure as Code and cloud infrastructure automation.
  • Experience using GitHub or similar platforms for source control, code review, and CI/CD automation.
  • Experience with monitoring, observability, incident response, production recovery, and operational readiness.
  • Experience building dashboards, alerts, and monitoring solutions using Grafana or comparable observability platforms.
  • Experience developing automation using Python, Go, Java, Bash, PowerShell, or similar technologies.
  • Experience using AI-assisted engineering tools with appropriate validation, governance, and human oversight.
  • Working knowledge of SRE practices including service-level objectives, post-incident improvement, runbooks, and toil reduction.
  • Strong collaboration, communication, problem-solving, and operational decision-making skills.

Preferred Qualifications

  • Hands-on experience with Google Cloud Platform, including Cloud Run and Google Compute Engine.
  • Experience creating reusable Infrastructure as Code components and integrating infrastructure changes with deployment workflows.
  • Experience supporting cloud-native, Linux-based, or enterprise platform environments.
  • Experience with Grafana administration, Prometheus-compatible monitoring, OpenTelemetry, log aggregation, and alerting technologies.
  • Experience with GitHub Actions, security scanning, policy controls, and software supply-chain practices.
  • Familiarity with containerization and Kubernetes concepts.
  • Experience implementing or supporting AIOps capabilities that improve operational efficiency and service reliability.
  • Experience modernizing enterprise applications and supporting highly available, resilient cloud environments.

In addition to the salary range, this role is eligible for bonus or incentive opportunities.

Similar roles