WEX

Manager TAG and Encompass SRE

WEX Chicago, Illinois, United States · $106K–$131K/yr

Software Development · 5,001-10,000 employees

20 h ago
sre Mid (2-5 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will lead and mentor a global team of Site Reliability Engineers while driving critical operations, incident response, and strategic reliability projects. Additionally, you will monitor system health and collaborate with development teams to implement automation and observability improvements.

What they look for

Site Reliability Engineering Team Management Python Bash Go Grafana ELK Stack Splunk Docker Kubernetes CI/CD Incident Response Cloud Platforms Infrastructure as Code Observability Automation

Requirements

The role requires at least 3 years of experience in system administration, DevOps, or SRE, along with 2 years of experience in people management. Proficiency in scripting languages like Python, Bash, or Go and experience with containerization and monitoring tools are essential.

Benefits

Health insurance Dental insurance Vision insurance Retirement savings plan Paid time off Health savings account Flexible spending accounts Life insurance Disability insurance Tuition reimbursement

Full description

About the Team & Role We are looking for a highly motivated and high-potential mid-level Site Reliability Engineer Manager (SRE) to join our team and help drive meaningful business impact while accelerating your growth as a reliability leader.

This is an exciting time to be part of the SRE evolution at WEX. Our complex systems and diverse product portfolio power a wide variety of customer businesses, generating rich operational and telemetry data across platforms and environments. As we scale, the need for resilient, observable, and efficient systems is greater than ever.

As a mid-level SRE Manager, you’ll play a key role in shaping and building the next generation of WEX’s reliability practices, platforms, and tools. You’ll contribute to driving improvements across monitoring, alerting, performance optimization, incident response, and capacity management. You’ll also help reduce operational toil through automation and proactive engineering, while partnering closely with engineering and product teams to embed reliability into everything we build.

We take a modern approach to engineering—leveraging agile practices, a product-oriented mindset, and a strong focus on innovation, including the use of AI and automation to enhance observability and response.

You’ll face meaningful challenges that have significant impact, and be part of a high-performing team with experienced engineers and leaders ready to support your continued development.

If you're passionate about reliability, solving hard problems, and growing fast in a high-impact environment, this is a great opportunity for you!

How you'll make an impact

  • You will lead, mentor, and elevate a globally diverse team of Site Reliability Engineers—driving critical operations, incident response, and strategic projects while championing the ongoing career development of your team.
  • Monitor and manage system health, availability, and performance.
  • Develop and maintain automation tools for system provisioning, monitoring, and alerting.
  • Participate in on-call rotations and respond to system alerts and incidents.
  • Collaborate with development teams to implement reliability-focused features.
  • Improve observability and logging for troubleshooting issues.
  • Follow IT security policies and compliance requirements.

Experience you'll bring 

  • 3+ years of experience in system administration, DevOps, or SRE roles.
  • 2+ years of experience managing people or teams
  • 2+ years of experience leading projects
  • Proficiency in scripting and automation using Python, Bash, or Go.
  • Experience with monitoring and logging (Grafana, ELK stack, Splunk, etc.).
  • Knowledge of containerization and orchestration (Docker, Kubernetes).
  • Understanding of CI/CD pipelines and version control systems.
  • Understanding of monitoring tools such as Prometheus, Grafana, or Splunk.
  • Strong problem-solving skills and a willingness to learn.

Preferred Qualification

  • Hands-on experience with cloud platforms (AWS, Azure, or GCP).
  • Familiarity with infrastructure as code (Terraform, Ansible, CloudFormation).
  • Knowledge of incident response processes and SLAs.
  • Experience with developing AI based solutions.
  • Ability to troubleshoot and resolve performance bottlenecks.
  • Strong communication skills and ability to work across teams.
  • Experience supporting systems that manage B2B payments, expense reimbursements, or virtual card creation.
  • Familiarity with monitoring and alerting in systems that are eventually consistent.
  • Awareness of key compliance frameworks (e.g., PCI-DSS, SOX) and how they apply to observability and automation.

The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non-sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well-being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section.

Pay Range: $106,000.00 - $130,600.00

Similar roles