Senior Site Reliability Engineer
MetaRouter $160K–$180K/yr
Data Infrastructure and Analytics · 11-50 employees
About the role
You will design, build, and operate cloud infrastructure while automating manual processes to ensure system reliability and scalability. Additionally, you will manage deployment pipelines, maintain observability tools, and mentor engineers to improve delivery velocity.
What they look for
Requirements
Candidates must have 6+ years of experience in SRE, DevOps, or platform engineering with hands-on experience in public cloud environments. Proficiency in container orchestration, infrastructure as code, and systematic troubleshooting of distributed systems is required.
Benefits
Full description
About Us
MetaRouter is a customer-data streaming platform for large, security-conscious enterprises. The MetaRouter platform dramatically reduces latency and bloat by providing server-side integration with third-party marketing, analytics, and data storage/transport tools. By enabling customers to access our SaaS tool, access a private PaaS architecture, or deploy fully within their own private cloud, customers can dramatically simplify and centralize their customer-data pipelines and maintain full control over security and compliance.
The MetaRouter platform, serving as a full-spectrum customer data collection, modification, and delivery platform, is a collection of many, varied microservices. Our client and server-side ingestion and identity libraries, ETL applications, and configuration and monitoring UIs serve to give data teams control over the shape and substance of their customer-behavior data, all while powering the flexibility and freedom to be creative with their architecture.
About The Role
As a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and operational tooling end to end. You will automate what is manual, instrument what is opaque, and make our deployments repeatable across a growing number of isolated customer environments.
We run dedicated, private environments per customer, so the interesting problems here are about repeatability, automation, and observability at scale. Experience with these patterns matters more than familiarity with any particular cloud, orchestrator, or monitoring vendor.
Core Responsibilities
- Design, build, and operate the cloud infrastructure that supports our applications and internal operations, from provisioning through decommissioning.
- Own deployment automation and release tooling, and improve the safety and speed of getting changes to production.
- Manage upgrades across infrastructure and the supporting software our applications depend on.
- Build and maintain dashboards, logs, metrics, and alerting so problems are caught early and alerts mean something.
- Troubleshoot infrastructure and application issues in production, and drive fixes through to resolution.
- Scale systems sustainably through automation, and push for changes that improve both reliability and delivery velocity.
- Handle identity and access management and single sign-on across platforms and services.
- Ensure infrastructure and applications meet compliance requirements.
- Work with customers to determine and implement custom infrastructure requirements.
- Participate in code reviews to ensure infrastructure, applications, and supporting services follow best practice.
- Improve and maintain infrastructure and process documentation.
- Mentor engineers earlier in their careers and pair with product teams to spread reliability practice.
- Participate in the on-call rotation.
Qualifications and Experience
- 6+ years in an SRE, DevOps, or platform engineering role operating production systems.
- Hands-on experience with at least one major public cloud provider.
- Solid experience configuring, maintaining, and troubleshooting container orchestration clusters.
- Working fluency with infrastructure as code, configuration management, containerization, version control, and CI/CD tooling.
- Strong scripting and automation skills, plus the ability to debug and optimize application code when needed.
- Experience building dashboards, metrics, and alerts in a modern observability platform, including its query language.
- Expertise troubleshooting large-scale distributed systems, with a systematic problem-solving approach.
- Solid understanding of Unix/Linux operating systems and networking fundamentals.
- Effective written and verbal communication, and comfort prioritizing a wide variety of tasks in a fast-paced environment.
- Familiarity with agile methodologies.
Employment Details
Job Type: Full Time Location: Fully Remote (US)
Benefits
- Health / Dental / Vision insurance
- 401(k)
- Unlimited vacation policy
- Fully remote (US)
Similar roles
-
Senior Application Support Engineer / Site Reliability Engineer (SRE)
DTCC Boston, Massachusetts, United States
-
Site Reliability Engineer - Privilege & Access Management
JPMorgan Chase & Co. Plano, Texas, United States
-
Senior Site Reliability Developer
Vena Solutions Canada · CA$123K–CA$167K/yr
-
Manager TAG and Encompass SRE
WEXWEXUS Berwyn, Illinois, United States · $106K–$131K/yr
-
CTIO - Site Reliability Engineer- Senior Associate
PwC Birmingham, Alabama, United States · $55K–$187K/yr
-
Sr. SRE
Pura Pleasant Grove, Utah, United States