Senior Database Reliability Engineer
Jobgether United States
Internet Marketplace Platforms · 11-50 employees
About the role
You will own the reliability, performance, and availability of production databases across EHR, data, and AI workloads. This includes managing observability, automating safeguards, and partnering with engineering teams to optimize database architecture and query quality.
What they look for
Requirements
The role requires 6+ years of experience in database reliability engineering or administration with deep expertise in MySQL and AWS Aurora. Candidates must possess strong skills in automation, observability tools, and experience working within HIPAA-compliant environments.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Database Reliability Engineer based in the United States.
This is a high-impact engineering role responsible for the reliability, performance, availability, and cost efficiency of production databases powering a modern healthcare technology platform.You will take end-to-end ownership of database operations across EHR, data, and AI workloads, with a strong focus on preventing issues rather than simply reacting to them.The role combines deep database engineering with observability, automation, incident response, architecture, and cloud cost optimization.You will partner closely with engineering teams to improve query quality, establish safe migration practices, and ensure scalable database architectures.You will also play a key role in protecting sensitive healthcare information within a HIPAA-compliant environment.This is an opportunity to influence platform architecture, engineering standards, and operational practices as the organization continues to scale.The ideal candidate is a database expert who enjoys solving complex production problems and building systems that remain reliable under growing workloads.
\n
Accountabilities:
- Own the reliability, performance, availability, and operational health of production databases running on AWS Aurora MySQL across EHR, Data, and AI workloads.
- Manage database observability end to end, maintaining the metrics and alerting pipeline into Datadog and integrating after-hours database alerts into the DevOps on-call rotation.
- Establish automated safeguards to identify and terminate long-running or runaway queries and provide immediate visibility into database activity across instances.
- Investigate database performance issues by analyzing execution plans, optimizing queries, re-indexing where appropriate, and reducing unnecessary database load and latency.
- Review application and EHR queries before production release, serving as a performance gate to prevent inefficient workloads from reaching production.
- Educate engineering teams on database hygiene, query optimization, and practices that improve reliability and performance.
- Create and maintain comprehensive database runbooks so first responders can resolve incidents quickly and consistently.
- Architect and continuously optimize Aurora reader topologies and read-routing strategies based on actual workload requirements.
- Own database cost efficiency through instance right-sizing, reserved-capacity and Savings Plan strategies, and appropriate storage tiering.
- Manage replication health, backups, restore testing, failover procedures, and disaster recovery capabilities.
- Establish safe, repeatable standards for database schema changes and migrations across engineering teams.
- Partner with platform architecture stakeholders to evaluate and determine the best technical approach for new and existing database workloads.
- Handle protected health information responsibly while maintaining database operations within a HIPAA-compliant environment.
Requirements:
- 6+ years of experience in database reliability engineering, database administration, database engineering, or a closely related discipline, including ownership of production systems at scale.
- Deep expertise in MySQL, including query optimization, execution-plan analysis, indexing strategies, and replication; hands-on Aurora MySQL experience is strongly preferred.
- Experience operating large, multi-reader Aurora or RDS clusters at terabyte scale, including read-routing and connection-management strategies.
- Strong knowledge of database observability and monitoring technologies such as Datadog, Percona Monitoring and Management (PMM), Performance Insights, Prometheus, or Grafana.
- Proficiency with Python, Bash, or a comparable scripting language for automation, along with practical experience using infrastructure-as-code tools such as Terraform.
- Strong AWS operational knowledge, particularly RDS/Aurora, database sizing, reserved capacity, storage options, and cloud cost optimization.
- Experience with production on-call rotations, incident response, troubleshooting, and creation of operational runbooks.
- Strong understanding of database reliability, security, backup, recovery, failover, and disaster-recovery practices.
- Excellent communication and collaboration skills, with the ability to educate and influence engineers across multiple teams.
- Experience with cloud data warehouses such as Snowflake, Databricks, or Redshift is a plus, particularly where analytical workloads can be moved away from OLTP databases.
- Experience working with HIPAA-regulated or other compliance-heavy environments involving sensitive data is advantageous.
- Familiarity with automated query-remediation approaches such as Percona Toolkit, pt-kill, statement timeouts, or custom query termination systems is beneficial.
- Application-side experience, particularly with PHP, is a plus for effective collaboration on query and application performance.
- Knowledge of database and cloud security best practices, agile methodologies, and fast-paced engineering environments is valued.
Benefits:
- Competitive salary and compensation package.
- Remote/hybrid working environment.
- Potential equity compensation based on outstanding performance.
- Flexible PTO.
- Company-sponsored lunches.
- Company-paid disability and life insurance.
- Company-paid family and medical leave.
- Medical, dental, and vision insurance.
- Discounted pet insurance.
- FSA/DCA and commuter benefits.
- 401(k) retirement plan.
- Credits toward online fitness classes and gym memberships.
- Access to an HQ recovery suite featuring a cold plunge, sauna, and shower.
- Opportunity to work on meaningful healthcare technology and solve complex infrastructure challenges.
- Collaborative environment focused on smart, sustainable work rather than simply working longer hours.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1