F

Site Reliability Engineer (SRE) - Database Focus

FyerX

Technology, Information and Media · 11-50 employees

19 h ago
Remote sre Senior (5-10 yrs) Full-time Contractor
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design and operate highly available, scalable database clusters, optimizing query performance and managing replication, backups, and disaster recovery. Automate database maintenance and monitoring, resolve production performance issues, and enforce data security and governance controls.

What they look for

Database Reliability Engineering Database Administration PostgreSQL MySQL Aurora Database Performance Optimization High Availability Database Replication Disaster Recovery Python Go Bash Terraform Prometheus Grafana Datadog

Requirements

Requires 5–9 years of systems engineering or database administration experience, including at least 4 years managing high-scale cloud-native databases, and strong knowledge of database engines, automation, telemetry, and transaction mechanics. A listed database certification is mandatory; experience with large-scale migrations and distributed databases or CDC tools is preferred.

Full description

This is a remote position.

Site Reliability Engineer (SRE) - Database Focus

Job Details

  • Employment Type: Contract
  • Work Mode: Remote
  • Location: Offshore
  • Total Experience Required: 5 to 9 years
  • Relevant Experience Required: 4+ years of dedicated database reliability engineering or database administration (DBA) experience in high-throughput cloud environments
  • Mandatory Certification: AWS Certified Database - Specialty, Certified Postgres Professional, or Oracle Database Administration Certified Professional

Job Summary

We are seeking an experienced Database Site Reliability Engineer (DB SRE) to oversee the stability, performance, and scaling limits of our high-volume transaction databases. The ideal candidate will build automated cluster management tools, orchestrate sub-millisecond query optimizations, manage multi-region replication layers, and design fault-tolerant configurations to ensure maximum database uptime and zero data loss.

Key Responsibilities

  • Design and maintain highly available database cluster topologies across public clouds, leveraging automated scaling, sharding mechanisms, and read-replica strategies.
  • Optimize sub-millisecond database query performance, conducting rigorous index evaluations, identifying locking contentions, and rewrites of inefficient SQL scripts.
  • Orchestrate automated backup and disaster recovery validation sweeps, configuring point-in-time recovery (PITR) parameters and multi-region failover tests.
  • Build automated telemetry dashboards and proactive alerts utilizing monitoring frameworks (e.g., Prometheus, Grafana, Datadog) to track database health metrics (IOPS, CPU utilization, connections).
  • Write robust automation scripts using Python, Go, or Bash to manage recurring database maintenance routines, schema migration deployments, and resource adjustments.
  • Diagnose and remediate production database performance bottlenecks, resolving replication lags, connection pool limits, memory usage leaks, and deadlocks.
  • Enforce data safety and governance protocols, configuring encryption-at-rest policies, fine-grained access parameters, and compliance masking routines to secure sensitive records.

Requirements

  • 5 to 9 years of core systems engineering or database administration experience, with 4+ dedicated years actively managing high-scale, cloud-native relational databases (e.g., PostgreSQL, MySQL, Aurora) or NoSQL databases.
  • Strong technical mastery of database engine configurations, connection poolers (e.g., PgBouncer), infrastructure automation (Terraform), and system telemetry structures.
  • Deep structural understanding of write-ahead logging (WAL), isolation levels, transaction mechanics, distributed storage bounds, and network latency impacts.
  • Mandatory certification: AWS Database Specialty, Certified Postgres Professional, or Oracle Database Admin Certified Professional.

Preferred Qualifications

  • Prior experience implementing large-scale live database data migrations with minimal production operational windows.
  • Familiarity with distributed database engines or streaming message queues (e.g., CockroachDB, Kafka, Debezium) for real-time change data capture (CDC).

Similar roles