SRE - Database Reliability (All Levels)
Casap San Francisco, California, United States
Software Development · 11-50 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
You will own the production database layer end-to-end, focusing on reliability, performance, and security for Aurora PostgreSQL and DynamoDB. This includes managing backups, migrations, observability, and on-call incident response within a regulated fintech environment.
What they look for
Requirements
Candidates must have deep experience running PostgreSQL in production, including query optimization and major-version upgrades. You should also be proficient with AWS data stores, infrastructure-as-code tools like Terraform, and working under compliance frameworks like PCI DSS or SOC 2.
Full description
Location(s)
This is a hybrid role based out of our San Francisco, CA or New York, NY office. Our team works in-office time on Mondays, Wednesdays, and Fridays.
About the Role
We're hiring an SRE to own our production databases end to end, weighted toward database reliability rather than general infrastructure.
Casap builds dispute-resolution software for banks and credit unions, and we onboard new financial-institution clients every month. Our data layer runs on Aurora PostgreSQL today, with OpenSearch and a TimescaleDB metrics store beside it, and part of it is moving to DynamoDB. You'll make sure that data is backed up, restorable, fast and upgraded on time, and you'll carry it safely through the migration and after it.
Platform engineering already covers infrastructure as code, CI and networking, so you'll spend most of your time on the data layer itself. You'll join the Foundations team and work in a regulated environment (PCI DSS and SOC 2).
What You'll Do
- Aurora PostgreSQL in production: availability, performance, capacity planning and major-version upgrades.
- Backups and restores: backup policy, point-in-time recovery, and scheduled restore tests with
- measured recovery time and data-loss targets.
- The DynamoDB migration, from the data side: access-pattern and data-model review, backfill
- and dual-write plans, validation, cutover and rollback.
- Database access and security: least-privilege roles, credential rotation, audit logging and
- encryption, in line with PCI DSS and SOC 2.
- Data-layer observability: slow queries, connection saturation, replication lag and storage
- growth, with alerting in Datadog.
- On-call for data-layer incidents: response, follow-through and written post-incident reviews.
- Runbooks clear enough that the rest of engineering can operate the databases without you.
What You’ve Done
- You've run PostgreSQL in production and been the person accountable when it broke.
- You have deep PostgreSQL skills: query planning, indexing, vacuum and bloat, connection pooling, replication and major-version upgrades.
- You've done real backup and restore work: point-in-time recovery, restore drills and measured recovery times.
- You know data stores on AWS, including Aurora PostgreSQL (Serverless v2) and DynamoDB data modeling (access patterns, key design, capacity).
- You're comfortable with Terraform or OpenTofu, CI pipelines and scripting in Python, Go or Bash.
- You've worked under PCI DSS, SOC 2 or a similar regime, or you're ready to.
- You write clearly: runbooks, migration plans and incident reviews that other people can act on.
Nice to Have
- TimescaleDB in production
- A completed migration from a relational database to DynamoDB
- Datadog database monitoring
- Experience in payments, banking, or another fintech setting
Similar roles
-
Site Reliability Engineer, Data Center Infrastructure
SpaceX Bastrop, Texas, United States
-
Sr. Site Reliability Engineer, Platform Infrastructure
SpaceX Bastrop, Texas, United States
-
Principal Site Reliability Engineer
UKG Seattle, Washington, United States · $184K–$265K/yr
-
Software Engineer, SRE and Production Engineering - DGX Cloud
NVIDIA Santa Clara, California, United States · $184K–$288K/yr
-
Production Site Reliability Engineer
Charles Schwab Inc. Omaha, Nebraska, United States · $120K–$155K/yr
-
Site Reliability Engineer (Associate, Experienced, or Senior)
Boeing Berkeley, Missouri, United States · $99K–$217K/yr