Site Reliability Engineer
Luma Financial Technologies New York, New York, United States · $120K–$145K/yr
Financial Services · 51-200 employees
About the role
The Site Reliability Engineer will manage AWS infrastructure, Kubernetes clusters, and CI/CD pipelines to ensure platform reliability and scalability. They will also lead incident response efforts and implement resilience strategies such as disaster recovery and multi-region architecture.
What they look for
Requirements
Candidates must have at least 5 years of experience in Site Reliability or Software Development Engineering. Proficiency in Java, Python, Bash, Go, AWS, Terraform, and Kubernetes is required, along with a strong focus on system resilience and automation.
Full description
About the role
At Luma, our Site Reliability Engineer (SRE) team keeps our platform reliable, secure, and lightning fast. They own everything from AWS infrastructure and Kubernetes clusters to CI/CD pipelines, monitoring, and alerting. If you’re passionate about tackling big challenges, automating at scale, and making systems more resilient, we’d love to have you on the team.
What you'll do
- Collaborate with product engineering teams to design and build the infrastructure their services run on.
- Keep our Kubernetes clusters on AWS EKS running smoothly, secure, and ready to scale.
- Design and deliver resilience strategies that cover multi-region architecture, backups, disaster recovery, and failover.
- Automate infrastructure with Terraform and Infrastructure-as-Code, reducing manual effort and human error.
- Help teams ship faster by improving CI/CD pipelines and deployment practices.
- Monitor performance and reliability using modern observability tools.
- Support on-call rotations and lead incident response with a focus on long-term fixes.
Qualifications
- 5+ years of applicable experience in Site Reliability or Software Development Engineering required
- Bachelor’s degree in Computer Science, Software Engineering or related concentration highly preferred
- You code to solve problems and are comfortable in the following languages: Java, Java, Python, Bash and Go.
- You have strong experience with AWS (RDS, CloudFront, IAM, VPCs), Terraform, and Kubernetes.
- You are resilience focused, with experience designing and running systems that remain dependable during failures and recover seamlessly.
- You have hands-on experience improving and operating CI/CD pipelines (e.g., CircleCI, GitHub Actions, or similar) to help teams ship faster with confidence.
- You stay calm under pressure, bringing incident response expertise and strong root-cause analysis skills.
- Most importantly, you are a team player who brings clear communication, strong collaboration, and a mindset of continuous improvement.
Luma Financial Technologies is unable to sponsor work visas or provide employment-based immigration sponsorship for this position, now or in the future. Applicants must be legally authorized to work in the United States without current or future sponsorship.
Any misrepresentation, including regarding work authorization or sponsorship needs during the application or interview process will result in disqualification from consideration for this role or termination of employment.
#LI-Hybrid, #LI-SS1
Similar roles
-
Site Reliability Engineer III (DBA)
Backblaze External Website United States · $125K–$150K/yr
-
Site Reliability Engineer II (AI Platform)
OpenTable Toronto, Ontario, Canada · CA$110K–CA$130K/yr
-
Site Reliability Engineer
Apple Hyderabad, Telangana, India
-
Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)
Encora Perímetro Urbano Santiago de Cali, Valle del Cauca, Colombia
-
Senior Site Reliability Engineer
Salesforce Dublin, Leinster, Ireland
-
Sr Staff Site Reliability Engineer, AI Infrastructure
d-Matrix Santa Clara, California, United States · $175K–$265K/yr