Senior Site Reliability Engineer
Felix Technologies, Inc. Ciudad de México, Mexico
Financial Services · 51-200 employees
About the role
You will manage and optimize infrastructure on GCP and GKE while automating provisioning and configuration tasks. Additionally, you will define SLOs, manage incident responses, and collaborate with cross-functional teams to embed reliability and security into the development lifecycle.
What they look for
Requirements
Candidates must have 4+ years of experience in SRE or Platform Engineering with strong proficiency in GCP, Kubernetes, and scripting languages like Go or Python. A solid understanding of monitoring, observability, and cloud security best practices is essential for this role.
Benefits
Full description
About Us
At Félix, we are building the indispensable financial companion for Latinos in the US. We combine an AI-powered, conversational-first interface with real-time financial infrastructure to make cross-border money movement as easy as sending a text. Starting with fast, affordable remittances powered by AI and crypto rails, we are expanding into credit, savings, and wallet services to support the complete immigrant financial journey. Our ambition is to deliver a white-glove financial experience with the simplicity of a conversation to a community the traditional financial system has historically overlooked.
We are a hyper-growth Series C company, backed by over $300 million in funding from top-tier global investors, including Andreessen Horowitz, QED, Castle Island, Switch Ventures, HTwenty, Monashees, General Catalyst Customer Value Fund. This isn't just about the numbers; it's a testament to the trust our investors have in our vision and our team. Additionally, Félix was selected as an “Endeavour Entrepreneur” and was a recipient of the CrossTech Fintech Startups Award.
Joining Félix means you will be part of a team building a legacy, a company that will outlive us all. This is a rare opportunity to apply your skills to a deeply meaningful mission—serving a community that has been underserved for too long because we are obsessed with our customers. We get things done with urgency and focus, driven by extreme ownership over our impact. We collaborate without ego, fostering radical transparency and fierce loyalty so we can grow together. Because we aim for insanely great rather than just good enough, we stay insatiably curious—always experimenting, building the future today, and delivering a product that truly makes our users' lives better.
About the Role
As a Senior Site Reliability Engineer you will be a critical part of our Platform Engineering team. You will be responsible for creating automations to provide velocity to product engineering . This is a hands-on role for a builder who is passionate about shifting security left and empowering developers to ship secure code, quickly and confidently. You will be instrumental in maturing our DevSecOps practices, building out our automations.
Responsibilities
- Manage and optimize our infrastructure on Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
- Automate provisioning and configuration using Terraform, Helm, and scripting languages such as Go, Python, and Bash.
- Build, maintain, and improve monitoring and alerting systems using OpenTelemetry standards
- Participate in on-call rotations, incident response, and post-mortem analyses, ensuring rapid recovery and continuous learning from failures.
- Define and track SLOs/SLIs and error budgets to monitor service health and performance.
- Implement cloud security best practices to protect sensitive data and maintain the integrity of our systems.
- Collaborate across Engineering, Security, and Product teams to embed reliability and automation in every phase of development and deployment.
- Contribute to GKE cost optimization and resource management strategies to enhance efficiency and control operational spend.
Requirements
- 4+ years of experience as a SRE/Platform Engineer.
- Strong hands-on experience with GCP and GKE.
- Proficiency in Kubernetes (architecture, deployments, networking, and troubleshooting).
- Solid programming or scripting skills in Go, Python, or Bash.
- Proficiency with Docker and Linux
- Experience with Terraform
- Experience with Helm
- Experience with GitHub Actions
- Strong understanding of monitoring and observability using Prometheus, Grafana, and logging frameworks.
- Familiarity with incident management, on-call operations, and post-mortem processes.
- Knowledge of network fundamentals (TCP/IP, DNS, Load Balancing).
- Experience with PostgreSQL or distributed databases.
- Awareness of FinOps and cloud cost management principles.
- Excellent problem-solving, communication, and collaboration skills, with a proactive mindset.
- GCP certifications, such as Professional DevOps Engineer or Cloud Architect.
- Certified Kubernetes Administrator (CKA).
- Experience in FinOps, cloud security, or regulated industries.
- Familiarity with PagerDuty or similar incident management tools.
- Background implementing SLOs/SLIs and error budgets in production environments.
- These are the applicable requisites, although equivalent competencies in any of the above will also be considered.
What We Offer
- Competitive salary
- Initial stock options grant
- Annual performance bonus
- Health, dental, and vision plans
- Remote work environment, although we have offices in Miami and México City and would love to work in hybrid model if you are up to it.
- Continuous learning opportunities
- Unlimited PTO
- Paid parental leave
- Empowering opportunities for growth in a dynamic entrepreneurial environment
Equal Opportunity Employer
At Félix, we are committed to providing equal employment opportunities to all qualified employees and applicants without regard to race, religion, nationality, sex, sexual orientation, gender identity, age, or disability. This policy applies to all terms and conditions of employment, including recruitment, hiring, placement, promotion, training, compensation, benefits, and termination.
Want to learn more about our privacy practices? Check out our Privacy Policy.
Similar roles
-
Lead Site Reliability Engineer
JPMorgan Chase & Co. Jersey City, New Jersey, United States · $157K–$215K/yr
-
Site Reliability Engineer
Schonfeld New York, New York, United States · $150K–$225K/yr
-
Site Reliability Engineer II
Axon Boston, Massachusetts, United States · $116K–$185K/yr
-
Senior II Site Reliability Engineer Lead
Akamai Cambridge, Massachusetts, United States · $146K–$264K/yr
-
Principal Site Reliability Engineer
MetaRouter $180K–$250K/yr
-
Senior Lead Site Reliability Engineer
JPMorgan Chase & Co. Jersey City, New Jersey, United States · $176K–$260K/yr