Senior SRE - BeReal
BeReal · Paris, Ile-de-France, France
Software Development · 51-200 employees
About the role
You will design, implement, and optimize infrastructure for scalability and reliability while maintaining SRE practices like SLIs and SLOs. Additionally, you will partner with development squads to improve service performance and automate operational workflows.
What they look for
Requirements
The role requires strong knowledge of Kubernetes, distributed systems, and at least one major cloud provider like GCP. You must have solid experience applying SRE practices and a pragmatic, ownership-driven mindset for managing production incidents.
Benefits
Full description
About BeReal
At BeReal, we are dedicated to authenticity in social media. By encouraging users to share unfiltered moments, we foster genuine connections and celebrate real life. We are now an international team of 100+ and have 40M+ monthly active users. Backed by Voodoo, our team is fully focused on scaling BeReal into an iconic social network used by hundreds of millions.
The Infrastructure team provides the backbone that powers the company’s growth, ensuring the scalability, efficiency, and reliability of our platform. We design and operate our infrastructure on GCP. Working hand in hand with developers, we enable teams to ship fast and efficiently while maintaining a strong focus on costs and performance. Our mission is to create a developer-friendly, cost-effective, and highly automated infrastructure that supports innovation at scale.
Role
- Apply and help maintain SRE practices across your scope, including SLIs, SLOs, error budgets, incident management, and postmortem processes
- Design, implement, and optimize infrastructure for availability, scalability, reliability, and cost efficiency
- Contribute to our observability stack, improving monitoring, alerting, logging, and distributed tracing
- Automate infrastructure and operational workflows (e.g., Terraform, Terragrunt, Kubernetes)
- Support FinOps initiatives, helping build tools and insights to optimize cloud costs
- Partner closely with development squads to improve service reliability, performance, and operational excellence
- Contribute to architectural decisions and help apply best practices for building resilient distributed systems
- Share knowledge with other Infrastructure engineers, helping raise the bar on reliability and operational excellence
- Analyze performance bottlenecks and work on solutions such as scaling strategies, service optimizations, and system debugging
Profile
- Strong knowledge of Kubernetes
- Experience with high traffic, distributed systems architectures, and related tools (service discovery, config/secret management, etc.)
- Strong knowledge of one Cloud provider (AWS or GCP preferred)
- Solid experience applying SRE practices (SLOs, incident management, observability, reliability engineering)
- Strong operational mindset with experience managing production incidents and driving reliability improvements
- Comfortable collaborating with and supporting other engineers, with the ability to weigh in on technical decisions
- Ownership-driven – If something isn't working, you don't wait for instructions; you improve it
- Pragmatic and impact-oriented – You balance reliability, delivery speed, and business priorities
- Performance vs cost-conscious – You make decisions that align with both technical excellence and financial sustainability
Our Stack
- Operator: Kubernetes
- CI/CD: Argocd, Github actions
- Cloud provider: GCP
- Monitoring: Datadog
- Infra as code: Terraform / Terragrunt
- Languages: golang / node
- Datastores: Spanner / PostgreSQL / Redis
Benefits
- Competitive salary based on experience
- Swile Lunch voucher
- Gymlib (100% covered by Voodoo)
- Premium healthcare coverage with SideCare, 100% covered for you and your family
- Wellness activities in our Paris office