About the role
You will own the end-to-end cloud infrastructure, managing it as code and ensuring robust GitOps delivery. Additionally, you will oversee observability, cloud security, and cost management while building internal tooling to support scientific and ML teams.
What they look for
Requirements
The role requires 2-3+ years of platform or infrastructure engineering experience with deep expertise in AWS and infrastructure-as-code. Strong production experience with Kubernetes, GitOps, and CI/CD pipelines is essential for success in this position.
Benefits
Full description
About AQEMIA
AQEMIA is a drug invention company dedicated to creating entirely new medicines to address major unmet medical needs. At the core of our mission is QEMI, our proprietary molecule-invention platform, which uniquely combines cutting-edge science with advanced technology. Powered by physics-based modeling, statistical mechanics, and generative AI, QEMI allows our teams to design novel drug candidates from first principles.
What makes AQEMIA different is our commitment to true innovation: our research is dedicated to the invention of new molecular entities, not the refinement of existing ones. We focus on inventing never-before-seen molecules, without relying on experimental data, and advancing them into a growing pipeline of proprietary programs and strategic partnerships with leading pharmaceutical companies.
Our most advanced preclinical programs are currently in vivo optimization, targeting diseases still waiting for effective treatments, offering our teams the opportunity to work on science that can make a real difference in people’s lives.
For more information, visit AQEMIA.com, our WTTJ Page, and our LinkedIn.
About our Team
AQEMIA brings together a diverse, multidisciplinary team of 80+ professionals based in Paris and London. Our scientists and engineers, including chemists, physicists, machine learning experts, and software engineers, work side by side to push the boundaries of early-stage drug discovery.
This close collaboration across disciplines is central to our approach, enabling us to tackle complex scientific challenges from first principles and translate cutting-edge ideas into novel therapeutic candidates. At AQEMIA, team members are encouraged to contribute their expertise, learn from one another, and play an active role in shaping the future of drug invention.
About our Engineering Department
The Engineering team (~12 people) builds and scales the technical foundations that power Aqemia’s drug discovery engine.
Bringing together expertise across software engineering, cloud infrastructure, data engineering, site reliability engineering (SRE), and scientific computing, the team designs and operates robust, secure, and high-performance platforms that enable Aqemia’s scientific and AI teams to experiment, train, deploy, and run models at scale.
Their work spans data infrastructure, scientific compute systems, cloud operations, orchestration, observability, CI/CD, developer tooling, and platform scalability. By transforming cutting-edge research into reliable and scalable systems, the team accelerates the discovery of new medicines while ensuring operational excellence and an outstanding developer experience.
The role
As our SRE, you'll own Aqemia's cloud platform end to end - infrastructure-as-code, GitOps delivery, observability, security and FinOps - for a company where nothing is deployed by hand and everything runs through Git.
The scale here is different from a typical product company: bursty, large-scale parallel scientific computation, GPU fleets to plan and optimize, and some of the company's most valuable data to protect, all while keeping cost under control.
You'll join a small team with full ownership of the stack, meaning your decisions on architecture and tooling directly shape how fast research moves.
As the platform evolves, from today's GitOps pipelines toward productionized MLOps and autonomous discovery workflows, you'll grow with it, taking on more scope rather than staying in a fixed lane.
\n
Responsibilities
- Own AWS infrastructure end to end, built and maintained as code with OpenTofu and Terragrunt - no manual changes, no exceptions.
- Operate and evolve Kubernetes workloads via GitOps (ArgoCD, Helm, Kustomize), and drive adoption of standardized infrastructure patterns across teams.
- Contribute to the platform's reliability practice: own observability (metrics, logs, alerting), respond to incidents, and run blameless postmortems through to completed action items.
- Manage cloud cost as a shared responsibility - producing the monthly cost report, maintaining the cost allocation model, and partnering with teams to plan capacity ahead of large GPU compute campaigns.
- Set and enforce security posture across the platform: patching, vulnerability follow-up, and least-privileged access management.
- Build internal tooling and CI/CD pipelines that reduce friction for engineering, ML and scientific teams, and make the platform approachable to non-infrastructure users.
- Shape platform architecture and long-term strategy through technical reviews and sprint planning, sharing knowledge across infrastructure and DevOps topics.
Qualifications
- Strong platform/infrastructure engineering background, with 2-3+ years of experience post-degree.
- Deep hands-on expertise in AWS and infrastructure-as-code (Terraform/OpenTofu, Terragrunt).
- Strong production experience with Kubernetes and GitOps delivery (ArgoCD, Helm, Kustomize).
- Experience building and maintaining CI/CD pipelines (GitHub Actions or GitLab).
- Experience with cloud security practices and least-privileged access management.
Nice-to-have
- MLOps experience - training/inference pipelines, model lifecycle, workflow orchestrators.
- GPU capacity planning - autoscaling GPU fleets, spot strategies, quota management.
- Exposure to AI-driven or data-intensive workflows.
- Experience with another cloud provider beyond AWS (e.g. GCP).
Our recruitment process
- First discussion with our Talent Acquisition
- Hiring Manager’s interview: you’ll meet directly with your future manager
- Technical assessment of your skills in a deep-dive interview with the team
- Cultural fit interview with our co-founder and COO, Emmanuelle
- Final interview with our co-founder and CEO, Maximillien
\nWhy Join Us?
At AQEMIA, we work for a mission: joining us means having your own impact on changing the way drugs are discovered, and helping to shape the direction of our fast-growing company and team.
Expanding Drug Discovery Pipeline : Focused on critical therapeutic areas like Oncology, CNS, Immuno-inflammation... with in vivo proof of concept/patent stage programs. Collaborations with top Pharma, including a $140M Sanofi deal.
World-Class Interdisciplinary Team : work alongside exceptional talent at the intersection of technology and life sciences. Our teams combine deep expertise in AI, physics-based modeling, biology, and medicinal chemistry to push the boundaries of innovation.
DeepTech Recognition: AQEMIA is proud to be part of the French Tech 120 and France 2030, highlighting our role as a key player in Europe’s DeepTech ecosystem.
Prime Location with Flexibility : Our offices are located in the heart of Paris and London (King’s Cross), with flexible work arrangements including up to two remote days per week.
Strong Financial Backing : $100M raised from leading European and International investors
Similar roles
-
Especialista em Infraestrutura | SRE
XP Inc. São Paulo, São Paulo, Brazil
-
Site Reliability Engineer (m/f/x)
Tipico Karlsruhe, Baden-Württemberg, Germany
-
Site Reliability Engineer
Razorpay Software Private Limited Bengaluru, Karnataka, India
-
Senior Software Engineer, Site Reliability Engineering
Google New York, New York, United States · $174K–$252K/yr
-
Software Engineer III, Site Reliability Engineering, Vertex AI
Google Warsaw, Masovian Voivodeship, Poland
-
Senior Staff Software Engineer, Site Reliability Engineering
Google Sunnyvale, California, United States · $262K–$364K/yr