Nucleome Therapeutics Ltd

Senior Cloud Data Engineer

Nucleome Therapeutics Ltd · United Kingdom

Biotechnology · 11-50 employees

20 h ago
Senior (5-10 yrs) Full-time United Kingdom
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design, build, and operate the cloud infrastructure for a genomics data platform, focusing on scalability, performance, and cost-effectiveness. Additionally, you will manage infrastructure-as-code repositories and implement observability solutions to ensure pipeline health and service reliability.

What they look for

AWS Kubernetes Python Pulumi Terraform Data Engineering CI/CD GitOps Observability PostgreSQL S3 SageMaker Data Pipelines Cloud Infrastructure Networking Distributed Tracing

Requirements

Candidates must hold a degree in computer science or a related quantitative discipline and possess extensive experience in designing production cloud infrastructure on AWS. Strong programming skills in Python and experience with Kubernetes, CI/CD workflows, and infrastructure-as-code tools are essential.

Full description

Overview of the role:

We are looking for an experienced Cloud Data Engineer to join our expanding computational team. Working in collaboration with engineering, bioinformatics, and drug discovery colleagues, you will design, build, and operate the cloud infrastructure that underpins our genomics data platform; from orchestration and compute through to storage, search, and machine learning. Your work will focus on scaling this platform reliably as data volumes and pipeline complexity grow, and on maturing our observability, security, networking and engineering operational practices so the team can move fast with confidence.

Key responsibilities:

  • Drive the design and evolution of our AWS-based platform. Including EKS cluster capacity, networking and ingress, and the data services the platform depends on: RDS PostgreSQL, S3 data Lakehouse, search engines, and SageMaker. You will identify bottlenecks, propose architectural improvements, and implement changes to keep pipelines and services performant and cost-effective as demand increases.
  • You will identify opportunities to scale existing and develop new data pipelines which the lab team rely on to generate the critical data needed for analysis on drug target viability, making your work directly impactful to therapeutic development priorities.
  • Own and extend our Pulumi based infrastructure-as-code repositories. You will work closely with the engineering team on technical and architectural decisions, taking ownership of implementation while collaborating on strategy, standards, and review. Together you will iterate on platform design and ensure consistent, repeatable deployments across environments.
  • Build on our existing observability stack (monitoring, logging, tracing) to deliver deeper visibility into pipeline health, service performance, and infrastructure utilisation. You will design and implement alerting and dashboards, extend instrumentation across Dagster jobs, Kubernetes workloads, and AWS services and introduce distributed tracing where it adds clear operational value. You will also contribute to cost visibility and optimisation (e.g. compute sizing, spot/on-demand strategy, storage lifecycle).

Required Qualifications and Skills:

  • BSc, MSc or PhD in computer science, engineering, or a related quantitative discipline. Experience in biotech, pharma, or other data-intensive scientific environments. Solid programming skills (Python preferred). Experience with CI/CD pipelines, GitOps or similar deployment workflows, and modern DevOps practices.
  • Extensive hands-on experience designing and operating production cloud infrastructure on AWS, including Kubernetes, VPC networking, IAM, S3, RDS, and containerised workloads. Experience with infrastructure-as-code tools (Pulumi or Terraform etc.) and Helm-based deployments.
  • Experience scaling data platforms and batch/orchestrated workloads: autoscaling (Karpenter, HPA), mixed compute strategies (spot, on-demand, Fargate), and performance/cost trade-offs for large-scale data processing. Experience in data modelling and implementation in relational and non-relational databases.
  • Strong practical experience with observability stacks; metrics, logging, and alerting. Familiarity with distributed tracing and an ability to implement end-to-end observability for data pipelines.
  • Strong communication, organisational, and time management skills, with the ability to explain complex infrastructure concepts to both technical and non-technical audiences. Comfortable working independently and as part of a multidisciplinary team in a fast-moving environment.
  • Detail oriented in implementation, alongside a clear view of the bigger picture, planning horizons, operational risk, and business impact.