Railsware

Senior DevOps Engineer

Railsware

Software Development · 201-500 employees

3 h ago
Remote devops Senior (5-10 yrs) Full-time
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own and evolve the production infrastructure on AWS while leading a strategic initiative to migrate specific workloads to bare-metal infrastructure. This role involves managing multi-region systems, ensuring high reliability, and coordinating technical migration plans end-to-end.

What they look for

DevOps AWS Terraform Kubernetes CI/CD GitHub Actions Docker Linux Networking Python Shell Scripting Ansible Cloudflare Prometheus Grafana Postgres Kafka

Requirements

The ideal candidate has 5+ years of experience in DevOps with deep expertise in AWS, Terraform, and container orchestration. You must be comfortable managing multi-quarter infrastructure initiatives and possess strong skills in Linux networking and automation.

Full description

We are looking for a Senior DevOps Engineer to join our team!

You will own and evolve Mailtrap's production infrastructure — a multi-region AWS platform powering high-volume email sending and testing. You will keep it reliable, secure, and cost-efficient day to day.

You will also lead a key strategic initiative: planning and executing a partial migration of selected workloads from AWS to rented bare-metal / colocation infrastructure to cut costs — a hybrid-by-design effort, not a full cloud exit.

You will decide what moves and what stays, design the target platform (likely Kubernetes or a similar orchestrator on bare metal), and own the migration end to end, from business case to cutover.

Our engineering team thrives in an agile, continuously improving, and automation-oriented environment. We value ongoing evolution, objective evaluation of our processes, and taking action to make things better.

Must have skills

  • 5+ years DevOps/platform (senior), production ownership
  • Deep AWS (VPC, ECS, IAM, RDS/ElastiCache or equivalents, networking)
  • Terraform at scale (modules, remote state / Terraform Cloud)
  • CI/CD with GitHub Actions, including automated deploys to production (e.g. blue/green deployments)
  • Containers (Docker); comfortable operating services on ECS or K8s
  • Strong Linux networking (DNS, TLS, load balancing, firewalls, VPN/hybrid)
  • Proven experience planning and executing migration to on-prem, colo, or private cloud (partial/hybrid OK — full exit not required)
  • Comfortable owning a multi-quarter infra initiative: TCO, design, vendors, cutover, rollback
  • Comfortable automating operational tasks with shell and/or Python
  • Fluent English (both spoken and written)

Strongly preferred

  • Colo/bare-metal ops: IPAM (NetBox), Ansible, image-based provisioning (Packer or equivalent), HAProxy/Nginx
  • Replacing managed AWS services with self-hosted (Postgres HA, Redis/Valkey, Kafka, OpenSearch)
  • Email infrastructure (MTA, SMTP, IP reputation, DKIM/SPF/DMARC) — Halon or similar
  • Cloudflare (DNS/WAF/Access)
  • Observability beyond CloudWatch (Prometheus/Grafana/Loki or equivalent)
  • Prior work with multi-region SaaS or EU data residency
  • Cost-driven architecture / FinOps mindset

Would be a plus

  • GCP (BigQuery/certificates)
  • Comfortable reading Ruby or Go (used in our services and tooling)
  • AWS Certificates

Responsibilities day-to-day

  • Operate and evolve AWS multi-account / multi-region infra
  • Terraform modules/workspaces
  • Ensure safe infrastructure changes across network, storage, and services, with zero-downtime deployments.
  • ECS services, blue/green deploys, Docker image pipelines
  • Reliability: CloudWatch/PagerDuty/Sentry, capacity, cost tags
  • Security baseline: IAM, secrets (SSM), Cloudflare edge rules
  • Partner with engineers on release automation and production readiness
  • Maintain the hybrid estate (AWS and rented bare metal) as one operable platform

Migration leadership

  • Build the business and technical case for what moves off AWS vs what stays
  • Design the rented bare-metal / colo landing zone (compute, network, storage, observability, secrets)
  • Produce migration waves, dependency maps, cutover/rollback plans
  • Stand up hybrid connectivity and dual-run periods; shift traffic safely (e.g. via Cloudflare)
  • Replace or re-home managed services where it pays off (compute, queues, search, cache, MTA nodes)
  • Coordinate colo/vendors, timelines, and eng teams; report progress and risk

Similar roles