MWDN

Platform & SRE Engineer

MWDN Kyiv, Ukraine

IT Services and IT Consulting · 51-200 employees

Yesterday
Remote sre Senior (5-10 yrs) Full-time Contractor Ukraine
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Platform & SRE Engineer will own the operational lifecycle of Linux-based runtime hosts and manage AWS infrastructure. They will also establish SRE and observability practices to ensure platform health, reliability, and performance.

What they look for

Linux systems administration AWS SRE Infrastructure engineering Python Bash Observability Grafana Prometheus OpenTelemetry Systemd Networking Security hardening Automation Cloud infrastructure Incident response

Requirements

Candidates must have 8+ years of experience in DevOps, SRE, or Linux systems engineering with strong hands-on skills in AWS and Linux administration. Proficiency in automation using Python or Bash and a deep understanding of SRE principles are essential.

Benefits

Legal assistance Professional development English courses Team-building events 29 days of PTO 10 paid recovery days Financial support for independent contractors Dedicated HR support

Full description

MWDN is a global IT outstaffing company with 23+ years of experience that connects exceptional tech talent with leading companies across Israel, the USA, Great Britain, and Western Europe. We offer opportunities to work on international products in a stable and professional environment.

Why does MWDN rock?

Here’s what you can expect when you join MWDN:

  • Security: We carefully vet our clients to minimize risks and ensure reliability and timely payments-no fraud or unpleasant surprises.
  • Career support: If a project isn’t the right fit, we support you and actively help find new opportunities that match your skills and career goals.
  • Legal assistance: We provide guidance on legal matters, including opening and managing your independent contractor or sole proprietorship status, taxes, and related processes.
  • Professional development: We offer English courses and professional growth opportunities, as well as team-building events.

Why choose us? MWDN is ranked among the top 5 IT employers in our region according to DOU. We take pride in our transparency and strong commitment to our team. Curious to learn more? See what our employees say about working with us on DOU.

What is your new project?

Domain: AI Infrastructure / Real-Time Data Processing

Client Location: Israel

Company size: 10 - 51

A fast-growing product company building next-generation infrastructure for production AI systems. The team is focused on high-performance, low-latency distributed technologies that operate directly within real-time data flows and AI workloads. The product addresses complex engineering challenges related to runtime reliability, distributed processing, networking, observability, and large-scale system performance.

What makes this project exciting?

The client is looking for a Platform & SRE Engineer to join the engineering team and take ownership of the infrastructure and operational foundation behind our distributed runtime environments.

This is a hands-on engineering role combining Linux systems engineering, AWS infrastructure, machine-image management, deployment automation, observability, security hardening, and SRE practices.

A key part of the role is owning the lifecycle of Linux-based runtime hosts — from AMI creation and machine provisioning through operating-system configuration, system services, runtime deployment, host tuning, upgrades, security hardening, monitoring, and operational troubleshooting.

You will maintain and evolve our internal host management and deployment tooling, ensuring runtime environments can be deployed, configured, upgraded, diagnosed, and operated reliably across cloud and, in the future, customer-hosted environments.

In parallel, you will help establish and own our SRE and observability practices, ensuring we can measure platform health and availability, identify failures quickly, and operate against clearly defined reliability objectives.

What makes you a great fit

  • 8+ years of experience in DevOps, SRE, Platform Engineering, Linux Systems Engineering, or Infrastructure Engineering.
  • Strong Linux systems administration and troubleshooting skills, including systemd, processes, networking, filesystems, permissions, kernel configuration, and system-level debugging.
  • Strong hands-on AWS experience, particularly with EC2, networking, IAM, storage, and production Linux workloads.
  • Strong experience with AWS AMIs, including building, configuring, hardening, validating, versioning, and maintaining machine images for production environments.
  • Strong experience with Linux server hardening and production security practices, including least privilege, service isolation, user and permission management, secure system configuration, and attack-surface reduction.
  • Hands-on experience with Grafana, OpenTelemetry, Prometheus, Mimir,
  • ClickHouse, Fluent Bit, or comparable observability technologies.
  • Experience building production monitoring, dashboards, metrics, logging, and alerting solutions.
  • Understanding of SRE principles, including SLIs, SLOs, availability, incident response, and root-cause analysis.
  • Experience writing automation and operational tooling using Python and Bash.
  • Strong understanding of networking fundamentals including TCP/IP, routing, DNS, TLS, and secure connectivity.
  • Experience automating machine provisioning, configuration, deployment, and upgrades.
  • Ability to independently troubleshoot complex issues spanning infrastructure, operating systems, networking, and application runtime.
  • Strong ownership mindset and ability to take responsibility for systems from deployment through production operation.

Nice to Have

  • Experience with Linux performance tuning including CPU pinning, NUMA, hugepages, IRQ affinity, and NIC tuning.
  • Experience with high-performance or low-latency networking environments.
  • Experience with DPDK or other user-space networking technologies.
  • Experience operationally integrating FPGA or other hardware accelerators into Linux environments.
  • Experience supporting software deployed in customer-managed or on-prem environments.
  • Experience with PKI, certificate management, TLS/mTLS, and machine identity.
  • Experience with Infrastructure as Code such as Terraform.
  • Experience with GitHub Actions or similar CI/CD systems.
  • Experience with vulnerability management and security/compliance initiatives such as SOC 2.
  • Experience working in startup or high-growth engineering environments.

Your day-to-day in this position

Linux Platform & Host Operations

  • Own the operational lifecycle of Linux-based runtime hosts.
  • Maintain and evolve internal host configuration, deployment, and management tooling.
  • Manage Linux system configuration and services using systemd and Linux-native tooling.
  • Automate host provisioning, configuration, upgrades, validation, and recovery.
  • Configure and troubleshoot host-level networking and runtime environments.
  • Work with performance-sensitive Linux configurations including CPU affinity, hugepages, IRQ affinity, kernel parameters, and NIC configuration.
  • Troubleshoot complex issues across the operating system, networking, runtime software, and hardware boundaries.

AWS & Machine Image Management

  • Build, package, validate, maintain, and troubleshoot AWS AMIs used for production runtime environments.
  • Own and improve AMI creation and release processes.
  • Automate machine provisioning and image lifecycle management.
  • Operate and troubleshoot AWS EC2-based runtime infrastructure.
  • Manage the integration between machine images, runtime software, system configuration, and deployment automation.
  • Support operational integration with hardware accelerators deployed in AWS environments.
  • Help evolve deployment mechanisms toward customer-hosted and on-prem environments.

Linux Security & Server Hardening

  • Define and implement Linux server-hardening standards for production environments.
  • Harden operating-system configuration, system services, users, permissions, networking, and runtime processes.
  • Apply least-privilege principles across services and host-level components.
  • Ensure services run with appropriate users, permissions, capabilities, and filesystem access.
  • Reduce attack surface by disabling unnecessary services, ports, packages, and privileges.
  • Implement secure configuration of SSH, systemd services, kernel parameters, logging, and host networking.
  • Automate and validate hardening as part of machine-image creation and deployment.
  • Work with CISO to remediate vulnerabilities and continuously improve host security posture.

Observability & SRE

  • Build and maintain production dashboards covering infrastructure, hosts, services, and overall platform health.
  • Design and maintain actionable alerts for service degradation, infrastructure failures, resource saturation, and availability issues.
  • Define and implement SLIs and SLOs for critical platform components.
  • Establish availability and uptime measurements used to track reliability and support SLA commitments.
  • Identify gaps in application and infrastructure telemetry and work with engineering teams to introduce the required metrics.
  • Improve monitoring of latency, errors, traffic, resource utilization, saturation, and system health.
  • Participate in incident response, root-cause analysis, and reliability improvements.
  • Continuously improve monitoring coverage and reduce alert noise.

Observability Infrastructure

  • Operate and evolve centralized metrics, logging, and telemetry infrastructure.
  • Deploy, configure, and maintain technologies such as Grafana, OpenTelemetry,
  • Prometheus-compatible systems, Mimir, ClickHouse, and Fluent Bit.
  • Maintain reliable telemetry collection from distributed runtime environments.
  • Build and maintain dashboards, alerting rules, telemetry pipelines, and operational diagnostics.
  • Troubleshoot observability pipelines from collection through storage, querying, visualization, and alerting.
  • Ensure observability infrastructure itself is reliable, scalable, and operationally maintainable.

Platform Automation & DevOps

  • Build automation and operational tooling using Python, Bash, and Linux tooling.
  • Improve release, deployment, upgrade, and rollback workflows.
  • Support CI/CD and infrastructure automation.
  • Reduce manual operational work through automation-first solutions.
  • Contribute to broader DevOps and infrastructure initiatives as needed.
  • Work closely with software, networking, security, and hardware engineering teams.

Why work with us?

  • People-first management with minimal bureaucracy
  • A friendly company culture, proven by employees who choose to return
  • Flexible working hours
  • 29 days of PTO (18 working days per year pluse all national holidays)
  • 10 paid recovery days
  • Full financial and legal support for independent contractors
  • Free English classes, with native speakers or Ukrainian teachers
  • Dedicated HR support

Our next steps

✅ Intro call with a Recruiter — ✅ Technical Interview — ✅ Interview with CTO and Head of Engineering — ✅ HR Interview — ✅ Offer

Requirements

null

Similar roles