Jobgether

Senior DevOps Engineer, AI Platform

Jobgether Canada

Internet Marketplace Platforms · 11-50 employees

7 h ago
Remote devops Senior (5-10 yrs) Full-time Canada
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will build and operate cloud infrastructure to support AI platforms, web applications, and backend services. This includes managing Kubernetes environments, automating infrastructure provisioning, and ensuring production reliability and scalability.

What they look for

Kubernetes Azure Kubernetes Service Oracle Kubernetes Engine Terraform Helm CI/CD Python Cloud Networking Observability PostgreSQL Redis RabbitMQ Docker ArgoCD Infrastructure as Code Linux

Requirements

Candidates must have 7+ years of professional experience in DevOps or SRE with strong hands-on expertise in Kubernetes and Microsoft Azure. Proficiency in infrastructure automation, cloud networking, and supporting distributed systems is essential.

Benefits

Full-time remote work Work on modern AI platforms High-impact role Ownership over production infrastructure Exposure to modern technologies Collaborative environment Professional growth

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer, AI Platform based in Canada.

As a Senior DevOps Engineer, you’ll build and operate the cloud infrastructure powering AI platforms, web applications, APIs, and backend services. You’ll translate technical designs into secure, scalable, reliable, and production-ready environments across modern cloud platforms. The role combines deep Kubernetes expertise with cloud networking, infrastructure automation, CI/CD, and observability. You’ll support AI workloads including LLM gateways, agent runtimes, RAG pipelines, and asynchronous processing services. You’ll work closely with AI engineers, application developers, and architects while independently owning infrastructure delivery and operations. Production reliability, incident response, scalability, security, and cost optimization will be central to your impact. This is an opportunity to help create reusable platform capabilities that enable engineering teams to deliver sophisticated services faster and more consistently.

\n

Accountabilities

  • Translate application and platform technical designs into reliable, secure, scalable, and production-ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments, primarily using Azure Kubernetes Service and Oracle Kubernetes Engine.
  • Build and support infrastructure for AI workloads, including LLM gateways, Python-based agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
  • Design and manage cloud networking, including virtual networks, subnets, routing, NAT, load balancers, DNS, TLS, private connectivity, firewalls, network policies, ingress, egress, and service-to-service communication.
  • Operate infrastructure supporting web applications, APIs, databases, caches, queues, scheduled jobs, microservices, and event-driven workloads.
  • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
  • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and Infrastructure as Code practices.
  • Implement comprehensive observability across infrastructure and applications through metrics, logs, distributed tracing, dashboards, alerts, health checks, and service-level objectives.
  • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
  • Create reusable infrastructure patterns, templates, and operational practices that enable engineering teams to launch services efficiently and consistently.
  • Determine required cloud resources, Kubernetes configurations, namespaces, scaling models, identities, secrets, and supporting services for new workloads.
  • Provision and operate dependencies such as PostgreSQL, Redis, RabbitMQ, storage systems, and other shared platform services.
  • Establish CI/CD workflows covering builds, testing, container publishing, deployment, validation, and rollback.
  • Define operational runbooks, capacity monitoring, dashboards, alerts, and health checks before production launches.
  • Own infrastructure delivery through UAT and production while partnering with architects and engineers to resolve technical design trade-offs.

Requirements

  • 7+ years of professional experience in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, or a closely related discipline.
  • Strong hands-on experience operating production Kubernetes environments, including expertise in networking, scheduling, storage, autoscaling, security, and troubleshooting.
  • Strong Microsoft Azure experience, particularly with AKS, networking, identity, storage, and monitoring; Oracle Cloud Infrastructure experience is preferred.
  • Deep understanding of cloud networking concepts, including virtual networks, subnets, routing, NAT, load balancing, private networking, DNS, TLS, firewalls, ingress, and egress.
  • Proven experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
  • Experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
  • Hands-on knowledge of databases, caching, and messaging technologies such as PostgreSQL, Redis, RabbitMQ, or equivalent platforms.
  • Experience implementing production observability using tools such as OpenTelemetry, Grafana, Prometheus, Sentry, or cloud-native monitoring solutions.
  • Strong Linux, systems administration, and production troubleshooting capabilities.
  • Working knowledge of Python, particularly backend services built with frameworks such as FastAPI.
  • Familiarity with at least one additional programming language such as C#, Java, Go, JavaScript, or TypeScript.
  • Strong understanding of HTTP/HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead-letter queues, and asynchronous processing.
  • Ability to read application logs and stack traces and diagnose infrastructure and application issues involving latency, memory, CPU, connections, and dependencies.
  • Strong cross-functional communication skills and the ability to independently execute technical designs while engaging architects and application engineers when needed.
  • Experience supporting AI or machine learning platforms, LLM gateways, agent runtimes, RAG pipelines, or MCP services is highly desirable.
  • Familiarity with Cloudflare, Envoy, ArgoCD, GitOps, and OpenTelemetry would be an advantage.
  • Experience building reusable infrastructure platforms for high-scale SaaS or customer-facing applications is a plus.
  • Strong experience operating distributed systems using technologies such as RabbitMQ, Redis, and PostgreSQL is beneficial.

Benefits

  • Full-time, fully remote position available across Canada.
  • Opportunity to work on infrastructure supporting modern AI platforms, agent systems, RAG workloads, and cloud-native applications.
  • High-impact role with significant ownership over production infrastructure, reliability, scalability, and platform engineering.
  • Opportunity to work across Microsoft Azure and Oracle Cloud Infrastructure.
  • Exposure to modern technologies including Kubernetes, Terraform, Helm, ArgoCD, Docker, OpenTelemetry, and GitOps.
  • Collaborative environment working closely with AI engineers, application engineers, architects, and platform teams.
  • Opportunity to create reusable infrastructure capabilities that accelerate engineering delivery.
  • Professional growth through hands-on work with large-scale distributed systems and emerging AI infrastructure.
  • Flexible remote-first working environment designed to support collaboration across distributed teams.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Similar roles