Charger Logistics Inc

Site Reliability Engineer

Charger Logistics Inc India

Truck Transportation · 501-1,000 employees

6 h ago
Remote sre Senior (5-10 yrs) Contractor India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Site Reliability Engineer will maintain the reliability, scalability, and availability of the production platform while managing Kubernetes clusters and containerized workloads. They will also automate operational tasks, lead incident responses, and collaborate with development teams on performance tuning and capacity planning.

What they look for

Kubernetes Docker Terraform Helm Python Go Bash Prometheus Grafana PostgreSQL CI/CD Linux Networking Microservices Cloud Computing Observability

Requirements

Candidates must have 6–8 years of experience in SRE or DevOps roles with strong hands-on expertise in Kubernetes, Docker, and cloud platforms. Proficiency in infrastructure as code, scripting languages like Python or Go, and a solid understanding of Linux and networking fundamentals are required.

Benefits

Healthcare Benefit Package

Full description

Charger Logistics Inc. is a leading asset-based transportation company with over 20 years of experience delivering innovative logistics solutions. We have evolved into a world-class transport provider and continue to expand across North America.

We're looking for a Site Reliability Engineer to keep our platform fast, stable and always available. Our systems run 24/7 to support trucking and logistics operations across North America. They're built on more than 700 .NET Core microservices running on Kubernetes with PostgreSQL. You'll work closely with our development, DevOps and QA teams in Canada to improve reliability, automate operations and respond to production issues.

Please note this is a Contract opportunity for potential to convert in a permanent role..

Responsibilities:

  • Keep our production platform reliable, scalable and highly available.
  • Define and track SLIs, SLOs and error budgets for critical services.
  • Build and improve monitoring, alerting, logging and tracing using Prometheus, Grafana, Loki, Jaeger and OpenTelemetry.
  • Manage and tune Kubernetes clusters and containerized workloads.
  • Automate repetitive operational work using Python, Go or Bash.
  • Build and maintain infrastructure as code with Terraform and Helm.
  • Support and improve CI/CD pipelines for safe, frequent deployments, including blue/green and canary rollouts.
  • Take part in an on-call rotation. Lead incident response, troubleshoot production issues and run blameless post-mortems.
  • Monitor and tune PostgreSQL performance, backups and high availability with the development team.
  • Plan capacity, test resilience and help reduce cloud costs.
  • 6–8 years in SRE, DevOps or production infrastructure roles.
  • Strong hands-on experience with Kubernetes and Docker in production.
  • Experience with at least one major cloud platform (AWS, Azure or GCP).
  • Hands-on with monitoring and observability tools (Prometheus, Grafana, ELK/Loki, Jaeger or similar).
  • Infrastructure as code with Terraform, Helm or similar.
  • Scripting or coding skills in Python, Go or Bash.
  • Strong Linux and networking fundamentals (DNS, load balancing, TCP/IP, TLS).
  • Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins or Azure DevOps.
  • A solid understanding of microservices and distributed systems.
  • Experience handling production incidents and writing root-cause analyses.

Nice to have

  • Service mesh experience (Istio, Consul or Linkerd)
  • PostgreSQL administration and performance tuning
  • Familiarity with .NET Core applications in production
  • Messaging systems such as Kafka or RabbitMQ
  • Certifications such as CKA, CKAD, or AWS/Azure DevOps Engineer
  • Experience in logistics, transportation or another 24/7 operations environment is a plus
  • Competitive Salary
  • Healthcare Benefit Package
  • Career Growth

Similar roles