CloudFactory

Senior SRE

CloudFactory Canada

Technology, Information and Internet · 1,001-5,000 employees

7 h ago
Remote sre Senior (5-10 yrs) Full-time Canada
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design and build scalable infrastructure while developing systems and pipelines to support automation and reliability. You will also collaborate with cross-functional teams to maintain production system health and communicate technical issues to stakeholders.

What they look for

Python Docker Kubernetes GCP AWS Terraform CI/CD Infrastructure as Code Site reliability engineering Observability Automation Monitoring Alerting System architecture Deployment pipelines

Requirements

Candidates must have at least 5 years of experience in production infrastructure and proficiency in Python, Docker, Kubernetes, and Terraform. A degree in Computer Science, Engineering, or a related quantitative field is required.

Benefits

Market competitive salary Quarterly variable compensation Hybrid working model Comprehensive medical cover Group life insurance Personal development and growth opportunities

Full description

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale. 

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Senior SRE, you will design and build scalable infrastructure, working closely with cross-functional teams to develop systems and pipelines that support the automation, reliability, and scalability of our production environments. You will bring a high degree of autonomy to designing new infrastructure components and applying site-reliability practices across our systems, while communicating complex technical issues clearly to stakeholders across the business. This is an exciting opportunity to make a real impact while working alongside talented people from developing nations.

Please note: This is a full-time, fixed-term employee position with an expected duration of 6 months.

Responsibilities:

Infrastructure design and automation

  • Design and implement new core infrastructure components with a high degree of autonomy.
  • Optimize and improve existing systems and operations, such as deployment pipelines, environment provisioning, and high-throughput batch jobs.
  • Use Infrastructure as Code (IaC) tools, such as Terraform, to manage and scale complex infrastructure.

CI/CD automation

  • Develop CI/CD pipelines to automate build, test, deployment, and monitoring processes.
  • Create and manage multi-step CI/CD pipelines, including environment setup and artifact handling.

Reliability and availability

  • Support the reliability, availability, and performance of production systems, applying site-reliability practices across the infrastructure.
  • Set up monitoring, alerting, and observability tooling to maintain visibility into system health.

Collaboration and communication

  • Collaborate closely with software engineers, product, and business stakeholders on the design and delivery of infrastructure and deployment systems.
  • Communicate complex technical issues clearly to stakeholders from technical and non-technical backgrounds alike.

Must-have skills (required)

  • 5+ years of experience building and operating infrastructure in production environments.
  • Fluent in Python, with strong experience writing production-ready code.
  • Experience with Docker and Kubernetes.
  • Knowledgeable about cloud platforms such as GCP or AWS.
  • Experience using Infrastructure as Code (IaC) tools such as Terraform.
  • Experience using CI/CD platforms to automate build, test, and deployment pipelines.
  • Comfortable applying site-reliability principles, such as availability, observability, and automation, across production systems.

Academic and professional requirements

  • Degree in Computer Science, Engineering, or another quantitative or computational field, or equivalent practical experience.

Nice-to-have skills (preferred)

  • Familiarity with monitoring tools such as Prometheus or Grafana.
  • Experience with configuration management tools (e.g., Ansible, Chef, Puppet).
  • Exposure to multi-cloud or hybrid-cloud environments.
  • Great Mission and Culture
  • Meaningful Work
  • Market competitive salary
  • Quarterly variable compensation
  • Hybrid Working Model
  • Comprehensive medical cover 
  • Group life insurance
  • Personal development and growth opportunities

At CloudFactory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Join us today and be part of our mission to connect people and technology for a better world! Apply now and bring your whole, authentic self to work—we can’t wait to meet you!

Similar roles