Senior DevOps Engineer
Lean Solutions Group Colombia
Technology, Information and Internet · 1,001-5,000 employees
About the role
You will design, build, and maintain scalable cloud infrastructure specifically optimized for AI and machine learning workloads. This involves collaborating with data scientists and AI engineers to manage CI/CD pipelines, infrastructure as code, and system observability.
What they look for
Requirements
The role requires at least 5 years of experience in cloud infrastructure and DevOps, with a strong preference for AWS expertise. Candidates must also possess at least 2 years of practical experience working with AI/ML systems and containerization technologies.
Full description
Company Overview:
Lean Tech is a rapidly expanding organization situated in Medellín, Colombia. We pride ourselves on possessing one of the most influential networks within software development and IT services for the entertainment, financial, and logistics sectors. Our corporate projections offer a multitude of opportunities for professionals to elevate their careers and experience substantial growth. Joining our team means engaging with expansive engineering teams across Latin America and the United States, contributing to cutting-edge developments in multiple industries.
Position Title: Senior DevOps / Infrastructure Engineer (AI/ML Focus) Category: Infrastructure & Cloud Operations Seniority: Senior
Location: LATAM (Remote)
What you will be doing: We are looking for a Senior DevOps/Infrastructure Engineer to design, build, and maintain the cloud infrastructure that powers our AI and machine learning solutions. Sitting at the intersection of infrastructure engineering and applied AI, you will work closely with AI engineers, data scientists, and product teams to ensure our infrastructure is scalable, secure, and optimized for AI workloads. This position requires a strong cloud infrastructure background alongside practical experience building and operating AI/ML systems beyond consuming AI tools.
Key Responsibilities:
- Design, implement, and maintain scalable, secure, and highly available cloud infrastructure, primarily on AWS.
- Build and manage CI/CD pipelines to support fast, reliable deployments across all environments.
- Provision and manage infrastructure using Infrastructure as Code (Terraform, CloudFormation, or similar).
- Support infrastructure needs specific to AI/ML workloads, including GPU resources, model training pipelines, and inference environments.
- Collaborate with AI/ML engineers on the infrastructure required to train, deploy, and monitor machine learning models.
- Implement observability practices (logging, monitoring, alerting) across services and pipelines.
- Own security and compliance practices for cloud resources, including secrets management, IAM policies, and network configuration.
- Troubleshoot and resolve infrastructure and deployment issues across production and non-production environments.
- Continuously evaluate and improve infrastructure reliability, cost efficiency, and performance.
- Document architecture decisions, runbooks, and operational procedures.
Required Skills & Experience:
- 5+ years of experience in cloud infrastructure, DevOps, or a closely related field.
- Proven hands-on experience with a major cloud provider (AWS strongly preferred).
- 2+ years of experience working with AI or machine learning tools and systems on the development side (building, training, deploying, or operating ML models/AI pipelines).
- Strong experience with Infrastructure as Code (Terraform, CloudFormation, Pulumi, or similar).
- Solid background in containerization and orchestration (Docker, Kubernetes, ECS, or similar).
- Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar).
- Strong understanding of networking, security, and identity management in cloud environments.
- Experience with monitoring and observability tools (CloudWatch, Grafana, Datadog, Prometheus, or similar).
- Comfortable working in Linux environments and scripting in Python, Bash, or similar languages.
- Strong communication skills with the ability to work cross-functionally with engineering, data science, and product teams.
Nice to Have Skills:
- Experience with MLOps practices and tools (MLflow, SageMaker, Kubeflow, or similar).
- Familiarity with vector databases, model serving frameworks, or LLM infrastructure.
- Experience with serverless architectures (Lambda, Fargate).
- Background supporting data pipelines or data engineering workflows.
- Relevant certifications (AWS Certified Solutions Architect, AWS Certified DevOps Engineer, or similar).
Similar roles
-
DevOps/ Platform Engineer (m/w/d)
PAUL Tech AG Mannheim, Baden-Württemberg, Germany
-
(9700) DevOps Engineer
IAMUS Consulting Fort Meade, Maryland, United States · $122K–$161K/yr
-
DevOps Engineer I (Remote)
Businessolver United States · $75K–$115K/yr
-
DevOps Engineer - SE II
Keywords Studios Pune, Maharashtra, India
-
Platform / DevOps Engineer (m/f/d) — CI/CD & Developer Environments
Autonomous Teaming Solutions ATS GmbH Munich, Bavaria, Germany
-
DevOps Track Sr.Engineer
HEXAWARE India