Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
You will own the design, development, and maintenance of the in-house AIOps, ML, and LLM platform across cloud and on-premise environments. Additionally, you will build and operate production ML/LLM workflows while troubleshooting complex issues across infrastructure and application layers.
What they look for
Requirements
Candidates must have strong software engineering fundamentals and practical experience with the ML/LLM lifecycle, including data pipelines and model deployment. Proficiency in Python, Linux, Kubernetes, and networking is required, along with experience in building production-grade MLOps/LLMOps workflows.
Full description
Ensign is hiring !
Key Responsibilities
- Own the design, development, maintenance, and evolution of the in-house AIOps / ML / LLM platform, including related cloud and on-premise Kubernetes solutions.
- Translate client, security, compliance, and internal requirements into practical platform designs with cross-functional teams.
- Build and operate production ML / LLM workflows, including retraining, deployment, inference serving, monitoring, rollback, and optimisation.
- Troubleshoot production issues across application, infrastructure, networking, Linux, Kubernetes, and ML serving layers.
Qualifications / Requirements
- Strong software/platform engineering fundamentals, including system design, API design, distributed systems, scalability, reliability, observability, authentication/authorization, testing, and maintainable code design.
- Practical understanding of the ML / LLM lifecycle, including data pipelines, model training/retraining, evaluation, experiment tracking, deployment, monitoring, and production feedback loops.
- Strong development experience in Python, with working proficiency in Go and C++ for reading, debugging, maintaining, and extending existing production codebases.
- Strong Linux, networking, and Kubernetes fundamentals, including production troubleshooting, service connectivity, ingress, resource limits, workload debugging, and deployment operations.
- Experience designing, deploying, and operating production platforms on AWS, Azure, GCP, or on-premise environments.
- Experience building CI/CD, automation, and MLOps / LLMOps workflows for production ML / LLM systems.
- Strong communication skills and ability to work with AI, deployment, infrastructure, and security teams.
Good to Have
- Deep experience operating Kubernetes in bare-metal, air-gapped, or restricted on-premise environments.
- Experience with MLflow, Kubeflow, vLLM, TensorRT, TGI, or similar ML / LLM platform tools.
- Exposure to TypeScript / React or Java-based services.
Similar roles
-
Principal AI Engineer, Machine Learning Operations
BlackSky United States · $185K–$215K/yr
-
Mechanical Engineer, Annapurna Labs, Machine Learning Hardware
Amazon Austin, Texas, United States · $136K–$184K/yr
-
Machine Learning Engineer (Model Dev)
Artera United States · $140K–$180K/yr
-
Machine Learning Engineer | Senior
Jobgether Brazil
-
Researcher- Applied AI and Machine Learning
Cohu San Diego County, California, United States · $125K–$175K/yr
-
Senior Machine Learning Engineer
Checkr San Francisco, California, United States · $176K–$244K/yr