E

Senior Machine Learning Ops Engineer

Ensign Infosecurity Singapore, Singapore, Singapore

Jul 31
machine-learning Senior (5-10 yrs) Full-time Singapore
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will own the design, development, and maintenance of the in-house AIOps, ML, and LLM platform across cloud and on-premise environments. Additionally, you will build and operate production ML/LLM workflows while troubleshooting complex issues across infrastructure and application layers.

What they look for

Machine Learning LLM Kubernetes Python Go C++ Linux Networking Cloud Computing CI/CD MLOps LLMOps System Design Distributed Systems Observability API Design

Requirements

Candidates must have strong software engineering fundamentals and practical experience with the ML/LLM lifecycle, including data pipelines and model deployment. Proficiency in Python, Linux, Kubernetes, and networking is required, along with experience in building production-grade MLOps/LLMOps workflows.

Full description

Ensign is hiring !

Key Responsibilities

- Own the design, development, maintenance, and evolution of the in-house AIOps / ML / LLM platform, including related cloud and on-premise Kubernetes solutions.

- Translate client, security, compliance, and internal requirements into practical platform designs with cross-functional teams.

- Build and operate production ML / LLM workflows, including retraining, deployment, inference serving, monitoring, rollback, and optimisation.

- Troubleshoot production issues across application, infrastructure, networking, Linux, Kubernetes, and ML serving layers.

Qualifications / Requirements

- Strong software/platform engineering fundamentals, including system design, API design, distributed systems, scalability, reliability, observability, authentication/authorization, testing, and maintainable code design.

- Practical understanding of the ML / LLM lifecycle, including data pipelines, model training/retraining, evaluation, experiment tracking, deployment, monitoring, and production feedback loops.

- Strong development experience in Python, with working proficiency in Go and C++ for reading, debugging, maintaining, and extending existing production codebases.

- Strong Linux, networking, and Kubernetes fundamentals, including production troubleshooting, service connectivity, ingress, resource limits, workload debugging, and deployment operations.

- Experience designing, deploying, and operating production platforms on AWS, Azure, GCP, or on-premise environments.

- Experience building CI/CD, automation, and MLOps / LLMOps workflows for production ML / LLM systems.

- Strong communication skills and ability to work with AI, deployment, infrastructure, and security teams.

Good to Have

- Deep experience operating Kubernetes in bare-metal, air-gapped, or restricted on-premise environments.

- Experience with MLflow, Kubeflow, vLLM, TensorRT, TGI, or similar ML / LLM platform tools.

- Exposure to TypeScript / React or Java-based services.

Similar roles