Machine Learning Operations Engineer
IPolarity LLC · Hanover Township, New Jersey, United States · $114K–$125K/yr
Staffing and Recruiting · 51-200 employees
About the role
The MLOps Engineer will design, build, and maintain scalable machine learning deployment pipelines and infrastructure. They are responsible for monitoring production models for drift, performance, and operational costs while ensuring robust governance and security.
What they look for
Requirements
Candidates must have 8+ years of experience in AI infrastructure and 3+ years supporting production machine learning platforms. A Bachelor's or Master's degree in a technical field is required along with strong proficiency in Python, SQL, and cloud-based ML tools.
Full description
Machine Learning Operations Engineer (Rate:$60 ) Location: Newark NJ
We are seeking a Machine Learning Operations (MLOps) Engineer to join our team. The MLOps Engineer will be responsible for building and maintaining the infrastructure that enables reliable deployment, monitoring, governance, and continuous improvement of production machine learning systems across enterprise client environments.
What You'll Do:
- Design, build, and maintain scalable machine learning deployment pipelines.
- Develop standardized model registries, artifact repositories, data versioning, and reproducible ML environments.
- Build automated evaluation pipelines for production machine learning models.
- Implement automated data quality monitoring including profiling, anomaly detection, validation, quarantine, and alerting.
- Develop automated retraining workflows, promotion gates, rollback capabilities, and audit trails.
- Monitor production environments for model drift, latency, prediction quality, infrastructure performance, and operational costs.
- Troubleshoot production machine learning issues and lead incident response activities.
- Build CI/CD pipelines supporting enterprise AI applications.
- Collaborate closely with Data Scientists and client engineering teams to deploy and maintain AI solutions.
- Ensure governance, security, lineage, reproducibility, and audit readiness across machine learning platforms.
Who You Are:
- 8+ years - Passionate about building reliable AI infrastructure at enterprise scale.
- Experienced deploying and maintaining production machine learning systems.
- Strong analytical and troubleshooting skills.
- Fast learner with attention to detail.
- Excellent communication and collaboration skills.
- Comfortable working with both software engineering and data science teams.
Education: Bachelor's or Master's degree in Computer Science, Software Engineering, Data Engineering, or a related technical field.
Related Work Experience: 3+ years supporting production machine learning platforms or cloud infrastructure.
Technical Skills:
- Advanced SQL
- Python
- Apache Spark / PySpark
- AWS SageMaker (Databricks, Azure ML, or Vertex AI experience is a plus)
- Kubernetes
- CI/CD pipelines
- Infrastructure as Code (Terraform, CloudFormation, or similar)
- Model monitoring and ML observability tools
- Data validation and automated testing frameworks
- Statistics related to monitoring, model drift, and performance evaluation
• • Git and modern DevOps practices