Senior Machine Learning Engineer, Speech & LLM Training Data
Propio Overland Park, Kansas, United States
Translation and Localization · 201-500 employees
About the role
The role involves defining data roadmaps and building scalable pipelines for multilingual speech, translation, and multimodal LLM training. You will manage end-to-end data workflows including curation, annotation, model training, and performance evaluation.
What they look for
Requirements
Candidates must have a Bachelor's or Master's degree in a technical field and at least 5 years of experience in ML engineering or speech/audio data workflows. Proficiency in Python, SQL, and cloud-based ML infrastructure is required, along with experience in speech-processing tasks and large-scale dataset management.
Full description
Description
Propio Language Services is one of the top 5 providers high-quality, real-time multilingual interpretation, translation, and localization services, operating at 9-figure scale across healthcare, legal, and other industries. We are driven by a passion for cutting-edge technology and exceptional service, building seamless experiences that bridge communication gaps across languages, cultures, and modalities.
Propio is hiring a Senior Machine Learning Engineer, Speech & LLM Training Data to transform large volumes of multilingual conversational audio into high-quality training and evaluation datasets. This hands-on role owns audio processing, dataset curation, annotation and QA workflows, model training, and evaluation for our multilingual speech, translation, and conversational AI systems.
Key Responsibilities:
- Define the data roadmap for multilingual speech, translation, multimodal LLMs, and conversational AI.
- Build audio-processing pipelines covering resampling, channel handling, VAD, diarization, language identification, transcription, alignment, and quality filtering.
- Build dataset pipelines for cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning, and lineage.
- Design annotation guidelines, QA rubrics, golden datasets, and reviewer workflows.
- Build evaluation datasets, analyze model failures, and translate performance gaps into targeted data improvements.
- Run training, fine-tuning, post-training, and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows, and synthetic data generation.
- Productionize secure, traceable, and reproducible data and ML workflows on AWS.
Requirements
Qualifications:
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, Computational Linguistics, or a related field, or equivalent practical experience.
- 5+ years of experience in ML engineering, speech/audio ML, ML data engineering, NLP, or LLM training-data workflows.
- Strong hands-on experience with Python, SQL, Linux, Git, and Docker.
- Experience training or evaluating models using PyTorch, Hugging Face, or comparable ML frameworks.
- Experience with FFmpeg and audio-processing libraries such as TorchCodec, torchaudio, librosa, or equivalent tools.
- Experience with speech-processing tasks such as VAD, diarization, ASR, forced alignment, language identification, and audio-quality analysis.
- Experience with Databricks/Spark, Parquet/Arrow, and large-scale dataset pipelines.
- Working knowledge of AWS S3, SageMaker, Glue, Step Functions, IAM, and KMS.
- Experience with an annotation platform such as Labelbox, Label Studio, Scale AI, Prodigy, Argilla, or custom internal tooling.
- Experience with experiment tracking and data versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake, or LakeFS.
- Experience with multilingual speech, translation, annotation workflows, and evaluation datasets.
Preferred Qualifications:
- Experience with multilingual telephony, healthcare, interpretation, or call-center audio.
- Experience with tools such as Silero VAD, pyannote, WhisperX, NeMo, Kaldi, or equivalent speech technologies.
- Experience with distributed processing or training using Ray, PySpark, or similar frameworks.
- Experience with HIPAA, PHI/PII redaction, and secure data governance.
- Experience with low-resource languages, accents, dialects, and code-switching.
- Experience with synthetic data, active learning, weak supervision, or LLM-as-judge evaluation.
#LI-JS1
Similar roles
-
Senior/Staff Engineer, Machine Learning - Online Mapping
Nuro Mountain View, California, United States · $194K–$291K/yr
-
Staff Machine Learning Engineer, Mapping
Waymo Mountain View, California, United States · $251K–$310K/yr
-
Senior Engineering Manager, Machine Learning
Checkr San Francisco, California, United States · $268K–$315K/yr
-
Software Engineering Manager (Machine Learning)
Cedar United States · $196K–$247K/yr
-
Machine Learning Scientist II, Drug Discovery Analytics
Revolution Medicines Redwood City, California, United States · $202K–$238K/yr
-
Senior Machine Learning Engineer
Air Pittsburgh, Pennsylvania, United States