Machine Learning Engineer - Voice AI & Generative Music (Part-time)
MWDN Kyiv, Ukraine
IT Services and IT Consulting · 51-200 employees
About the role
The engineer will own the end-to-end workflow for singing voice conversion, text-to-speech model deployment, and Spanish lyrics generation. This involves fine-tuning models, optimizing inference, and integrating solutions into the existing production platform.
What they look for
Requirements
Candidates must have 3+ years of experience in deep learning, specifically in audio or NLP, with strong proficiency in PyTorch. Native or professional proficiency in Spanish is required for quality assurance and lyrics evaluation.
Benefits
Full description
MWDN is a global IT outstaffing company with 23+ years of experience that connects exceptional tech talent with leading companies across Israel, the USA, Great Britain, and Western Europe. We offer opportunities to work on international products in a stable and professional environment.
Why does MWDN rock?
Here’s what you can expect when you join MWDN:
- Security: We carefully vet our clients to minimize risks and ensure reliability and timely payments - no fraud or unpleasant surprises.
- Career support: If a project isn’t the right fit, we support you and actively help find new opportunities that match your skills and career goals.
- Legal assistance: We provide guidance on legal matters, including opening and managing your independent contractor or sole proprietorship status, taxes, and related processes.
- Professional development: We offer English courses and professional growth opportunities, as well as team-building events.
Why choose us? MWDN is ranked among the top 5 IT employers in our region according to DOU. We take pride in our transparency and strong commitment to our team. Curious to learn more? See what our employees say about working with us on DOU.
What is your new project?
Domain: Music Technology / Voice AI / Generative AI
An AI-native music label developing advanced tools for producing and releasing Spanish-language music. Artificial intelligence serves as the foundation of its production process, powering artist-consistent audio generation, vocal cloning, custom text-to-speech, and generative lyrics.
The company has built its own end-to-end music production pipeline covering audio data preparation, model training, vocal conditioning, inference, quality evaluation, and deployment. The environment is hands-on, fast-moving, and focused on turning sophisticated AI models into release-ready music products rather than experimental prototypes.
What makes this project exciting?
Our client is the technology-driven music label where artificial intelligence powers the entire content-production process. The company develops artist-specific audio, creates custom vocal models, and operates a proprietary system designed to produce and release music faster than conventional labels.
Each track moves through an internally developed workflow that covers audio data preparation, model training, vocal conditioning, generation, quality control, and production deployment.
We are looking for an Applied Machine Learning Engineer specializing in audio and generative AI to take ownership of this workflow. This is a practical engineering position rather than a research, analytics, or prompt-only role. The ideal candidate has already built and launched complete ML systems, taking them from raw datasets and training experiments through optimized inference and production use.
What makes you a great fit
- 3+ years of experience training and deploying deep learning models in production, preferably within audio or NLP.
- Hands-on experience in at least two of the following areas: Voice cloning or Singing Voice Conversion (SVC), Text-to-Speech model fine-tuning, LLM fine-tuning using techniques such as LoRA, DPO, and RAG.
- Strong proficiency in PyTorch.
- Practical experience with GPU cloud infrastructure, such as RunPod or AWS, as well as Docker and serverless inference.
- A rigorous, evidence-based approach to model evaluation, including benchmarks, ablation studies, and blind testing.
- Native or strong professional proficiency in Spanish, required for lyrics evaluation and voice quality assurance.
Nice to Have
- A music background or experience with music-production tools and workflows, including stems, MIDI, and DAWs.
- Experience with singing-voice synthesis solutions such as ACE Studio, ACE-Step, RVC, so-vits-svc, or similar technologies.
- Previous experience working with licensed celebrity or artist voices and consent-based voice AI.
Your day-to-day in this position
1.Custom Singing Voice Model
- Fine-tune and improve an existing singing-voice conversion and cloning model, increasing its current 70–75% fidelity baseline to production-level quality of 90% or higher in blind listening tests.
- Reproduce expressive vocal characteristics beyond basic timbre, including vibrato, falsetto, dynamics, and both spoken and sung delivery.
- Evaluate different improvement strategies, including expanding the training dataset, testing alternative base models, and assessing enterprise APIs that offer no-training guarantees.
2. Text-to-Speech Model
- Take ownership of an existing validated LoRA fine-tune based on Chatterbox Multilingual for a Spanish-speaking voice.
- Ensure accurate reproduction of a Mexican accent and correct pronunciation of sounds and letters such as J, Ñ, and X.
- Containerize the model and deploy it as a serverless inference endpoint using RunPod or a similar platform.
- Integrate the endpoint with the existing web platform through an API.
- Train a second version using clean studio recordings to eliminate the remaining output instability.
3. ALMA — Spanish Lyrics Model
- Develop a proprietary Spanish-language lyrics-generation model, with experience in regional Mexican genres considered a strong advantage.
- Build the solution using an open-weight LLM, LoRA fine-tuning, preference optimization through DPO, and RAG based on a curated content corpus.
- Establish clear evaluation criteria and organize quality assessment with native Spanish speakers.
- Integrate the completed model into the existing frontend.
Why work with us?
- People-first management with minimal bureaucracy
- A friendly company culture, proven by employees who choose to return
- Flexible working hours
- 29 days of PTO (18 working days per year pluse all national holidays)
- 10 paid recovery days
- Full financial and legal support for independent contractors
- Free English classes, with native speakers or Ukrainian teachers
- Dedicated HR support
Our next steps
✅ Intro call with a Recruiter — ✅ Intro call with a client — ✅ Final client interview— ✅ Offer
Requirements
null
Similar roles
-
Machine Learning Engineer - Multimodal Intelligence
Apple Sunnyvale, California, United States
-
Senior Machine Learning Engineer (W/M/X)
Ubisoft Saint-Mandé, Ile-de-France, France
-
Staff Machine Learning Scientist/Engineer
Wayve Sunnyvale, California, United States · $370K–$419K/yr
-
Senior Machine Learning Engineer, Ads Response Prediction
Instacart Wasaga Beach, Ontario, Canada · CA$180K–CA$190K/yr
-
AI / Machine Learning Engineer
Lynx Madrid, Community of Madrid, Spain
-
Vice President, AI / Machine Learning Data Engineer
BNY Manchester, England, United Kingdom