Data Scientist
Ford Motor Company Chennai, Tamil Nadu, India
Motor Vehicle Manufacturing · 10,001+ employees
About the role
Develop and maintain AI pipelines involving text, image, and audio modalities while managing the full lifecycle of Gen AI models. Implement real-world conversational agents and generative tools at scale using advanced NLP and prompt engineering techniques.
What they look for
Requirements
Requires a Bachelor’s or Master’s degree in Computer Science, Engineering, or related fields with strong experience in NLP, LLMs, and cloud infrastructure. Candidates must demonstrate proficiency in Python, containerization tools, and real-world deployment of transformer-based models.
Full description
ML/DL Skills:
• High familiarity in the use of DL theory/practices in NLP applications
• Comfort level to code in ADK, A2A, AgentSkills, Ontology, Huggingface, LangGraph, LangChain, Chainlit, Tensorflow and/or Pytorch, Scikit-learn, Numpy and Pandas
• Comfort level to use two/more of open source NLP modules like SpaCy, TorchText, fastai.text, farm-haystack, and others
NLP Skills:
• Knowledge in fundamental text data processing (like use of regex, token/word analysis, spelling correction/noise reduction in text, segmenting noisy unfamiliar sentences/phrases at right places, deriving insights from clustering, etc.,)
• Have implemented in real-world BERT/or other transformer fine-tuned models (Seq classification, NER or QA) from data preparation, model creation and inference till deployment
Python Project Management Skills
• Familiarity in the use of Docker tools, pipenv/conda/poetry env
• Comfort level in following Python project management best practices (use of setup.py, logging, pytests, relative module imports,sphinx docs,etc.,)
• Familiarity in use of Github (clone, fetch, pull/push,raising issues and PR, etc.,)
Cloud Skills and Computing:
• Use of GCP services like BigQuery, Cloud function, Cloud run, Cloud Build, VertexAI,
• Good working knowledge on other open source packages to benchmark and derive summary
• Experience in using GPU/CPU of cloud and on-prem infrastructures
• Skillset to leverage cloud platform for Data Engineering, Big Data and ML needs.
Deployment Skills:
• Use of Dockers (experience in experimental docker features, docker-compose, etc.,)
• Familiarity with orchestration tools such as airflow, Kubeflow
• Experience in CI/CD, infrastructure as code tools like terraform etc.
• Kubernetes or any other containerization tool with experience in Helm, Argoworkflow, etc.,
• Ability to develop APIs with compliance, ethical, secure and safe AI tools.
UI:
• Good UI skills to visualize and build better applications using Gradio, Dash, Streamlit, React, Django, etc.,
• Deeper understanding of javascript, css, angular, html, etc., is a plus.
Data Engineering:
• Skillsets to perform distributed computing (specifically parallelism and scalability in Data Processing, Modeling and Inferencing through Spark, Dask, RapidsAI or RapidscuDF)
• Ability to build python-based APIs (e.g.: use of FastAPIs/ Flask/ Django for APIs)
• Experience in Elastic Search and Apache Solr is a plus, vector databases.
Similar roles
-
Business Data Scientist, gTech Users and Products
Google Dublin, Leinster, Ireland · $138K–$197K/yr
-
Senior Product Data Scientist, Consumer Payments
Google Singapore
-
Environmental Data Scientist
Uni Systems Ispra, Lombardy, Italy
-
Senior Data Scientist, Member Lifecycle
Vinted Vilnius, Vilnius County, Lithuania · €55K–€75K/yr
-
Senior Data Scientist / Analytics Manager - Pharmaceutical domain
WNS Global Services Gurgaon, Haryana, India
-
Tax Technology and Transformation - Advanced Technologies - Data Scientist - Staff
EY New York, New York, United States · $77K–$143K/yr