Ford Motor Company

Data Scientist

Ford Motor Company Chennai, Tamil Nadu, India

Motor Vehicle Manufacturing · 10,001+ employees

6 h ago Closes in 4d
data-scientist Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Develop and maintain AI pipelines involving text, image, and audio modalities while managing the full lifecycle of Gen AI models. Implement real-world conversational agents and generative tools at scale using advanced NLP and prompt engineering techniques.

What they look for

Natural Language Processing Large Language Models Generative AI Python Tensorflow Pytorch Docker Kubernetes GCP BigQuery VertexAI Spark Transformers Prompt Engineering CI/CD Data Engineering

Requirements

Requires a Bachelor’s or Master’s degree in Computer Science, Engineering, or related fields with strong experience in NLP, LLMs, and cloud infrastructure. Candidates must demonstrate proficiency in Python, containerization tools, and real-world deployment of transformer-based models.

Full description

ML/DL Skills:

• High familiarity in the use of DL theory/practices in NLP applications

• Comfort level to code in ADK, A2A, AgentSkills, Ontology, Huggingface, LangGraph, LangChain, Chainlit, Tensorflow and/or Pytorch, Scikit-learn, Numpy and Pandas

• Comfort level to use two/more of open source NLP modules like SpaCy, TorchText, fastai.text, farm-haystack, and others

NLP Skills:

• Knowledge in fundamental text data processing (like use of regex, token/word analysis, spelling correction/noise reduction in text, segmenting noisy unfamiliar sentences/phrases at right places, deriving insights from clustering, etc.,)

• Have implemented in real-world BERT/or other transformer fine-tuned models (Seq classification, NER or QA) from data preparation, model creation and inference till deployment

Python Project Management Skills

• Familiarity in the use of Docker tools, pipenv/conda/poetry env

• Comfort level in following Python project management best practices (use of setup.py, logging, pytests, relative module imports,sphinx docs,etc.,)

• Familiarity in use of Github (clone, fetch, pull/push,raising issues and PR, etc.,)

Cloud Skills and Computing:

• Use of GCP services like BigQuery, Cloud function, Cloud run, Cloud Build, VertexAI,

• Good working knowledge on other open source packages to benchmark and derive summary

• Experience in using GPU/CPU of cloud and on-prem infrastructures

• Skillset to leverage cloud platform for Data Engineering, Big Data and ML needs.

Deployment Skills:

• Use of Dockers (experience in experimental docker features, docker-compose, etc.,)

• Familiarity with orchestration tools such as airflow, Kubeflow

• Experience in CI/CD, infrastructure as code tools like terraform etc.

• Kubernetes or any other containerization tool with experience in Helm, Argoworkflow, etc.,

• Ability to develop APIs with compliance, ethical, secure and safe AI tools.

UI:

• Good UI skills to visualize and build better applications using Gradio, Dash, Streamlit, React, Django, etc.,

• Deeper understanding of javascript, css, angular, html, etc., is a plus.

Data Engineering:

• Skillsets to perform distributed computing (specifically parallelism and scalability in Data Processing, Modeling and Inferencing through Spark, Dask, RapidsAI or RapidscuDF)

• Ability to build python-based APIs (e.g.: use of FastAPIs/ Flask/ Django for APIs)

• Experience in Elastic Search and Apache Solr is a plus, vector databases.

Similar roles