Data Engineer
Ford Motor Company Chennai, Tamil Nadu, India
Motor Vehicle Manufacturing · 10,001+ employees
About the role
Design and maintain scalable data pipelines using Google Cloud Platform to process high-volume telematics and vehicle health data. Ensure data quality and reliability through automated testing, monitoring, and proactive incident management.
What they look for
Requirements
Requires 3-7+ years of hands-on data engineering experience with a strong background in GCP, Python, and SQL. Candidates must hold a Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field.
Full description
The ODIA Data Engineering team is the backbone of Ford’s Vehicle Care Platform (VCP) and Ford APP. We are responsible for ingesting, transforming, and modeling vast volumes of connected vehicle telematics, Vehicle health alerts, and Repair Order Data.
As a Data Engineer on this team, you will build and maintain mission-critical batch and near-real-time data pipelines (running on cadences from 15 minutes to daily batches), build alert-to-repair-order correlation models, drive data quality frameworks (Great Expectations / DQM), and support multi-market international rollouts (North America, Europe, IMG).
Responsibilities
Data Pipeline Architecture & Engineering
- Design and maintain scalable Airflow DAGs and data pipelines (GCP Cloud Composer, Dataproc, BigQuery) processing multi-million record datasets across Tier 1 and Tier 2 analytics layers.
- Build idempotent, fault-tolerant data ingestion logic capable of graceful retry recovery, backfilling, and strict execution-date time-windowing.
- Develop complex business transformations and correlation engines, such as Alert-to-Repair-Order attribution, VIN-to-Dealer PA code mappings, and multi-market localization logic.
- Optimize BigQuery performance through partitioning, clustering, materialized/authorized views, and query cost optimization.
Data Quality, Reliability & Observability
- Implement and expand automated data quality suites using Great Expectations and custom DQM (Data Quality Management) dashboards to proactively catch schema drift, data anomalies, and orphan records.
- Instrument centralized logging and monitoring using GCP Cloud Logging, Cloud Monitoring, and Dynatrace for end-to-end pipeline visibility.
- Participate in production support, incident management, and Root Cause Analysis (RCA) for DAG failures, ensuring high availability and data SLAs.
Compliance, Security & Platform Operations
- Ensure strict data privacy compliance across international markets (e.g., CCPA, GDPR, CNIL, ICO), managing customer consent state pipelines and regulatory alert deletion/response flows.
- Contribute to Business Continuity and Disaster Recovery (BCP/DR) plans and support compliance audits (ISO certifications).
- Maintain CI/CD automation pipelines for data artifacts using GitHub Actions, SonarQube, and infrastructure as code.
Cross-Functional Collaboration
- Collaborate closely with VCP software engineering teams, Enterprise Data Architecture (CDW/CDM), Product Managers, and global market stakeholders.
- Produce clear technical documentation, runbooks, and data dictionaries on Confluence.
Qualifications
Required Qualifications:
- Experience: 3–7+ years of hands-on data engineering experience building production-grade data pipelines.
- Cloud & Big Data: Strong working experience with Google Cloud Platform (GCP)—specifically BigQuery, GCS, Cloud Composer / Apache Airflow, and Dataproc.
- Programming & SQL: Expert-level proficiency in SQL (complex joins, window functions, query optimization) and Python (data manipulation, ETL scripting, API integration).
- Pipeline Orchestration: Deep understanding of Apache Airflow best practices (DAG design, custom operators, sensors, idempotent task execution, backfills, dynamic task mapping).
- Data Modeling: Demonstrated experience with data warehousing concepts (Tiered architecture / medallion architecture, dimensional modeling, star/snowflake schemas).
- Education: Bachelor’s or Master’s degree in Computer Science, Data Engineering, Information Technology, or equivalent practical experience.
Tech Stack Snapshot
- Cloud Platform: Google Cloud Platform (GCP) — BigQuery, Cloud Storage (GCS), Dataproc, Cloud Composer / Apache Airflow, Cloud Logging.
- Languages & Querying: Python, SQL (GoogleSQL/BigQuery SQL), PySpark / Spark.
- Data Quality & Testing: Great Expectations, DQM Dashboards, automated data validation suites.
- DevOps & Tools: Git, GitHub Actions / Tekton, SonarQube, Docker, Hoppscotch / Postman, Jira, Confluence.
Similar roles
-
ML Data Engineer (m/f/d) - Sensor Data & Pipelines
Autonomous Teaming Solutions ATS GmbH Munich, Bavaria, Germany
-
(Senior) Data Engineer with AI - Freelance
Netguru Poland · €62K–€77K/yr
-
Senior Data Engineer
Great Eastern Cuenca, Azuay, Ecuador
-
Data Engineer
Booking Experts Enschede, Overijssel, Netherlands · €49K–€83K/yr
-
Data Engineer
Quicklizard Petah Tikva, Center District, Israel
-
Stage - Data Engineer - Aeroline - Toulouse
Sopra Steria Toulouse, Occitania, France