Autofleet

Senior Data Engineer

Autofleet Bengaluru, Karnataka, India

Software Development · 51-200 employees

21 h ago
Remote data-engineer Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design, implement, and maintain scalable data pipelines for batch and real-time processing while owning the backend data infrastructure. You will also collaborate with cross-functional teams to translate data requirements into robust solutions for machine learning and analytics.

What they look for

Data Engineering Python SQL Spark Google Cloud Platform ETL/ELT Airflow Docker Kubernetes Data Pipelines Distributed Systems Machine Learning Data Modeling CI/CD Data Infrastructure

Requirements

The role requires 4+ years of experience in backend data engineering or infrastructure-focused software development. Candidates must be proficient in Python and SQL, with a strong background in cloud-native environments and distributed data processing.

Full description

We are making the future of Mobility come to life starting today.

At Autofleet we support the world’s largest vehicle fleet operators and transportation providers to optimize existing operations and seamlessly launch new, dynamic business models - driving efficient operations and maximizing utilization.

At the heart of our platform lies the data infrastructure, driving advanced machine learning models and optimization algorithms. As the owner of data pipelines, you'll tackle diverse challenges spanning optimization, prediction, modeling, inference, transportation, and mapping.

As a Senior Data Engineer, you will play a key role in owning and scaling the backend data infrastructure that powers our platform—supporting real-time optimization, advanced analytics, and machine learning applications.

  • Design, implement, and maintain robust, scalable data pipelines for batch and real-time processing using Spark, and other modern tools.
  • Own the backend data infrastructure, including ingestion, transformation, validation, and orchestration of large-scale datasets.
  • Leverage Google Cloud Platform (GCP) services to architect and operate scalable, secure, and cost-effective data solutions across the pipeline lifecycle.
  • Develop and optimize ETL/ELT workflows across multiple environments to support internal applications, analytics, and machine learning workflows.
  • Build and maintain data marts and data models with a focus on performance, data quality, and long-term maintainability.
  • Collaborate with cross-functional teams including development teams, product managers, and external stakeholders to understand and translate data requirements into scalable solutions.
  • Help drive architectural decisions around distributed data processing, pipeline reliability, and scalability.
  • 4+ years in backend data engineering or infrastructure-focused software development.
  • Proficient in Python, with experience building production-grade data services.
  • Solid understanding of SQL 
  • Proven track record designing and operating scalable, low-latency data pipelines (batch and streaming).
  • Experience building and maintaining data platforms, including lakes, pipelines, and developer tooling.
  • Familiar with orchestration tools like Airflow, and modern CI/CD practices.
  • Comfortable working in cloud-native environments (AWS, GCP), including containerization (e.g., Docker, Kubernetes).
  • Bonus: Experience working with GCP
  • Bonus: Experience with data quality monitoring and alertingֿ
  • Bonus: Experience with Snowflake, DBT, Flink, Kafka
  • Bonus: Strong hands-on experience with Spark for distributed data processing at scale.
  • Degree in Computer Science, Engineering, or related field.

Similar roles