Caterpillar Inc.

Software Engineer - Data Engineering

Caterpillar Inc. · Bangalore, Karnataka, India

Machinery Manufacturing · 10,001+ employees

13 h ago
Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain scalable data pipelines on AWS while optimizing data warehousing solutions using Snowflake. Collaborate with cross-functional teams to ensure data quality, security, and performance across all data processing stages.

What they look for

Data Engineering AWS Snowflake Python SQL Data Pipelines Data Warehousing Graph Databases Vector Databases CI/CD GitHub Actions Azure DevOps Agile Data Modeling Performance Tuning Cloud Computing

Requirements

Requires 5-8 years of experience in data engineering with advanced proficiency in Python, SQL, and AWS cloud services. Candidates must have a strong understanding of Snowflake architecture and experience with graph and vector data modeling.

Full description

Career Area:

Technology, Digital and Data

Job Description:

Your Work Shapes the World at Caterpillar Inc.

When you join Caterpillar, you're joining a global team who cares not just about the work we do – but also about each other.  We are the makers, problem solvers, and future world builders who are creating stronger, more sustainable communities. We don't just talk about progress and innovation here – we make it happen, with our customers, where we work and live. Together, we are building a better world, so we can all enjoy living in it.

Job Summary We are looking for a highly motivated and experienced Data Engineer to join our data engineering team. The ideal candidate will have a strong background in building scalable data pipelines using the AWS cloud stack and extensive hands-on experience with Snowflake. Proficiency in Python and SQL, along with graph and vector database technologies, is essential. This role requires strong problem-solving abilities and a proactive mindset to deliver efficient, scalable, and reliable data solutions. ________________________________________ Key Responsibilities •    Design, develop, and maintain scalable data pipelines on AWS using services such as S3, Glue, Lambda, Redshift, and EMR. •    Build and optimize data warehousing solutions using Snowflake, including performance tuning and data modeling. •    Write efficient and reusable code in Python and SQL for data transformation and processing. •    Collaborate with cross-functional teams, including data scientists, analysts, and business stakeholders, to understand data requirements. •    Monitor, troubleshoot, and improve pipeline performance and reliability. •    Ensure data quality, integrity, and security across all stages of the pipeline. •    Participate in code reviews, architecture discussions, and continuous improvement initiatives. ________________________________________ Required Qualifications •    5-8 years of experience in data engineering or related roles. •    Strong hands-on experience with AWS cloud services, including data and AI workloads. •    Deep understanding of Snowflake architecture, performance tuning, and best practices. •    Advanced proficiency in Python and SQL for data pipelines, transformations, and services. •    Strong understanding of graph and vector data modelling concepts and their practical applications. •    Experience developing and operating CI/CD pipelines (GitHub Actions)and cloud native deployments. •    Experience working with Azure DevOps (AzDO) boards for backlog management in Agile environments. •    Excellent analytical and problem-solving skills. •    Strong communication and collaboration abilities. •    Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. ________________________________________ Nice to Have skills •    Knowledge of the NVIDIA ecosystem and its applications in data and AI. •    Hands on experience with Vector (Milvus or similar) and Graph (Neo4J or similar) databases ________________________________________ Preferred Qualifications •    Experience with orchestration tools such as AWS Step Functions. •    Familiarity with data governance and compliance practices. •    Exposure to real-time data processing frameworks (e.g., Kafka, Spark Streaming).

This position requires working onsite five days a week. 

Relocation is available for this position.

Posting Dates:

August 7, 2026 - August 20, 2026

Caterpillar is an Equal Opportunity Employer.  Qualified applicants of any age are encouraged to apply

Not ready to apply? Join our Talent Community.