Chubb

Data Engineer - Regional

Chubb San Pedro Garza García, Nuevo León, Mexico

Insurance · 1,001-5,000 employees

8 h ago
data-engineer Mid (2-5 yrs) Full-time Mexico
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design, build, and maintain scalable batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks. Additionally, you will partner with data scientists and stakeholders to model data and ensure high data quality through monitoring and validation.

What they look for

Python SQL PySpark Databricks ETL/ELT Delta Lake Data Modeling Data Engineering Git CI/CD Data Pipelines Performance Optimization Data Warehousing Streaming Workflow Orchestration

Requirements

Candidates must have 3+ years of experience in data engineering with strong proficiency in Python, SQL, and Databricks. A solid understanding of data warehousing, performance tuning, and software engineering practices like version control is required.

Full description

Data Engineer

About the role

We're looking for a Data Engineer to build and maintain the pipelines that power analytics and data products across Chubb. You'll design ETL/ELT workflows on Databricks, turn raw source data into reliable, well-modeled datasets, and keep those pipelines fast and cost-efficient as data volumes grow.

This role suits someone who is comfortable owning a pipeline end to end — from ingestion through transformation to the tables analysts and data scientists actually query.

What you'll do

- Design, build, and maintain batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks. - Model data across raw, cleansed, and curated layers (medallion architecture) with Delta Lake. - Ingest data from a range of sources — relational databases, APIs, files, and event streams — including incremental and change data capture patterns. - Tune Spark jobs and SQL queries for performance and cost: partitioning, file sizing and compaction, caching, join strategies, and shuffle reduction. - Build data quality checks, validation rules, and monitoring so problems are caught before downstream consumers see them. - Orchestrate and schedule workflows (Databricks Workflows, Airflow, or similar), with proper retry, alerting, and dependency handling. - Apply software engineering practices to data work: version control, code review, testing, and CI/CD for pipeline deployments. - Partner with analysts, data scientists, and business stakeholders to translate requirements into usable data models. - Document pipelines, data lineage, and design decisions.

Required qualifications

- [3]+ years of experience in a data engineering or comparable role. - Strong Python for data processing, automation, and pipeline development. - Advanced SQL: complex joins, window functions, aggregations, and query optimization. - Hands-on experience with Databricks and PySpark in a production environment. - Demonstrated experience designing and operating ETL/ELT pipelines at scale. - Practical knowledge of performance optimization — able to diagnose a slow or expensive job and explain what you changed and why. - Solid understanding of data warehousing and modeling concepts (dimensional modeling, slowly changing dimensions, normalization trade-offs). - Experience with Git and collaborative development workflows.

Nice to have

- Delta Lake internals: OPTIMIZE, Z-ordering, liquid clustering, time travel, VACUUM. - Databricks features such as Unity Catalog, Delta Live Tables / Lakeflow Declarative Pipelines, Auto Loader, or Databricks SQL. - Cloud platform experience ([AWS / Azure / GCP]) and its storage and compute services. - Streaming experience with Structured Streaming, Kafka, or Event Hubs. - Infrastructure as code (Terraform) and CI/CD pipelines for data workloads. - dbt or similar transformation frameworks. - Databricks certification (Data Engineer Associate or Professional). - Familiarity with data governance, access control, and PII handling.

Similar roles