Data Engineer - Regional
Chubb San Pedro Garza García, Nuevo León, Mexico
Insurance · 1,001-5,000 employees
About the role
You will design, build, and maintain scalable batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks. Additionally, you will partner with data scientists and stakeholders to model data and ensure high data quality through monitoring and validation.
What they look for
Requirements
Candidates must have 3+ years of experience in data engineering with strong proficiency in Python, SQL, and Databricks. A solid understanding of data warehousing, performance tuning, and software engineering practices like version control is required.
Full description
Data Engineer
About the role
We're looking for a Data Engineer to build and maintain the pipelines that power analytics and data products across Chubb. You'll design ETL/ELT workflows on Databricks, turn raw source data into reliable, well-modeled datasets, and keep those pipelines fast and cost-efficient as data volumes grow.
This role suits someone who is comfortable owning a pipeline end to end — from ingestion through transformation to the tables analysts and data scientists actually query.
What you'll do
- Design, build, and maintain batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks. - Model data across raw, cleansed, and curated layers (medallion architecture) with Delta Lake. - Ingest data from a range of sources — relational databases, APIs, files, and event streams — including incremental and change data capture patterns. - Tune Spark jobs and SQL queries for performance and cost: partitioning, file sizing and compaction, caching, join strategies, and shuffle reduction. - Build data quality checks, validation rules, and monitoring so problems are caught before downstream consumers see them. - Orchestrate and schedule workflows (Databricks Workflows, Airflow, or similar), with proper retry, alerting, and dependency handling. - Apply software engineering practices to data work: version control, code review, testing, and CI/CD for pipeline deployments. - Partner with analysts, data scientists, and business stakeholders to translate requirements into usable data models. - Document pipelines, data lineage, and design decisions.
Required qualifications
- [3]+ years of experience in a data engineering or comparable role. - Strong Python for data processing, automation, and pipeline development. - Advanced SQL: complex joins, window functions, aggregations, and query optimization. - Hands-on experience with Databricks and PySpark in a production environment. - Demonstrated experience designing and operating ETL/ELT pipelines at scale. - Practical knowledge of performance optimization — able to diagnose a slow or expensive job and explain what you changed and why. - Solid understanding of data warehousing and modeling concepts (dimensional modeling, slowly changing dimensions, normalization trade-offs). - Experience with Git and collaborative development workflows.
Nice to have
- Delta Lake internals: OPTIMIZE, Z-ordering, liquid clustering, time travel, VACUUM. - Databricks features such as Unity Catalog, Delta Live Tables / Lakeflow Declarative Pipelines, Auto Loader, or Databricks SQL. - Cloud platform experience ([AWS / Azure / GCP]) and its storage and compute services. - Streaming experience with Structured Streaming, Kafka, or Event Hubs. - Infrastructure as code (Terraform) and CI/CD pipelines for data workloads. - dbt or similar transformation frameworks. - Databricks certification (Data Engineer Associate or Professional). - Familiarity with data governance, access control, and PII handling.
Similar roles
-
Big Data Engineer
TP-Link Systems Inc. Irvine, California, United States
-
Data Engineer, GTM
Anthropic San Francisco, California, United States · $320K–$405K/yr
-
Senior Data Engineer
Babylist Canada
-
Data Engineer, People Team
Changi Airport Group Singapore
-
Senior Data Engineer
Fresenius Medical Care Bengaluru, Karnataka, India
-
DATA ENGINEER SR
Alicorp S.A.A. Lima, Lima, Peru