Capgemini

Azure Data Engineer

Capgemini Bogota, RAP (Especial) Central, Colombia

IT Services and IT Consulting · 10,001+ employees

20 h ago
Remote data-engineer Senior (5-10 yrs) Full-time Colombia
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design, develop, and maintain scalable data pipelines and architectures using Azure cloud services to support BI, ML, and AI initiatives. Implement robust ETL/ELT processes and ensure data quality, security, and observability across distributed processing workflows.

What they look for

Azure Databricks Azure Data Factory Python PySpark SQL ETL/ELT Azure DevOps Data Lakehouse FastAPI PyTest REST APIs Git Azure Functions Azure Logic Apps Data Governance CI/CD

Requirements

Requires a Bachelor's degree in Computer Science or a related field with proficiency in Python, PySpark, and the Azure data ecosystem. Candidates should have experience with modern data architectures, distributed computing, and Agile methodologies.

Full description

At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.

Job Description

Your Role:

We are seeking a highly skilled Senior Data Engineer to design, develop, optimize, and maintain scalable data solutions in cloud environments. You will be responsible for building robust pipelines, data processing architectures, and analytics platforms that support Business Intelligence, Machine Learning, and Artificial Intelligence initiatives.

  • Design, develop, and maintain scalable, high-performance data pipelines.
  • Build ETL/ELT processes using Python, PySpark, SQL, Azure Databricks, and Azure Data Factory.
  • Implement integration mechanisms for structured and unstructured data from multiple sources.
  • Develop storage and processing solutions based on Data Lake and Lakehouse architecture.
  • Ensure data quality, consistency, security, and availability.
  • Optimize distributed processing workflows for large volumes of data.
  • Implement data governance, monitoring, and observability standards.
  • Design and implement solutions using Azure Databricks, Azure Data Factory, Azure Functions, Azure Storage Accounts, Azure Key Vault, and Azure Logic Apps.
  • Automate deployments via CI/CD pipelines using Azure DevOps.
  • Monitor and resolve incidents related to data processes in production environments.
  • Apply security and access management best practices in cloud environments.
  • Implement secure authentication, authorization, and secret management mechanisms.
  • Collaborate with Data Science teams to prepare, transform, and make data available for analytical models.
  • Participate in Machine Learning, Artificial Intelligence, and NLP initiatives.
  • Implement and support data pipelines for forecasting and predictive analytics models.

Your Profile:

  • Bachelor’s degree in systems engineering, Computer Science, Software Engineering, Data Science, or a related field.
  • Python, PySpark, FastAPI, PyTest.
  • Microsoft Azure: Azure Databricks (Databricks Notebooks, Databricks Workflows), Azure Data Factory, Azure Functions, Azure Key Vault, Azure Storage Accounts, Azure Logic Apps, Azure DevOps.
  • REST APIs.
  • SQL.
  • Git.

Desirable:

  • Web Scraping. • Microsoft Excel for exploratory analysis and data validation.
  • Data and Analytics Knowledge.
  • Design and implementation of modern data architectures.
  • Massive data processing and distributed computing.
  • Data Warehousing and Data Lakes.
  • Machine Learning Engineering fundamentals.
  • Business Intelligence and enterprise analytics.
  • Forecasting and time-series analysis.
  • Credit and financial risk modeling.
  • Financial analysis and performance metrics.
  • Cash flow analysis and modeling.
  • Experience working with Agile methodologies and Jira.

#LI-DC10

#LI-Remote

Capgemini is a global business and technology transformation partner, helping organizations to accelerate their dual transition to a digital and sustainable world, while creating tangible impact for enterprises and society. It is a responsible and diverse group of 340,000 team members in more than 50 countries. With its strong over 55-year heritage, Capgemini is trusted by its clients to unlock the value of technology to address the entire breadth of their business needs. It delivers end-to-end services and solutions leveraging strengths from strategy and design to engineering, all fueled by its market leading capabilities in AI, generative AI, cloud and data, combined with its deep industry expertise and partner ecosystem.

Similar roles