O

Data Engineer LATAM

onebeat Rio de Janeiro, Rio de Janeiro, Brazil

Software Development · 51-200 employees

10 h ago
Remote data-engineer Mid (2-5 yrs) Full-time Brazil
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will be responsible for designing, developing, and maintaining a scalable open-source data platform including Data-Lakehouse and ETL processes. You will also collaborate cross-functionally to optimize data pipelines and ensure the reliability of the data infrastructure.

What they look for

SQL Data Modeling Apache Iceberg ClickHouse Temporal Python Scala Java AWS Microservices Apache Spark ETL ELT Data Lakehouse Data Governance Data Infrastructure

Requirements

Candidates must have 3+ years of experience in data engineering with strong proficiency in SQL and Python or Scala/Java. Experience with data lakehouse architectures, columnar databases, and orchestration tools like Temporal is required.

Full description

We are seeking an experienced Data Engineer. The ideal candidate is self-motivated, a multitasker, and a demonstrated team player. You will be responsible for designing, developing, managing, and maintaining our open-source data platform, including our Data-Lakehouse (S3, Apache Iceberg, and ClickHouse), ETL processes, and orchestration tool (Temporal Workflow).

What You Will Do

● Develop a scalable data platform integrating multiple sources for easy access.

● Design and enhance data tools (orchestration, governance, Data-Lakehouse, BI, etc.).

● Ensure smooth operation of data systems for analysts, scientists, and engineers.

● Optimize data pipelines (ingestion, processing, and output) in a microservices environment.

● Build, maintain, and monitor ETL/ELT processes and orchestrate workflows using Temporal.

● Troubleshoot and improve the performance, scalability, and reliability of the data infrastructure (S3, Apache Iceberg, ClickHouse).

● Collaborate cross-functionally with data scientists, analysts, and backend engineers to understand data needs and deliver solutions.

● Implement and champion data quality, governance, and security best practices across the platform.

Requirements

● 3+ years of experience as a Data Engineer or in a similar data infrastructure role.

● Strong proficiency in SQL and hands-on experience with data modeling.

● Experience with data lake/lakehouse architectures (e.g., Apache Iceberg, S3, or similar).

● Experience with analytical / columnar databases (e.g., ClickHouse or similar).

● Experience building and orchestrating ETL/ELT pipelines (e.g., Temporal, Airflow, or similar).

● Strong programming skills in Python and/or Scala/Java.

● Experience working within a microservices architecture and cloud environments (AWS preferred).

● Self-motivated, strong multitasking skills, and a demonstrated team player.

● Excellent communication skills and the ability to work both independently and collaboratively.

● Hands-on experience with Apache Spark (or similar technologies) for large-scale data processing.

● Professional proficiency in written and spoken English.

● Note: this role is focused on batch data processing (not real-time streaming).

Nice to Have

● Experience working with and contributing to open-source data platforms and tools.

● Familiarity with BI and visualization tools (e.g., Superset, Looker, Tableau, Metabase, or similar).

● Experience with containerization and orchestration (Docker, Kubernetes).

● Experience with infrastructure-as-code and CI/CD practices.

● Experience with AWS EMR and running Apache Spark workloads in a cloud environment.

● Experience leveraging AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to boost engineering productivity.

Similar roles