Principal Data Engineer
Lula · Cape Town, Western Cape, South Africa
Banking · 51-200 employees
About the role
The Principal Data Engineer will guide the data team through modernization efforts while designing and maintaining complex data pipelines. They will collaborate with cross-functional teams to build robust data systems and mentor junior and senior engineers.
What they look for
Requirements
Candidates must have extensive experience with Snowflake, dbt, and Azure cloud platforms, along with advanced SQL and Python proficiency. A strong background in data modeling, CI/CD practices, and performance optimization for large-scale data systems is required.
Full description
About Lula
Lula is on a mission to simplify business funding for South African SMEs. We build fast, digital-first fintech tools that make cash flow management seamless so business owners can focus on doing what they love. We’re a fast-growing, tech-driven team that values curiosity, collaboration, and high impact over red tape. Speaking of love, we’re looking for Lulas who love to make a difference to join our team and change the game.
Culture Code
- We Embrace Curiosity - We continuously seek better ways to deliver value with a solutions-over-problems mindset.
- We win as One - We collaborate, build strong relationships and value diverse perspectives
- We’re Driven by Purpose - We are passionate and committed to delivering the best products and services for SMEs
- We Execute with Ambition - We set ambitious goals, embrace challenges, and deliver with focus and determination.
ROLE PURPOSE
The role requires an experienced, highly technical Principal Data Engineer to join and help guide our established data team through the next phase of our data modernization journey. Building on the strong foundation our team has already laid, you will work directly alongside them, providing deep technical expertise, mentoring talent, and diving into the fine details of building and optimizing high-performance data systems.
In this role, you will bridge strategic vision and hands-on execution. Collaborating closely with Solutions Architects and cross-functional engineering teams, you will design, ingest, and maintain complex data pipelines across a fast-paced FinTech ecosystem. The ideal candidate thrives on building robust data solutions from the ground up, enjoys solving intricate data management challenges, and is excited to roll up their sleeves to elevate our overall data capability.
KEY RESPONSIBILITIES
- Work as part of a multi-disciplinary data focussed team (Data, Analytics, Machine Learning Engineers)
- Collaborate with broader data functionality across the business (Analysts, Data Scientists)
- Develop greenfield projects using our Azure and Snowflake platforms
- Assemble data sets that meet functional and non-functional business requirements
- Develop and maintain operational and analytical data systems
- Create and maintain robust ELT/ETL products for batch, micro-batch and near real-time data pipelines using Airbyte, Event Grid, Azure Data Factory, Kafka or similar tools
- Follow test-driven development practices
- Demo work to both technical and non-technical stakeholders
- Create documentation and training material for the solutions being delivered
- Guide and mentor senior and junior team members
OUR TECH STACK
- Azure (Functions, databases, blob storage)
- DBT
- Snowflake
- Airflow
- Airbyte
- Event Grid
THE SKILLS AND EXPERIENCE WE’RE LOOKING FOR
- Strong hands-on experience with Snowflake, including warehouse/resource management, semi-structured data handling (VARIANT, JSON), and performance/cost optimisation.
- Experience with dbt (Core or Cloud) - models, tests, macros, snapshots, and documentation generation.
- Experience using schemachange, Flyway, Liquibase, or a similar tool to manage database change control as code.
- Proficient in Jinja templating for building dynamic, reusable SQL/config.
- Advanced working knowledge of SQL (DDL, DML, JSON, XML) and extensive experience managing incremental/batch loading methodologies (CDC, CT, CDC-style watermarking).
- Proven experience building ingestion pipelines from diverse source types: APIs, flat files, relational/NoSQL databases, Azure Table Storage, and web scraping.
- Skilled and experienced in the Azure (or AWS/GCP) cloud platform.
- Advanced understanding of relational data structures, including keys, constraints, and triggers.
- Experience with performance tuning and optimisation of RDBMS and/or cloud data warehouses.
- Experience with relational and NoSQL database technologies (MS SQL Server, MongoDB, CosmosDB, etc.).
- Ability to design and implement conceptual, logical and physical data models that support organisational needs.
- Solid understanding and experience in data modeling, data management and governance methodologies.
- Good understanding of data-related frameworks, methodologies, and patterns.
- Strong analytic skills working with structured, semi-structured and unstructured data sets.
- Proficiency in Python (preferred), Java, or Scala.
- Practical experience applying generative AI / LLM tools to engineering or data workflows (e.g. Copilot, ChatGPT, Claude, or similar).
- Experience implementing CI/CD pipelines through technologies such as GitLab, Azure DevOps, etc.
- Experience deploying data systems in an Infrastructure-as-Code (IaC) manner, preferably using Terraform.
- Strong ability to produce high-quality technical documentation as a routine part of delivery.
- Experience supporting and working with cross-functional teams in a dynamic environment.
- Communicates effectively with both technical and non-technical stakeholders.