principal33

Data Engineer with strong Python Spark

principal33 Helsinki, Uusimaa, Finland

IT Services and IT Consulting · 201-500 employees

4 h ago
Remote python Senior (5-10 yrs) Full-time Finland
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, build, and operate scalable batch and streaming data pipelines using Python, PySpark, and Azure Databricks. Manage data architecture, orchestration, and CI/CD processes while ensuring high data quality and performance.

What they look for

Python PySpark Azure Databricks Data pipelines SQL Delta Lake Unity Catalog Terraform GitHub Actions Azure DevOps GitOps CI/CD Data modeling API integration Cloud infrastructure Data quality

Requirements

Requires at least 5 years of experience in data engineering, ideally within the energy or trading sectors. Candidates must possess strong proficiency in Python, Spark, and cloud-based data platform management.

Benefits

Medical insurance Holiday flat in Valencia Gifts for special occasions Anniversary gifts Team-building events End of year celebrations Day off on your birthday Community and social initiatives Access to Udemy German language courses Personal and professional growth

Full description

Your mission

Data pipelines and integration

  •  Design, build and operate scalable batch, micro-batch and streaming data pipelines using Python, PySpark and Azure Databricks.
  •  Integrate internal and external data sources, including REST APIs, GraphQL, WebSocket and gRPC interfaces, databases, files, event streams and third-party data feeds.
  •  Develop robust ingestion solutions for structured, semi-structured and unstructured data, including JSON, CSV, Parquet, Delta and API-based payloads.
  •  Build reliable web scraping and data acquisition components where APIs or managed integration mechanisms are unavailable.
  •  Implement pagination, throttling, retries, exponential backoff, checkpointing, schema evolution and recovery patterns for external integrations.
  •  Design pipelines that support idempotent processing, reprocessing, controlled backfills and graceful recovery from partial failures.
  •  Develop and maintain batch and streaming patterns for market, weather, fundamental and time-series data.

  Software engineering

  •  Develop modular, reusable and testable Python and PySpark components rather than relying on monolithic notebooks.
  •  Apply object-oriented and functional design principles appropriately to data-focused development.
  •  Structure solutions as maintainable software projects with clear separation between source code, configuration, tests, deployment assets and notebooks.
  •  Write clean, readable and well-documented code using type hints, meaningful interfaces and appropriate design patterns.
  •  Build automated unit, integration, contract and data-quality tests and incorporate them into delivery pipelines.
  •  Conduct code reviews and promote engineering standards covering readability, testability, security, performance and maintainability.
  •  Package reusable functionality as Python modules or wheels where appropriate.
  •  Troubleshoot complex issues across source systems, APIs, processing logic, infrastructure and production runtime environments.

  Data architecture and modelling

  •  Design maintainable data models that support analysts, traders, reporting solutions and downstream data products.
  •  Implement Lakehouse and Medallion architecture patterns across Bronze, Silver and Gold layers.
  •  Preserve raw data appropriately while applying cleansing, validation, standardisation and business transformations in downstream layers.
  •  Design solutions for schema evolution, data retention, lineage and reproducible processing.
  •  Apply sound data architecture principles across operational, analytical, event-based and time-series workloads.
  •  Optimise data layouts, partitioning, joins, file sizes, caching and Spark execution plans for performance and cost.

  Orchestration and DataOps

  •  Design, schedule and operate workflows using Databricks Workflows and Astronomer.
  •  Implement dependency management, parameterisation, environment-specific configuration and controlled promotion across development, test and production environments.
  •  Define operational runbooks and support effective diagnosis, recovery and problem management.
  •  Monitor pipeline health, freshness, completeness, performance and data-quality indicators.
  •  Use production-safe release patterns, including controlled rollouts, rollback and validation where appropriate.

  GitOps, CI/CD and Infrastructure as Code

  •  Manage all production code through Git using clear branching, pull-request and review practices.
  •  Build and maintain automated CI/CD pipelines using GitHub Actions and/or Azure DevOps.
  •  Deploy Databricks jobs, pipelines and application artefacts using Databricks Asset Bundles or equivalent approved mechanisms.
  •  Provision and configure relevant cloud and Databricks resources through Terraform.
  •  Treat application code, infrastructure, data pipeline definitions and operational configuration as version-controlled artefacts.
  •  Apply automated validation, security scanning and testing before production deployment.
  •  Contribute to reusable pipeline templates, engineering standards and platform automation.

  Reliability, performance and cost efficiency

  •  Engineer solutions for availability, recoverability, scalability and predictable operational behaviour.
  •  Optimise Spark workloads through appropriate partitioning, built-in Spark functions, efficient joins, adaptive execution and avoidance of unnecessary shuffles or UDFs.
  •  Select suitable compute models and cluster configurations based on workload characteristics.
  •  Apply cost-awareness to pipeline design, compute sizing, scheduling, storage and data-retention decisions.
  •  Monitor resource consumption and identify opportunities to reduce processing times and cloud costs without compromising reliability or data quality.
  •  Balance immediate delivery requirements with sustainable architecture and long-term maintainability.

  Data quality, governance and security

  •  Implement automated data validation, schema checks, null checks, referential-integrity controls and business quality rules.
  •  Detect and manage schema drift and unexpected changes in source data.
  •  Use Delta Lake and Unity Catalog capabilities to support data lineage, access control, metadata and governance.
  •  Ensure secrets and credentials are handled securely using approved secret-management mechanisms and managed identities.
  •  Maintain technical documentation, metadata and operational information for assigned data products.
  •  Collaborate with data governance, architecture, security and platform teams to ensure alignment with enterprise standards.

  Collaboration and delivery

  •  Work closely with traders, analysts, data scientists, software engineers, product owners and platform teams to translate business requirements into robust technical solutions.
  •  Communicate design decisions, risks, dependencies and technical trade-offs clearly to technical and non-technical stakeholders.
  •  Contribute reusable components, templates, documentation and engineering guidelines for the wider data community.
  •  Work effectively in a distributed, international and cross-functional environment.

Your profile

An experienced data engineer with at elast 5 years of experience, ideally in the energy sector and/or trading. Be able to operate fundamental power-price forecasting models for short- to mid-term trading and large amounts of data. Be able to work fully remote in a collaborative environment with interdisciplinary teams. Good communication skills and professional behaviour. What we offer

  • Medical insurance: Your health, and your family's, is our top priority. You're fully covered.
  • Holiday flat in Valencia: Dreaming of sunny days in Spain? Our company flat is ready for your next getaway.
  • Gifts for special occasions: We love celebrating you. Expect thoughtful surprises on Easter, Women's Day, Father's Day, and more.
  • Anniversary gifts: Your 1st, 5th, and 10th work anniversaries are milestones worth celebrating, and we mark each of them with a special gift.
  • Team-building events: We value connection beyond work, offering engaging team experiences in great locations to inspire collaboration and fun.
  • End of year celebrations: We wrap up each year with a special celebration, filled with joy, laughter, and unforgettable moments.
  • Day off on your birthday: Your special day is yours to enjoy, no work required.
  • Community and social initiatives: We bring people together through activities like Bring Your Kids to the Office Day, donation drives, tree planting, sporting events, decorating the Christmas tree with your colleagues, celebrating new office openings, Principal33's anniversary, and more.
  • Access to Udemy: Growth matters, so you get unlimited access to thousands of courses. Develop your skills anytime, anywhere.
  • German language courses: Dedicated courses to help you build your German skills, supporting both personal and professional growth.
  • Personal and professional growth: We support your development through masterclasses with experienced trainers, and we're open to investing in courses and training that help you grow.

Email address

About us

At principal33, we are driven by innovation, collaboration, and a commitment to excellence. Our mission is to deliver impactful solutions that empower our clients and inspire our teams. We are dedicated to cultivating a work environment that attracts top talent, drives innovation, and values diverse perspectives. Join us and become part of an organization that prioritizes growth, creativity, and meaningful impact every day.

Similar roles