Ciklum

Data Engineer

Ciklum · Kyiv, Ukraine

IT Services and IT Consulting · 1,001-5,000 employees

4 h ago
Remote Mid (2-5 yrs) Full-time Ukraine
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Develop, maintain, and monitor robust data pipelines to process AST outputs, knowledge graph structures, and vector embeddings. Collaborate with engineering teams to normalize data and ensure consistency across automated documentation generation workflows.

What they look for

Python SQL Neo4j Qdrant Tree-sitter Data pipelines Git Docker CI/CD Unit testing Integration testing Data transformation Vector search Knowledge graphs AST parsing Markdown

Requirements

Requires 3+ years of commercial data engineering experience with proficiency in Python, SQL, and data manipulation frameworks. Candidates must have hands-on experience with graph databases, vector search engines, or static code parsing.

Benefits

Medical insurance Mental health support Financial consultations Legal consultations Udemy access Language courses Company-paid certifications Internal events Flexible work environment

Full description

Ciklum is looking for a Data Engineer to join our team full-time in Ukraine.

We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.

About the role:

As a Data Engineer, become a part of a cross-functional development team engineering experiences of tomorrow.

The Legacy Code Semantic Documentation Project is a 26-week enterprise initiative for a global industrial automation leader. The primary objective is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented ~400K LOC codebase (spanning IEC 61131-3 languages and ANSI C/C++).

Operating within a dedicated, zero-data-egress secure tenant, the technical architecture combines Tree-sitter AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral family) with NLI-based validation. Delivery is structured across 7 work packages executing over a 6-month period, governed by strict contractual KPIs: ≥95% code coverage, ≥92% NLI-verified factual precision, and ≥90% SME validation acceptance. All outputs align with EU Cyber Resilience Act (EU CRA) requirements for SBOM and source-level traceability.

Responsibilities:

  • Data Pipeline Execution: Develop, test, and maintain robust data pipelines that process AST outputs, knowledge graph structures, and vector embeddings
  • Knowledge Base Ingestion: Write data processing scripts to ingest, normalize, and update code dependency graphs in Neo4j and vector indexes in Qdrant
  • Data Cleansing & Transformation: Parse raw source code structures and raw metadata into structured Markdown assets and standardized JSON/RAG inputs
  • Pipeline Monitoring & Debugging: Monitor execution throughput, resolve batch processing errors, and ensure data state consistency across pipeline runs
  • Collaboration: Work closely with Senior Data Engineers and AI/ML Engineers to optimize data retrieval speeds and pipeline efficiency

Requirements:

  • Professional Experience: 3+ years of commercial Data Engineering experience building and maintaining production data pipelines
  • Core Technical Proficiency: Solid proficiency in Python, SQL, and standard data manipulation frameworks
  • Hands-on Stack Exposure: Direct working knowledge of graph databases (Neo4j), vector search engines (Qdrant), or static code parsing frameworks (Tree-sitter)
  • Engineering Best Practices: Proficiency with Git version control, Docker containerization, unit/integration testing for data pipelines, and CI/CD workflows
  • Problem-Solving & Detail Orientation: Strong analytical skills with a focus on data accuracy, schema consistency, and output validation
  • Language: Professional proficiency in English (B2+/C1)

What’s in it for you?

  • Strong community: Work alongside top professionals in a friendly, open-door environment
  • Growth focus: Take on large-scale projects with a global impact and expand your expertise
  • Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications
  • Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies
  • Flexibility: Enjoy radical flexibility – work remotely or from an office, your choice
  • Care: We’ve got you covered with company-paid medical insurance, mental health support, and financial & legal consultations

About us:

At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress.

As one of Ukraine’s largest IT companies and a top employer recognized by Forbes, we’ve spent over 20 years delivering meaningful tech solutions. We proudly support diverse talent and military veterans, recognizing their unique skills and perspectives they bring to shaping the future.

Explore, empower, engineer with Ciklum!

Interested already? We would love to get to know you! Submit your application. We can’t wait to see you at Ciklum.

#LI-NV1