Jobgether

Senior Software Engineer - Semantic Data Lake

Jobgether India

Internet Marketplace Platforms · 11-50 employees

7 h ago
Remote Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design and implement core infrastructure and control-plane services to orchestrate data lifecycles and ensure data reliability. Additionally, you will develop self-service tools and automate governance capabilities to enable domain teams to manage trusted data products.

What they look for

Data engineering Software engineering SQL Python Scala AI engineering Data modeling Apache Airflow Kubernetes Terraform Data governance RAG systems Metadata management Observability Cloud-native infrastructure Semantic modeling

Requirements

Candidates must have 4–8 years of experience in software or data engineering with strong proficiency in SQL and Python or Scala. You should also possess practical knowledge of AI-native architectures, data orchestration frameworks, and cloud-native infrastructure practices.

Benefits

Competitive compensation Comprehensive benefits Remote-office flexibility

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer - Semantic Data Lake based in India.

Join a forward-looking data engineering team building the foundations of an AI-native enterprise data platform. In this role, you will help transform raw enterprise data into trusted, reusable, semantically meaningful business assets. You will design and evolve the control plane that orchestrates data operations, governance, quality, lineage, and observability at scale. The role combines modern software engineering, data platform architecture, AI engineering, and DevOps practices. You will work closely with data scientists, domain experts, and product stakeholders to turn complex business concepts into reliable data products. You will also help shape how AI-assisted development and agentic architectures are adopted across the engineering organization.

\n

Accountabilities:

  • Design and implement the core infrastructure and control-plane services that orchestrate data lifecycles from ingestion and transformation through distribution and consumption.
  • Build proactive reliability capabilities, including observability frameworks, automated quality gates, validation mechanisms, and operational controls that embed data trust into the platform.
  • Develop self-service portals, developer tooling, and reusable templates that enable domain teams to independently create, manage, and consume trusted data products.
  • Automate governance and security capabilities, including end-to-end data lineage, metadata management, role-based access control, PII protection, auditability, and compliance controls.
  • Contribute to AI-native data architecture, including context and retrieval-augmented generation (RAG) systems, agentic workflows, semantic models, and technologies that support intelligent data consumption.
  • Use AI coding assistants such as Claude, GitHub Copilot, Cursor, and similar tools to accelerate development, testing, refactoring, documentation, and data exploration while applying strong engineering judgment to validate generated outputs.
  • Help establish effective AI-native engineering practices by sharing prompts, development patterns, workflows, and responsible approaches to AI-assisted coding and documentation.
  • Collaborate with domain experts, data scientists, product stakeholders, and adjacent platform teams to translate business concepts into scalable frameworks, semantic models, classifications, KPIs, scoring algorithms, and business rules.
  • Implement traceable and reproducible data logic while maintaining strong standards for data modeling, semantic clarity, documentation, governance, and responsible use of AI-generated artifacts.
  • Partner with ingestion, master data management, and data product teams to ensure seamless integration across the broader data platform.

Requirements:

  • 4–8 years of experience in software engineering or data engineering, with a strong focus on data transformation, modeling, analytics platforms, or data-intensive applications.
  • Strong SQL skills and proficiency in at least one general-purpose programming language, such as Python or Scala.
  • Hands-on experience using AI development tools such as Claude, GitHub Copilot, Cursor, or comparable solutions as an integral part of the software development workflow.
  • Practical knowledge of AI-native architectures, including LLM-driven applications, RAG systems, agentic workflows, vector databases, graph databases, context management, and ontologies. Experience with frameworks such as LangGraph or CrewAI is valuable.
  • Familiarity with modern AI engineering practices, including prompt design, Spec-Driven Development (SDD), AI-assisted development, and AI-supported code review.
  • Strong experience with data orchestration and reliable pipeline engineering, particularly Apache Airflow, with an understanding of data validation and quality frameworks such as Great Expectations.
  • Experience with Infrastructure as Code and platform engineering practices, including Terraform, Kubernetes, CI/CD automation, and scalable cloud-native infrastructure.
  • Experience implementing enterprise data governance and metadata management using DataHub or similar platforms, including lineage, cataloging, access controls, and security policies.
  • Knowledge of observability and incident-management practices, with experience using technologies such as Grafana and automated alerting or incident-response integrations.
  • A self-service mindset, with experience building developer portals, internal platforms, or tools that reduce operational dependencies and enable engineering teams to work autonomously.
  • Strong understanding of data quality practices, including validation, enrichment, schema enforcement, business-rule encoding, and automated controls.
  • Strong analytical, problem-solving, and communication skills, with the ability to collaborate effectively across engineering, data, product, and business teams.
  • A demonstrated commitment to traceability, reproducibility, semantic clarity, and building data models that are reliable, reusable, and trusted by their users.

Benefits:

  • Competitive, market-aligned compensation.
  • Comprehensive benefits designed to support employees' personal and professional well-being.
  • Opportunities to work on modern AI-native data platforms and large-scale enterprise technology.
  • Exposure to advanced data engineering, semantic modeling, agentic AI, RAG, governance, and platform engineering practices.
  • A collaborative, cross-functional environment focused on continuous innovation and engineering excellence.
  • Opportunities to contribute to emerging AI-assisted development practices and influence how modern engineering workflows evolve.
  • Remote-office flexibility for the role based in India.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1