WEXWEXUS

Senior Software Engineer - Semantic Data Lake

WEXWEXUS Bengaluru, Karnataka, India

Software Development · 5,001-10,000 employees

20 h ago
Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design and implement the core infrastructure of the Control Plane to orchestrate the entire data lifecycle from ingestion to distribution. Additionally, you will build observability frameworks and automated quality gates to ensure data trust and platform reliability.

What they look for

Data Engineering Software Engineering SQL Python Scala Airflow Terraform Kubernetes CI/CD Data Modeling LLM Architectures RAG Systems Data Governance DataHub Observability Agentic Workflows

Requirements

Candidates should have 4-8 years of experience in data or software engineering with strong proficiency in SQL and Python or Scala. You must demonstrate experience as an AI-native engineer with hands-on knowledge of LLM-driven architectures and modern DevOps practices.

Full description

About the team / Role Senior Software Engineer, Semantic Data Team

WEX is reimagining its enterprise data platform with a powerful goal: transforming raw data into semantically meaningful, reusable, and trusted business assets. As a Senior Software Engineer on the Semantic Data Team, you'll contribute to designing, building, and maintaining control plane for Semantic layer and AI Native data platforms supporting our core 360 data objects (Customer, Fleet, Health and Payments)

An embedded DataOps backbone and self-service control plane ensures governance, lineage, observability, and operational efficiency at scale. Together, these elements empower WEX to modernize responsibly—continuously delivering business value while evolving toward a future where intelligence and automation are built on a foundation of governed, high-quality data.

This team is at the heart of WEX's DaaS platform where the platform integrates DataOps automation, control plane services, and governance tooling to uphold quality and trust at scale, enabling trusted, production-grade semantic models through automated controls, audit validation, and operational alignment.

Wex is highly invested in building AI native data platform. The Agentic Control Plane is the runtime layer that determines how an AI agent interprets a user request, assembles the right context, applies the right rules and constraints, uses approved tools, and drives the next step.

We're looking for an AI-native engineer: someone who builds with modern AI coding tools (Claude, Copilot, Cursor, and similar) and Spec-Driven Development (SDD) as a core part of their daily workflow. You'll use these tools to accelerate development and apply solid engineering judgment to ship production-grade, trustworthy data assets.

How you'll make an impact

  • Engineer the Operational Backbone: You will design and implement the core infrastructure of the Control Plane, serving as the central execution engine that orchestrates the entire data lifecycle—from ingestion and transformation to final distribution.
  • Shift from Reactive to Proactive: You will transform platform reliability by building sophisticated observability frameworks and automated quality gates, ensuring that data trust is engineered into the system rather than inspected after the fact.
  • Drive Self-Service Autonomy: You will eliminate operational bottlenecks by developing intuitive portals and "golden templates," empowering domain teams to autonomously build, manage, and consume trusted data products.
  • Champion Platform Governance: You will automate critical compliance and security policies, including end-to-end data lineage, role-based access control (RBAC), and PII protection, making security an inherent feature of our data objects.
  • Prepare for AI-Native Scale: You will build the foundational data architecture required for our AI future, contributing to the development of the context and RAG-based systems that will power the next generation of data-driven intelligence.
  • Leverage AI coding assistants (Claude, Copilot, Cursor, and similar) to accelerate development—drafting transformation logic, generating tests, refactoring pipelines, exploring datasets, and producing semantic documentation—while critically reviewing AI output for correctness and alignment with business rules.
  • Share patterns, prompts, and workflows that help the team get more leverage out of AI tooling, contributing to AI-native engineering practices across the Semantic Data Team.
  • Work closely with domain experts, data scientists, and product stakeholders to translate business concepts into framework that supports development
  • Implement logic for classifications, KPIs, scoring algorithms, and business rules, ensuring traceability and data lineage.
  • Follow and contribute to standards for data modeling, documentation, and governance within the semantic layer—including responsible, auditable use of AI-generated code and artifacts.
  • Collaborate across teams to integrate with ingestion, MDM, and data product layers.

Experience you'll bring

  • 4–8 years of experience in data engineering or software engineering with a focus on data transformation, modeling, or analytics platforms.
  • Strong proficiency in SQL and at least one general-purpose language such as Python or Scala.
  • Demonstrated experience as an AI-native engineer—using tools like Claude, GitHub Copilot, Cursor, or similar as a regular part of your development workflow.
  • AI-Native Architecture: You are passionate about the future of data intelligence and have hands-on experience with LLM-driven architectures. You are comfortable designing RAG systems, building agentic workflows (e.g., LangGraph, CrewAI), and working with vector/graph databases to manage context and ontology.
  • Familiarity with modern AI engineering practices such as prompt design, Spec-Driven Development (SDD), and AI-assisted code review.
  • Advanced Data Orchestration & Pipeline Reliability: You possess deep expertise in building and managing complex orchestration systems (specifically Airflow) and are obsessed with reliability. You know how to engineer "quality-first" pipelines using frameworks like Great Expectations to ensure data trust at scale.
  • Infrastructure as Code (IaC) & Platform Engineering: You are fluent in the modern DevOps stack—including Terraform, Kubernetes, and CI/CD automation—allowing you to build self-healing, scalable infrastructure that minimizes operational toil.
  • Enterprise Governance & Metadata Management: You have a proven track record of implementing automated governance, specifically utilizing DataHub or similar platforms to manage end-to-end lineage, cataloging, and RBAC policies that make security inherent to the data object.
  • Observability & Incident Management: You understand that "the platform is the product." You bring experience building comprehensive observability frameworks (e.g., Grafana, xMatters integration) that move teams from reactive firefighting to proactive, automated alerting.
  • Self-Service Mindset: You advocate for "developer autonomy" over manual intervention. You have experience building intuitive portals and developer tools that enable domain teams to build, manage, and consume trusted data products independently.
  • Solid understanding of data quality practices—including validation, enrichment, schema enforcement, and business rule encoding.
  • Comfort operating in a collaborative, cross-functional environment, balancing business logic with platform scalability.
  • A proven track record for traceability, reproducibility, and semantic clarity—you build data models others can trust and reuse.