Senior Data Engineer - Data Platform
Jobgether India
Internet Marketplace Platforms · 11-50 employees
About the role
Architect and build scalable data transformation and warehouse systems capable of processing billions of records. Design production-grade patterns for LLM-powered data extraction, enrichment, and validation while ensuring high data quality and performance.
What they look for
Requirements
Requires 5+ years of experience in building production-grade data platforms with end-to-end architectural responsibility. Must have expert-level skills in Python, SQL, Snowflake, dbt, and Airflow, along with proven experience deploying LLMs in production pipelines.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer - Data Platform based in India.
This is a hands-on senior engineering role focused on building a modern, AI-native data platform at massive scale. You’ll architect the transformation and warehouse layer that turns billions of raw records into accurate, trusted B2B data. You’ll design production-grade systems for LLM-powered extraction, enrichment, validation, and entity resolution. The role combines deep technical ownership with strong architectural judgment around performance, reliability, cost, and data quality. You’ll work across data engineering, AI, analytics, and product teams in a lean, highly autonomous environment. Your work will establish reusable standards and infrastructure that directly influence the quality of data delivered to customers globally.
\n
Accountabilities
- Architect and build scalable transformation, warehouse, and data-processing systems capable of handling billions of records across multiple markets.
- Design production patterns for LLM-powered extraction, enrichment, entity resolution, and semantic validation, including structured outputs, retries, fallbacks, and reusable abstractions.
- Develop evaluation frameworks with labelled datasets, scoring mechanisms, judge calibration, regression suites, and measurable precision/recall targets for AI-powered data workflows.
- Establish clear standards for when deterministic rules, SQL, data contracts, and dbt tests should be used versus when LLM-based validation is appropriate.
- Own warehouse architecture and optimization, including dbt transformation layers, Snowflake performance, clustering, materialization strategies, access controls, and cost management.
- Build reliable orchestration and infrastructure patterns using Airflow and AWS, with appropriate recovery, dependency management, scheduling, and infrastructure-as-code practices.
- Implement observability across AI and data pipelines by tracking model versions, prompts, costs, latency, decisions, and operational alerts.
- Develop matching, deduplication, embedding, and retrieval solutions to support entity resolution across company and people data.
- Provide technical leadership in architecture discussions with data quality, sourcing, analytics, product, and other cross-functional teams.
- Balance technical excellence with product impact, making informed decisions about when to optimize, refactor, redesign, or ship.
Requirements
- 5+ years of experience building and owning production-grade data platforms in business-critical environments, preferably with end-to-end architectural responsibility.
- Proven experience working with billions of rows and designing systems, pipelines, and quality controls that remain reliable and performant at significant scale.
- Demonstrated experience deploying LLMs within production data pipelines for extraction, enrichment, or validation, including structured outputs, versioned prompts, and labelled evaluation datasets.
- Experience designing evaluation or test harnesses used by other engineers, including scorers, regression suites, labelled datasets, and judge-calibration approaches.
- Strong architectural judgment when deciding between deterministic data-quality approaches and LLM-based semantic evaluation.
- Expert-level Python and SQL skills, with strong knowledge of performance optimization, concurrency, and large-scale data transformations.
- Extensive experience with dbt and Snowflake, including modular transformation architecture, business logic, query optimization, clustering, RBAC, and cost management.
- Extensive production experience with Airflow, including orchestration, dependencies, recovery strategies, and cost-aware scheduling.
- Solid AWS experience across services such as S3, Lambda, Glue, ECS, and RDS.
- Practical experience using AI coding tools such as Claude Code, Cursor, or equivalent, with evidence of real production work delivered through these tools.
- Strong product mindset and engineering judgment, with an ability to connect platform decisions to data quality and customer outcomes.
- Experience with evaluation frameworks, LLM tracing, embeddings, vector search, fuzzy matching, Spark/PySpark, streaming technologies, B2B data, or data privacy and compliance would be highly valued.
- Experience in startup or scale-up environments where you have established technical standards and architectures rather than simply inherited them is a plus.
Benefits
- Fully remote position with the flexibility to work from India.
- Opportunity to work on greenfield data and AI infrastructure at significant scale.
- High technical ownership, with the ability to shape architecture, engineering standards, and production practices.
- Exposure to frontier challenges involving LLM evaluation, AI-powered data pipelines, entity resolution, observability, and large-scale data quality.
- Lean, senior engineering environment with minimal hierarchy and strong individual ownership.
- Fast-paced delivery culture with frequent releases and end-to-end responsibility for the technology stack.
- Competitive base salary with meaningful equity participation.
- Opportunity to build systems that directly influence the quality and reliability of customer-facing data.
- Inclusive environment that values diverse perspectives, collaboration, professional growth, and technical excellence.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Data Engineer - Splunk and Azure
Bosch Group Bengaluru, Karnataka, India
-
Data Engineer
BID Operations Shenzhen, Guangdong Province, China
-
Data Engineer
Prodigal Bengaluru, Karnataka, India
-
Lead Data Engineer
Bristlecone Noida, Uttar Pradesh, India
-
Quantitative Data Engineer
Qube Research & Technologies Hong Kong, Hong Kong Island, Hong Kong S.A.R.
-
Data Engineer
CipherHealth United States