Principal Data Engineer
HG Insights Pune, Maharashtra, India
Software Development · 51-200 employees
About the role
You will own the end-to-end architecture of the data platform, managing everything from raw data ingestion to the delivery of finished datasets. Additionally, you will build quality control layers, improve entity matching systems, and optimize compute costs across the infrastructure.
What they look for
Requirements
The role requires over 15 years of experience in building production-grade data engineering systems, with at least 5 years in a staff or principal capacity. Deep expertise in Spark, Databricks, SQL, and AWS ecosystems is essential, along with strong software engineering discipline.
Benefits
Full description
About HG HG Insights is the pioneer of Revenue Growth Intelligence. For more than a decade, we have delivered comprehensive, AI- driven datasets on B2B buyers, technology adoption, IT spend, and buyer intent, sourced from billions of data points. Today, we are a trusted partner to Fortune 500 technology companies, hyperscalers, and innovative B2B vendors seeking precise go-to-market analytics and decision-making. Through an evolving suite of AI agents that incorporate first-party data and buyer signals, HG Insights enables AI-powered GTM automation across sales, marketing, RevOps, and data analytics teams, modernizing GTM execution from strategy through activation.
About the Role This is a senior individual contributor role on the team that builds and runs HG's data platform: everything between a raw vendor file landing in our lake and a finished dataset arriving in a customer's hands. It includes the transformation stack that turns messy external signals into product-grade company attributes, the matching systems that decide which company a piece of evidence belongs to and the internal applications our analysts use to curate and correct the result.
You will be based in Pune, working with engineers at our Pune, US & Brazil locations.
What you’ll do
- Own the architecture of the data platform end to end, and make the calls on build vs. buy, batch vs. incremental, and where each workload belongs.
- Build the quality layer that catches silent failure: contracts at every handoff, freshness and completeness monitoring, and pipelines that quarantine bad data rather than publish it.
- Improve entity matching and attribution — and make match decisions explainable to the customers who depend on them.
- Make our release cycle boring. Recurring deliveries should not depend on people watching them.
- Treat compute cost as a first-class engineering metric across Spark, orchestration, and the serving tier.
What we are looking for
- 15+ years building production grade data engineering systems, with a minimum of 5 years in a staff or principal scope.
- Deep Spark and Databricks expertise at multi-terabyte scale, including the operational side of lakehouse table formats — merges, compaction, small files, schema evolution.
- Strong SQL and dimensional modelling, and the judgment to know when to break the rules.
- Experience with performant and scalable OLTP setups (MySQL/Postgres).
- Airflow orchestration (DAGs, operators, sensors) and integration with Spark/Databricks.
- Proven experience in AWS ecosystems (EC2, S3, EMR).
- Hands-on entity resolution, fuzzy matching, or record linkage at scale. This is central to what we do.
- Python/Scala/Java, with real software discipline: testing, CI/CD, code review, infrastructure as code is a plus.
- Experience in Docker - Kubernetes, Terraform is a plus.
- Experience in integrating AI first implementations in traditional data engineering setups will greatly help in shaping our future designs.
- Experience with machine learning pipelines (Spark MLlib, Databricks ML) for predictive analytics.
- Knowledge of data governance frameworks and compliance standards (GDPR, CCPA).
Preferred Qualifications
B2B firmographic, technographic or intent data; streaming and CDC; warehouse-to-lakehouse migration experience.
What we offer/Benefits We take care of our people fully and genuinely. Every HG Insights employee receives comprehensive healthcare coverage for themselves and their families, a monthly wellness reimbursement and access to Able To professional mental health coaching for stress, anxiety and depression support as per the regions. For your financial future, US employees benefit from a 401(k) with up to 5% company matching and all UK employees are enrolled in a pension scheme. Every employee also receives stock options because the people who build this company should share in its success. We invest in your growth through an annual learning reimbursement and a genuine commitment to developing your career at every stage. And through our Culture Club, you will always have a real voice in shaping the kind of workplace we build together.
Equal Opportunity Employer HG Insights is proud to be an Equal Employment Opportunity employer. We do not accept unsolicited resumes from recruiters or employment agencies. In the event of a recruiter or agency submitting a resume or candidate without a signed agreement being in place, we explicitly reserve the right to pursue and hire such candidates without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted directly to hiring managers, are deemed to be the property of HG Insights.
Similar roles
-
Senior Data Engineer
Knowit Poland Warsaw, Masovian Voivodeship, Poland
-
Data Engineer III - Data Platform
Zinnia - Employee Referral Gurugram, Haryana, India
-
Associate Director / Director Data Engineer
Weekday AI Bengaluru, Karnataka, India
-
Data Engineer
AEGEAN Spata, Attica, Greece
-
Data Engineer - OESIS Framework
OPSWAT Ho Chi Minh City, Vietnam
-
GCP Data Engineer
Capco Bengaluru, Karnataka, India