Innovaccer Analytics

4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

Innovaccer Analytics Noida, Uttar Pradesh, India

Hospitals and Health Care · 1,001-5,000 employees

4 h ago
data-engineer Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will build and operate high-volume data ingestion pipelines using Spark and Trino to transform raw healthcare data into actionable insights. Additionally, you will manage table maintenance, optimize query performance, and implement data quality checks within the Lakehouse architecture.

What they look for

Apache Spark Apache Iceberg Trino SQL Python Java Data engineering AWS Airflow CI/CD Parquet Data pipelines Distributed SQL engines Schema handling Performance tuning

Requirements

The role requires a degree in Computer Science or a related field and at least 5 years of experience in data engineering. Candidates must possess strong SQL skills, proficiency in Apache Spark, and experience with distributed SQL engines like Trino or Presto.

Benefits

Generous leaves Parental leave Sabbatical Health insurance Care program Financial assistance

Full description

Engineering at InnovaccerWith every line of code, we accelerate our customers' success, turning complex challenges into innovative solutions. Collaboratively, we transform each data point we gather into valuable insights for our customers. Join us and be part of a team that's turning dreams of better healthcare into reality, one line of code at a time. Together, we're shaping the future and making a meaningful impact on the world.

About the RoleAs a Senior Software Engineer on the Lakehouse team, you will build and operate the data pipelines at the heart of Innovaccer's on-premise platform: Spark ingestion jobs landing raw healthcare data into Apache Iceberg, and Trino SQL transforms building the layered tables thatpower analytics and applications. You will work hands-on across the full pipeline surface, from file validation and quarantine at ingestion to query performance and table health in serving.

A Day in the Life● Build Spark ingestion jobs that land high-volume raw files into Iceberg tables with schema handling, bad-record quarantine, and idempotent batch replay. ● Develop and operate Trino SQL transform pipelines across data layers: validation and typing, business-rule transforms, MERGE-based deduplication, and aggregate builds. ● Port existing warehouse SQL workloads to Trino and Spark SQL dialects, and validate results against source outputs. ● Automate Iceberg table maintenance: compaction, snapshot expiry, and orphan-file cleanup as scheduled workflows. ● Tune query and pipeline performance: partitioning strategy, file sizing, statistics, and resource-group behavior. ● Instrument pipelines with data-quality checks, reconciliation reports, and alerting, and participate in per-dataset validation during rollout phases.

● B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.

● 5+ years of data engineering experience building production pipelines at scale.

● Strong SQL skills and hands-on experience with Apache Spark for batch processing.

● Experience with Trino/Presto (or a comparable distributed SQL engine) and open table formats: Iceberg preferred, Delta Lake or Hudi acceptable.

● Working knowledge of S3-compatible object storage and columnar file formats (Parquet).

● Experience with workflow orchestration tools (Airflow or equivalent) and CI/CD for data pipelines.

● Professional development experience with Python and/or Java.

● Healthcare data formats (HL7, CCDA, claims files) and regulated-environment experience are pluses.

Here’s What We Offer

  • Generous Leaves: Enjoy generous leave benefits of up to 40 days.
  • Parental Leave: Leverage one of industry's best parental leave policies to spend time with your new addition.
  • Sabbatical: Want to focus on skill development, pursue an academic career, or just take a break? We've got you covered.
  • Health Insurance: We offer comprehensive health insurance to support you and your family, covering medical expenses related to illness, disease, or injury. Extending support to the family members who matter most.
  • Care Program: Whether it’s a celebration or a time of need, we’ve got you covered with care vouchers to mark major life events. Through our Care Vouchers program, employees receive thoughtful gestures for significant personal milestones and moments of need.
  • Financial Assistance: Life happens, and when it does, we’re here to help. Our financial assistance policy offers support through salary advances and personal loans for genuine personal needs, ensuring help is there when you need it most.

Innovaccer is an equal-opportunity employer. We celebrate diversity, and we are committed to fostering an inclusive and diverse workplace where all employees, regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, marital status, or veteran status, feel valued and empowered.

Disclaimer: Innovaccer does not charge fees or require payment from individuals or agencies for securing employment with us. We do not guarantee job spots or engage in any financial transactions related to employment. If you encounter any posts or requests asking for payment or personal information, we strongly advise you to report them immediately to our HR department at px@innovaccer.com. Additionally, please exercise caution and verify the authenticity of any requests before disclosing personal and confidential information, including bank account details.

About InnovaccerInnovaccer builds software that helps hospitals, clinics, and health insurance companies make sense of all the scattered data they deal with every day — patient records, appointments, insurance claims, lab results — and turns it into something useful and actionable.

Think of a typical hospital: patient information often lives in a dozen different systems that don't talk to each other well. This causes delays in scheduling, slows down doctors, and makes it harder to catch health problems early. Innovaccer's platform connects all of that data together, and increasingly, uses AI "agents" to actually do some of the manual work for healthcare teams — things like scheduling appointments, preparing insurance paperwork, or drafting clinical notes — so that staff can spend more time with patients and less time on admin work.

Some of the largest healthcare systems in the US — including CommonSpirit Health, Atlantic Health, and Banner Health — use Innovaccer's platform today.

The company is also investing heavily in this direction: in June 2026, Innovaccer signed a multi-year partnership with AWS to run its AI agents at a much larger scale, using AWS's cloud AI tools (Amazon Bedrock) and healthcare-specific data infrastructure (AWS HealthLake). This is a strong signal of where the platform — and this engineering team — is headed next. For more information, visit www.innovaccer.com and check us out on YouTube, Glassdoor, LinkedIn, Instagram, and the Web.

Similar roles