Jobgether

Principal Azure Data Engineer (12 to 15 yrs) (Python / PySpark / SQL/ DataBricks)

Jobgether India

Internet Marketplace Platforms · 11-50 employees

23 h ago
Remote python Principal (10+ yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Develop, optimize, and support large-scale Databricks data pipelines while ensuring high-quality, trusted data models for healthcare products. Collaborate with business stakeholders and clinical leaders to translate requirements into scalable data platform capabilities and AI/ML-enabled solutions.

What they look for

Azure Databricks Python PySpark SQL Data Engineering Data Modeling Healthcare Data Cloud Computing Data Pipelines Machine Learning Artificial Intelligence Software Development Lifecycle Data Quality Metadata Modeling Technical Leadership

Requirements

Requires 12-15 years of data engineering experience with a strong background in healthcare datasets and cloud platforms like Azure. Candidates must possess advanced skills in Python, PySpark, and SQL, along with a solid understanding of data architecture and software development best practices.

Benefits

Remote work opportunities Flexible working hours Stock options Mentorship Professional development opportunities

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Azure Data Engineer (12 to 15 yrs) (Python / PySpark / SQL/ DataBricks) based in India.

This is a senior technical role focused on building and evolving large-scale data platforms that support healthcare products and improve patient outcomes. You’ll work with business-critical healthcare data across diverse sources, transforming complex information into reliable, trusted datasets and data products. The role combines hands-on engineering with architectural thinking across Azure, Databricks, Python, PySpark, and SQL. You’ll partner closely with product teams, data engineers, business stakeholders, and clinical leaders to translate needs into scalable solutions. You’ll also play a key role in data quality, unified data models, AI/ML enablement, and platform evolution. This is an opportunity to make a tangible impact in a data-rich healthcare environment while working independently and driving engineering best practices.

\n

Accountabilities

  • Develop new Databricks pipelines and enhance, optimize, and support existing production data pipelines.
  • Work with business-critical healthcare data while maintaining close collaboration with business stakeholders to ensure reliability and usability.
  • Apply data engineering best practices to create high-quality, highly available, intuitive, and trusted data models.
  • Define and implement new entities within unified data models to support evolving business and product requirements.
  • Partner with business customers to gather requirements and translate them into new datasets and data-platform capabilities.
  • Identify, investigate, and resolve data-quality issues proactively to improve the reliability and user experience of data products.
  • Extract, transform, combine, and manage data from diverse and heterogeneous healthcare data sources.
  • Design, implement, and support platforms that enable efficient ad-hoc access to large-scale datasets.
  • Develop data and metadata models that support machine learning, artificial intelligence, analytics, and downstream data products.
  • Support data integration activities associated with acquisitions and other organizational initiatives.
  • Apply disciplined software development practices across documentation, coding standards, code reviews, source control, deployment, testing, and operations.
  • Provide technical leadership through independent problem-solving, proactive identification of opportunities, and continuous improvement of data engineering practices.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, or a related field; a Master’s degree in a healthcare-related discipline is preferred.
  • 12–15 years of overall data engineering or related technology experience, including substantial experience working with healthcare data.
  • 10+ years of experience working with diverse healthcare datasets such as claims, encounters, ADT, FHIR, EHR, or comparable clinical and healthcare data.
  • 8+ years of experience working with cloud-based services across Azure, AWS, or GCP, with strong hands-on exposure to Microsoft Azure preferred.
  • Advanced practical skills in Python, PySpark/Spark, and SQL, including strong debugging capabilities for complex, business-critical data workloads.
  • Strong experience with Databricks pipelines, including developing, enhancing, troubleshooting, and supporting production environments.
  • Solid understanding of database structures, principles, practices, and data architecture.
  • Experience with normalized, dimensional, star-schema, and snowflake-schema data modelling approaches.
  • Demonstrated understanding of the full software development lifecycle, including documentation, coding standards, code reviews, source control, deployment, testing, and operational support.
  • Experience building scalable data platforms and integrating data from multiple heterogeneous sources.
  • Ability to model data and metadata for machine learning and AI use cases.
  • Strong written and verbal communication skills, with the ability to collaborate effectively with technical teams, business stakeholders, and clinical leaders.
  • Naturally curious and analytical mindset, with the ability to understand what data represents, how it is used, and how it can create downstream value.
  • Proactive approach to identifying data-quality problems and improving common data use cases.
  • Ability to work independently, take ownership of complex technical challenges, and operate effectively in a fast-paced environment.
  • Databricks or Microsoft Azure certification is a plus.

Benefits

  • Remote work opportunities from India.
  • Flexible working hours.
  • Stock options and participation in the organization’s long-term growth.
  • Benefits that go beyond statutory requirements.
  • Friendly, trusting, and collaborative work environment.
  • Opportunities to collaborate with experienced technology professionals.
  • Access to mentorship and ongoing learning and professional development opportunities.
  • Opportunity to work on large-scale healthcare data and technology initiatives with meaningful impact on patient outcomes.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Similar roles