Vi

Senior Data Engineer

Vi Tel-Aviv, Tel-Aviv District, Israel

Wellness and Fitness Services · 11-50 employees

Sep 03
data-engineer Senior (5-10 yrs) Full-time Israel
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design, build, and operate scalable data pipelines and reusable data models to support the company's AI platform. Additionally, you will ensure data quality, reliability, and production performance while collaborating cross-functionally to solve complex architectural challenges.

What they look for

Python SQL Data Engineering ETL/ELT Apache Spark AWS Data Modeling Data Pipelines Distributed Systems Data Quality Observability CI/CD Docker Infrastructure as Code Apache Iceberg

Requirements

Candidates must have 5+ years of professional experience in data engineering with strong proficiency in Python, SQL, and distributed processing frameworks like Apache Spark. You should also possess deep experience with AWS services, modern data lakehouse architectures, and CI/CD practices.

Full description

  • Vi is an enterprise AI platform for health enterprises - healthcare, biopharma, and wellness. We deploy agentic AI and predictive models into production environments where the output drives next best actions for patients, care teams, and operations to deliver ROI and improve health outcomes.
  • We are looking for a Senior Data Engineer to build and scale the data foundation behind Vi's platform and products. You will own complex data end-to-end - from ingestion and transformation through modeling, quality, observability, and production delivery.
  • This is a hands-on senior IC role for a strong builder who can solve difficult data problems independently, set a high technical bar, and collaborate closely with engineering, DS, and product. You will turn large, fragmented datasets into reliable, reusable capabilities that power every Vi product.

Responsibilities

  • Build and own scalable data pipelines- Design, implement, and operate robust pipelines for high-volume structured and unstructured data, with validation, monitoring, lineage, and recovery built in.
  • Scale the platform for growth- A key near-term initiative is re-architecting the system to support a significantly larger customer base. You will own performance and cost-efficiency across pipelines and services, keeping reliability and operating costs under control as the platform scales.
  • Build across the stack- This is not a pipelines-only role. You will also write backend services and some frontend, including the internal backoffice the team runs on. We hire builders, not narrow specialists.
  • Own the core data tables- Own schema design and evolution, data contracts, and the modeling standards the team follows — naming, shared dimensions, normalization, documentation. Be accountable when a table is wrong, late, or drifting.
  • Level up the team's data work- Pair with and advise software engineers and data scientists on Spark, SQL, and modeling, and help turn notebook-grade code into production-grade pipelines.
  • Partner cross-functionally- Translate product, client, compliance, and business requirements into clear technical designs and dependable production systems.

Requirements

  • Spark at scale- You have tuned real Spark jobs for performance and cost — skew, shuffle, partitioning, memory, spill — run pipelines over TB-scale or billions of rows in production, and can reason about the physical execution plan, not just write DataFrame code.
  • 5+ years of professional experience building and owning production systems.
  • Strong Python and SQL, with maintainable, tested production code.
  • Strong software engineering fundamentals across the stack. You can own backend services and pick up frontend when the work needs it — not a pipelines-only specialist.
  • AI-first way of working- You build with AI in your day-to-day development, using it to move faster and raise the quality of what you ship.
  • Deep experience designing and operating ETL/ELT pipelines, data models, and distributed data-processing systems.
  • Comfortable advising and pairing with other engineers and data scientists on data work.
  • Strong AWS experience: S3, Glue, EMR, Athena, and related compute and orchestration services.
  • Experience with modern data lakehouse or warehouse architectures. Apache Iceberg is a strong advantage.
  • Experience with workflow orchestration (Airflow or similar), CI/CD, Docker, Git, and infrastructure as code such as AWS CDK and CloudFormation.
  • Strong understanding of data quality, schema evolution, lineage, observability, privacy, security, and access controls. Experience with regulated or sensitive data, such as healthcare / PHI, is an advantage.
  • High comfort in a fast-moving environment with incomplete requirements, high ownership, and a strong sense of urgency.

Advantages

null

Why choose Vi?

null

Why choose Muuv?

null

Similar roles