Srijan Technologies PVT LTD

Technical Lead - Data Engineer

Srijan Technologies PVT LTD Gurgaon, Haryana, India

IT Services and IT Consulting · 1,001-5,000 employees

13 h ago
data-engineer Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Lead the development of scalable ETL/ELT pipelines and data architectures across multi-cloud platforms like Databricks and Snowflake. Manage data integration, orchestration, and CI/CD workflows while mentoring team members on best practices in data governance and MLOps.

What they look for

Data Engineering Databricks Snowflake AWS Azure Python SQL Apache Airflow Airbyte CI/CD Docker Kubernetes MLOps Data Architecture ETL/ELT REST APIs

Requirements

Requires 5+ years of experience in data engineering with mandatory expertise in data lakes, warehouses, and complex data processing paradigms. Proficiency in Python, SQL, and cloud-based data platforms is essential, along with experience in DevOps and AI-assisted engineering.

Full description

Lead Data Engineer

Overview

We are looking for a Lead Data Engineer who combines hands-on multi-platform expertise with strong leadership in data architecture, pipelines, and CI/CD. This role requires a versatile engineer with deep technical skills across modern data platforms (such as Databricks, Snowflake, AWS, and Azure), an understanding of MLOps/DevOps practices, and the ability to guide a high-performing team in building scalable, production-ready data solutions. You will not be limited to a single platform but will leverage a diverse toolkit to solve complex data challenges.

Key Responsibilities

· Pipeline & Architecture: Lead hands-on development of scalable ETL/ELT pipelines, data models, and integration frameworks to process high-volume (billions of records) structured and unstructured retail data.

· Multi-Platform Engineering: Design, develop, and optimize data processing applications across multiple platforms, including Databricks (Spark/Delta Lake), Snowflake, AWS, or Azure.

· Data Integration & Orchestration: Build and manage robust data pipelines using Apache Airflow for orchestration and Airbyte for seamless data integration and movement.

· Data Processing: Architect and implement robust solutions for Change Data Capture (CDC), large-scale batch processing, and low-latency real-time/streaming data processing.

· API Management: Work extensively with external APIs for data ingestion, as well as design, create, and manage internal REST APIs to serve data to downstream applications and users.

· AI-Augmented Deliverables: Actively leverage AI assistants to conceptualize, design, and accelerate the development of data pipelines and everyday engineering tasks.

· DevOps & CI/CD: Own and evolve CI/CD pipelines (Git workflows, automated testing, release cycles, secrets management, documentation). Guide DevOps-oriented deployments utilizing Dockerized applications, Kubernetes orchestration, and monitoring/logging tools (Splunk, Datadog, Dynatrace).

· MLOps Alignment: Collaborate with Data Scientists on data readiness for ML projects and ensure alignment with ML lifecycle stages (data prep, feature engineering, model deployment).

· Governance & Leadership: Establish and enforce best practices in data governance, data quality, metadata, and security. Mentor team members through peer reviews, knowledge sharing, and technical leadership.

· Innovation: Stay ahead of industry trends in MLOps, observability, and GenAI, introducing relevant tools and practices.

Required Skills & Experience

· Experience: 5+ years of experience in Data Engineering.

· Data Lakes & Warehouses: Mandatory expertise in designing, building, and managing large-scale Data Warehouses and Data Lakes from the ground up.

· Data Processing Paradigms: Extensive, hands-on experience working with Change Data Capture (CDC) mechanisms, complex batch processing, and real-time/streaming data processing.

· Platform Expertise: Proven expertise in more than one major cloud data platform/ecosystem (e.g., Databricks, Snowflake, AWS Analytics, Azure Data Engineering).

· SQL Mastery: Advanced proficiency in writing, optimizing, and debugging complex SQL queries for large-scale data processing and analytics.

· Programming: Strong programming skills in Python (async, threading, decorators, advanced I/O).

· APIs: Strong proficiency in interacting with third-party APIs and hands-on experience creating and managing REST APIs (using frameworks like FastAPI, Flask, or similar).

· Tooling: Deep hands-on experience with workflow orchestration (Apache Airflow) and data integration platforms (Airbyte).

· AI-Assisted Engineering: Mandatory capability to use AI coding assistants and tools to design pipelines, write code, and enhance day-to-day productivity.

· Data Architecture: Experience with data modeling (e.g., Delta Lake or Snowflake architecture) and scalable ETL/ELT design.

· DevOps/CI/CD: Hands-on experience with Git-based CI/CD (GitLab preferred) and a working knowledge of Docker & Kubernetes for deployment and scaling.

· MLOps: Understanding of MLOps concepts including data preparation, model lifecycle, registries, and monitoring.

· Soft Skills: Strong problem-solving skills with the ability to design for scale and performance, coupled with excellent collaboration, communication, and leadership skills.

Good to Have

· Customer Data Platform (CDP): Experience working with, building, or implementing CDPs to unify customer data across systems.

· Experience in the retail domain or other large-scale data-heavy environments.

· Familiarity with streaming frameworks (Kafka, Spark Streaming, etc.).

· Agentic Pipeline Development: Experience or strong interest in building agentic pipelines using LLMs for dynamic data orchestration and automation.

· Knowledge of model observability tools and ML deployment pipelines.

· Exposure to GenAI concepts (vector embeddings, vector databases, RAG).

Similar roles