Data Engineer AI Java/Python/Spark
Infotree Global Solutions Capon Bridge, West Virginia, United States
Staffing and Recruiting · 1,001-5,000 employees
About the role
You will design, build, and maintain scalable data and AI platforms while driving technical initiatives across teams. The role involves developing production-ready software, implementing large-scale data pipelines, and creating GenAI applications using modern engineering practices.
What they look for
Requirements
Candidates must have at least 5 years of professional software or data engineering experience with strong proficiency in Python or Java. You are required to have hands-on experience with distributed data processing, cloud-native technologies, and LLM-based application development.
Full description
Role Overview:
We are looking for a Senior/Lead Data & GenAI Engineer to design, build and maintain scalable, production-grade data and AI platforms.
The role combines software engineering, distributed data processing, cloud-native technologies and Generative AI. You will work across teams, drive technical initiatives and build reusable libraries and frameworks that enable reliable, scalable and testable systems.
Key Responsibilities:
- Develop, test and maintain high-quality, production-ready software.
- Design and implement large-scale data pipelines and distributed processing systems.
- Build scalable cloud-native services and platforms using modern engineering practices.
- Provide technical leadership for cross-team initiatives and complex engineering projects.
- Design and develop reusable libraries, frameworks and platform components.
- Optimize distributed data processing workloads for performance, scalability and reliability.
- Work with data platforms including Databricks, Apache Spark and Snowflake.
- Develop and deploy applications using Python and/or Java.
- Build and operate containerized workloads using Kubernetes and cloud-native technologies.
- Design and implement GenAI/LLM-based applications and services.
- Work with frameworks such as LangChain and LangGraph for LLM orchestration and agentic workflows.
- Collaborate with data scientists, software engineers, architects and product teams.
- Establish engineering best practices around testing, observability, reliability and deployment.
Required Experience:
- 5+ years of professional software/data engineering experience.
- Strong hands-on experience with Python and/or Java.
- Strong experience with Apache Spark and distributed data processing.
- Experience with Databricks and/or modern lakehouse platforms.
- Experience with Snowflake or comparable cloud data warehouses.
- Practical experience with Kubernetes and cloud-native technologies.
- Experience designing and maintaining large-scale data pipelines.
- Strong understanding of distributed systems, scalability and production engineering.
- Experience developing ML/AI or GenAI applications.
- Experience with LLM-based applications, RAG, AI agents or LLM orchestration.
- Familiarity with LangChain, LangGraph or similar GenAI frameworks.
- Strong software engineering fundamentals including testing, code quality and system design.
Nice to Have:
- Experience with AWS, Azure or GCP.
- Experience with streaming technologies such as Kafka.
- Experience with Delta Lake / Lakehouse architecture.
- Experience building RAG pipelines and vector-search solutions.
- Experience with LLM evaluation, observability and productionization.
- Experience with AI agents, tool calling and multi-step workflows.
- Experience building internal developer platforms, frameworks or reusable engineering libraries.
- Experience leading cross-functional or cross-team technical initiatives.
Ideal Candidate Profile:
The strongest candidate is not purely a Data Engineer and not purely an ML Engineer.
We are looking for someone who combines:
Software Engineering + Data Engineering + Cloud/Platform Engineering + GenAI
Typical backgrounds may include:
- Senior Data Engineer
- Lead Data Engineer
- Senior Software Engineer – Data
- Data Platform Engineer
- Senior Cloud Data Engineer
- AI/ML Platform Engineer
- Senior ML Engineer with strong data engineering experience
- GenAI Engineer with strong distributed-data/platform experience
- Data & AI Architect / Technical Lead
Core Technology Stack:
Languages: Python, Java
Data: Apache Spark, Databricks, Snowflake, Delta Lake
Cloud/Platform: Kubernetes, Docker, AWS/Azure/GCP, cloud-native technologies
GenAI/ML: LLMs, RAG, LangChain, LangGraph, AI agents, vector search
Engineering: Distributed systems, APIs, CI/CD, automated testing, observability, scalability
Similar roles
-
Staff Data Engineer
Realtor.com Careers Austin, Texas, United States
-
Software Data Engineer
Apple Cupertino, California, United States
-
Senior Data Engineer
Blend360 Montevideo, Montevideo, Uruguay
-
Data Engineer Manager level
Blend360 Montevideo, Montevideo, Uruguay
-
Lead Data Engineer - Auctions & Outcomes
Kargo New York, New York, United States · $200K–$230K/yr
-
Ingénieur Data / Data Engineer
mthree Recruiting Portal Montreal, Quebec, Canada