Forward Deployed Data Engineer - Houston
Indicium AI Houston, Texas, United States · $180K–$230K/yr
Information Technology & Services · 501-1,000 employees
About the role
You will act as a technical anchor embedded in client accounts to modernize legacy data ecosystems into cloud-native Lakehouses. You will lead architecture design, manage data pipelines, and guide engineering squads to deliver high-stakes software solutions.
What they look for
Requirements
Candidates must possess extensive hands-on experience in data engineering, specifically with Databricks, Python, SQL, and infrastructure-as-code tools like Terraform. Strong communication skills and the ability to bridge the gap between technical execution and executive business strategy are essential.
Benefits
Full description
The Opportunity
As a Forward Deployed Data Engineer (FDE) on our US Launch Team, you will act as the principal technical anchor embedded directly within high-stakes enterprise client accounts (focusing heavily on Financial Services and Healthcare). You won't just be writing specs or advising from afar—you will be on the front lines in client environments, bridging the gap between C-suite data strategy and hands-on production code.
Working at the tip of the spear alongside our elite partners at Databricks and Anthropic, you will lead architecture design, modernize brittle legacy ecosystems into cloud-native Lakehouses, build agentic AI data harnesses, and guide integrated "SWAT" pods to ship working software in weeks rather than months.
Key Responsibilities
• Embedded Technical Leadership: Serve as the primary technical authority on client engagements. Partner directly with client VP/CTO stakeholders to scope architectures, map data domains, and turn complex requirements into execution-ready engineering plans.
• Hands-On Lakehouse Modernization: Roll up your sleeves to write and optimize production PySpark, Databricks SQL, and Delta Live Tables (DLT)—actively refactoring complex legacy logic (Informatica, PL/SQL, legacy stored procedures) into Medallion Architectures.
• Agentic AI & Data Infrastructure: Architect deterministic, secure data "harnesses" that integrate LLM workflows (Anthropic/Claude) into traditional enterprise data pipelines safely and cost-effectively.
• DataOps & IaC Ownership: Write modular Terraform scripts to provision cloud environments (AWS/GCP), manage governance and lineage via Unity Catalog, and enforce CI/CD rigor across projects.
• Pod & Delivery Guidance: Partner with our LATAM-based nearshore engineering squads (600+ experts) as the US technical lead—conducting code reviews, debugging performance bottlenecks, and maintaining high engineering standards.
• Technical Unblocking & Escalation: Act as the ultimate technical safety net. If a pipeline breaks or a deployment stalls at 2:00 AM, you have the hands-on depth to jump into the code, fix the issue, and keep client delivery on track.
Qualifications & Requirements
Technical Mastery & Execution
• Extensive Hands-On Data Engineering: Deep proficiency in Python and advanced SQL with a proven track record of shipping production data software.
• Deep Databricks Platform Expertise: Hands-on experience building, scaling, and tuning Databricks Lakehouse environments (Delta Lake, Unity Catalog, PySpark, DLT, Workflows).
• Modern Stack & IaC Discipline: Expert-level experience with dbt (Core or Enterprise) for data modeling, Terraform for Infrastructure as Code, and Git/Airflow for orchestration and CI/CD.
• Legacy Migration Experience: Demonstrated background migrating enterprise client workloads from legacy platforms (Teradata, Netezza, Informatica, Oracle) to modern cloud infrastructure.
Client Presence & Startup Grit
• The "Player-Coach" Mindset: Equal comfort pitching target architectures to a client CTO and debugging a failing PySpark job or dbt macro alongside junior engineers.
• Executive Communication: Ability to translate messy business requirements into clean architectural patterns and speak authoritatively across both technical and business functions.
• Scrappy Nation-Builder: High autonomy and resilience—thriving in a fast-paced, Series A growth environment without relying on rigid corporate playbooks or large support structures.
Nice-to-Haves
• Prior experience in forward-deployed, technical consulting, or ProServe roles at top-tier agencies or hyper-growth vendors.
• Pragmatic experience building or deploying LLM/GenAI orchestration frameworks (LangChain, LlamaIndex, Anthropic API) into production pipelines.
• Official certifications in Databricks, AWS, GCP, or dbt.
The anticipated base salary range for this role is $180,000 - $230,000. In addition to base pay, this position may be eligible for an annual discretionary bonus. An individual's final salary offer will be determined based on a variety of factors, including geographic location, experience, specialized skills, and qualifications. This compensation range is subject to updates or modifications at the company’s discretion
Similar roles
-
Senior Data Engineer, Economy
Roblox San Mateo, California, United States · $243K–$295K/yr
-
Fabric Senior Data Engineer
EXL Pune, Maharashtra, India
-
Cloud Data Engineer Senior
Paradigma Digital - Nuestras ofertas de Empleo Pozuelo de Alarcón, Community of Madrid, Spain
-
Data Engineer Senior
Paradigma Digital - Nuestras ofertas de Empleo Pozuelo de Alarcón, Community of Madrid, Spain
-
Data Engineer
EXL Atlanta, Georgia, United States
-
Senior Data Engineer
Clarity Innovations Tampa, Florida, United States · $77K–$180K/yr