Sr. Full-Stack Engineer, Data Systems
Datavations New York, New York, United States · $150K–$170K/yr
Software Development · 51-200 employees
About the role
You will own the end-to-end domain of product taxonomy and attribute systems, including the LLM-based extraction engine and data transformation layers. Additionally, you will manage the orchestration, reliability, and application development for both internal and customer-facing data platforms.
What they look for
Requirements
Candidates must have at least 5 years of experience in Cloud Infrastructure, SRE, or Platform Engineering with strong expertise in AWS and IaC. Proficiency in Python, SQL, dbt, and full-stack development with TypeScript/Next.js is required, along with experience in AI/LLM-based systems.
Benefits
Full description
About Datavations
Datavations is a leading New York-based data and AI software specializing in the $2.3 trillion dollar building materials industry. Datavations gives building materials and home improvement manufacturers real-time, store-level visibility into pricing, assortment, and inventory across major retailers, including Home Depot, Lowe’s, and Menards. Manufacturers use this data to make sharper decisions about pricing, distribution, and how they show up on the shelf.
About the role
Product taxonomy and attributes form the foundation of every insight Datavations delivers to our customers. We are looking for an engineer to own this domain end-to-end—driving both the data systems that organize the market and the applications that enable our teams and customers to interact with them. This is a hands-on role with significant architectural latitude for someone who wants to take full ownership of a system that already powers the business.
WHAT YOU'LL OWN
- The attribute extraction engine — LLM-based extraction at production scale, its configuration model, its quality gates, and its cost profile.
- The transformation layer for this domain — the dbt models that turn extracted values into published attributes: standardization, and the override model that lets human judgment reliably beat the machine. Built on our ClickHouse warehouse alongside the data platform team.
- Orchestration and reliability for these pipelines — the Dagster jobs, sensors, and schedules that run taxonomy and attribute processing, including run monitoring, retries, alerting, and recovery tooling. Reliability is not a separate team here; for this domain, it is this seat.
- The applications — internal and customer-facing. You own the screens people actually use, not only the services behind them: the internal platform our teams run taxonomy and attributes from, and the customer-facing views of this data as they move onto it. React/Next.js front end, backend services and APIs, application architecture, deployment and CI/CD.
- The write path — every change validated, logged, and visible before it ships, with approval where it matters.
- Data quality and observability — automated testing on the data itself, plus statistical detection of what goes wrong quietly: outliers, drift, a retailer that stopped updating, a value that flips between runs.
- Engineering standards for the domain — documentation, tests, runbooks, and review culture as the system and the team around it scale.
BUILDING WITH AI, SPECIFICALLY
- A meaningful part of this platform is AI, and not bolted on the side: an LLM extraction engine already running at scale, a rules engine that produces a confidence score per item, and an in-app assistant that answers questions in natural language and proposes changes for review. We are looking for someone who has built this kind of system properly, not someone who has called a completion endpoint.
- Evaluation before assertion — golden sets, regression suites that run when a prompt changes, and a defensible answer to “did that make it better?” Quality you can measure, not quality you claim.
- Prompts and rules as versioned data — stored, diffable, testable, auditable, with a clear record of what changed and what it did. Not constants in a file that move on deploy.
- Tool-calling and MCP — exposing internal systems to agents safely: scoped, read-first, audited. It is how our teams will query and operate the platform, and how we already work internally.
- Cost and latency as design constraints — model tiering, caching, batching, circuit breakers. Knowing what a run costs before it runs, and why a small model first is usually the right answer.
WHAT YOU'LL WALK INTO
- A modernization already in motion, not a greenfield and not a rescue: a production system that has grown fast with the business, a detailed technical map of it, an engaged leadership team, a modern monorepo, and a new internal platform in its first release — with full air cover to do it properly. You own the taxonomy and attribute systems end to end; the warehouse, ingestion, and shared infrastructure are owned alongside you by the data platform team.
REQUIREMENTS
- At least 5 years of experience in Cloud Infrastructure, Site Reliability Engineering (SRE), or Platform Engineering.
- Strong expertise with Terraform and using infrastructure as code (IaC).
- Experience with CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI/CD, ArgoCD).
- Strong expertise in architecting solutions using AWS services.
- Experience in Python, SQL, dbt, and modern orchestration tools (e.g., Dagster, Airflow, Prefect).
- Proficiency with a columnar warehouse (e.g., ClickHouse, Snowflake, BigQuery).
- Experience with full-stack development (TypeScript, Next.js or similar).
- Experience building production data pipelines and maintaining observability/reliability.
- Experience with AI/LLM-based systems, including evaluation and versioning of prompts.
- Experience with data visualization tools (e.g., Tableau, Power BI) and project management tools (e.g., Jira).
YOU'RE A FIT IF
- You think like a product person as well as an engineer — you form opinions about what the people using this should be able to do, and would rather understand the problem behind a request than build the ticket as written.
- You build with AI daily — both in what you ship and in how you work. AI-assisted development is a habit, not a novelty.
- You write things down by default, and can take ownership of an existing codebase without needing its original author in the room.
- Bonus: retail, e-commerce, or point-of-sale data.
YOU'RE NOT A FIT IF
- You want a greenfield with no history, or “it ran green” is your definition of done.
Preferred Location: NY, Chicago, Dallas, Cincinnati, Minneapolis
Why Join Datavations
- Impact at Scale: Influence a $2.3 trillion industry by shaping how data science accelerates ROI for major manufacturers.
- Autonomy & Growth: Enjoy the freedom to experiment with new technologies and see your ideas realized in production.
- Collaborative Culture: Work alongside a supportive team that values positivity, proactive ownership, and continuous learning.
Our values
- Customer Obsession: We are integrated in the industry with a customer-first obsession.
- Proactive Ownership: We have agency for our actions and embody an entrepreneurial mindset.
- Bias for Action: We maintain momentum with a can-do attitude; we value progress over perfection.
- Eager to Grow: We stay curious and hungry to learn, viewing every failure with humility as an opportunity to grow.
- Foster Collaboration: We work across boundaries with quiet competence, welcoming diverse perspectives and offering help unselfishly.
Compensation
The average range for this role is $150,000 - $170,000 depending on experience, skills, and alignment with the role’s responsibilities. This range reflects our current national expectations for qualified candidates. Exceptional candidates based in our NYC office may be considered for a higher range. Total compensation may also include equity, performance bonuses, and a comprehensive benefits package.
We’re committed to paying competitively and equitably, and we regularly review our compensation structures to ensure they align with the market and support our values