Staff Software Engineer, Data & AI
Holly New York, New York, United States · $202K–$262K/yr
Technology, Information and Internet · 2-10 employees
About the role
You will own the technical direction and health of the Data Platform, designing architecture for ingestion, processing, and AI-powered data systems. You will lead cross-team initiatives to turn messy public data into clean, canonical datasets that product teams rely on.
What they look for
Requirements
Candidates must have 8+ years of experience in production software engineering with a focus on Data or AI platforms. You should be proficient in data modeling, SQL, and building scalable ETL pipelines, with a strong ability to lead architecture across teams.
Benefits
Full description
About the Role
We're hiring a Staff Software Engineer, Data & AI to own the technical direction of Holly's Data Platform. Local governments publish enormous amounts of public information salary schedules, job classifications, MOUs (Labor Union Agreements), budgets but it's scattered across thousands of websites and buried in messy formats: scanned PDFs, inconsistent HTML, spreadsheets, and everything in between. Turning that chaos into clean, searchable, trustworthy data is one of our biggest product challenges and deepest moats.
You'll set direction for the Data Platform: ingestion and processing pipelines, canonical datasets, semantic search and retrieval, and inference APIs that product teams build on. AI is part of the infrastructure, not a layer added later. You'll establish how Holly uses models for extraction, classification, normalization, matching, and search, and how we evaluate and operate those systems in production.
This is a hands-on Staff role with domain-wide scope. You'll set direction across ingestion, processing, modeling, search, inference, quality, and product consumption; resolve architecture decisions that span teams; and make other engineers more effective on the platform. You'll work directly with the founders and partner across Product and Engineering to decide which Data Platform investments matter most.
If you've worked with large, high-volume data and love the challenge of taming messy real-world inputs into something people can rely on, we'd love to talk.
What You'll Do
You'll own the technical health of Holly's Data Platform, lead cross-team decisions, and build the data and AI systems and standards the rest of the product depends on.
Own and Build the Data Platform
- Own the data platform end-to-end from raw public sources to clean, canonical datasets the product consumes
- Design the architecture, schemas, and standards for how Holly ingests, models, and trusts its data
- Partner with the founders to identify, scope, and de-risk our highest-leverage data initiatives
- Set technical direction for the data domain across teams
Build Ingestion & Normalization Pipelines
- Build systems that collect large volumes of public government data from thousands of local-government sources across the web
- Turn messy, heterogeneous inputs scanned PDFs, inconsistent HTML, spreadsheets into structured, normalized data (parsing, extraction, OCR, dedupe, entity resolution, schema mapping)
- Where it adds leverage, incorporate LLM-assisted extraction and embeddings into the pipeline
- Build for freshness, reliability, and scale so data stays current and trustworthy
Model & Serve Data for the Product
- Design canonical data models and domain schemas that product engineers build on
- Expose clean, versioned, well-documented datasets the main app can reliably consume
- Own data quality, validation, lineage, and observability so downstream teams can trust what they're building on
Own Search & AI Infrastructure
- Set the architecture for semantic search and retrieval over large, changing government datasets
- Design inference APIs for extraction, classification, matching, and other AI-powered data processing
- Establish evals, tracing, fallbacks, and cost controls for model-backed systems
- Create reusable interfaces that let product teams ship Data & AI capabilities safely
Make It Automated & Intelligent
- Evolve pipelines from manual/one-off toward automated, self-healing, monitored systems
- Establish data-quality checks, alerting, and standards that keep the platform reliable as it grows
- Establish patterns other engineers can use to collect, validate, and serve data safely
- Lead cross-team architecture work and make the system easier for others to extend
How We Work
Six principles drive how we build:
- Work on What Matters, Default to No every yes has a cost, so we save them for what moves the business and spend time on what matters.
- Question Everything, Be Opinionated titles don't settle arguments, the better case wins. Feel empowered to push back to everyone from the Head of Engineering to one of the founders.
- Obsess Over Craft quality first, and we don't trade it for a date. If you wouldn't put your name on it, it shouldn'''t end up in the codebase.
- Own It End to End if you build it, you own it: to production, in tests, and when it breaks. With great power comes great responsibility, with autonomy comes responsibility to make sure you own your work.
- Ship Small, Ship Often the smallest thing that stands on its own, kept reversible. Small ships compound, are easier to review and easier to fix if there are issues.
- Automate the Hurt, Not the Itch automate the recurring pain, the Toil aka things you do repeatedly that waste time, not the one-off annoyances or what seems ''"fun''" to automate.
What You'll Have
We'd love to talk if you're a Staff engineer who has led a Data or AI Platform across teams and knows how to turn messy real-world inputs into dependable product capabilities.
The Essentials
- Have 8+ years building and shipping production software, or equivalent experience, including sustained Staff-level scope across a Data or AI Platform
- Have led the architecture and operation of production data systems used by multiple teams
- Have worked with large-scale, high-volume data ideally where lots of sources, users, or records make volume and reliability matter
- Are strong at data modeling and SQL, with experience designing schemas that others build on (Postgres a plus)
- Have built and owned ETL/ELT pipelines that handle messy, heterogeneous, real-world inputs (scraped data, PDFs, HTML, spreadsheets)
- Have led production search, retrieval, inference, or AI-assisted data-processing systems
- Know how to evaluate model quality and operate model-backed systems when outputs are probabilistic
- Bring a strong data-quality mindset validation, testing, monitoring, lineage, and reliability are core to how you work
- Turn ambiguous company problems into technical direction, sequenced work, and clear ownership
- Make other engineers better through architecture, mentorship, and reusable standards
- Are pragmatic about tooling and comfortable working in (or ramping quickly into) a modern TypeScript/Postgres codebase
Bonus Points
- Open source contributions to or maintainer of a widely used tool.
- Experience with large-scale web scraping / crawling, document extraction (OCR), or LLM-assisted parsing
- Experience with embeddings / vector search or supporting ML/AI data workflows
- Experience with analytical/columnar or warehouse stacks (ClickHouse, BigQuery, Snowflake) and/or streaming pipelines
- Experience in government, public sector, or civic tech
- Prior early-stage startup experience
Don't meet every bullet? Apply anyway. If you're strong on most of this and excited about the work, we want to hear from you we'll help you ramp on the rest.
What You'll Get
- Domain ownership. Own the technical health and direction of the data systems that define how the product works.
- Cross-team influence. Make the high-leverage calls on data architecture, standards, and how we scale, then help teams execute them.
- Commitment to Open Source. We are big believers in supporting open source, and provide a monthly day of Open Source where you can work on your favorite tool. In addition to internal hackathons and other projects
- Direct access. Work directly with the founders, with autonomy to drive major initiatives end-to-end.
- High-impact scope. Build the data foundation the entire product depends on, and see its power features customers rely on quickly.
- Public-service impact. Your work improves how local governments operate, helping millions of Americans access public-service careers.
- Competitive package. $202k-$262k base, 0.25-0.60% equity (L4), comprehensive health benefits (platinum plan with vision and dental), 401(k), paid parental leave, and a professional development stipend.
Ready to Join Us? A few important notes:
Location: This is an onsite role based out of our New York City HQ, four days a week (typically Mondays Thursdays), with some flexibility depending on the role and the candidate. Candidates must reside in New York or be able to commute to our NYC office. Applicants must be authorized to work in the U.S. without requiring sponsorship.
Work Philosophy: We'''re an early-stage startup serving government clients with hard deadlines. There may be occasional off-hours work around launches or critical issues (rare and typically planned). We value flexibility and trust you to manage your schedule while maintaining a high bar for responsiveness and customer outcomes.
We're excited to build with you
Team Holly 🌆
www.hollygov.com
Holly is committed to building a diverse company and working with the broadest talent pool possible. We encourage applications from all races, religions, national origins, genders, sexual orientations, gender identities, gender expressions, and ages, as well as veterans and individuals with disabilities. If you need a reasonable accommodation during the application or interview process, let us know.