Senior Data Engineer, Blockchain data and/or NLP pipelines
Inca Digital, Inc. United States
Software Development · 51-200 employees
About the role
Design, build, and operate end-to-end batch and incremental ETL/ELT pipelines to transform raw data into production-grade APIs and dashboards. Collaborate with cross-functional teams to translate intelligence requirements into scalable products while ensuring data quality and reproducibility.
What they look for
Requirements
Requires advanced proficiency in Python and SQL with proven experience building production ETL/ELT pipelines using modern orchestrators. Candidates must demonstrate deep expertise in either blockchain ledger data or NLP/LLM production text processing.
Benefits
Full description
About Us: Inca Digital is a veteran-owned data and intelligence company specializing in digital-asset analytics for exchanges, financial institutions, regulators, and blockchain ecosystems. Our technology and expertise provide clarity across crypto markets, tracking blockchain transactions, liquidity movements, and illicit finance, helping clients identify risk, enhance transparency, and improve decision-making. Inca’s infrastructure fuses structured and unstructured data from blockchains, exchanges, social networks, and financial markets. The result is a powerful analytics engine that supports ecosystem monitoring, market surveillance, and counter–illicit finance intelligence across digital-asset networks. Inca operates as a fast-paced, nimble, global, and remote technology company. We leverage an asynchronous‑first workflow, try to minimize time spent on meetings, and believe in open debate and logic over authority. Work alongside some of the sharpest minds in the world, including intelligence analysts, defense veterans, data engineers, quant researchers, linguists, and more.
Domain
- Blockchain data: direct experience treating ledger data as data. You've pulled from RPC endpoints, archive nodes, or indexers; decoded logs and ABIs; handled reorgs and chain-specific quirks; and understand UTXO vs. account-based models in practice.
- NLP / LLM: you've shipped production text processing — extraction, classification, embeddings and vector search, or LLM-in-the-loop enrichment — with a clear view on evaluation, cost, and failure modes.
We expect real depth in one and working competence in the other; tell us which is which in your application.
What You'll Own
- Design, build, and operate end-to-end batch and incremental ETL/ELT pipelines, turning raw data into production-grade APIs, dashboards, and automated alerts.
- High-throughput data acquisition from diverse sources: relational databases, third-party and vendor APIs, flat files, blockchain nodes and indexers, and social and web text at scale.
- NLP and LLM-assisted processing over unstructured text — entity and claim extraction, classification, deduplication, and enrichment feeding downstream risk models.
- Model and query data across relational, object, and graph (Neo4j) stores, choosing the right one for the access pattern rather than defaulting to a favorite.
- Provenance and reproducibility: raw captures immutable, datasets versioned, and any published finding reproducible as of the date it was made.
- Data quality as a first-class deliverable: validation, lineage, reconciliation, and freshness monitoring.
- Partner with the Head of R&D, engineers, data scientists, and our investigations team to translate intelligence requirements into scalable client-facing products.
- Contribute to technical design and architecture in a written, argued design process.
What We Require
- Advanced Python and SQL.
- Proven experience building and operating production ETL/ELT pipelines under an orchestrator (Airflow, Dagster, Prefect, Step Functions, or equivalent), including backfills, idempotency, retries, schema drift, and on-call for your own pipelines.
- Hands-on with PostgreSQL/RDS and object storage (S3).
- Serving data to consumers: building and versioning APIs (Kong OSS, FastAPI, or similar) and/or feeding dashboards.
- Docker, Git, CI/CD (GitHub Actions, AWS CodeBuild/CodePipeline), and IaC (Terraform, CDK, or equivalent).
Nice to Have
- Streaming and near-real-time architectures (Kafka, AWS SQS/SNS).
- Document stores (MongoDB) and analytical engines or warehouses (ClickHouse, DuckDB, Snowflake, Athena).
- Graph data modeling and querying (Neo4j/Cypher or comparable).
- Transformation and data-quality tooling (dbt, Great Expectations, Soda, OpenLineage).
- Vector stores (pgvector, Qdrant, OpenSearch).
- Go; Kubernetes.
- Domain:
- Blockchain ledger data, smart contracts, or web3 languages (Solidity, Vyper, Go, Bitcoin Script)
- NLP techniques, LLM integrations, or model feature engineering
- Financial services, trading venues, or regulatory frameworks (SEC, CFTC, FinCEN)
- OSINT, dark web, or social platform data collection
Why join us?
- We operate in a fully remote, high-trust environment and prioritize impact and delivered results over hours logged.
- We offer competitive compensation and healthcare stipend.
- We invest in our people through childcare stipends and dedicated employee assistance resources.
- Join a thought leader at the intersection of national security and digital assets, working with cutting-edge tech that defines the field.
Curious about how we build at Inca Digital? Hear directly from our Head of R&D on how our team builds the infrastructure that powers Inca's intelligence, merging financial, natural language, and blockchain data into custom dashboards and automation tools built around each client's needs.
🎥 Watch on YouTube: Inside Inca | The Engine Behind Inca's Intelligence
Inclusion & Equal Opportunity: As a veteran-owned company, Inca Digital thrives on diverse perspectives, specifically across race, gender, neurodivergence, and veteran status, to solve complex data challenges. We are an Equal Opportunity Employer; all qualified applicants receive consideration without regard to protected status, including disability or national origin.
Similar roles
-
Data Engineer
Green River Data Analysis Vermont, United States · $120K–$135K/yr
-
Data Engineer Mid
Bluetab, an IBM Company San Isidro, Lima, Peru
-
Data Engineer Jr
Bluetab, an IBM Company Bogota, Capital District, RAP (Especial) Central, Colombia
-
Senior Data Engineer
Sporty Group São José da Laje, Alagoas, Brazil
-
Senior Data Engineer
Tastewise Tel Aviv, Tel-Aviv District, Israel
-
Senior Software Engineer (Data Engineer)
NielsenIQ Chennai, Tamil Nadu, India