Modash

Senior Product Data Engineer- Data Insights (remote, Europe)

Modash Capital City of Prague, Prague, Czechia · €100K–€130K/yr

Software Development · 51-200 employees

5 h ago
Remote data-engineer Senior (5-10 yrs) Full-time Czechia
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will own the end-to-end development of data products, transforming raw social media data into high-quality, actionable signals for customers. This involves building production pipelines, implementing automated quality checks, and leveraging LLMs to enrich datasets at scale.

What they look for

PySpark Spark Data engineering SQL Python AWS Data quality LLMs Airflow Entity matching System design Data enrichment Infrastructure as code Pulumi Data pipelines

Requirements

You must have solid experience as a data engineer working with large-scale, complex datasets and strong proficiency in Spark and Python. The role requires a senior engineer capable of owning full features from architecture and implementation to release and iteration.

Benefits

Stock options Flexible hours Unlimited paid vacation Personal development support Regular offsites

Full description

Hey, I'm Tadas. I'm hiring a Senior Product Data Engineer for Data Insights, the team that builds Modash's data products.

Modash helps brands find, understand, and work with creators on Instagram, TikTok, and YouTube. More than 2,700 companies, including Stanley 1913, Sennheiser, and NordVPN, use us to run and grow their creator partnerships.

Data Insights turns raw social media data into the datapoints customers buy: which brands a creator has worked with, how to contact them, where they're based, and which other creators are like them. Most of the work is taking messy public data across 200M+ profiles and making it accurate enough that a brand will act on it.

In this role you'll own datapoints end to end, from the raw signal to what the customer sees, and you'll be judged on their quality.

Why we're hiringAt Modash, the data is the product customers pay for. Brands use our collaboration, contact, and location data to decide which creators to work with, and our largest customers take it in regular bulk deliveries through our API.

We need a senior engineer who can take one of these datapoints from a rough idea to a production pipeline and keep improving it after launch. Right now that means building our own creator-location data, using LLMs to find and validate creator contact emails, and matching sponsored posts to the right brand and its parent company.

You'll join a team of four data engineers who work closely with four backend engineers. You'll make important technical decisions yourself, with strong teammates to back you up, and customers will see what you build.

For a feel of how we build software, read our Engineering Blog.

What you'll own1. Build the datapoints customers pay for.

You'll turn raw posts, bios, and profile data into signals: brand collaborations, contact details, creator location, related creators, and links between one creator's accounts on different platforms.

2. Make data quality measurable.

Every datapoint we ship has automated quality checks. You'll decide what good looks like for your datapoint, measure it per platform against real coverage numbers, and catch regressions before customers do.

3. Take projects from idea to production.

You'll shape the problem, scope the work, design the architecture, write the code, ship it, and learn from how it performs. We expect senior engineers to own the outcome, even when the problem arrives half-defined.

4. Use LLMs where they pay off.

We run LLM enrichment at large scale. You'll decide where an LLM beats a rule or a simpler model, and work out what it costs across 200M+ profiles before it ships.

What the day-to-day looks likeMost weeks you're on one project for days at a time, with some of your time going to pipeline issues and code review. A typical week might look like this:

  • Monday. You spot-check your datapoint's output and find a quality problem. You write a short plan for fixing it and get feedback from another data engineer before you start.
  • Tuesday. You build the change in PySpark and test it. It's too slow on the full dataset, so you profile the job and fix it.
  • Wednesday. A pipeline failed overnight. You find the cause, ship a fix, and rerun it. In the afternoon you review a teammate's PR, then get back to your project.
  • Thursday. You compare results before and after the change on each platform. One number moves more than expected, and you dig in until you can explain it.
  • Friday. You merge, watch the first production run, and check the quality reports. You answer a question from a backend engineer about the data, then plan next week.

Every meeting needs a reason, and we protect time for focused work. You'll have a short standup, pair with people when it helps, and get long stretches to plan, build, and ship.

What you've done before• Built data systems at meaningful scale. You have solid experience as a data engineer and have worked with large, complex datasets in production.

  • Worked deeply with Spark. PySpark, Scala, or Databricks all count. We use PySpark, but strong Spark fundamentals matter more than the exact flavour. Here is the blog post that helps explain the kind of work we do.
  • Turned messy data into signals people rely on. Entity matching, classification, deduplication, or enrichment: you've built reliable output from data that started out incomplete, inconsistent, or wrong.
  • Treated data quality as part of the job. You've defined quality checks, tracked coverage and accuracy, and handled a quality drop like a bug.
  • Owned full features. You've taken a feature through planning, scoping, architecture, implementation, release, and iteration.
  • Written production Python and SQL. You care about code quality, system design, and maintainability.
  • Worked with orchestration and cloud infrastructure. Experience with Airflow or AWS Step Functions, and with services such as EMR, Glue, Athena, DynamoDB, S3, Kinesis, Lambda, or ECS, will help you get moving quickly.
  • Worked using agentic development. You use coding agents like Cursor in your daily work, give them clear context, and review what they write as carefully as a teammate's PR.
  • Worked autonomously without working alone. You make progress with incomplete information, say what you think, ask for feedback, and help teammates do better work.

Curiosity about the creator economy helps too, but we'll get you up to speed.

Our stack• AWS, with Pulumi for infrastructure as code

  • Spark on EMR, mostly PySpark
  • Iceberg tables on S3, queried through Glue and Athena
  • PostgreSQL on Aurora
  • Airflow
  • Great Expectations for data quality checks
  • LLM models for data enrichment
  • DynamoDB, Kinesis, Lambda, and ECS
  • GitHub, Notion, Linear, Cursor

The interview processWe move quickly and can finish the process in under a week:

  • Intro chat
  • Two technical interviews: an agentic PySpark coding challenge and a system design session
  • Team fit and project presentation
  • Culture and alignment conversation with our CEO, Avery Shrader

What we offer• Fully remote in Europe 🏠 Work from wherever you do your best work.

  • Compensation. Your compensation is made up of salary and stock options. As we're growing fast, the stock option package is especially significant. Annual salary range is €100,000 to €130,000. We hire across Europe, so the exact number depends on your location, employment type, skills, and experience.
  • Flexible hours ⏱ We care about outcomes, not when you log on.
  • Unlimited paid vacation 🌴 Take the time you need to stay rested and do great work.
  • Personal development support 🧠 Courses, books, and conferences are on us.
  • Real ownership 💡 Take meaningful customer problems from ambiguity to impact.
  • Regular offsites ✈️ We’re remote-first, but we make time to connect in person.

And a little more about us...

Founded in 2018 by a high-school dropout and a Canadian (yes, we’re also shocked it’s going so well), Modash is building a suite of tools that help brands scale partnerships with online content creators.

2,700+ companies like Stanley 1913, Sennheiser, and NordVPN already use Modash to manage and scale their influencer marketing work. And we're just getting started. Over the coming decade, brand investment in creators will continue to boom, and Modash will be at the centre of it all.

Modash is here to stay. We have 8-figures in ARR across two products, a $12M series A investment, and we are default alive. We are building a company that will still be here in 20 years; not rushing towards an exit.

We’re almost 100 people distributed across 20+ countries, operating with a fast, async-first culture. If you join Modash, you’ll be surrounded by people who truly want to be the greatest at their craft. People who make you better. Interesting people too, who have done everything from building solar cars, to hanging out with Metallica and Bon Jovi.

Come join us. Be great, do great things, create great memories, all while making a great impact. Do it.

Similar roles