About the role
You will own the end-to-end data infrastructure, including job aggregation, ingestion, enrichment, and user profile modeling. Additionally, you will collaborate with AI/ML teams to improve matching quality and build scalable data pipelines.
What they look for
Requirements
The role requires 5+ years of professional experience in data or backend engineering with strong proficiency in Python and SQL. Candidates must have experience with cloud infrastructure, data modeling, and designing production-grade data pipelines.
Full description
About Jobgether
Jobgether is building the job search platform for experienced professionals navigating the global remote job market.
We aggregate hundreds of thousands of opportunities from companies around the world, structure and enrich this data, and combine it with user profiles and preferences to help people identify the opportunities and companies where they have the strongest fit.
Data is at the core of our product. The quality of our job inventory, user data, enrichment systems and matching models directly determines the quality of the experience we provide.
We are a fully remote, international team focused on building a smarter and more effective way to search for work.
\n
What you will own As our Senior Data Engineer, you will own the data infrastructure behind Jobgether's job marketplace and matching engine.
Your responsibility will span the complete data lifecycle: from discovering and collecting jobs, to cleaning and enriching them, capturing high-quality user data, and making sure both sides of the marketplace can be matched accurately.
This is a highly hands-on role for someone who enjoys building systems, solving messy data problems and taking ownership of data quality at scale.
1/ Job Aggregation & Ingestion
Own and continuously improve the systems responsible for collecting job opportunities from thousands of companies and external sources.
You will:
• Design and maintain scalable job aggregation and scraping pipelines.
• Improve coverage, freshness and reliability of our job inventory.
• Build systems to detect failed sources, missing jobs and ingestion anomalies.
• Manage deduplication and job lifecycle management.
• Improve the scalability and resilience of our aggregation infrastructure.
• Monitor the health and performance of the entire ingestion ecosystem.
2/ Data Quality & Enrichment
Raw job data is only the starting point. You will be responsible for transforming heterogeneous job information into reliable, structured data that can power search, matching and product experiences.
This includes:
• Normalizing job titles, locations, companies, skills, seniority and employment information.
• Improving classification and taxonomy systems.
• Developing automated quality controls and anomaly detection.
• Designing enrichment pipelines using deterministic systems, external data and AI/LLMs.
• Defining and tracking data-quality metrics across the job inventory.
• Identifying systematic quality issues and building solutions rather than manual fixes.
3/ User Data & Profiles
You will also own the data layer on the candidate side of the marketplace.
You will work on:
• Structuring user profiles from CVs, onboarding data, preferences and product interactions.
• Improving how skills, experience, seniority, job preferences and career signals are represented.
• Designing reliable data models that can evolve as Jobgether collects richer career information.
• Ensuring user data is consistent, usable and available to our matching and personalization systems.
• Maintaining strong privacy and data-governance standards.
4/ Matching Quality
Ultimately, our data exists to create better matches between people and opportunities.
You will work closely with Product and AI/ML teams to continuously improve matching quality.
This includes:
• Building the data foundations used by our matching algorithms.
• Identifying missing or unreliable signals affecting match quality.
• Creating datasets and evaluation frameworks to measure matching performance.
• Monitoring match quality across roles, geographies and user segments.
• Helping productionize new scoring, ranking and AI/ML systems.
• Connecting user behavior and outcomes back into the matching system to continuously improve recommendations.
5/ Data Platform & Engineering
Across these areas, you will:
• Design scalable and maintainable data architectures.
• Build and operate reliable ETL/ELT pipelines.
• Improve observability, monitoring and alerting.
• Optimize database performance and data access patterns.
• Maintain clear documentation of data models, pipelines and architecture.
• Contribute to technical decisions around our broader data stack.
• Troubleshoot production issues and take ownership from diagnosis through resolution.
What we are looking for
We're looking for someone with strong data-engineering fundamentals and a product mindset. Someone who's comfortable with messy, real-world data and takes ownership of the results, making sure the pipeline produces best-in-class data to power the product.
Core Requirements
• 5+ years of professional experience in Data Engineering, Backend Engineering or a closely related field.
• Strong professional experience with Python.
• Deep experience designing and operating production data pipelines.
• Strong SQL skills and experience with relational databases such as PostgreSQL or MySQL.
• Experience working with NoSQL databases such as MongoDB.
• Strong understanding of data modeling, schemas and large-scale data transformation.
• Experience with cloud infrastructure, preferably AWS.
• Experience designing monitoring, observability and data-quality systems.
• Strong understanding of APIs, web data ingestion and distributed systems.
• Ability to investigate complex data problems and identify their root causes.
• Strong written and spoken English.
Strong Plus
Experience in one or several of these areas would be particularly relevant:
• Large-scale web scraping or data aggregation.
• Search engines, marketplaces or recommendation systems.
• Job, talent or HR data.
• Entity resolution and deduplication.
• Taxonomies, classification and semantic data enrichment.
• LLM-based data extraction and enrichment.
• Machine Learning pipelines and MLOps.
• Ranking or recommendation systems.
• Docker, Kubernetes and CI/CD environments.
• Data privacy and GDPR-compliant architectures.
This role will suit you particularly well if:
• You like owning a problem end-to-end rather than maintaining one small part of a system.
• Messy data problems interest you more than perfect datasets.
• You naturally investigate why data is wrong, not only why a pipeline failed.
• You think about the product impact of the infrastructure you build.
• You are comfortable making architectural decisions while remaining highly hands-on.
• You prefer automating recurring problems instead of creating manual processes.
• You enjoy working in an environment where there is a lot to build and improve.
Why join Jobgether
High ownership: You will own one of the most critical systems at Jobgether and have significant influence over its architecture.
Real scale: Your systems will process and enrich a large, constantly changing global job inventory and millions of candidate data points.
Direct product impact: Improvements in your work translate directly into better job discovery, matching and recommendations for our users.
Technical challenges: Aggregation, entity resolution, classification, enrichment, search and matching create genuinely complex data problems to solve.
Remote by design: Jobgether is a fully remote company with an international team and a culture focused on accountability and outcomes.
\n
Similar roles
-
Data Engineer H/F
Scalian Paris, Ile-de-France, France
-
Experienced Data Engineer
EY Greece Thessaloniki, Macedonia and Thrace, Greece
-
Data Engineer
EY Greece Thessaloniki, Macedonia and Thrace, Greece
-
Lead Data Engineer
JPMorgan Chase & Co. Bengaluru, Karnataka, India
-
Data Engineer
Reale Mutua Assicurazioni Turin, Piedmont, Italy · €39K/yr
-
Data Engineer
Ayming Lisbon, Portugal