Data Scientist
Oxylabs Vilnius, Vilnius County, Lithuania · €42K–€82K/yr
IT Services and IT Consulting · 201-500 employees
About the role
You will develop and maintain entity resolution techniques and ML-based extraction models to transform unstructured data into structured outputs. Additionally, you will automate scraping workflows and monitor applied ML models to ensure data accuracy and reliability at scale.
What they look for
Requirements
The role requires 4+ years of experience as a Data Scientist, Data Analyst, or Data Engineer with a strong background in ML, NLP, and data modeling. Proficiency in Python, SQL, and Spark is essential, along with strong structural thinking and communication skills.
Benefits
Full description
We’re a team of 500+ professionals who develop cutting-edge proxy and web data scraping solutions for thousands of the world’s best known businesses, including Fortune 500 companies.
What’s in store for you:
You’ll be developing complex products with high coding standards, maintaining our own infrastructure, handling petabytes of data, and solving challenges on a daily basis. We got you covered with a team of strong professionals to support you, a well-built tech stack, and loads of ownership.
The team waiting for you:
Our Engineering team manages a powerful data platform that has unlocked a growing backlog of applied ML work. As a Data Scientist, you will take dedicated ownership of this backlog, driving initiatives like entity resolution, LLM-based structured extraction, and topic modeling. We are rapidly scaling our source acquisition and ingestion workflows, making this the perfect time to join and shape our capabilities. By turning unstructured data into structured, product-ready outputs, you will solve the core challenge of ensuring data accuracy and reliability at scale.
\n
In this role, you will:
- Develop and maintain entity resolution techniques to match and link records across sources with inconsistent identifiers.
- Prototype, build, and evaluate LLM and ML-based extraction and inference models that turn unstructured or free-text data into structured outputs.
- Automate scraping logic and source queue generation using data-driven models to streamline and scale source acquisition workflows.
- Build, iterate on, and monitor applied ML models to proactively identify data or concept drift.
- Enrich and enhance datasets with product-ready calculated fields and derived metrics.
- Assist in building automated or ad-hoc QA validation processes to validate model output accuracy, reliability, and consistency.
Your skills & experiences
- 4+ years of previous experience as a Data Scientist, Data Analyst, or Data Engineer, with a strong background in data modeling, ML, and NLP.
- Excellent programming skills in Python, proficiency in SQL, and hands-on experience with Spark.
- Deep, structural thinking with the ability to communicate clearly and align with various stakeholders.
- A collaborative, self-driven mindset with a keen attention to detail and a propensity to dig into deeper layers to inspire improvements.
Nice to Have:
- Experience with Dagster, dbt, Superset, or Trino.
- Previous experience working closely in a team with data engineers.
- Strong business acumen and an understanding of how insights convert to value and unlock new revenue.
- Excellent written and spoken English.
Tech stack:
- SQL,
- Python,
- Spark,
- Dagster,
- dbt,
- Superset,
- Trino.
Salary & Benefits:
- Gross salary: 3500 - 6800 EUR/month. Keep in mind that we are open to discuss a different salary based on your skills and competencies.
- Growth & Learning: 40+ internal learning options, external conferences, mentorship, and year-round knowledge-sharing.
- Health & Well-being: Private health insurance, psychotherapy, on-site well-being consultants, 24/7 gym access, and a wellness app.
- Celebration & Community: Team events, an overseas workation, quarterly team-building budgets, and plenty of ways to mark milestones together.
Up for the challenge? Let’s talk!
\nUp for the challenge? Let’s talk!
Similar roles
-
Data Scientist
EXL London, England, United Kingdom
-
Data Scientist
TECDATA ENGINEERING Madrid, Community of Madrid, Spain
-
Senior Data Scientist
Ignite IT Suitland, Maryland, United States
-
Data Scientist
Bluetab, an IBM Company San Isidro, Lima, Peru
-
Data Scientist, AI Ongoing Delivery
Braze London, England, United Kingdom
-
Data Scientist II (AI Deployment)
Braze São Paulo, Brazil