Tango

ML Engineer - Tabular Data & Experimentation

Tango Warsaw, Masovian Voivodeship, Poland

Consumer Services · 201-500 employees

13 h ago
Senior (5-10 yrs) Full-time Poland
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The role focuses on building end-to-end machine learning models on tabular and panel data while designing rigorous experiments to validate performance. You will translate complex business questions into ML formulations and determine the most effective methods for offline and online evaluation.

What they look for

Machine Learning Tabular Data Experimentation A/B Testing Python SQL Causal Inference Gradient Boosting LLM Feature Engineering Statistical Analysis Data Modeling Drift Monitoring Recommender Systems Power Analysis

Requirements

Candidates must have 5+ years of applied machine learning experience with a deep understanding of tabular data, causal inference, and statistical experimentation. Proficiency in production-grade Python and SQL is required, along with the ability to justify the use of specific models like GBDTs or LLMs.

Benefits

Stock options Competitive salary Medical insurance Lunch budget Parking Multisport card

Full description

Tango is a successful, market leader, a live-streaming Platform with 450+ Million registered users, in an industry projected to reach $240 BILLION in the next couple of years.

The B2C platform, based on the best-quality global video technology, allows millions of talented people around the world to create their own live content, engage with their fans, and monetize their talents.

Tango live stream was founded in 2018 and is powered by 500+ global employees operating in a culture of growth, learning, and success!

The Tango team is a vigorous cocktail of hard workers, creative brains, energizers, geeks, overachievers, athletes, and more. We push the limits to bring our app from “one of the top” to “the leader”.

The best way to describe Tango's work style is not to use the word “impossible”. We believe that success is a thorny path that runs on sleepless nights, corporate parties, tough releases, and of course our users' smiles (and as we are a LIVE app, we truly get to see our users all around the world smiling right in front of us in real-time!).

Do you want to join the party?

Responsibilities

The core of this role is building ML on tabular and panel data, and designing the experiments that tell you whether it works. Most of the interesting questions are about experimentation: how to design an A/B that's worth running, how to extract a credible answer from offline data when an online experiment isn't feasible, and how to know when offline analysis is enough and when it isn't. This applies whether the underlying model is a gradient boosting model, a classical recommender, or an LLM-based component. LLMs are one tool among several here - chosen on merits against the alternatives, and held to the same experimental rigor as anything else the team ships.

  • End-to-end ML on tabular and panel data: feature engineering, validation strategy including time-aware splits, gradient boosting and related methods, calibration, drift monitoring.
  • Designing A/B tests that answer real questions: hypothesis design, power and MDE, handling peeking and multiple testing, variance reduction (CUPED and similar), interference and network effects.
  • Extracting credible answers from offline data when online experimentation isn't feasible - DiD, synthetic control, IV, uplift modeling - and choosing the right method for the situation rather than the most fashionable one.
  • Knowing when offline evaluation is sufficient and when it isn't, and being willing to defend that judgment.
  • Designing evaluations for LLM-based components the team owns: offline metrics, online proxies, drift detection, power analysis, guardrail metrics - with the same rigor as classical ML.
  • Translating business questions into ML formulations: metrics, loss, constraints, the trade-offs that actually matter to the product.

Requirements

  • 5+ years in applied ML with real product impact, not Kaggle-only.
  • Deep working knowledge of tabular and panel data: temporal leakage, non-stationarity, the ways these problems differ from i.i.d.
  • Solid statistics and experimentation - comfortable designing a test from scratch, and equally comfortable pushing back when someone wants to "just look at the p-value."
  • Causal inference at the depth where you can pick a method, justify it, and explain what it doesn't tell you.
  • Production-grade Python - maintainable code, not just research notebooks.
  • SQL at production-analytics level; able to get to the right sample independently, including non-trivial joins, window functions, and reasoning about query plans.
  • A clear view of when an LLM is the right tool for a problem and when a GBDT or classical method is, and the ability to defend either choice with evaluations and cost.

Nice to have

  • Building or substantially reworking a production recommender - two-tower, GBDT ranking, rerankers, candidate generation and ranking at scale.
  • LLM-augmented retrieval and ranking: semantic retrieval, LLM rerankers, embedding-based candidate generation.
  • Causal ML tooling (DoubleML, EconML).
  • MLflow, W&B, feature stores.
  • LLM-as-judge methodology and a working understanding of its failure modes.
  • Prompt optimization as an empirical practice (DSPy-style or hand-rolled).

What we offer:

  • Stock options grant (we’re a Silicon Valley Company)
  • Competitive salary
  • Medical insurance for you and 75% off for your relatives
  • On-site position with 4 days at the office and 1 day WFH
  • Budget for lunch
  • Parking
  • Multisport card
  • Cheerful team spirit and fun office atmosphere

If this sounds like you, apply and help empower live entertainers and creators to build independent businesses around their live talents.

#LI-YT1