About the role
You will own the quality and automation frameworks for a high-scale, low-latency streaming Customer Data Platform. This involves building end-to-end test suites, validating data pipelines, and ensuring performance and governance standards are met.
What they look for
Requirements
Candidates must have 6+ years of experience in software quality engineering with a strong focus on building test automation frameworks. Proficiency in programming languages like Python or Java and experience with data-intensive systems like Kafka and Kubernetes are essential.
Full description
This is a remote position.
Owning quality for a high-scale, low-latency streaming Customer Data Platform deployed on infrastructure we run ourselves (on-prem, Kubernetes), not managed cloud services. The platform ingests from thirty or more source systems at hundreds of thousands of events per second, resolves customer identity, computes customer attributes in real time, and emits governed signals to external destinations. You will build and own the automation that proves this works: API and end-to-end test frameworks, streaming and data-pipeline validation, latency and load testing against explicit SLOs, and governance testing that proves consent, data-protection, and tenant-isolation rules actually hold at runtime. Testing here is evidence-producing, not exploratory only — test outcomes gate whether an artifact is allowed to go live.
Requirements
- 6+ years in software quality engineering, with the majority of your work in test automation rather than manual testing
- Demonstrated ownership of test automation frameworks you built or substantially re-architected — not only writing test cases against someone else's framework (core requirement)
- Ability to write high-quality code in Python, Java, TypeScript, or equivalent languages, including framework structure: base classes, shared utilities, configuration handling, reporting, and test data management
- Strong API test automation experience (REST Assured, Karate, pytest, Postman/Newman, Playwright API, or similar) — test organization, authentication handling, response and schema validation, and contract testing between services
- End-to-end UI automation experience (Playwright, Cypress, Selenium, or similar) — framework architecture, Page Object Model or equivalent, handling dynamic elements, and parallel execution
- Demonstrated ability to diagnose and eliminate flaky tests — proper waits, test isolation, deterministic setup and teardown, and root-cause investigation rather than blanket retries
- Performance and load testing experience (JMeter, k6, Locust, Gatling, or similar) — designing load scenarios, identifying bottlenecks, and validating throughput and latency against stated targets
- Strong SQL for data validation — joins, aggregations, window functions, and source-versus-target reconciliation queries over large datasets
- Practical experience testing data pipelines and ETL/streaming jobs — schema validation, source-to-destination comparison at row and value level, completeness and duplication checks, and sampling strategies for datasets too large to compare in full
- Experience testing event-driven systems and message queues (Kafka or equivalent) — event and payload validation, ordering guarantees, delivery semantics (at-least-once versus exactly-once-effective), idempotency, consumer lag, and offset behaviour
- Experience testing stateful stream processing (Apache Flink or similar) — correctness after restart from checkpoint, savepoint-based upgrades, state restoration, late and out-of-order event handling, and behaviour under backpressure
- Ability to validate end-to-end latency against an explicit SLO — measuring p95 and p99 across pipeline stages, and distinguishing in-scope platform processing from excluded external calls such as third-party API round-trips and destination acknowledgement
- Experience integrating tests into CI/CD pipelines — trigger configuration, test stages, artifact and report publication, and failure gates that block promotion
- Disciplined test data management — setup and cleanup strategy, inter-test dependency handling, and environment-specific data; experience working with synthetic, generated, or masked datasets where production data cannot be used
- Ability to produce structured, auditable test evidence: test results tied to the exact artifact version and configuration under test, so that results are traceable and cannot be silently reused after the artifact changes
- Experience testing data governance and privacy controls — consent enforcement, data classification and usage rules, tokenization and masking, and verifying that raw sensitive identifiers do not appear in downstream topics, logs, or exports
- Experience testing identity resolution or entity matching — deterministic and probabilistic matching outcomes, merge and split behaviour, lifecycle state transitions, and identifier reassignment scenarios
- Experience testing multi-tenant systems — verifying tenant and workspace isolation, role- and attribute-based access control, and absence of cross-tenant data leakage
- Experience validating replay, backfill, and reconciliation — comparing recomputed values against originally emitted values and confirming that reruns do not re-trigger external side effects
- Experience validating analytical or lakehouse data (Apache Paimon, Iceberg, Delta Lake, or similar, queried through Trino, Spark, or equivalent) — comparing emitted streams against materialized tables for completeness and correctness
- Experience with failure-injection and resilience testing — node, broker, cache, or service loss; verifying recovery path, data integrity after recovery, and absence of duplicate or lost records
- Experience testing on Kubernetes-deployed platforms — namespace-scoped environments, Helm-based deployments, pod and job lifecycle, and access to logs and metrics for diagnosis
- Working familiarity with observability tooling for test diagnosis (Grafana, Prometheus, OpenSearch or equivalent log search, distributed tracing)
- Experience validating AI/ML or LLM outputs is an advantage — approaches to testing non-deterministic responses, accuracy and regression measurement, and detection of unsupported or fabricated output
- Rigorous defect discipline — reproducible defect reports with evidence, clear severity and triage judgement, and ownership of a maintained regression suite
- Domain exposure to any of the following is an advantage: customer data platforms or customer 360 systems, telecom or other high-volume transactional systems, real-time systems with latency SLAs, AdTech or MarTech platforms, and data analytics or dashboard systems
- Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)
Similar roles
-
Web QA & Automation Engineer
Atera Tel-Aviv, Tel-Aviv District, Israel
-
QA Engineer
Bluewhite Tel Aviv, Tel-Aviv District, Israel
-
Divisional QC / QA Manager
Saulsbury Houston, Texas, United States
-
QA Automaticien·ne Playwright - Nantes (F/H)
MOBIAPPS Nantes, Pays de la Loire, France · €40K–€48K/yr
-
QA Engineer
DKB Code Factory Valencia, Valencian Community, Spain
-
QA Auditor
Ascent Aviation Services Marana, Arizona, United States