About the role
You will own the quality and automation frameworks for a high-scale, low-latency streaming Customer Data Platform. This involves building end-to-end test suites, validating data pipelines, and ensuring performance and governance standards are met.
What they look for
Requirements
Candidates must have 6+ years of experience in software quality engineering with a strong focus on building test automation frameworks. Proficiency in programming languages like Python or Java and experience with data-intensive systems like Kafka and Kubernetes are essential.
Full description
Owning quality for a high-scale, low-latency streaming Customer Data Platform deployed on infrastructure we run ourselves (on-prem, Kubernetes), not managed cloud services. The platform ingests from thirty or more source systems at hundreds of thousands of events per second, resolves customer identity, computes customer attributes in real time, and emits governed signals to external destinations. You will build and own the automation that proves this works: API and end-to-end test frameworks, streaming and data-pipeline validation, latency and load testing against explicit SLOs, and governance testing that proves consent, data-protection, and tenant-isolation rules actually hold at runtime. Testing here is evidence-producing, not exploratory only — test outcomes gate whether an artifact is allowed to go live.
Requirements
- 6+ years
in software quality engineering, with the majority of your work in test automation rather than manual testing
- Demonstrated
ownership of test automation frameworks you built or substantially re-architected — not only writing test cases against someone else's framework (core requirement)
- Ability to
write high-quality code in Python, Java, TypeScript, or equivalent languages, including framework structure: base classes, shared utilities, configuration handling, reporting, and test data management
- Strong API
test automation experience (REST Assured, Karate, pytest, Postman/Newman, Playwright API, or similar) — test organization, authentication handling, response and schema validation, and contract testing between services
- End-to-end
UI automation experience (Playwright, Cypress, Selenium, or similar) — framework architecture, Page Object Model or equivalent, handling dynamic elements, and parallel execution
- Demonstrated
ability to diagnose and eliminate flaky tests — proper waits, test isolation, deterministic setup and teardown, and root-cause investigation rather than blanket retries
- Performance
and load testing experience (JMeter, k6, Locust, Gatling, or similar) — designing load scenarios, identifying bottlenecks, and validating throughput and latency against stated targets
- Strong SQL
for data validation — joins, aggregations, window functions, and source-versus-target reconciliation queries over large datasets
- Practical
experience testing data pipelines and ETL/streaming jobs — schema validation, source-to-destination comparison at row and value level, completeness and duplication checks, and sampling strategies for datasets too large to compare in full
- Experience
testing event-driven systems and message queues (Kafka or equivalent) — event and payload validation, ordering guarantees, delivery semantics (at-least-once versus exactly-once-effective), idempotency, consumer lag, and offset behaviour
- Experience
testing stateful stream processing (Apache Flink or similar) — correctness after restart from checkpoint, savepoint-based upgrades, state restoration, late and out-of-order event handling, and behaviour under backpressure
- Ability to
validate end-to-end latency against an explicit SLO — measuring p95 and p99 across pipeline stages, and distinguishing in-scope platform processing from excluded external calls such as third-party API round-trips and destination acknowledgement
- Experience
integrating tests into CI/CD pipelines — trigger configuration, test stages, artifact and report publication, and failure gates that block promotion
- Disciplined
test data management — setup and cleanup strategy, inter-test dependency handling, and environment-specific data; experience working with synthetic, generated, or masked datasets where production data cannot be used
- Ability to
produce structured, auditable test evidence: test results tied to the exact artifact version and configuration under test, so that results are traceable and cannot be silently reused after the artifact changes
- Experience
testing data governance and privacy controls — consent enforcement, data classification and usage rules, tokenization and masking, and verifying that raw sensitive identifiers do not appear in downstream topics, logs, or exports
- Experience
testing identity resolution or entity matching — deterministic and probabilistic matching outcomes, merge and split behaviour, lifecycle state transitions, and identifier reassignment scenarios
- Experience
testing multi-tenant systems — verifying tenant and workspace isolation, role- and attribute-based access control, and absence of cross-tenant data leakage
- Experience
validating replay, backfill, and reconciliation — comparing recomputed values against originally emitted values and confirming that reruns do not re-trigger external side effects
- Experience
validating analytical or lakehouse data (Apache Paimon, Iceberg, Delta Lake, or similar, queried through Trino, Spark, or equivalent) — comparing emitted streams against materialized tables for completeness and correctness
- Experience
with failure-injection and resilience testing — node, broker, cache, or service loss; verifying recovery path, data integrity after recovery, and absence of duplicate or lost records
- Experience
testing on Kubernetes-deployed platforms — namespace-scoped environments, Helm-based deployments, pod and job lifecycle, and access to logs and metrics for diagnosis
- Working
familiarity with observability tooling for test diagnosis (Grafana, Prometheus, OpenSearch or equivalent log search, distributed tracing)
- Experience
validating AI/ML or LLM outputs is an advantage — approaches to testing non-deterministic responses, accuracy and regression measurement, and detection of unsupported or fabricated output
- Rigorous
defect discipline — reproducible defect reports with evidence, clear severity and triage judgement, and ownership of a maintained regression suite
- Domain
exposure to any of the following is an advantage: customer data platforms or customer 360 systems, telecom or other high-volume transactional systems, real-time systems with latency SLAs, AdTech or MarTech platforms, and data analytics or dashboard systems
- Familiarity
using AI tools for development and debugging (Claude, Cursor, Codex)
Similar roles
-
Team Member - Offsite Audit, Concurrent Audit & QA| Mumbai
CSB Bank Mumbai, Maharashtra, India
-
Manufacturing Floor & QA Supervisor - HOLT Manufacturing
HOLT Group Waco, Texas, United States
-
QA Lab Relief III - Visalia - Shift Flex
Ventura Coastal, LLC Visalia, California, United States · $51K/yr
-
QA Lab Relief II - Visalia - Shift Flex
Ventura Coastal, LLC Visalia, California, United States · $49K/yr
-
Senior Manager, QA Compliance
Sarepta Therapeutics Andover, Massachusetts, United States · $136K–$170K/yr
-
Senior QA Engineer / Analyst
Perseus Group, Constellation Software Jacksonville, Florida, United States