Jobgether

Senior Data Product Engineer

Jobgether Brazil

Internet Marketplace Platforms · 11-50 employees

7 h ago
Remote Senior (5-10 yrs) Full-time Contractor Brazil
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Architect and scale distributed web-scraping systems to collect large volumes of data from internet sources. Develop robust data pipelines and maintain cloud infrastructure to transform raw data into structured datasets for analytics.

What they look for

Python Web scraping Data engineering AWS PostgreSQL Airflow GraphQL Hasura Docker Kubernetes Scrapy Playwright Selenium Puppeteer Data pipelines Cybersecurity

Requirements

Requires strong hands-on experience with Python, web-scraping technologies, and AWS cloud services. Candidates must demonstrate the ability to build reliable data pipelines and manage anti-scraping mechanisms at scale.

Benefits

Remote work Flexible working hours Flexible, self-managed vacation time Sick leave Personal days Public holidays Paternity leave Maternity leave Study leave Moving days Company-provided equipment Training in technology and engineering Access to books and technical talks Continuous learning opportunities In-house English classes Continuous feedback Career development sessions Internal events and team activities Birthday day off

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Product Engineer based in Brazil.

This is a senior-level opportunity to help transform a research-driven data collection platform into a scalable, production-grade data product for cyber risk analytics. You will design and operate distributed web-scraping systems capable of collecting large volumes of data from thousands of internet sources. The role spans data collection, data engineering, infrastructure, and data product development. You’ll turn raw web data into reliable, structured, and well-documented datasets that can power SaaS products and analytics. Working closely with Data Science and engineering teams, you’ll help establish robust pipelines, APIs, and data access patterns. You’ll have significant ownership over technical decisions, reliability, observability, and continuous improvement. The position is fully remote across LATAM, with flexible working hours and a contractor setup paid in USD.

\n

Accountabilities

  • Architect, build, and scale distributed web-scraping systems capable of reliably collecting data from thousands of online sources.
  • Develop robust approaches for handling anti-scraping and bot-protection mechanisms, including proxy rotation, CAPTCHAs, rate limiting, IP blocking, browser fingerprinting, and headless browsing.
  • Build and maintain data extraction workflows using technologies such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, and CSS selectors.
  • Refactor and productionize existing Python data-collection pipelines, strengthening reliability, observability, error handling, retry mechanisms, monitoring, and alerting.
  • Build schedulable, containerized ingestion workflows using Airflow or comparable orchestration technologies.
  • Design PostgreSQL schemas, views, partitioning strategies, and efficient data-access patterns for processed web data.
  • Develop and maintain cloud infrastructure and data services using AWS technologies including S3, EKS, IAM/IRSA, Parameter Store, and ECR.
  • Build and maintain the data-access layer using GraphQL and Hasura, while collaborating with Data Science and SaaS teams on reliable API contracts.
  • Support the integration of existing ML components and ensure dependable movement of data from collection and ingestion through consumption.
  • Take end-to-end ownership of the data product, contribute to code reviews, and continuously improve engineering practices, architecture, and system reliability.

Requirements

  • Strong hands-on experience designing and operating advanced web-scraping systems at scale.
  • Demonstrated experience working with anti-scraping and bot-protection mechanisms, including proxy pools, headless browsers, rate limiting, IP blocking, fingerprinting, or similar techniques.
  • Strong professional experience with Python and data processing, including production-grade pipeline development.
  • Extensive experience with web-scraping and DOM-parsing technologies such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.
  • Hands-on AWS experience, particularly with S3 and boto3; familiarity with EKS, IAM/IRSA, Parameter Store, and ECR is highly valuable.
  • Experience building reliable, schedulable data pipelines with Airflow or an equivalent orchestration platform.
  • Strong knowledge of PostgreSQL, SQL, relational database fundamentals, schema design, and data-access patterns.
  • Ability to take end-to-end ownership of technical solutions, make sound architectural decisions, and work effectively with a high degree of autonomy.
  • Experience participating in code reviews and maintaining strong software engineering standards.
  • Familiarity with GraphQL, Hasura, Redis, Elasticsearch, Docker, Kubernetes, Helm, or GitHub Actions is an advantage.
  • Experience with cybersecurity data, NLP/ML pipelines, or data-as-a-product environments is a plus.
  • English proficiency at B2 level or higher, with the ability to communicate effectively with technical and cross-functional stakeholders.

Benefits

  • 100% remote work across LATAM.
  • Flexible working hours.
  • Contractor engagement with compensation paid in USD.
  • Flexible, self-managed vacation time.
  • Sick leave, personal days, and public holidays.
  • Paternity and maternity leave.
  • Study leave and moving days.
  • Company-provided equipment and work materials.
  • Training in technology and engineering best practices.
  • Access to books, technical talks, and continuous learning opportunities.
  • In-house English classes.
  • Continuous feedback and 1:1 career development sessions.
  • Internal events and team activities.
  • Birthday day off.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1