A

Software Engineer, Data Infrastructure and Acquisition

Analogy Group India

Technology, Information and Media · 2-10 employees

7 h ago
Remote Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design and maintain data acquisition infrastructure while building AI-powered, self-healing scrapers for public data. Develop resilient pipelines and agentic workflows to ensure data accessibility and system observability.

What they look for

Python JavaScript TypeScript Data Infrastructure Web Scraping AWS Django Prefect AI Agents Workflow Orchestration Data Pipelines Observability System Architecture Headless Browsers Data Validation Distributed Systems

Requirements

Requires a Bachelor's or Master's degree in Computer Science or Engineering and 5+ years of professional software engineering experience. Candidates must possess strong Python skills and practical experience with web scraping, data pipelines, and AI agent development.

Full description

Position: Software Engineer, Data Infrastructure and Acquisition

Analogy Group builds cutting-edge AI political intelligence tools. We work with advocacy organizations, researchers, and companies working on the most important problems in the world—from climate to democracy to AI policy. We use LLMs to make sense of messy data, enabling these organizations to move faster with more informed strategy.

We are hiring a Software Engineer focused on Data Infrastructure and Acquisition to help us design the infrastructure and build scrapers for public data. This role is a mix of backend engineering, agentic AI engineering, and web scraping. We want to make government data accessible and useful to help fight corruption, protect democracy, and make citizens able to understand their government, and your work will directly contribute to this mission.

This is a fully remote role for engineers based in India. We offer highly competitive, top-of-market compensation and are looking for candidates with a strong record of technical excellence, whether developed at a top engineering program, a high-performing technology company, an ambitious startup, or through exceptional independent work.

What You’ll Do:

  • Design, build, and maintain our data acquisition infrastructure using technologies such as Prefect, AWS, and Django.
  • Develop clear, durable abstractions for scraping patterns and infrastructure, keeping the system manageable as we scale to thousands of data sources.
  • Build AI-powered infrastructure for agentic, self-healing scrapers.
  • Develop agents capable of creating new scrapers from established templates and validating their output.
  • Build and test resilient scrapers for messy, difficult public data sources, including PDFs, videos, forms, and legacy government portals.
  • Improve observability, retries, data validation, and failure recovery across our pipelines.
  • Open-source selected parts of our work and write about what we learn.

Minimum Requirements:

  • Bachelor's or Master's degree in Computer Science, Engineering, or similar.
  • 5+ years of professional software engineering experience, with strong Python skills and working knowledge of JavaScript or TypeScript.
  • Experience designing and operating complex data pipelines.
  • Experience with workflow orchestration platforms such as Temporal, Prefect, Dagster, or Airflow.
  • Practical web-scraping experience, including headless browsers, retries, rate limiting, proxy rotation, and changing or unreliable source websites.
  • Experience using coding agents such as Codex or Claude Code, including an understanding of what makes agent workflows reliable.
  • Experience developing AI agents, agent tooling, evaluation systems, skills, or execution harnesses.
  • Strong judgment about abstractions, testing, observability, and long-term system maintainability.

Nice to Have

  • Experience extracting structured information from PDFs, scanned documents, audio, or video.
  • Experience operating distributed systems or data infrastructure on AWS.
  • Experience with Django and Prefect.
  • Contributions to open-source software or published technical writing.
  • An interest in government transparency, public-interest technology, politics, or civic data.