【Full Remote - ¥15M】Data Engineer – AI Data Platform / Annotation & Data Management
GAIDOOR株式会社 Minato, Japan
Staffing and Recruiting · 2-10 employees
About the role
You will design and operate the data management and annotation infrastructure to support the development of proprietary large-scale LLMs and VLM technologies. This includes building data pipelines, managing annotation workflows, and collaborating with AI researchers to improve model accuracy.
What they look for
Requirements
Candidates must have experience in data collection, preparation, or annotation workflows, along with proficiency in Python or SQL. Business-level Japanese communication skills are required to discuss technical development topics.
Benefits
Full description
One of the leading AI companies we work closely with is looking for a Data Engineer to build and improve the data foundation behind the development of advanced Generative AI, LLM, and VLM technologies.
The company develops proprietary large-scale Japanese-language LLMs from scratch and provides enterprise AI solutions used by more than 300 companies, including major Japanese enterprises.
In this position, you will be responsible for designing and operating the data management and annotation infrastructure that supports AI model development.
The role covers the entire data lifecycle, including data collection, dataset creation, annotation design and quality management, data pipelines, metadata management, and analysis of model outputs.
You will also have the opportunity to develop and customize internal annotation tools specifically designed for the company’s AI development tasks.
This is an excellent opportunity for a Data Engineer who wants to work closer to AI/ML development and directly contribute to improving the accuracy and quality of proprietary LLM and VLM models.
About the Company
The company is a Japan-based AI technology company with strong expertise in Generative AI and LLM technologies.
Rather than simply utilizing external AI models, the company has the technical capability to develop large-scale Japanese-language LLMs entirely in-house.
In 2024, the company released a Japanese-specialized LLM with 100 billion parameters. Its technology is designed specifically for Japanese-language and business use cases, with a strong focus on accuracy, security, and enterprise applications.
The company currently operates enterprise AI products including an AI agent for the manufacturing industry and a platform supporting enterprise AI implementation.
Its solutions have already been adopted by more than 300 companies, including approximately 30% of Nikkei 225 companies.
In 2024, the company raised JPY 4.5 billion in Series D funding, bringing its total funding to approximately JPY 8.8 billion. The organization currently has approximately 150 employees.
Job Responsibilities
・Design annotation standards and workflows for training and evaluation data
・Manage annotation operations and data quality
・Develop, customize, and operate annotation tools
・Collect data from customer datasets, public sources, and other relevant sources
・Build, organize, and maintain datasets for AI model development
・Review model outputs and analysis results and identify initial accuracy or data-quality issues
・Provide feedback to engineers and researchers based on data and model analysis
・Design, develop, and operate data pipelines covering data collection, preprocessing, and storage
・Manage metadata across datasets and data pipelines
・Support data security and personal information protection requirements
・Collaborate closely with engineers and AI researchers to continuously improve data quality and model performance
Development Environment
TypeScript / Vue.js / Node.js / Python / Docker / Terraform / AWS / Azure
Why Join This Role?
・Work directly on the data foundation behind proprietary LLM and VLM development
・Influence model accuracy by improving the quality of training and evaluation data
・Take ownership from annotation-standard design through internal tooling and operations
・Build datasets and data pipelines used directly by AI engineers and researchers
・Work with document, image, and drawing data in advanced AI applications
・Collaborate closely with AI engineers and researchers
・Work on technology already being deployed at hundreds of major enterprises
・Work fully remotely from anywhere in Japan
Requirements
Required Skills
All of the following are required:
・Experience designing and operating data collection, data preparation, or annotation workflows
・Basic experience with data processing and data pipeline development using SQL, Python, or similar technologies
・Ability to continuously manage and maintain high-quality datasets
・Active use of AI coding tools such as GitHub Copilot, Cursor, Claude Code, or similar tools in daily development
・Business-level or higher Japanese communication skills, including the ability to discuss technical development topics in Japanese
・Ability to reside and work in Japan
Preferred Skills
・Experience developing or customizing annotation tools, including web application development
・Experience building datasets or designing evaluation processes for Machine Learning projects
・Experience building data platforms on AWS, GCP, or Azure
・Experience with annotation of images, drawings, documents, or other unstructured data
Benefits
Benefits
Compensation
Expected Annual Salary:
¥7,500,000 - ¥15,000,000
Salary Review:
Twice a year
Location
Tokyo / Full Remote Work Available Within Japan
Working Style
・Primarily full remote
・Work from anywhere within Japan regardless of region
・Flex-time system
・Core hours: 10:00 - 14:00
・Average monthly overtime: approximately 10 - 20 hours
Holidays & Leave
・123 annual holidays
・Five-day workweek (Saturdays, Sundays, and national holidays off)
・Year-end and New Year holidays
・Paid annual leave
・Maternity leave
・Childcare leave
Benefits
・Full social insurance coverage
・Corporate defined contribution pension plan
・Up to ¥240,000 per year in financial support for professional development and AI-related activities
・Transportation expenses
・Regular health checkups
・Influenza vaccinations
・Company-provided PC
・Five-day onboarding program after joining
Similar roles
-
Data Engineer Working Student
Cheil Germany GmbH Eschborn, Hesse, Germany
-
HSO International - Data Engineer
HSO Group B.V. Reims, Grand Est, France
-
Data Engineer - Senior
Cummins Pune, Maharashtra, India
-
Data Engineer (m/f/d)
T-Systems Iberia Granada, Andalusia, Spain
-
Senior Data Engineer
myKaarma Noida, Uttar Pradesh, India
-
AI Data Engineer Intern
Aumovio Budapest, Central Hungary, Hungary