Associate Engineer/Engineer, SRE
Rakuten Viki Singapore
Entertainment Providers · 51-200 employees
About the role
The role involves pioneering agentic infrastructure by designing autonomous systems and building internal platforms to abstract infrastructure complexity. You will also contribute to performance and reliability engineering while partnering with development teams to implement SRE best practices.
What they look for
Requirements
Candidates must have a Bachelor's degree in a technical field and at least 2 years of experience in SRE, DevOps, or backend development. Proficiency in Linux, containerization, cloud platforms, and software development is required.
Full description
Job Description:
Rakuten International oversees 7 businesses with over 4,000 employees globally. The brand is recognized for its leadership and innovation in e-commerce, digital content, advertising, entertainment and communications, bringing the joy of discovery and access to more than 1 billion members across the world. Our teams deliver on the company’s mission to delight merchants and customers through innovation, optimism, and teamwork.
Rakuten Viki is a global entertainment streaming platform that specializes in Asian content. Our platform enables millions of viewers to discover and enjoy primetime shows and movies, subtitled in over 150 languages. Headquartered in San Mateo, California, we also have offices in Singapore, Seoul, and Shanghai, ensuring a strong global presence and a deep connection to the heart of Asian entertainment. Our platform is home to a large and loyal community of fans who share a passion for Asian culture and entertainment. Join us in our mission to bridge cultures and connect the world to Asian entertainment. At Rakuten Viki, we offer a chance to be part of a global community that celebrates culture, creativity, and connection.
We are in search of a Associate Engineer/Engineer, SRE to join our team and support our business growth. This role will be based in Singapore and reporting to SRE Manager.
About the SRE Team:
The Site Reliability Engineering (SRE) team at Viki builds and operates the platform that powers Viki’s large-scale, distributed systems. We develop and maintain services that power Viki's API and business intelligence, as well as make architecture changes to keep them scalable, reliable, and secure. Our scope is evolving into a Platform and Performance Engineering model, where we are building a sophisticated framework to run agents at scale within secure, policy-bounded, and sandboxed environments. We run our systems on GCP with GKE and our media pipeline on AWS. We also use Spinnaker, Cloudbuild, Datadog, PostgreSQL, RabbitMQ and Redis, to name a few tools.
Our team is currently focused on architecting a high-performance internal platform that elevates developer velocity and system efficiency. We are leading the charge in autonomous infrastructure, building the specialized environments required to run complex processes safely and at scale. We are also committed to driving innovation that boosts productivity across the entire Engineering organization.
Key Responsibilities:
- Pioneering Agentic Infrastructure: Shift away from manual infrastructure management by designing and implementing "blueprints" and skills that enable AI agents to autonomously provision, configure, and maintain our systems.
- Platform Engineering & Abstraction: Build and refine internal platforms that abstract infrastructure complexity, allowing developers to self-serve while maintaining high standards of security and reliability.
- Performance & Reliability Engineering: Drive/Contribute initiatives to optimize system performance and reliability, moving beyond reactive maintenance to proactive, automated performance tuning.
- Architectural Influence: Partner with development teams to instill SRE best practices, ensuring that security, scalability, and cost-efficiency are "baked in" to the architecture from the start.
- Autonomous Lifecycle Management: Evolve our operational model by developing agents to handle iterative tasks such as middleware upgrades, reducing manual toil and increasing deployment velocity.
- Developer Productivity: Build the next generation of internal tooling that enables developers to move faster.
- Operational Excellence: Participate in the on-call rotation to ensure platform availability. Use the insights gained from incidents to refine our agentic workflows and automated guardrails.
- Governance & Continuous Improvement: Monitor key system metrics (performance, cost, and security) and translate these insights into new agentic skills or policy improvements to continuously harden our platform.
- Proactive Security: Integrate security into the automated development cycle by building agents that perform vulnerability assessments, monitor for risks, and execute remediation strategies at scale.
Required Qualifications:
- Bachelor’s Degree in Computer Science, Engineering, or an equivalent technical field.
- Core Experience: 2+ years in SRE, DevOps role or backend development, with a proven track record of building scalable, robust systems that deliver high-impact global services. Fresh graduates with a strong foundation in software development are also encouraged to apply and will be considered.
- Software Engineering: Proficiency in software development, with the ability to write clean, maintainable code to solve complex infrastructure challenges.
- Technical Foundation: Fundamentals in Linux/Unix operating systems and core networking principles.
- Cloud & Containerization: Practical experience with Docker and Kubernetes (GKE/EKS) and familiarity with either GCP or AWS.
- Tooling: Proficiency of Infrastructure as Code (IaC), CI/CD pipelines, and observability best practices.
- Security Mindset: Understanding of application and infrastructure security fundamentals.
- Innovation & Agents: Familiarity with agent-based architectures and a passion for building intelligent autonomous systems.
- Ownership & Collaboration: A proactive problem-solver who takes full ownership of resolution and thrives in cross-functional environments to drive effective, scalable solutions.
Rakuten provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type. Rakuten considers applicants for employment without regard to race, color, religion, age, sex, national origin, disability status, genetic information, protected veteran status, sexual orientation, gender, gender identity or expression, or any other characteristic protected by federal, state, provincial or local laws.
Five Principles for Success Our worldwide practices describe specific behaviors that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
Always improve, Always Advance - Only be satisfied with complete success - Kaizen Passionately Professional - Take an uncompromising approach to your work and be determined to be the best Hypothesize - Practice - Validate – Shikumika - Use the Rakuten Cycle to succeed in unknown territory Maximize Customer Satisfaction - The greatest satisfaction for our teams is seeing their customers smile Speed!! Speed!! Speed!! - Always be conscious of time - take charge, set clear goals, and engage your team
Similar roles
-
Senior Site Reliability Engineer
Commonwealth Bank Sydney, New South Wales, Australia
-
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
Software Mind Montreal, Quebec, Canada
-
Consultant Specialist(SRE)
HSBC Guangzhou City, Guangdong Province, China
-
Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA
Qualcomm Riyadh, Riyadh, Saudi Arabia
-
Embedded Site Reliability Engineer
Finning Fort McMurray, Alberta, Canada
-
Senior Site Reliability Engineer - Azure Storage
Microsoft Redmond, Washington, United States · $120K–$261K/yr