GoTymeX

DevOps Engineer (Databricks-based Data as a Product Platform)

GoTymeX Ho Chi Minh City, Vietnam

Software Development · 501-1,000 employees

7 h ago
Remote devops Senior (5-10 yrs) Full-time Vietnam
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The DevOps Engineer will design, build, and maintain robust CI/CD pipelines to automate the deployment and operation of a Databricks-based Data as a Product platform. They will also manage platform infrastructure as code, implement GitOps practices, and integrate observability and AI tools to ensure platform reliability and efficiency.

What they look for

CI/CD Terraform AWS Python Bash Databricks GitOps Datadog Infrastructure as Code Observability Unity Catalog Git CloudWatch IAM Automation DORA metrics

Requirements

Candidates must have senior-level experience in CI/CD pipeline development, infrastructure as code using Terraform, and cloud environment management on AWS. A bachelor's or master's degree in Computer Science or a related field is required, along with strong scripting skills in Python or Bash.

Benefits

Meal allowance Parking allowance Premium health care Professional development Overseas training opportunities Performance bonus Annual leave Sick leave

Full description

We are seeking a DevOps Engineer to own the CI/CD pipelines and infrastructure-as-code practices that power our Data as a Product (DaaP) Platform, built on Databricks (AWS). This role is primarily focused on designing and building robust CI/CD pipelines for deploying and operating our Databricks-based DaaP Platform — automating deployments, managing environments through config-as-code, and integrating with observability, AI, and cloud tooling (Datadog, Claude/Anthropic, AWS native services) to give data and engineering teams a fast, reliable, and self-service platform experience.

Main responsibilities:

  • Design, build, and maintain CI/CD pipelines as the core focus of this role — enabling reliable, repeatable deployment of the Data Platform and its workloads/jobs/infrastructure on Databricks and AWS
  • Manage platform environments (workspaces, clusters, jobs, permissions, and related AWS resources) as code using Terraform and/or config-as-code approaches
  • Implement GitOps practices for infrastructure and platform configuration, ensuring changes are version-controlled, reviewed, and auditable
  • Build and maintain integrations between the Data Platform and third-party platforms (Datadog for observability/alerting, Claude/Anthropic for AI-assisted workflows, and other SaaS/AWS services)
  • Automate provisioning, scaling, and lifecycle management of AWS resources supporting the Data Platform (networking, IAM, S3, VPC endpoints, etc.)
  • Develop monitoring, logging, and alerting for pipeline health, cost, and platform reliability
  • Define and track key engineering metrics (e.g., DORA metrics — deployment frequency, lead time for changes, change failure rate, time to restore service — plus review/cycle time) to measure and improve platform and DevOps efficiency
  • Partner with data engineering and platform teams to translate manual/ad-hoc processes into automated, code-first workflows
  • Continuously improve deployment velocity, environment consistency, and developer experience across dev/staging/prod

Must have:

  • Education: Bachelor's or Master's in Computer Science, IT, or related field
  • Seniority: Senior-level, hands-on experience designing and owning CI/CD pipelines in a cloud environment
  • CI/CD (core focus): Strong, proven experience building and operating CI/CD pipelines with GitHub Actions, GitLab CI, CodePipeline/CodeBuild, or similar, including GitOps branching strategies
  • IaC: Strong proficiency in Terraform and config-as-code approaches for managing cloud/platform environments
  • AWS: Working knowledge of core AWS services relevant to a data platform — IAM, VPC/networking, S3, Lambda, CloudWatch, Secrets Manager
  • Scripting: Proficient in Python and/or Bash for automation and tooling
  • Observability: Experience integrating platforms with Datadog (or similar) for monitoring, dashboards, and alerting
  • Engineering metrics: Familiarity with DORA metrics and related delivery/quality metrics (deployment frequency, lead time, change failure rate, MTTR, review/cycle time) to drive platform and team improvement
  • AI tooling integration: Familiarity with integrating LLM/AI platforms (e.g., Claude/Anthropic APIs) into engineering workflows
  • Version control: Strong Git practices and experience working in a collaborative, code-reviewed environment
  • Soft skills: Clear communication, self-directed problem-solving, comfort working across data engineering and platform teams

Nice to have:

  • Databricks: Experience managing Databricks on AWS — workspace administration, Unity Catalog, cluster policies, job orchestration, Databricks Asset Bundles (DABs)
  • Experience with Unity Catalog governance and data product/data mesh concepts
  • Experience in banking, fintech, or other regulated industries
  • AWS certification(s) (Solutions Architect, DevOps Engineer – Professional)
  • Databricks certification (Data Engineer or Platform Administrator)
  • Experience with workflow orchestration tools (n8n, Step Functions, Airflow)

At GoTymeX, opportunities are here for the taking. If you want to be part of our purpose and live and lead through our values, we can offer exciting development opportunities through expanded lateral roles, stretch assignments, or people leadership.

Some of our benefits:

  • Meal and parking allowances are covered by the company.
  • Full benefits and salary rank during probation.
  • Insurances such as Vietnamese labor law and premium health care for you and your family.
  • SMART goals and clear career opportunities (technical seminar, conference, and career talk) - we focus on your development.
  • Values-driven, international working environment, and agile culture.
  • Overseas travel opportunities for training and work-related.
  • Internal Hackathons and company events (team building, coffee run, etc.).
  • Pro-Rate and performance bonus.
  • 15-day annual + 3-day sick leave per year from the company.
  • Work-life balance 40-hr per week from Mon to Fri.

Similar roles