Weekday AI

DevOps Engineer

Weekday AI Bengaluru, Karnataka, India

Technology, Information and Internet · 11-50 employees

3 h ago
devops Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The role involves designing, deploying, and maintaining scalable AWS cloud infrastructure while ensuring system reliability and observability. You will be responsible for managing CI/CD pipelines, troubleshooting Linux-based systems, and collaborating with engineering teams to optimize operational efficiency.

What they look for

DevOps Linux AWS Prometheus Grafana Cloud CI/CD Infrastructure Automation Observability Cloud Infrastructure Networking Bash Python Git Docker Kubernetes System Reliability

Requirements

Candidates must have at least 2 years of professional experience in DevOps or a related cloud engineering role. Strong hands-on proficiency with Linux, AWS services, Prometheus, and Grafana is required, along with scripting skills in Python or Bash.

Full description

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟭𝟮𝟯𝟴𝟬𝟬𝟬 - 𝗥𝘀 𝟮𝟬𝟲𝟰𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟭𝟮.𝟯𝟴-𝟮𝟬.𝟲𝟰 𝗟𝗣𝗔)

Experience: 2+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are looking for a hands-on and technically strong DevOps Engineer to build, maintain, and improve reliable, scalable, and secure cloud infrastructure and deployment environments. The role will focus on Linux, AWS, Prometheus, and Grafana Cloud, with responsibility for infrastructure automation, monitoring, observability, deployment processes, system reliability, and production support.

The ideal candidate will have strong troubleshooting skills, a practical understanding of cloud infrastructure, and the ability to work closely with software engineering and other technical teams to improve application reliability and operational efficiency.

KEY RESPONSIBILITIES• Design, deploy, configure, and maintain scalable AWS cloud infrastructure across development, staging, and production environments.

  • Administer and troubleshoot Linux-based servers and systems, including performance, availability, security, and resource utilisation.
  • Support cloud services across compute, networking, storage, databases, IAM, and other AWS components.
  • Implement and maintain infrastructure automation and configuration-management practices.
  • Build and maintain reliable CI/CD pipelines to automate application build, testing, deployment, and release processes.
  • Configure and manage Prometheus for infrastructure and application monitoring, metrics collection, and alerting.
  • Develop and maintain Grafana Cloud dashboards, visualisations, alerts, and observability solutions.
  • Monitor system health, application performance, resource utilisation, availability, and service-level indicators.
  • Investigate production incidents, identify root causes, and implement permanent corrective actions.
  • Troubleshoot Linux, networking, application deployment, infrastructure, and cloud-related issues.
  • Improve system reliability through automation, proactive monitoring, capacity planning, and performance optimisation.
  • Implement appropriate security controls across AWS infrastructure, Linux systems, access management, and deployment environments.
  • Collaborate with software engineers, QA, architects, and other technical teams to improve deployment and operational processes.
  • Maintain infrastructure documentation, operational runbooks, monitoring standards, and troubleshooting procedures.
  • Support backup, disaster recovery, high-availability, and business-continuity requirements.
  • Identify opportunities to reduce operational overhead through automation and standardisation.
  • Participate in production releases, incident response, maintenance activities, and continuous improvement initiatives.
  • Stay current with AWS services, DevOps practices, cloud-native technologies, observability tools, and infrastructure automation.

WHAT MAKES YOU A GREAT FIT• 2+ years of professional experience in DevOps, Cloud Engineering, Site Reliability Engineering, Infrastructure Engineering, or a related role.

  • Strong hands-on experience administering and troubleshooting Linux environments.
  • Good practical experience with AWS cloud services and cloud infrastructure management.
  • Strong understanding of AWS compute, networking, storage, IAM, monitoring, and security concepts.
  • Hands-on experience with Prometheus for metrics collection, monitoring, and alerting.
  • Practical experience with Grafana Cloud, including dashboards, visualisations, alerts, and observability.
  • Experience building and maintaining CI/CD pipelines and automated deployment workflows.
  • Understanding of infrastructure-as-code and configuration-management practices.
  • Good knowledge of networking fundamentals, DNS, HTTP/HTTPS, TCP/IP, load balancing, and security concepts.
  • Strong troubleshooting and root-cause analysis skills across infrastructure and application environments.
  • Understanding of system reliability, availability, scalability, monitoring, and performance optimisation.
  • Experience with scripting or automation using Bash, Python, or similar technologies.
  • Familiarity with Git and modern software development and deployment workflows.
  • Exposure to Docker, Kubernetes, or other containerisation technologies will be an advantage.
  • Strong understanding of DevOps principles, automation, observability, and production operations.
  • Excellent communication and collaboration skills with the ability to work effectively with cross-functional engineering teams.
  • Proactive mindset with strong ownership of infrastructure reliability, operational excellence, and continuous improvement.

Similar roles