Weekday AI

Linux SME / SRE Engineer

Weekday AI Bengaluru, Karnataka, India · ₹2M–₹3M/yr

Technology, Information and Internet · 11-50 employees

5 h ago
sre Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Administer and maintain critical Linux production environments, including PCS/Pacemaker high-availability clusters, while ensuring reliability, availability, and performance. Resolve production incidents, perform root-cause analysis, support failover and maintenance activities, collaborate with stakeholders, and improve operational efficiency through SRE practices and automation.

What they look for

Linux Administration Red Hat Enterprise Linux (RHEL) PCS/Pacemaker Clustering High Availability Site Reliability Engineering (SRE) Production Support Cluster Troubleshooting Incident Management Root-Cause Analysis Monitoring Failover and Recovery ITIL Practices VMware Administration AWS Oracle Database Infrastructure Automation

Requirements

Requires 5–9 years of Linux administration, infrastructure engineering, SRE, or production support experience, including at least 4 years administering PCS/Pacemaker clusters and strong hands-on RHEL expertise. Candidates should have proven production troubleshooting and incident-resolution skills, strong communication and stakeholder collaboration abilities, and willingness to work rotational shifts; VMware, cloud, and Oracle experience are advantageous.

Full description

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟮𝟬𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟬𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟮𝟬-𝟯𝟬 𝗟𝗣𝗔)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India, Hyderabad, Telangana, India

Job Type: Full-time

We are looking for an experienced Linux SME / SRE Engineer with strong expertise in Core Linux Administration, RHEL, and PCS/Pacemaker clustering to support business-critical production environments.

The role focuses on maintaining highly available Linux infrastructure, resolving complex production issues, ensuring system reliability, and supporting clustered environments. The ideal candidate will have strong hands-on troubleshooting capabilities, a solid understanding of high-availability architectures, and the ability to work effectively with clients and technical stakeholders.

Key Responsibilities

  • Administer and support Linux-based production environments across critical infrastructure.
  • Perform day-to-day Core Linux administration, configuration, monitoring, maintenance, and troubleshooting.
  • Manage, monitor, configure, and troubleshoot PCS/Pacemaker high-availability clusters.
  • Ensure availability, reliability, stability, and performance of Linux infrastructure and clustered services.
  • Troubleshoot complex and critical production incidents and drive issues through to resolution.
  • Perform root-cause analysis and implement sustainable solutions for recurring infrastructure problems.
  • Monitor system and cluster health and proactively identify potential availability or performance issues.
  • Support failover, recovery, maintenance, and operational activities across high-availability environments.
  • Collaborate with clients, infrastructure teams, application teams, and other technical stakeholders on incidents and enhancements.
  • Participate in incident management, problem management, change management, and production maintenance activities.
  • Follow SRE practices for monitoring, reliability improvement, incident response, and operational efficiency.
  • Maintain technical documentation, operational procedures, troubleshooting guides, and support records.
  • Participate in rotational shifts to provide continuous production support.
  • Identify opportunities to automate repetitive infrastructure tasks and improve operational efficiency.
  • Support infrastructure changes, upgrades, patching, and maintenance activities in accordance with established processes.
  • Contribute to service reliability, availability, and continuous improvement initiatives.

What Makes You a Great Fit

  • 5–9 years of overall experience in Linux administration, infrastructure engineering, SRE, or production support, with a maximum of 10 years preferred.
  • Minimum 4 years of hands-on experience with PCS/Pacemaker cluster administration.
  • Strong expertise in Core Linux Administration and production infrastructure support.
  • Strong hands-on experience with RHEL (Red Hat Enterprise Linux).
  • Solid understanding of High Availability, clustering, failover, resource management, and cluster troubleshooting.
  • Proven experience supporting critical production environments with strict availability and reliability requirements.
  • Strong troubleshooting, debugging, root-cause analysis, and incident-resolution capabilities.
  • Experience working with production monitoring, incident management, and infrastructure maintenance processes.
  • Strong understanding of SRE and ITIL practices is desirable.
  • Excellent communication and client-facing skills with the ability to explain technical issues clearly to stakeholders.
  • Strong stakeholder-management and collaboration skills.
  • Ability to work effectively under pressure during critical production incidents.
  • Willingness to work in rotational shifts, including scheduled production-support coverage.
  • Experience with VMware administration is an advantage.
  • Exposure to AWS or other cloud platforms is desirable.
  • Knowledge of Oracle Database and its infrastructure dependencies is an advantage.
  • Strong ownership mindset with a focus on system reliability, operational excellence, and continuous improvement.

Similar roles