Prodapt

Tier 1 Production Support Engineer

Prodapt Hyderabad, Telangana, India

Technology, Information and Internet · 5,001-10,000 employees

20 h ago
support-engineer Mid (2-5 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Provide 24x7 production support for critical enterprise integration and event-driven applications by independently diagnosing and resolving incidents. Maintain operational runbooks, track performance metrics, and ensure effective communication during incident lifecycles.

What they look for

Kafka Confluent Azure Kubernetes Service Splunk PagerDuty Prometheus Grafana SQL Postgres Java Spring Boot React Python Linux Unix Incident Management

Requirements

Requires 3+ years of experience in application production support with hands-on expertise in Kubernetes, Kafka, and observability tools. Candidates must possess strong troubleshooting skills in a Linux environment and the ability to work in rotational 24x7 shifts.

Full description

Overview

  • 1) Messaging system. Kafka/confluent

2) Azure kubernetes services

3) Azure

Job Title

Tier 1 Application Production Support Specialist

Job Summary / About the Role

AT&T is seeking a technically strong and self-sufficient Tier 1 Application Production Support Specialist to provide 24x7 support for critical enterprise integration, messaging, and event-driven applications. This is not a pure monitoring or escalation role. The ideal candidate is expected to independently identify root causes, apply fixes, and resolve the majority of incidents without Tier 2 involvement. Tier 2 is engaged only for complex or architectural issues that go beyond Tier 1 resolution scope. The candidate must have solid working knowledge of the full technology stack and the ability to triage, diagnose, and remediate issues end-to-end.

Key Responsibilities

- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.

- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.

- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.

- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.

- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.

- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.

- Monitor application support mailboxes and manage operational follow-through.

- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.

- Share client profile data and usage reports on request.

- Maintain clear incident communication to clients even when Tier 2 is engaged.

- Track and report operational metrics including MTTR and ticket resolution trends.

- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.

Required Qualifications / Must-Have Skills

- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.

- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.

- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.

- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.

- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.

- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.

- Solid SQL/Postgres skills for data-level investigation and validation during incidents.

- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.

- Basic Python scripting capability for operational checks and quick-fix automation.

- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.

- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.

- Willingness to work in rotational 24x7 shifts.

Good-to-Have / Nice-to-Have

- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.

- Exposure to IBM Sterling Integrator integration flows for context during incident triage.

- Telecom or high-availability enterprise support experience.

- Experience with CI/CD-driven deployment pipelines in a support context.

Experience Level

Associate to Mid-Level (typically 3 to 5 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

Responsibilities

  • 1) Messaging system. Kafka/confluent

2) Azure kubernetes services

3) Azure

Job Title

Tier 1 Application Production Support Specialist

Job Summary / About the Role

AT&T is seeking a technically strong and self-sufficient Tier 1 Application Production Support Specialist to provide 24x7 support for critical enterprise integration, messaging, and event-driven applications. This is not a pure monitoring or escalation role. The ideal candidate is expected to independently identify root causes, apply fixes, and resolve the majority of incidents without Tier 2 involvement. Tier 2 is engaged only for complex or architectural issues that go beyond Tier 1 resolution scope. The candidate must have solid working knowledge of the full technology stack and the ability to triage, diagnose, and remediate issues end-to-end.

Key Responsibilities

- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.

- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.

- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.

- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.

- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.

- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.

- Monitor application support mailboxes and manage operational follow-through.

- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.

- Share client profile data and usage reports on request.

- Maintain clear incident communication to clients even when Tier 2 is engaged.

- Track and report operational metrics including MTTR and ticket resolution trends.

- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.

Required Qualifications / Must-Have Skills

- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.

- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.

- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.

- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.

- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.

- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.

- Solid SQL/Postgres skills for data-level investigation and validation during incidents.

- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.

- Basic Python scripting capability for operational checks and quick-fix automation.

- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.

- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.

- Willingness to work in rotational 24x7 shifts.

Good-to-Have / Nice-to-Have

- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.

- Exposure to IBM Sterling Integrator integration flows for context during incident triage.

- Telecom or high-availability enterprise support experience.

- Experience with CI/CD-driven deployment pipelines in a support context.

Experience Level

Associate to Mid-Level (typically 3 to 5 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

Requirements

  • 1) Messaging system. Kafka/confluent

2) Azure kubernetes services

3) Azure

Job Title

Tier 1 Application Production Support Specialist

Job Summary / About the Role

AT&T is seeking a technically strong and self-sufficient Tier 1 Application Production Support Specialist to provide 24x7 support for critical enterprise integration, messaging, and event-driven applications. This is not a pure monitoring or escalation role. The ideal candidate is expected to independently identify root causes, apply fixes, and resolve the majority of incidents without Tier 2 involvement. Tier 2 is engaged only for complex or architectural issues that go beyond Tier 1 resolution scope. The candidate must have solid working knowledge of the full technology stack and the ability to triage, diagnose, and remediate issues end-to-end.

Key Responsibilities

- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.

- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.

- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.

- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.

- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.

- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.

- Monitor application support mailboxes and manage operational follow-through.

- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.

- Share client profile data and usage reports on request.

- Maintain clear incident communication to clients even when Tier 2 is engaged.

- Track and report operational metrics including MTTR and ticket resolution trends.

- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.

Required Qualifications / Must-Have Skills

- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.

- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.

- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.

- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.

- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.

- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.

- Solid SQL/Postgres skills for data-level investigation and validation during incidents.

- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.

- Basic Python scripting capability for operational checks and quick-fix automation.

- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.

- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.

- Willingness to work in rotational 24x7 shifts.

Good-to-Have / Nice-to-Have

- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.

- Exposure to IBM Sterling Integrator integration flows for context during incident triage.

- Telecom or high-availability enterprise support experience.

- Experience with CI/CD-driven deployment pipelines in a support context.

Experience Level

Associate to Mid-Level (typically 3 to 5 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

Similar roles