Tier 1 Production Support Engineer
Prodapt Hyderabad, Telangana, India
Technology, Information and Internet · 5,001-10,000 employees
About the role
Provide 24x7 production support for critical enterprise integration and event-driven applications by independently diagnosing and resolving incidents. Maintain operational runbooks, track performance metrics, and ensure effective communication during incident lifecycles.
What they look for
Requirements
Requires 3+ years of experience in application production support with hands-on expertise in Kubernetes, Kafka, and observability tools. Candidates must possess strong troubleshooting skills in a Linux environment and the ability to work in rotational 24x7 shifts.
Full description
Overview
- 1) Messaging system. Kafka/confluent
2) Azure kubernetes services
3) Azure
Job Title
Tier 1 Application Production Support Specialist
Job Summary / About the Role
AT&T is seeking a technically strong and self-sufficient Tier 1 Application Production Support Specialist to provide 24x7 support for critical enterprise integration, messaging, and event-driven applications. This is not a pure monitoring or escalation role. The ideal candidate is expected to independently identify root causes, apply fixes, and resolve the majority of incidents without Tier 2 involvement. Tier 2 is engaged only for complex or architectural issues that go beyond Tier 1 resolution scope. The candidate must have solid working knowledge of the full technology stack and the ability to triage, diagnose, and remediate issues end-to-end.
Key Responsibilities
- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.
- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.
- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.
- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.
- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.
- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.
- Monitor application support mailboxes and manage operational follow-through.
- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.
- Share client profile data and usage reports on request.
- Maintain clear incident communication to clients even when Tier 2 is engaged.
- Track and report operational metrics including MTTR and ticket resolution trends.
- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.
Required Qualifications / Must-Have Skills
- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.
- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.
- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.
- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.
- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.
- Solid SQL/Postgres skills for data-level investigation and validation during incidents.
- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
- Basic Python scripting capability for operational checks and quick-fix automation.
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.
- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.
- Willingness to work in rotational 24x7 shifts.
Good-to-Have / Nice-to-Have
- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.
- Exposure to IBM Sterling Integrator integration flows for context during incident triage.
- Telecom or high-availability enterprise support experience.
- Experience with CI/CD-driven deployment pipelines in a support context.
Experience Level
Associate to Mid-Level (typically 3 to 5 years)
Location / Work Mode
Onsite (Hyderabad / Bangalore or designated AT&T location)
Responsibilities
- 1) Messaging system. Kafka/confluent
2) Azure kubernetes services
3) Azure
Job Title
Tier 1 Application Production Support Specialist
Job Summary / About the Role
AT&T is seeking a technically strong and self-sufficient Tier 1 Application Production Support Specialist to provide 24x7 support for critical enterprise integration, messaging, and event-driven applications. This is not a pure monitoring or escalation role. The ideal candidate is expected to independently identify root causes, apply fixes, and resolve the majority of incidents without Tier 2 involvement. Tier 2 is engaged only for complex or architectural issues that go beyond Tier 1 resolution scope. The candidate must have solid working knowledge of the full technology stack and the ability to triage, diagnose, and remediate issues end-to-end.
Key Responsibilities
- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.
- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.
- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.
- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.
- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.
- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.
- Monitor application support mailboxes and manage operational follow-through.
- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.
- Share client profile data and usage reports on request.
- Maintain clear incident communication to clients even when Tier 2 is engaged.
- Track and report operational metrics including MTTR and ticket resolution trends.
- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.
Required Qualifications / Must-Have Skills
- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.
- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.
- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.
- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.
- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.
- Solid SQL/Postgres skills for data-level investigation and validation during incidents.
- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
- Basic Python scripting capability for operational checks and quick-fix automation.
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.
- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.
- Willingness to work in rotational 24x7 shifts.
Good-to-Have / Nice-to-Have
- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.
- Exposure to IBM Sterling Integrator integration flows for context during incident triage.
- Telecom or high-availability enterprise support experience.
- Experience with CI/CD-driven deployment pipelines in a support context.
Experience Level
Associate to Mid-Level (typically 3 to 5 years)
Location / Work Mode
Onsite (Hyderabad / Bangalore or designated AT&T location)
Requirements
- 1) Messaging system. Kafka/confluent
2) Azure kubernetes services
3) Azure
Job Title
Tier 1 Application Production Support Specialist
Job Summary / About the Role
AT&T is seeking a technically strong and self-sufficient Tier 1 Application Production Support Specialist to provide 24x7 support for critical enterprise integration, messaging, and event-driven applications. This is not a pure monitoring or escalation role. The ideal candidate is expected to independently identify root causes, apply fixes, and resolve the majority of incidents without Tier 2 involvement. Tier 2 is engaged only for complex or architectural issues that go beyond Tier 1 resolution scope. The candidate must have solid working knowledge of the full technology stack and the ability to triage, diagnose, and remediate issues end-to-end.
Key Responsibilities
- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.
- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.
- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.
- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.
- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.
- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.
- Monitor application support mailboxes and manage operational follow-through.
- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.
- Share client profile data and usage reports on request.
- Maintain clear incident communication to clients even when Tier 2 is engaged.
- Track and report operational metrics including MTTR and ticket resolution trends.
- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.
Required Qualifications / Must-Have Skills
- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.
- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.
- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.
- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.
- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.
- Solid SQL/Postgres skills for data-level investigation and validation during incidents.
- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
- Basic Python scripting capability for operational checks and quick-fix automation.
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.
- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.
- Willingness to work in rotational 24x7 shifts.
Good-to-Have / Nice-to-Have
- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.
- Exposure to IBM Sterling Integrator integration flows for context during incident triage.
- Telecom or high-availability enterprise support experience.
- Experience with CI/CD-driven deployment pipelines in a support context.
Experience Level
Associate to Mid-Level (typically 3 to 5 years)
Location / Work Mode
Onsite (Hyderabad / Bangalore or designated AT&T location)
Similar roles
-
Designated Support Engineer II
Zscaler Bengaluru, Karnataka, India
-
Technology Support Engineer
AA New Zealand Auckland, Auckland, New Zealand
-
Technical Support Engineer
Megaport Brisbane, Queensland, Australia
-
Application Support Engineer
Accenture Mumbai City, Maharashtra, India
-
Technical Support Engineer II - Hybrid (Tampa)
FusionTek Tampa, Florida, United States · $62K–$73K/yr
-
Remote Technical Support Engineer II
FusionTek United States · $62K–$73K/yr