SDM SRE Senior Engineer
The TJX Companies, Inc. Hyderabad, Telangana, India
Retail · 10,001+ employees
About the role
Implement SRE practices with a focus on automation, observability, and AIOps to improve production reliability and reduce manual toil. Collaborate with cross-functional teams to modernize support services and ensure operational readiness across enterprise applications.
What they look for
Requirements
Requires a bachelor's degree in a technical field and 6+ years of experience in SRE, DevOps, or production support engineering. Candidates must have strong hands-on experience with Azure, automation tools, observability platforms, and scripting languages.
Benefits
Full description
TJX Companies
At TJX Companies, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our four global Home Offices, Distribution Centers or Retail Stores—TJ Maxx, Marshalls, Homegoods, Homesense, Sierra, Winners, and TK Maxx, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer.
Job Description:
Location: India IT Office, Hyderabad / India
What you’ll discover
- Inclusive culture and career growth opportunities
- Global IT Organization which collaborates across U.S., Canada, Europe, India and Australia
- Challenging, collaborative, and team-based environment
What you’ll do
The Infrastructure and Operations organization embodies the hub of lifecycle engineering at TJX, delivering, maintaining, and optimizing our technology portfolio at cloud scale. We are a service-oriented team aimed at providing extraordinary experiences to thousands of TJX associates, business partners, and application delivery teams across the portfolio.
We are seeking a Site Reliability Engineer – AI Operations to implement Site Reliability Engineering practices with a strong focus on automation, observability, AIOps, Gen AI, and self-healing operations. This role will help improve production reliability, reduce manual toil, accelerate incident response, and enable intelligent operational automation across enterprise applications and platforms.
The ideal candidate combines strong production operations experience with engineering, automation, observability, and AI-enabled operations capabilities. This role will work closely with application, infrastructure, cloud, security, operations, and leadership teams to strengthen reliability, improve operational efficiency, and modernize how production services are supported.
Key Responsibilities
Site Reliability Engineering & Operations
- Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.
- Support production reliability by improving availability, performance, scalability, and operational readiness of applications and platforms.
- Participate in incident response, troubleshooting, root cause analysis, postmortems, and corrective/preventive action planning.
- Drive shift-left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines.
- Support reliable application and platform operations in Azure Cloud.
Automation & Self-Healing
- Build and maintain automation scripts, runbooks, self-healing workflows, and operational tools to reduce manual effort and improve MTTR.
- Identify repetitive operational activities and automate them where appropriate.
- Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.
- Create and maintain reusable automation playbooks, operational procedures, and knowledge articles.
- Ensure automation activities follow change management, governance, security, and compliance processes.
Observability, Monitoring & Reliability Metrics
- Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.
- Work with monitoring and APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
- Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards.
- Analyze service health trends, alert patterns, and operational data to identify improvement opportunities.
- Improve alert quality, event correlation, and operational visibility across enterprise applications and platforms.
AIOps, Gen AI & Intelligent Operations
- Support AI-driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations.
- Contribute to the design and implementation of Gen AI and AIOps use cases that improve incident response and operational productivity.
- Explore opportunities to use AI-enabled operations to reduce toil, improve service reliability, and enhance support experiences.
- Ensure AI-enabled operations follow enterprise security, compliance, data privacy, and responsible AI guidelines.
- Collaborate with engineering, operations, and platform teams to evaluate emerging AI and automation capabilities.
Collaboration & Continuous Improvement
- Work closely with application, infrastructure, cloud, operations, security, and leadership teams.
- Support cross-functional initiatives focused on reliability, automation, service health, and operational excellence.
- Communicate technical issues, risks, and improvement opportunities clearly to stakeholders.
- Contribute to continuous improvement by identifying gaps in process, tooling, monitoring, and documentation.
- Support knowledge sharing and operational readiness across global teams.
What you’ll need
We seek creative, customer-focused individuals with strong SRE, DevOps, production operations, automation, observability, and AI-enabled operations experience. This role requires a continuous improvement mindset and the ability to balance innovation with reliability, stability, security, and compliance.
Education
- Bachelor’s degree in Information Technology, Computer Science, Engineering, or equivalent practical experience.
Experience
- 6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles.
- Hands-on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation.
- Experience supporting production applications and platforms in large-scale enterprise environments.
- Experience working across application, infrastructure, cloud, operations, security, and leadership teams.
- Experience working within ITIL-based operational environments.
Required Technical Skills
- Strong experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
- Experience with SLIs, SLOs, SLAs, operational KPIs, service health dashboards, and reliability metrics.
- Strong coding/scripting experience in one or more languages such as Python, Java, PowerShell, Go, or Shell scripting.
- Hands-on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate, Rundeck, or similar tools.
- Knowledge of AIOps, Gen AI, or AI-enabled operations use cases such as anomaly detection, log analysis, ticket classification, event correlation, and incident summarization.
- Experience integrating enterprise tools using APIs, webhooks, scripts, and automation workflows.
- Hands-on experience supporting or implementing solutions in Azure Cloud.
- Understanding of distributed systems, APIs, microservices, cloud-native applications, and reliability engineering principles.
- Strong troubleshooting, analytical, communication, and stakeholder management skills.
Preferred Qualifications
- Experience with Agentic AI, AI agents, or Gen AI-based automation for IT operations.
- Experience with Azure OpenAI, Microsoft Copilot Studio, LangChain, Semantic Kernel, vector databases, or RAG-based solutions.
- Experience with ITSM and incident response tools such as ServiceNow, Jira Service Management, PagerDuty, xMatters, Opsgenie, or similar platforms.
- Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code, GitOps, or DevSecOps.
- Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time-series analysis.
- Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls.
- Experience in the retail domain.
- Relevant certifications in Azure, DevOps, SRE, AI, or ITIL.
Key Competencies
- Strong analytical and problem-solving mindset.
- Automation-first and continuous improvement mindset.
- Strong communication and stakeholder management skills.
- Ability to influence without direct authority.
- Ability to work across product, engineering, operations, cloud, security, and leadership teams.
- Curiosity and willingness to explore AIOps, Gen AI, and Agentic AI capabilities.
- Ability to balance innovation with reliability, stability, security, and compliance.
- Customer-focused approach with strong ownership and accountability.
Working Expectations
This role is based in India and requires collaboration with global teams across U.S., Canada, Europe, India, and Australia. The candidate should be comfortable supporting flexible working hours as needed to engage with global stakeholders, participate in critical meetings, support escalations, and ensure seamless delivery across regions.
Join us
Join us and Discover Different at TJX.
In addition to our open door policy and supportive work environment, we also strive to provide a competitive salary and benefits package. TJX considers all applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, gender identity and expression, marital or military status, or based on any individual's status in any group or class protected by applicable federal, state, or local law. TJX also provides reasonable accommodations to qualified individuals with disabilities in accordance with the Americans with Disabilities Act and applicable state and local law.
Address:
Salarpuria Sattva Knowledge City, Inorbit Road
Location:
APAC Home Office Hyderabad IN
Similar roles
-
Staff Site Reliability Engineer
Anduril Industries Costa Mesa, California, United States · $191K–$253K/yr
-
Senior Cloud Architect (Network & SRE)
LivePerson New York, New York, United States · $150K–$160K/yr
-
Site Reliability Engineer
The Next Chapter W&S Nijverdal, Overijssel, Netherlands · €66K–€78K/yr
-
Staff Site Reliability Engineer, AI Foundations, F1 Query
Google San Jose, California, United States · $207K–$300K/yr
-
Site Reliability Engineer - 7 Month Contract
Orion Health Minot, North Dakota, United States
-
Senior Application Support Engineer (SRE)
DTCC Tampa, Florida, United States · $75K–$150K/yr