The Home Depot

Software Engineering Manager - Reliability Engineering, Store Systems (Remote)

The Home Depot · Flórida, Paraná, Brazil · $140K–$240K/yr

Retail · 10,001+ employees

Yesterday
Remote Senior (5-10 yrs) Full-time Brazil
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Software Engineering Manager will lead the Store Systems Reliability Engineering team to ensure platform resilience, performance, and security through automation and incident management. They are responsible for mentoring junior engineers, driving systemic problem resolution, and establishing Service Level Objectives to support critical engineering services.

What they look for

Reliability Engineering Incident Management Root Cause Analysis Service Level Objectives Automation Infrastructure as Code Cloud Computing Observability Team Leadership Mentoring Chaos Engineering CI/CD Python Go Java System Architecture

Requirements

Candidates must have at least 5 years of relevant work experience and a bachelor's degree or equivalent in a related field. Proficiency in cloud infrastructure, programming languages like Python or Go, and modern observability tools is required for this leadership role.

Full description

With a career at The Home Depot, you can be yourself and also be part of something bigger.

Position Purpose:

As a Software Reliability Engineer Manager on the Store Systems RE team, you will ensure the resilience, performance, and security of our Store System and related application, working alongside partner Dev and Enablement teams supporting our critical Engineering Experience services. As a subject matter expert in Reliability Principles and Practices, you will act as an anchor for our Store Systems RE team, mentoring junior engineers — leading incident triage, root cause analysis, driving no-repeat resolutions of systemic problems through ownership of blameless postmortems. Your mission is to ensure that reliability is engineered into our platforms through automation, rigorous change, incident, problem management, and destructive testing, while establishing and enforcing Service Level Objectives (SLOs) that let product teams build and run customer-facing workloads on highly available, paved-path solutions.

Key Responsibilities:

  • 30% Strategy & Planning:
  • Looks across teams with a focus on alignment and dependencies
  • Gains a thorough understanding of infrastructure needs and guides teams to design infrastructure platforms that meet end user requirements
  • Translates product and project goals into infrastructure strategy and clearly communicates direction and priorities to teams and business partners
  • Determines value to the business of anticipated Systems Engineering efforts
  • Identifies goals, metrics, and appropriate analytics to measure the performance of Systems Engineering teams; continually makes recommendations and refinements on approaches based on learnings
  • Reviews recommended solutions and work of Systems Engineers to ensure alignment with company, stakeholder, and end user priorities
  • 20% Delivery & Execution:
  • Leads configuration, debugging, and support for infrastructure
  • Documents, reviews and ensures that all quality and change control standards are met
  • Leads field and corporate roll-outs of technology
  • Leads the stand up of necessary system software, hardware, and equipment (physical or virtual) to meet changing infrastructure needs
  • Creates and optimizes specifications for complex technology solutions
  • Provides regular status to leadership regarding progress of Systems Engineering efforts
  • Manages vendor relationships
  • Manages, reviews, and approves purchase requests for hardware and software
  • 20% Support & Enablement:
  • Removes roadblocks and obstacles that may impair Systems Engineers to help ensure efforts meet strategic, financial, and technical goals
  • Receives and prioritizes escalations and incoming requests from product teams and stakeholders
  • Guides the production of in-house documentation around solutions
  • Monitors tools and proactively helps teams struggling with systems issues
  • 30% People:
  • Provides leadership, mentoring, and coaching to Systems Engineering professionals
  • Attracts, retains, and develops top talent
  • Conducts annual and mid-year reviews, reviewing individual development plans and providing performance feedback
  • Fosters collaboration with team members to drive value, and identify and resolve impediments
  • Advocates for the end user and stakeholder by becoming associated with the product, empathizing with and understanding user needs
  • Guides more junior team members in strategy, alignment, analysis, and execution tasks
  • Participates in and contributes to learning activities around systems engineering core practices (communities of practice)

Direct Manager/Direct Reports:

  • Typically reports to the Systems Engineer Sr. Manager, Technology Director or Sr. Director.

Travel Requirements:

  • Typically requires overnight travel 5% to 20% of the time.

Physical Requirements:

  • Most of the time is spent sitting in a comfortable position and there is frequent opportunity to move about. On rare occasions there may be a need to move or lift light articles.

Working Conditions:

  • Located in a comfortable indoor area. Any unpleasant conditions would be infrequent and not objectionable.

Minimum Qualifications:

  • Must be eighteen years of age or older.
  • Must be legally permitted to work in the United States.
  • Must be legally permitted to work in the United States

Preferred Qualifications:

  • Technical Leadership: Guide and mentor teams of SREs, infrastructure engineers, and operations staff, fostering a culture of continuous learning and engineering excellence.
  • Strategic Planning: Define and execute the long-term reliability roadmap, aligning technical goals with broader business objectives and product delivery schedules.
  • Talent Development: Recruit, retain, and develop top engineering talent. Conduct performance reviews, establish clear career paths, and manage skill-gap training.
  • Budget & Resource Management: Oversee departmental budgets, cloud spending, and vendor negotiations. Optimize resource allocation to balance cost with system performance.
  • Service Level Management: Define, measure, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) across various product lines.
  • Incident & Crisis Management: Serve as the ultimate escalation point for high-severity incidents. Lead blameless post-mortems and ensure preventative action items are successfully implemented.
  • Capacity Planning: Anticipate system growth and scale infrastructure accordingly to prevent degradation during peak loads.
  • Chaos Engineering & Resilience: Champion chaos engineering practices and disaster recovery planning to proactively identify and mitigate systemic vulnerabilities.
  • Cloud & Infrastructure: Deep expertise in managing enterprise-scale environments across major cloud providers (Google Cloud Platform, AWS, or Microsoft Azure).
  • Infrastructure as Code (IaC): Advanced knowledge of provisioning and configuration management tools such as Terraform, Ansible, Chef, or Puppet.
  • Observability & Telemetry: Strong proficiency in modern monitoring, logging, and tracing stacks (e.g., Datadog, Prometheus, Grafana, Splunk, New Relic, ELK stack).
  • CI/CD & Automation: Experience overseeing robust continuous integration and deployment pipelines (e.g., Jenkins, GitLab CI, GitHub Actions) to ensure safe and rapid software releases.
  • Software Engineering Background: Proficiency in at least one or more programming languages (Python, Go, Java, or Bash) to guide automation efforts and participate in high-level architectural reviews.
  • Engineering Alignment: Partner closely with software development, QA, and security teams to build reliability into the software development lifecycle (SDLC) from day one ("Shift Left" reliability).
  • Stakeholder Communication: Translate complex technical metrics and incidents into clear, actionable insights for non-technical executive leadership.
  • Vendor Management: Evaluate, select, and manage relationships with third-party SaaS and infrastructure providers, ensuring compliance and optimal service delivery.
  • Security Integration: Collaborate with InfoSec to ensure infrastructure complies with industry regulations (e.g., PCI-DSS, SOC2, HIPAA) and corporate security policies.
  • Access & Identity Management: Advocate for and implement least-privilege access models and robust audit logging across critical systems.
  • Minimum Education:
  • The knowledge, skills and abilities typically acquired through the completion of a bachelor's degree program or equivalent degree in a field of study related to the job.

Preferred Education:

  • No additional education

Minimum Years of Work Experience:

  • 5

Preferred Years of Work Experience:

  • No additional years of experience

Minimum Leadership Experience:

  • None

Preferred Leadership Experience:

  • None

Certifications:

  • None

Competencies:

  • Attracts Top Talent: Attracting and selecting the best talent to meet current and future business needs
  • Balances Stakeholders: Anticipating and balancing the needs of multiple stakeholders
  • Builds Effective Teams: Building strong-identity teams that apply their diverse skills and perspectives to achieve common goals
  • Business Insight: Applying knowledge of business and the marketplace to advance the organization's goals
  • Collaborates: Building partnerships and working collaboratively with others to meet shared objectives
  • Communicates Effectively: Developing and delivering multi-mode communications that convey a clear understanding of the unique needs of different audiences
  • Develops Talent: Developing people to meet both their career goals and the organization's goals
  • Drives Engagement: Creating a climate where people are motivated to do their best to help the organization achieve its objectives
  • Drives Vision and Purpose: Painting a compelling picture of the vision and strategy that motivates others to action
  • Manages Ambiguity: Operating effectively, even when things are not certain or the way forward is not clear
  • Organizational Savvy: Maneuvering comfortably through complex policy, process, and people-related organizational dynamics
  • Situational Adaptability: Adapting approach and demeanor in real time to match the shifting demands of different situations

For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents:  

The pay range for this position is between $140,000.00 - $240,000.00