Oracle

Director, Reliability Engineering (Nashville, TN on-site)

Oracle Nashville, Tennessee, United States · $146K–$306K/yr

IT Services and IT Consulting · 10,001+ employees

4 d ago
Principal (10+ yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will lead the reliability engineering organization to define strategy, standards, and operating models for mission-critical data center infrastructure. This role involves partnering with cross-functional teams to improve asset performance, resilience, and lifecycle management across a global portfolio.

What they look for

Reliability engineering Data center operations Mission-critical infrastructure Strategic leadership Electrical systems Mechanical systems Controls and automation Root cause analysis FMEA Reliability-centered maintenance Asset lifecycle management Predictive maintenance Risk management Infrastructure resilience Technical program management Cross-functional leadership

Requirements

Candidates must have 10+ years of experience in engineering or critical facilities, with significant leadership experience in mission-critical environments. A bachelor's degree in a relevant technical discipline is required, along with deep expertise in reliability methodologies and infrastructure systems.

Benefits

Medical insurance Dental insurance Vision insurance Short term disability Long term disability Life insurance Supplemental life insurance Flexible spending accounts Pre-tax commuter and parking benefits 401(k) savings and investment plan Paid time off Paid holidays Paid sick leave Paid parental leave Adoption assistance Employee stock purchase plan Financial planning Group legal

Full description

>>This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle’s relocation policies. Candidates should expect a minimum of 25% travel, with additional travel as business needs require.<<

As Director of Building Automation, you will provide strategic and organizational leadership for the reliability engineering function supporting Oracle Cloud Infrastructure’s mission-critical data center portfolio. You will own the vision, operating model, engineering standards, and portfolio programs that improve infrastructure availability, maintainability, resilience, and lifecycle performance at scale.

This role requires significant hands-on and leadership experience within mission-critical environments. Direct data center experience is preferred. Candidates must have demonstrated experience supporting the electrical, mechanical, controls, and operational systems required to maintain continuous operations in mission-critical environments.

You will lead managers, engineers, analysts, and technical programs responsible for reliability engineering, asset performance, predictive maintenance, failure analysis, defect elimination, and lifecycle risk management across OCI's data center infrastructure.

You will partner with senior leaders across Data Center Operations, Engineering, Design, Construction, Commissioning, Automation, Procurement, and other infrastructure organizations to translate operational experience and engineering data into long-term reliability strategy. You will ensure lessons learned from incidents, equipment performance, maintenance activities, and portfolio trends result in durable improvements to standards, designs, operating practices, and investment priorities.

Success in this role requires the ability to operate at both strategic and technical levels—setting multi-year direction for the reliability organization while maintaining sufficient engineering depth and data center operational knowledge to challenge assumptions, assess complex infrastructure risks, and drive disciplined decision-making across a rapidly scaling global portfolio.

Responsibilities

Key Responsibilities

  • Lead the Reliability Engineering organization supporting multiple regions, sites, and infrastructure programs across OCI's mission-critical data center portfolio.
  • Define the multi-year reliability engineering strategy, organizational roadmap, operating model, and investment priorities required to improve data center infrastructure resilience and support OCI's continued growth.
  • Establish and govern portfolio-wide reliability engineering standards and methodologies, including FMEA/FMECA, RCA, Reliability-Centered Maintenance (RCM), criticality assessment, defect elimination, reliability growth, and continuous improvement practices.
  • Build and lead a high-performing organization of managers, engineers, analysts, and technical specialists, establishing clear accountability, technical expectations, career development, and succession plans.
  • Own portfolio-level programs that improve reliability across critical data center infrastructure, including electrical distribution, UPS systems, generators, mechanical cooling systems, controls, automation, and supporting facility systems.
  • Establish a comprehensive reliability measurement framework that provides leadership with visibility into asset health, failure trends, systemic risks, repeat events, corrective actions, maintenance effectiveness, and lifecycle exposure.
  • Define reliability KPIs, targets, governance mechanisms, and executive reporting that enable data-driven prioritization of operational and engineering investments.
  • Establish governance for corrective and preventive actions resulting from incidents, root cause analyses, audits, equipment failures, and reliability trend reviews, ensuring actions are completed, verified for effectiveness, and sustained.
  • Drive systematic identification and elimination of recurring and systemic failure modes across the data center portfolio rather than relying solely on site-specific remediation.
  • Sponsor the development and adoption of predictive and condition-based maintenance capabilities, including monitoring, analytics, automation, asset health modeling, and emerging technologies that improve early detection of equipment degradation and failure risk.
  • Partner with Data Center Operations leadership to continuously improve maintenance strategy, operational readiness, troubleshooting practices, procedures, failure response, and infrastructure risk management.
  • Provide reliability governance and technical leadership for commissioning, acceptance testing, operational handover, major maintenance, retrofits, capacity expansion, and infrastructure lifecycle decisions.
  • Establish portfolio approaches to asset lifecycle management, including equipment health, utilization, failure history, remaining useful life, obsolescence, spare parts strategy, replacement planning, and end-of-life risk.
  • Partner with Design, Construction, Engineering, and Procurement leadership to ensure lessons from operating facilities influence equipment specifications, design standards, redundancy strategies, maintainability requirements, vendor selection, and total cost of ownership.
  • Develop mechanisms to convert site-level events and engineering findings into portfolio-wide standards, design changes, maintenance improvements, and risk-reduction programs.
  • Lead technical and business reviews of significant reliability risks and provide clear recommendations regarding mitigation strategies, priorities, investment requirements, and residual operational risk.
  • Develop strong partnerships with equipment manufacturers, service providers, and technology partners to improve equipment performance, failure intelligence, serviceability, and long-term reliability.
  • Establish effective operating rhythms for the organization, including portfolio reviews, technical reviews, risk escalation, program governance, resource prioritization, and executive communications.
  • Represent Reliability Engineering in senior leadership discussions involving infrastructure risk, operational performance, capacity growth, capital planning, and long-term data center strategy.

Minimum Qualifications

  • 10+ years of progressive engineering, reliability, maintenance, critical facilities, or infrastructure experience, including significant experience directly supporting mission-critical environments.
  • Demonstrated experience working within mission-critical operations or engineering environments where infrastructure availability, redundancy, maintenance execution, and operational risk directly affect service continuity.
  • 5+ years of progressive leadership experience, including responsibility for engineering managers, senior technical professionals, or large multi-site technical organizations and programs.
  • Demonstrated technical knowledge of mission-critical infrastructure, including experience with electrical distribution, UPS systems, generators, mechanical cooling systems, controls/automation, and integrated facility operations.
  • Demonstrated experience developing and implementing reliability, maintenance, asset-management, or operational excellence strategies across mission-critical infrastructure.
  • Strong working knowledge of reliability engineering methodologies, including structured root cause analysis, FMEA/FMECA, RCM, criticality analysis, defect elimination, reliability metrics, and lifecycle risk management.
  • Demonstrated experience evaluating infrastructure failures, operational events, equipment performance, maintenance effectiveness, and systemic reliability risks within mission-critical environments.
  • Demonstrated ability to use operational and engineering data to identify systemic risks, establish priorities, and influence significant technical or business decisions.
  • Experience leading cross-functional initiatives involving Data Center Operations, Engineering, Design, Construction, Commissioning, Procurement, OEMs, vendors, and other technical stakeholders.
  • Experience establishing engineering governance, standards, KPIs, and management mechanisms across multiple sites, regions, or infrastructure programs in mission-critical environments.
  • Demonstrated ability to communicate complex technical risks, tradeoffs, and investment recommendations to senior and executive leadership.
  • Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Industrial Engineering, Systems Engineering, Reliability Engineering, or a related technical discipline; or equivalent relevant industry experience.

Skills and Competencies

  • Mission-Critical Technical Leadership: Strong understanding of mission-critical infrastructure, operating practices, redundancy, maintenance risk, failure modes, and the interdependencies between electrical, mechanical, controls, and operational systems.
  • Organizational Leadership: Ability to build, develop, and lead managers and senior technical professionals while establishing clear accountability and a strong engineering culture.
  • Reliability Strategy: Ability to translate data center infrastructure performance, business growth, and operational risk into a coherent multi-year reliability strategy and investment roadmap.
  • Technical Judgment: Ability to evaluate complex infrastructure reliability issues, challenge technical assumptions, understand operational consequences, and make decisions under uncertainty.
  • Systems Thinking: Ability to connect individual equipment failures and site-level events to systemic portfolio risks involving design, maintenance, operations, suppliers, processes, or organizational practices.
  • Data-Driven Decision Making: Ability to convert large volumes of operational, maintenance, failure, and asset data into actionable insights and investment priorities.
  • Executive Influence: Ability to communicate technical risk and recommendations clearly to senior leaders and build alignment across organizations with different objectives and priorities.
  • Operational Excellence: Strong commitment to disciplined execution, corrective-action rigor, standardization, measurable improvement, and sustained results.
  • Change Leadership: Ability to introduce and scale new engineering methods, technologies, processes, and operating models across a large and geographically distributed data center organization.
  • Talent Development: Demonstrated ability to develop engineering leaders and technical talent, establish career paths, strengthen organizational capability, and build succession depth.

Preferred Qualifications

  • Direct data center experience supporting mission-critical infrastructure and operations is preferred.
  • Experience leading reliability engineering, critical facilities engineering, or asset-management organizations within hyperscale, colocation, or large-scale enterprise data centers.
  • Experience supporting geographically distributed or global data center portfolios.
  • Deep technical expertise in one or more critical infrastructure domains, with broad working knowledge across electrical distribution, UPS, standby generation, mechanical cooling, controls/automation, and integrated facility operations.
  • Experience developing and scaling predictive maintenance, condition-based monitoring, failure trend analysis, asset health modeling, and equipment risk-ranking programs within data center environments.
  • Advanced knowledge of reliability, availability, and maintainability analysis; maintenance strategy optimization; spare parts planning; lifecycle modeling; and total cost of ownership.
  • Experience governing commissioning, operational acceptance, maintenance program design, and readiness of new or modified mission-critical data center infrastructure.
  • Experience with CMMS/EAM, DCIM, EPMS, BMS, monitoring, telemetry, analytics, and automation platforms used to manage critical data center infrastructure.
  • Experience developing portfolio-level KPI frameworks, executive dashboards, reliability reviews, and risk-governance mechanisms.
  • Experience partnering with OEMs and strategic suppliers to address systemic equipment issues, improve product reliability, and influence equipment roadmaps or specifications.
  • Experience incorporating operational lessons learned into engineering standards, design requirements, equipment specifications, and capital investment decisions.
  • Demonstrated experience managing organizational growth, workforce planning, resource prioritization, and technical capability development across geographically distributed teams.

Preferred Credentials / Certifications

  • Certified Maintenance & Reliability Professional (CMRP) preferred.
  • Certified Reliability Engineer (CRE) preferred.
  • ASQ, SMRP, or equivalent reliability, maintenance, engineering, or quality certifications are a plus.
  • Data center or critical-environment credentials, including relevant Uptime Institute training or certifications, are a plus.
  • OEM, controls, analytics, asset-management, or condition-monitoring training applicable to critical data center infrastructure is a plus.
  • Advanced training or certification in FMEA/FMECA, RCA, RCM, Lean, Six Sigma, or structured problem-solving methodologies is a plus.
  • Advanced technical or business degree is a plus.

Leadership Scope

  • The Director – Reliability Engineering is expected to operate beyond individual programs or regions and establish the mechanisms through which OCI manages reliability risk at portfolio scale. The role requires balancing immediate operational priorities with long-term infrastructure strategy and ensuring reliability engineering becomes an increasingly predictive, data-driven, and standardized capability.
  • The Director will be accountable for building an organization capable of identifying emerging reliability risks before they become widespread operational issues, converting lessons learned into durable improvements, and ensuring investments in maintenance, technology, engineering, and infrastructure are prioritized according to measurable risk and business impact.
  • This position requires a leader who understands the operational realities of mission-critical environments and can apply that experience to reliability strategy at scale. Direct data center experience is preferred.

Why Oracle Cloud Infrastructure?

Global impact at scale: Contribute directly to how mission-critical OCI data centers operate across regions and continents, influencing infrastructure reliability, security, sustainability, and long-term capacity growth.

Technically rigorous environment: Work alongside experienced engineers, automation specialists, and compliance teams in a rapidly scaling hyperscale cloud infrastructure, where disciplined execution and technical depth matter.

Culture built on operational excellence: Join an organization that values safety, process rigor, clear accountability, and continuous improvement as foundational to protecting uptime and customer trust.

Long-term career development: Benefit from internal mobility, role-based technical training, and development opportunities designed for professionals building long-term careers in cloud infrastructure and facilities operations.

#LI-SB36

Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9. 11 paid holidays 10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours. 11. Paid parental leave 12. Adoption assistance 13. Employee Stock Purchase Plan 14. Financial planning and group legal 15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted. As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.