Director Operational Resilience Engineering

Hybrid$142,800 – $304,200
Japan · United States · India

Job Description, Responsibilities & Requirements

About the Position

We are seeking a Director of Global Operational Resilience Engineering to lead reliability engineering efforts and improve datacenter reliability worldwide. The role is pivotal in applying reliability engineering methodologies to identify and eliminate recurring failure mechanisms, thereby enhancing the reliability, availability, resilience, and scalability of Microsoft's global datacenter fleet.

Responsibilities

  • Lead root cause analyses (RCA) and reliability engineering investigations into high-severity datacenter incidents to identify systemic failure mechanisms and drive permanent corrective actions.
  • Apply advanced reliability engineering methodologies, including Failure Modes and Effects Analysis (FMEA), Fault Tree Analysis (FTA), Weibull analysis, reliability modeling, and probabilistic risk assessment.
  • Perform incident forensics, statistical analysis, and trend analysis to identify recurring failure patterns, emerging risks, and opportunities for systemic improvement.
  • Develop and maintain fleet-wide reliability models that quantify operational risk, predict failure behavior, and prioritize mitigation efforts.
  • Embed RCA findings and reliability engineering insights into global operational standards, engineering requirements, maintenance strategies, and operating procedures.
  • Conduct systemic risk assessments across critical power, cooling, controls, and operational processes to improve resilience and availability.
  • Develop predictive metrics, leading indicators, and reliability dashboards that enable proactive identification and mitigation of operational risks.
  • Partner with Engineering, Regional Operations, Design, and Product teams to drive fleet-wide reliability improvements throughout the datacenter lifecycle.
  • Measure and validate the effectiveness of corrective actions using data-driven analysis to ensure sustainable elimination of recurring failure mechanisms.
  • Establish best practices, analytical frameworks, and continuous improvement methodologies that improve the reliability, availability, resilience, and scalability of the global datacenter fleet.
  • Embody our culture and values.

Requirements

Required Qualifications:

  • Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience OR
  • Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 7+ years technical engineering experience OR
  • Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 8+ years technical engineering experience.

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role.

Preferred Qualifications:

  • Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 7+ years technical engineering experience OR
  • Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 10+ years technical engineering experience OR
  • Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 12+ years technical engineering experience
  • 4+ years people management experience
  • Experience integrating telemetry, predictive analytics, and AI-driven tools into forensic processes
  • Prior experience in multinational environments
  • Professional certifications such as Certified Reliability Engineer (CRE), or equivalent.
  • Proven track record in Root Cause Analysis (RCA) and post-incident governance
  • Knowledge of risk assessment methodologies and containment strategies.
  • Familiarity with forensic frameworks, evidence handling, and reliability metrics.
  • Experience leading cross-functional teams and managing global programs
  • Advanced capability in failure analysis, data interpretation, and corrective action planning.
  • Deep understanding of industry standards for forensic engineering (e.g., ISO, IEC)

We Offer

  • Competitive salary range: USD $142,800 - $304,200 per year (depending on location)
  • Full-Time employment
  • Opportunity to work in a dynamic and innovative environment
  • Potential for career growth and professional development

About the Company

Microsoft Cloud Infrastructure and Operations (CO+I) is the engine that powers Microsoft's cloud services. The group is responsible for designing, building, and operating Microsoft’s global datacenters; managing the programmatic delivery of our critical infrastructure design, equipment procurement, construction delivery, infrastructure innovation, demand planning and capacity utilization of our unified infrastructure; and responsible for all operations needed to run the physical infrastructure. We focus on smart growth with an emphasis on automation, data-driven engineering, cost‐effectiveness, and environmental sustainability. We deliver the core infrastructure and foundational technologies for Microsoft's 200+ online businesses including Azure, Office 365, Bing, Xbox Live, Skype, and OneDrive. Our portfolio is built and managed by a team of subject matter experts working 24x7x365 to support services for more than 1 billion customers and 20 million businesses in over 90 countries worldwide.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances.

Job Details

Company name:
Microsoft Ukraine
Salary:
$142,800 – $304,200
Location:
Japan · United States · India
Employment Type:
Full-time
Work Mode:
Hybrid
Posted on TheJob:
Jul 13, 2026
Last checked:
Jul 13, 2026
Apply Now
© 2026 TheJob, Inc. All rights reserved.