Vice President, Major Incident Manager (Production Services Specialist II), Application Production Services & Engineering
Bank of America Singapore, Singapore
Banking · 10,001+ employees
About the role
The Production Services Specialist II leads end-to-end response activities for significant business-impacting incidents using disciplined major incident command practices. They are responsible for orchestrating cross-functional response teams, validating incident severity, and facilitating executive communications to ensure rapid service restoration.
What they look for
Requirements
Candidates must have 7+ years of experience in enterprise incident management, technology operations, or command center leadership. Strong proficiency in ITSM principles, executive-level communication, and the ability to make sound decisions under pressure are required.
Benefits
Full description
Job Description:
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.
We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.
Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!
Job Description:
The Production Services Specialist II provides advanced leadership across enterprise incident response, serving in Major Incident Commander and Major Incident Communications capacities for significant business-impacting events. The role orchestrates coordinated response efforts across technology domains, validates incident severity and business impact, drives time-sensitive decisions, and maintains clear communication with technical teams, business partners, operational leaders, and executive stakeholders.
This specialist is accountable for establishing structure during critical events, maintaining situational awareness, aligning resources to the highest-priority recovery actions, and ensuring escalation paths are engaged appropriately. The position requires strong judgment under pressure, executive-level communication skills, broad understanding of production technology operations, and the ability to translate complex technical conditions into concise business-relevant updates.
Responsibilities:
- Lead end-to-end response activities for significant business-impacting incidents using disciplined major incident command practices, including clear role assignment, bridge control, time-boxed recovery checkpoints, decision logging, and structured restoration governance.
- Establish incident command structures and coordinate cross-functional response teams.
- Direct prioritization, escalation, resource allocation, containment, workaround, and restoration activities during active major incidents while ensuring technical teams remain aligned to the highest-probability recovery path.
- Maintain situational awareness across business impact, operational impact, risk, dependencies, and recovery progress.
- Facilitate executive, stakeholder, and major incident communications throughout the incident lifecycle, ensuring updates are timely, business-relevant, risk-aware, and consistent with established communication cadence and escalation protocols.
- Validate incident severity, business impact assessments, escalation decisions, and major incident engagement criteria.
- Drive rapid stabilization while minimizing customer, associate, and business disruption.
- Ensure compliance with incident management standards, governance requirements, operational procedures, audit expectations, severity models, escalation criteria, and post-incident documentation requirements.
- Partner with technical teams to identify recovery strategies, validate service restoration, preserve incident timelines, support root cause analysis activities, and translate post-incident findings into actionable improvement opportunities.
- Monitoring triage orchestration for single channel impacting incidents to ensure that incidents are being managed appropriately and making adequate traction given impact and urgency and alert executive stakeholders of possible escalations.
- Perform succinct and comprehensive regional turnover for service continuity.
- Apply SRE-aligned concepts such as reliability engineering, alert hygiene, toil reduction, post-incident reviews, automation opportunities, service-level awareness, and continuous improvement to strengthen operational resiliency.
- Mentor junior incident management resources and develop command center operational capability through coaching, shadowing, simulation exercises, playbook reinforcement, communications review, feedback loops, and knowledge sharing across recurring incident patterns and command practices.
Required Skills:
- 7+ years in enterprise incident management, technology operations, command center leadership, or production services.
- Strong understanding of ITSM principles, major incident response, incident severity, business impact validation, and governance controls. Knowledgeable with Change Management and Problem Management.
- Demonstrated ability to direct cross-functional technical domains during high urgency service disruptions.
- Experience leading crisis and incident response as an Incident Commander in large-scale or high-impact environments.
- Executive-level written and verbal communication skills, including ability to produce concise incident updates and business impact summaries.
- Ability to make sound decisions under pressure while balancing restoration, risk, communications, and stakeholder expectations.
- Proficiency with Service Management (e.g. ServiceNow), collaboration, telemetry, and Microsoft 365 tools.
Desired Skills:
- ITIL Foundation, advanced ITSM certification, incident command, crisis management, business continuity, or operational resilience training that supports disciplined response leadership in high-impact production environments.
- Experience in financial services, regulated technology operations, enterprise production support, or global command center environments where auditability, risk management, resiliency, and stakeholder confidence are critical.
- Demonstrated exposure to post-incident reviews, root cause analysis support, problem management, corrective action tracking, governance forums, operational risk reviews, and continuous improvement practices that reduce repeat incidents and improve response quality.
- Experience mentoring incident managers, communications leads, operational coordinators, or junior command center resources through coaching, shadowing, playbook review, incident simulations, communications feedback, and readiness development.
- Working knowledge of observability, monitoring, telemetry, service health dashboards, alerting practices, and incident correlation methods used to accelerate triage, validate impact, and support restoration decisions.
- Familiarity with SRE, reliability engineering, service-level indicators, service-level objectives, automation, toil reduction, and operational excellence practices that strengthen resiliency and response maturity.
- Ability to influence without direct authority across technology, operations, risk, and business teams while maintaining composure, urgency, and clear accountability during ambiguous or rapidly evolving incidents.