VP, Site Reliability Engineering (SRE & Observability Platform)
FIS Global Jacksonville, Florida, United States
IT Services and IT Consulting · 10,001+ employees
About the role
The VP of SRE will lead the platform organization to improve observability and reduce production incidents using agentic engineering. They will define the enterprise strategy, build high-performing teams, and deliver a paved road for product engineers to meet reliability goals.
What they look for
Requirements
Candidates must have 15+ years of software engineering experience with at least 7 years in senior leadership roles managing SRE or platform organizations. Deep expertise in observability, AI/agentic systems, and cloud-native infrastructure is required, along with a bachelor's degree in Computer Science.
Full description
Position Type :
Full time Type Of Hire :
Experienced (relevant combo of work and education) Education Desired :
Bachelor of Computer Science
Job Description
About the role:
Production reliability is a board-level priority, and we are placing it in one hands-on leader. We are hiring a Vice President of Engineering to found and lead our Site Reliability Engineering (SRE) platform organization and own the mission end to end: raise observability maturity and materially reduce production incidents across the engineering estate. You will do it by making agentic engineering the core method of the function — using AI agents to run the platform itself, and putting AI agents in the hands of product engineers so they can meet the company's observability requirements and SRE goals with far less manual effort.
This is a builder's role with executive visibility. You will stand up a central platform team that treats reliability and observability as a product, deliver an aggressive multi-phase roadmap, and change how thousands of engineers instrument, operate, and take ownership of what they ship — in a large, complex, tech-debt-heavy environment where the winning strategy is paved roads over mandates. Central to that is a two-part AI mandate: making this function fully agent-native, and delivering AI agents to product engineers so they can fulfill their observability requirements and the company's SRE goals with far less manual effort.
You will report directly to the SVP of Platform Engineering and partner with leaders across engineering, product, and the executive team. The mission has a named executive sponsor and committed funding; your job is to turn that mandate into outcomes.
The AI mandate: two transformations you will lead:
Agentic engineering is not a side initiative in this role — it is central to how the function operates and to the value it delivers to the enterprise. You will be accountable for two connected transformations:
1. Transform this function to be fully agent-native
Re-found the platform organization around agentic engineering. Agents — not manual toil — should carry the load of instrumenting legacy code, investigating incidents, generating configuration, assisting on-call, and, over time, executing guarded remediation. You will build the agent control plane, guardrails, and evaluation harnesses that make this safe, and turn the function into the company's proof point for what disciplined, agent-first engineering looks like at scale.
2. Put AI agents in the hands of product engineers for observability and SRE outcomes
Deliver AI agents to product engineers as part of the paved road so they can meet the company's observability requirements and SRE goals with far less manual effort — agents that instrument their services, generate SLOs, dashboards, and alerts, investigate incidents, and assist on-call. These agents, and the reliability, evaluation, guardrails, and cost governance behind them, are how product teams hit reliability targets at scale. This role owns the SRE platform: the agents it provides serve observability and reliability, not product-feature development.
What you'll own:
- The SRE platform organization — a central platform team plus embedded/partner SREs, an enablement function, and a cross-cutting reliability champions network.
- The observability & reliability platform — telemetry pipeline (OpenTelemetry), metrics/logs/traces backends, dashboards and alerting, the SLO and error-budget system, incident management, and RCA/correlation.
- The paved road — shared instrumentation SDKs, templates, and dashboards/alerts/SLOs-as-code that make golden-signal observability near-automatic.
- The agentic engineering stack — internal agents for instrumentation, investigation/RCA, config generation, on-call, and guarded remediation — plus the agent control plane, guardrails, and evaluation that keep them safe.
- The agentic paved road for product engineers — AI agents delivered to product teams to instrument services, generate SLOs/dashboards/alerts, investigate incidents, and assist on-call — so they meet observability requirements and SRE goals with minimal manual effort, backed by guardrails and evaluation.
- Reliability & AI governance — error-budget policy, blameless incident and postmortem practice, AI safety and human-in-the-loop controls, and the metrics reported to senior leadership.
- The roadmap, budget, and vendor strategy — an 18-month phased plan, the operating budget, tooling selection, and cost governance for both telemetry and AI workloads.
Key responsibilities:
- Set the strategy and vision. — Own the enterprise strategy for reliability, observability, and agentic engineering. Define what “reliability as a product” and “agent-native engineering” mean here, and keep both tied to business outcomes.
- Build and lead the organization. — Recruit, structure, and grow a high-performing platform organization spanning SRE, platform, and AI engineering — hiring and developing senior, staff, and principal talent and the managers who lead them.
- Deliver the roadmap. — Execute the phased plan — mobilize and instrument, build foundations, scale and standardize the paved road with SLO coverage and error-budget policy, then bring proactive and agentic capabilities to production — on aggressive, overlapping timelines.
- Make the function agent-native. — Drive adoption of internal agents across the platform's own work, with least-privilege access, blast-radius limits, human-in-the-loop controls, and evaluation before any increase in autonomy.
- Put agents in product engineers' hands. — Deliver AI agents through the paved road that let product engineers meet observability requirements and SRE goals — instrumenting services, standing up SLOs, and resolving incidents — with far less manual effort, backed by guardrails and evaluation.
- Make reliability measurable. — Stand up SLOs and error budgets, modern incident management, and blameless postmortems; establish error-budget policy in partnership with leadership.
- Drive adoption across the estate. — Win teams over with paved roads and lighthouse wins, not mandates — through reliability reviews, enablement, office hours, and published before/after results.
- Own the tooling and its economics. — Select and evolve an industry-leading, OpenTelemetry-native and AI-native tool stack, with cost and cardinality governance built in from the start for both telemetry and inference.
- Operate as an executive partner. — Manage stakeholders across engineering and product, report progress and reliability/AI metrics to the CTO and executive team, and steward budget and headcount.
What success looks like:
- First 90 days — organization mobilized and sponsor alignment confirmed; reliability baseline established; 2–4 lighthouse services selected; instrumentation underway with the first golden-signal dashboards and SLOs live; the first internal agents piloted.
- By 12 months — the paved road is self-service and adopted by tier-1 teams; SLOs and error-budget policy are in effect; on-call load and alert noise are measurably down; the function is operating agent-first; and product engineers are using the SRE agents to instrument and meet SLOs with far less manual effort.
- By 18 months — proactive and guarded agentic capabilities are in production; the program hits its targets — a significant reduction in Sev1/Sev2 incidents, mean-time-to-resolution cut substantially, full SLO coverage on tier-1 services, and a healthier on-call — and the company's SRE goals are being met at scale through agents that product engineers rely on.
Required Qualifications:
- 15+ years in software engineering, including 7+ years in senior engineering leadership leading SRE, platform, infrastructure, or AI organizations at scale — including managing managers.
- A proven track record of improving reliability and reducing production incidents across a large, complex, multi-team estate (thousands of engineers and/or services).
- Demonstrated experience building and operating production AI and/or agentic systems at scale — with a working command of LLMs, agent frameworks and orchestration, retrieval, evaluation (evals), guardrails, and AI/agent observability.
- Experience establishing AI governance and safety for production agents — least-privilege access, human-in-the-loop controls, blast-radius limits, and evaluation-gated autonomy.
- Deep expertise in observability (OpenTelemetry; metrics, logs, traces), SLOs and error budgets, incident management, and cloud-native infrastructure (Kubernetes, infrastructure-as-code).
- Experience running an internal platform as a product — paved roads, developer experience, and adoption measured by usage, not decree — and enabling other teams to build on it.
- A demonstrated ability to drive org-wide change through influence in a large, matrixed organization rather than through mandate.
- Excellent executive communication — able to translate reliability and AI strategy into business terms and present to C-level stakeholders.
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
Preferred Qualifications:
- A track record of building internal AI/agent developer tooling that other engineering teams adopted at scale.
- Experience in financial services or another regulated, high-availability, high-compliance environment.
- Familiarity with a modern reliability and AI stack — e.g., OpenTelemetry, Prometheus/Grafana or a major observability SaaS, PagerDuty or Slack-native incident tooling, SLO tooling, and LLM/agent observability and evaluation platforms.
- A track record establishing error-budget policy and a genuinely blameless postmortem culture.
- Advanced degree in a relevant field.
Leadership competencies:
- Technical credibility — deep enough to earn the trust of senior engineers and make hard architecture, tooling, and AI calls, while leading through others.
- AI fluency & responsible innovation — moves fast on agentic capability while insisting on evaluation, guardrails, and safety.
- Systems thinking — sees the whole sociotechnical system — tooling, incentives, culture, and cost.
- Change leadership — moves a large organization through influence, evidence, and paved roads.
- Talent magnet — attracts, grows, and retains exceptional engineering and AI leaders and specialists.
- Bias for delivery — ships outcomes on aggressive timelines and is accountable for measurable results.
- Psychological safety — builds a blameless, high-trust culture where failure is analyzed, not punished.
Privacy Statement
FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.
EEOC Statement
FIS is an equal opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, marital status, genetic information, national origin, disability, veteran status, and other protected characteristics. The EEO is the Law poster is available here supplement document available here
For positions located in the US, the following conditions apply. If you are made a conditional offer of employment, you will be required to undergo a drug test. ADA Disclaimer: In developing this job description care was taken to include all competencies needed to successfully perform in this position. However, for Americans with Disabilities Act (ADA) purposes, the essential functions of the job may or may not have been described for purposes of ADA reasonable accommodation. All reasonable accommodation requests will be reviewed and evaluated on a case-by-case basis.
Sourcing Model
Recruitment at FIS works primarily on a direct sourcing model; a relatively small portion of our hiring is through recruitment agencies. FIS does not accept resumes from recruitment agencies which are not on the preferred supplier list and is not responsible for any related fees for resumes submitted to job postings, our employees, or any other part of our company.
#pridepass
Similar roles
-
Summer 2027 Site Reliability Internship
Tradeweb London, England, United Kingdom
-
Digital Site Reliability Engineer
Radisson Hotel Group Madrid, Community of Madrid, Spain
-
Site Reliability Engineer
WorldQuant Montevideo, Montevideo, Uruguay
-
Site Reliability Engineer II
Axon Boston, Massachusetts, United States · $116K–$165K/yr
-
Staff Site Reliability Engineer
Crunchyroll, LLC Los Angeles, California, United States · $210K–$263K/yr
-
Senior Site Reliability Engineer
Ciklum Kyiv, Ukraine