VetJobs

Site Reliability Engineer II - Tampa, FL

VetJobs Tampa, Florida, United States

Armed Forces · 51-200 employees

Yesterday
sre Mid (2-5 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Site Reliability Engineer will deliver end-to-end application and infrastructure service delivery to ensure reliable business operations. They will also partner with cross-functional teams to define SLOs/SLIs and manage production incidents using observability tools and AI-assisted triage.

What they look for

Site Reliability Engineering Java Spring Boot Apache Kafka Dynatrace Splunk Kubernetes Docker Jenkins Python SQL Oracle CI/CD Microservices Observability Incident Management

Requirements

Candidates must have at least 2 years of applied experience in software engineering and formal training in relevant concepts. Proficiency in large-scale infrastructure, microservices, CI/CD pipelines, and observability tools like Dynatrace and Splunk is required.

Full description

Job Description

ATTENTION MILITARY AFFILIATED JOB SEEKERS - Our organization works with partner companies to source qualified talent for their open roles. The following position is available to Veterans, Transitioning Military, National Guard and Reserve Members, Military Spouses, Wounded Warriors, and their Caregivers. If you have the required skill set, education requirements, and experience, please click the submit button and follow the next steps. All positions are onsite, unless otherwise stated.

Job Description: Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.

As a Site Reliability Engineer II at JPMorgan Chase within the within the Corporate and Investment Bank, Payments Technology Team, you will use technology to solve business problems and leverage software engineering best practices as we strive towards excellence. This role often works independently to execute small to medium projects, but you’ll also have the opportunity to collaborate with cross functional teams to continually improve your level of knowledge about JPMorgan Chase’s business and relevant technologies.

Job Responsibilities:

  • Deliver end-to-end application and/or infrastructure service delivery to enable reliable business operations across the firm.
  • Partner with cross-functional teams to define and maintain SLOs/SLIs and error budgets for key production services, proactively resolving issues before customer impact.
  • Uses enterprise-authorized AI capabilities within the work environment to speed up incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Own day-to-day operational processes including incident management, problem management (RCA), and change/event management (monitoring and alerting).
  • Monitor production environments for anomalies using standard observability tools; build and maintain alerting mechanisms to detect incidents early and minimize downtime.
  • Design and implement Dynatrace instrumentation to monitor performance, infrastructure health, and user experience; analyze trends, escalate/communicate as needed, and provide solutions to business and technology stakeholders.
  • Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes.

Required Experience

Required Qualifications, Capabilities, and Skills:• Formal training or certification on software engineering concepts and 2+ years applied experience ( NAMR/APAC – India/ LATAM/ Hong Kong)

  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows (e.g., troubleshooting support and runbook drafting) with strong validation habits and awareness of data sensitivity.
  • Ability to assess AI-assisted operational recommendations for correctness and risk, and apply appropriate controls to maintain resiliency, security, and auditability.
  • Large-scale application/infrastructure experience across on‑prem and public cloud environments, including strong understanding of core networking (DNS, TCP/IP, VPN, load balancing).
  • Hands-on observability expertise: production monitoring, distributed tracing, log analysis, and instrumentation using tools like Dynatrace, Geneos, and Splunk plus application logging frameworks.
  • Strong microservices engineering with Java and Spring Boot, including event-driven patterns using Apache Kafka.
  • Solid data layer skills: Oracle databases, GemFire/Geode distributed caching, Liquibase migrations, and strong SQL/database management experience.
  • DevOps and delivery proficiency: Jenkins-based CI/CD, Docker, Kubernetes, and modern automation practices for build, test, deploy, and release.
  • Scripting capability for automation and ops tasks across Python, Bash/PowerShell, plus Node.js scripting experience.

Preferred Experience

Preferred Qualifications, Capabilities, and Skills:• Experience with one or more general-purpose programming languages and/or automation scripting.

  • Experience working in banking and payment domains such as: Client Liquidity, High and low value payment and SWIFT.
  • Collaborate well and possess the ability to foster relationships effectively with diverse groups across geographies.
  • Ability to multi-task, handling complex requirements and is adaptable.

Similar roles