Z

Site Reliability Engineering SRE Leader with AI Experience

ZENITH INFOTEK LLC · United States

IT Services and IT Consulting · 51-200 employees

8 h ago
Remote Principal (10+ yrs) Contractor United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The SRE Leader will design and implement scalable, high-performance enterprise applications while leveraging AI and Generative AI technologies to improve system reliability. They will collaborate with cross-functional teams to build automation solutions, manage CI/CD pipelines, and ensure operational excellence through rigorous monitoring and capacity planning.

What they look for

Site Reliability Engineering Java Spring Boot AWS Microservices CI/CD DevOps Generative AI Splunk Datadog Dynatrace New Relic Capacity Planning Automation Cloud Computing Agile

Requirements

Candidates must possess a bachelor's degree and over 10 years of experience in enterprise Java application development and SRE principles. Strong proficiency in AWS cloud services, observability tools, and modern DevOps practices is required, with additional experience in AI technologies considered a plus.

Full description

Job Title: Site Reliability Engineering (SRE) Leader with AI Experience

Location: (Remote) Duration: 9+ Months (Contract) Interview: Video Interview

Job Summary

We are seeking an experienced Site Reliability Engineering (SRE) Leader with strong expertise in Java application development, cloud technologies, DevOps, and AI-driven solutions. The ideal candidate will lead initiatives focused on improving application reliability, scalability, automation, and operational excellence while leveraging modern AI and Generative AI technologies. This role requires close collaboration with engineering, infrastructure, and operations teams to build highly available, resilient, and high-performing enterprise applications.

Key Responsibilities

  • Design, develop, and implement scalable, reliable, and high-performance enterprise applications.
  • Collaborate with development, infrastructure, and operations teams to improve system reliability and availability.
  • Build and enhance automation solutions to streamline deployment, monitoring, and operational processes.
  • Develop and maintain CI/CD pipelines using Jenkins, CloudBees, or similar tools.
  • Monitor application health using observability platforms such as New Relic, Datadog, Dynatrace, and Splunk.
  • Analyze system performance, identify bottlenecks, and implement performance optimization strategies.
  • Design solutions for high availability, redundancy, auto-recovery, and fault tolerance.
  • Perform capacity planning and manage application performance against SLA/SLO objectives.
  • Lead Proof of Concepts (POCs) and successfully scale them into enterprise-wide implementations.
  • Implement best practices for application lifecycle management, monitoring, and continuous improvement.
  • Work with AWS cloud technologies to deploy and manage cloud-native applications.
  • Evaluate and integrate AI and Generative AI technologies into operational workflows.
  • Mentor engineering teams on SRE principles, DevOps practices, and reliability engineering.
  • Stay current with emerging technologies and recommend innovative solutions to improve platform stability and efficiency.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.

• 10+ years of experience in enterprise Java application development and runtime support.

  • Strong experience with:
  • Java
  • Spring Boot
  • Microservices Architecture
  • Java Enterprise Frameworks
  • Deep understanding of:
  • Site Reliability Engineering (SRE)
  • Reliability Engineering concepts
  • Automation
  • High Availability
  • Auto Recovery
  • Scalability
  • Redundancy
  • Experience with AWS Cloud services (AWS Certification preferred).
  • Hands-on experience with CI/CD tools such as Jenkins or CloudBees.
  • Strong knowledge of observability and monitoring tools including:
  • New Relic
  • Datadog
  • Dynatrace
  • Splunk
  • Experience with PCF (Pivotal Cloud Foundry), App Pilot, and Presto.
  • Strong understanding of application infrastructure, runtime environments, capacity planning, and SLA/SLO management.
  • Experience designing and scaling Proof of Concepts (POCs) into enterprise-grade solutions.
  • Familiarity with Agile methodologies and DevOps practices.
  • Excellent troubleshooting, analytical, and problem-solving skills.

Preferred Qualifications

  • Experience with Artificial Intelligence (AI) technologies.
  • Hands-on knowledge of Generative AI platforms and tools such as:
  • Google Cloud AI (Vertex AI/GHCP)
  • Claude
  • AWS Bedrock
  • Experience implementing AI-enabled operational automation.
  • AWS Professional or Specialty Certifications are a plus.

Technical Skills

  • Java
  • Spring Boot
  • Microservices
  • AWS Cloud
  • Jenkins
  • CloudBees
  • Splunk
  • Datadog
  • Dynatrace
  • New Relic
  • PCF (Pivotal Cloud Foundry)
  • App Pilot
  • Presto
  • CI/CD
  • DevOps
  • Site Reliability Engineering (SRE)
  • Capacity Planning
  • SLA/SLO
  • AI
  • Generative AI
  • AWS Bedrock
  • Claude
  • Google Cloud AI

Agile

This is a remote position.