Site Reliability Engineer II
Kalmbach Feeds Inc Upper Sandusky, Ohio, United States
Animal Feed Manufacturing · 1,001-5,000 employees
About the role
Operate and troubleshoot production and non-production Kubernetes clusters across on-premises data centers and Azure, including infrastructure, networking, storage, observability, and capacity. Participate in on-call incident response and disaster recovery, and reduce operational toil through automation, infrastructure as code, safer CI/CD, and reliability practices.
What they look for
Requirements
Candidates need at least three years of experience in SRE, platform, systems engineering, or DevOps, including hands-on production Kubernetes operations and data-center or equivalent infrastructure experience. Strong Linux, networking, monitoring, incident response, scripting, Git, and familiarity with CI/CD, GitOps, or infrastructure-as-code practices are required.
Full description
Site Reliability Engineer II
Hybrid Kubernetes & On-Premises Infrastructure | Hybrid role
Join a small SRE team at the start of an exciting build: we are shaping the next generation of our hybrid platform across company data centers and Microsoft Azure—not simply maintaining someone else’s established setup. You’ll have direct ownership and a real voice in the architecture, standards, and technologies we put into production. The platform will include production RKE2 clusters managed with Rancher, Azure Kubernetes Service (AKS), hybrid workloads, and disaster recovery, while storage, networking, and data-platform designs are still open to influence. This hands-on role spans racks to cloud, and what you help design will become what the company runs. You need not know every tool on day one, but should learn quickly and bring a practical, curious approach.
What You’ll Do
- Operate production and non-production RKE2/Rancher Kubernetes clusters, including upgrades, node lifecycle, networking, ingress, DNS, certificates, and capacity; troubleshoot control-plane, scheduling, CoreDNS, and memory (OOM) issues, and improve observability, alerting, and runbooks.
- Run Kubernetes on physical HPE servers and virtual machines in company data centers, including hardware, firmware, RAID, and out-of-band management; partner with infrastructure teams on Cisco networking, SAN/NVMe/object storage, failure domains, and capacity planning.
- Support AKS, Azure Container Registry (ACR), and connectivity between data centers and Azure. Implement, test, and document disaster-recovery plans, and verify that backups can be restored.
- Join the on-call rotation; investigate and respond to incidents, escalating to the Staff SRE when appropriate. Reduce recurring toil through automation, infrastructure as code (IaC), safer CI/CD, SLIs/SLOs, and attention to single points of failure.
What We’re Looking For
- 3+ years in SRE, platform, systems engineering, or DevOps, with hands-on experience operating and troubleshooting Kubernetes in production.
- Hands-on data-center, colocation, or equivalent experience with servers, virtualization, storage, and networking.
- Strong Linux fundamentals and working knowledge of TCP/IP, DNS, and TLS.
- Experience with monitoring, logging, alerting, and incident response; clear communication and a calm, methodical approach during outages.
- Scripting experience in Bash, Python, or Go; Git experience; and familiarity with CI/CD, GitOps, or IaC practices using any toolset.
Helpful but not required: RKE2/Rancher; AKS in a hybrid environment; Kyverno/OPA; 25/100GbE or Cisco Nexus; SAN/NVMe/object storage; stateful data platforms on Kubernetes; GPU/AI workloads; GitOps pipelines; Terraform or Ansible.
Why Join Us
Take meaningful ownership of a platform being built for its next chapter. You’ll help shape it from the ground up, influence foundational decisions, and see your work become the systems the company relies on—from physical servers to cloud. If you want to build, improve, and own real infrastructure rather than inherit a ticket queue, this is your opportunity.
Similar roles
-
ME00680-Site Reliability Engineer 3
Momentum Engineering, Inc. Annapolis Junction, Maryland, United States · $165K–$230K/yr
-
SRE
EX Squared Costa Rica
-
Site Reliability Engineer II
Abbott Los Angeles, California, United States · $82K–$141K/yr
-
Sr Mgr, Site Reliability Engineer (SRE)
The Walt Disney Company Orlando, Florida, United States · $175K–$215K/yr
-
Site Reliability Engineer
The Descartes Brazil
-
Senior Kafka SRE Engineer
Charles Schwab Inc. Austin, Texas, United States · $120K–$155K/yr