Infrastructure Site Reliability Engineer (w/ active Secret)
Critical Solutions Norfolk, Virginia, United States · $87K–$112K/yr
IT Services and IT Consulting · 11-50 employees
About the role
The SRE will develop and execute automated testing frameworks to ensure system resilience, performance, and reliability under various load conditions. They are responsible for managing infrastructure, coordinating production releases, and performing system maintenance across Windows, storage, and network environments.
What they look for
Requirements
Candidates must possess a U.S. citizenship, an active DoD Secret security clearance, and a minimum of a Bachelor's degree with 4+ years of experience. A DoD 8570.01 IAT Level II certification is also required to support classified environments.
Benefits
Full description
Infrastructure Site Reliability Engineer (w/ active Secret)
Location: Norfolk, VA Clearance: Secret Type: Full-time, Onsite Salary Range: $87,000 - $112,000
JOB DESCRIPTION
Critical Solutions is seeking a Infrastructure Site Reliability Engineer to support our federal customer.
The Site Reliability Engineers (SREs) will develop and execute tests focused on system resilience, performance under load, and failure scenarios. They will work in tandem with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions that validate system behavior under normal and stress conditions, ensuring our services are resilient and meet established service level objectives (SLOs). This work will contribute to the development of robust and scalable services that operate reliably in production.
Your responsibilities will include maintaining complex computer systems by writing code to automate software releases, monitor systems, and detect and fix problems before users even know there is an issue. You will use these skills to improve site performance and overall reliability.
The SRE Engineer role is responsible for supporting, migrating, automation and optimization of software development and deployment process, infrastructure as code, and contribute to the overall maturity of the Site Reliability Engineering program.
PRIMARY ROLES AND RESPONSIBILITIES:
- Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform.
- Coordinate and execute production release activities during designated maintenance windows, including scheduled off-hours.
- Windows Server & Endpoints: Administer patch distribution, baseline compliance, and reporting using Microsoft Endpoint Configuration Manager (MECM/SCCM).
- Storage Infrastructure: Coordinate, stage, and apply firmware, disk qualification packages, and system upgrades across NetApp ONTAP environments using Active IQ Unified Manager and OnCommand Insight (OCI).
- Network Infrastructure: Plan and execute non-disruptive software maintenance releases, SMUs, and image upgrades across Cisco NX-OS data center switches.
BASIC QUALIFICATIONS:
- Must be a U.S. citizen and possess an active DoD Secret Security Clearance or higher.
- B.S Degree and 4+ years of hands-on experience in enterprise systems administration and patch management across mixed infrastructure environments.
- Currently possess and ability to maintain an active DoD Secret security clearance
- Minimum of DoD 8570.01 IAT Level II Certification
- Must be able to support program execution in classified environments and access SIPRNet from an Agency location on short notice (local travel).
- Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred.
- Solid understanding of Linux/Unix and command-line tools.
- Proven experience executing centralized patch deployments using MECM / SCCM.
- Direct experience managing and updating NetApp ONTAP storage systems via Active IQ Unified Manager or OnCommand Insight.
- Demonstrated knowledge of network operating system updates and maintenance on Cisco NX-OS platforms.
- Experience in application administration, configuration, and integration.
- Familiarity with Agile development methodologies.
- Capable of working and communicating within a distributed team.
- Ability to work in a highly collaborative, forward-thinking, and innovation-driven environment
- Knowledge of Agile and DevSecOps/SRE concepts and best practices, with a desire to grow that knowledge
- Hand-on experience with Atlassian products (Jira, Confluence, Bitbucket, etc.).
- Experience creating JIRA and/or Azure DevOps workflows, projects, custom configurations
- Experience automating tasks using scripting languages such as Python or PowerShell.
- Working knowledge of the Risk Management Framework (RMF), DISA STIGs
- Working knowledge of formal IT Change Management workflows and ticketing systems (specifically HP Service Manager or similar enterprise ITSM platforms).
- Availability to support scheduled off-hours maintenance windows and production releases.
CERTIFICATION REQUIREMENT:
- Must possess one of the DoD 8570.01 IAT Level II Certification
PREFERRED QUALIFICATIONS:
- ITILv4, Scrum Master, or Agile SAFe certification(s) or applicable experience.
- Experience implementing scripts using PowerShell or Ansible Playbooks for patch rollout, system health checks, or reporting.
- Vendor-specific technical certifications (e.g., NetApp NCDA, Cisco CCNA, Microsoft Certified: Windows Server Hybrid Administrator).
- Prior experience supporting defense-oriented enterprise networks or environments with strict SLA and compliance mandates.
LOCATION:
- On-site 100%. Norfolk, VA
- Other acceptable work location: Jacksonville, FL, San Diego, CA and Honolulu, HI
Key Metrics of Success for the Team:
- Strict adherence to a 7-day vulnerability-to-production patching cadence across all assigned enterprise infrastructure assets.
- Comprehensive and regularly updated automated test coverage for all critical systems and infrastructure components.
ADDITIONAL INFORMATION:
CLEARANCE REQUIREMENT: Must possess an active DoD Secret Clearance. In addition, selected candidate must undergo background investigation (BI) and finger printing by the federal agency and successfully pass the preceding to qualify for the position. US CITIZENSHIP IS REQUIRED.
CRITICAL SOLUTIONS PAY AND BENEFITS:
Salary range $87,00 - $112,000. The salary range for this position represent the typical salary range for this job level and this does not guarantee a specific salary. Compensation is based upon multiple factors such as responsibilities of the job, education, experience, knowledge, skills, certifications, and other requirements.
BENEFIT SNAPSHOT: 100% premium coverage for Medical, Dental, Vision, and Life Insurance, Supplemental Insurance, 401K matching, Flexible Time Off (PTO/Holidays), Higher Education/Training Reimbursement, and more.
Similar roles
-
Senior Site Reliability Engineer (SRE)
Sequoia Connect Ciudad de México, Mexico
-
Summer 2027 SRE - Observability Engineering Internship
Tradeweb Jersey City, New Jersey, United States · $73K–$156K/yr
-
Senior Site Reliability Engineering- CTJ- Secret (Cleared Environments)
Microsoft Redmond, Washington, United States · $120K–$261K/yr
-
Senior Site Reliability Engineer
Aptean Alpharetta, Georgia, United States
-
Site Reliability Engineer
Cosm El Segundo, California, United States · $110K–$145K/yr
-
Senior Business Analyst, Site Reliability Engineering
CFA Institute Charlottesville, Virginia, United States · $95K–$130K/yr