Software Finder

DevOps Engineer

Software Finder 201 District, Virginia, United States

IT Services and IT Consulting · 201-500 employees

Jul 10
devops Mid (2-5 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design and maintain scalable AWS cloud infrastructure and CI/CD pipelines to support AI and data platforms. Implement observability, security best practices, and automate infrastructure using Terraform to ensure high system availability.

What they look for

AWS Terraform CI/CD Docker Kubernetes Python Bash Linux Monitoring Observability Git ECS CloudWatch Infrastructure as Code SRE Networking

Requirements

Requires a Bachelor's degree in Computer Science or a related field and 3+ years of experience in DevOps or SRE roles. Must have strong hands-on experience with AWS, containerization, and scripting in Python or Bash.

Full description

DevOps Engineer

Software Finder is seeking a skilled DevOps Engineer to build, manage, and optimize the infrastructure supporting our AI, automation, and data platforms. This role focuses on developing scalable, secure, and highly available cloud environments while ensuring reliable deployment and operation of production workloads.

The ideal candidate will have strong experience with AWS, infrastructure automation, CI/CD pipelines, containerization, monitoring, and platform security. Working closely with AI, Data Engineering, and Product teams, this individual will help improve infrastructure reliability, deployment efficiency, and overall system performance.

Key Responsibilities

  • Design, implement, and maintain scalable CI/CD pipelines to support reliable and efficient software deployments.
  • Build, manage, and optimize AWS cloud infrastructure, including compute, networking, storage, databases, and identity management.
  • Develop and maintain infrastructure as code using Terraform or equivalent tools to support reproducible and version-controlled environments.
  • Deploy, manage, and optimize containerized applications using Docker, Amazon ECS, or Kubernetes.
  • Implement monitoring, logging, alerting, and observability solutions to proactively identify and resolve production issues.
  • Support distributed and high-throughput workloads while maintaining system availability, reliability, and performance.
  • Implement infrastructure security controls, including access management, secrets management, vulnerability remediation, and system hardening.
  • Monitor cloud usage and optimize infrastructure costs without compromising performance or reliability.
  • Collaborate with AI and Data Engineering teams to deploy machine learning models, automation workflows, and data pipelines into production.
  • Participate in incident response, root cause analysis, and post-incident reviews.
  • Identify and implement improvements to infrastructure scalability, automation, and operational processes.
  • Maintain infrastructure documentation, deployment procedures, configuration records, and operational runbooks.
  • Support secure and consistent development, testing, staging, and production environments.
  • Participate in code reviews and follow established engineering standards and Git-based workflows.
  • Contribute to additional infrastructure, platform engineering, and DevOps initiatives as required.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
  • At least three years of experience in DevOps, site reliability engineering, platform engineering, or infrastructure engineering.
  • Strong hands-on experience with AWS services, including Amazon EC2, Amazon S3, Amazon RDS, AWS Lambda, AWS IAM, Amazon VPC, and Amazon CloudWatch.
  • Proven experience developing and managing infrastructure as code using Terraform or similar technologies.
  • Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI, or comparable platforms.
  • Hands-on experience with Docker and container orchestration technologies such as Amazon ECS or Kubernetes.
  • Strong scripting skills in Python, Bash, or both.
  • Solid understanding of Linux system administration, networking, and cloud infrastructure security.
  • Experience implementing monitoring and observability solutions using CloudWatch, Datadog, Prometheus, Grafana, or similar tools.
  • Proficiency with Git-based development workflows and code review practices.
  • Experience supporting production systems with distributed or high-throughput workloads.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent communication skills and the ability to collaborate with cross-functional engineering teams.
  • Experience supporting production AI/ML infrastructure, model serving, or inference environments is preferred.
  • Familiarity with vector databases and other AI infrastructure components is preferred.
  • Experience managing distributed workloads across multiple servers or regions is preferred.
  • Knowledge of relational and NoSQL databases at scale is preferred.
  • Experience with messaging or streaming platforms such as Amazon SQS or Apache Kafka is preferred.
  • Familiarity with cloud cost optimization, FinOps, and infrastructure efficiency initiatives is preferred.
  • Relevant AWS certifications, such as AWS Certified DevOps Engineer, AWS Certified SysOps Administrator, or AWS Certified Solutions Architect, are preferred.
  • Experience working in Agile or DevOps-driven software development environments is preferred.
  • Strong interest in automation, infrastructure reliability, and continuous improvement.

Similar roles