Senior Site Reliability Engineer
Ciklum Kyiv, Ukraine
IT Services and IT Consulting · 1,001-5,000 employees
About the role
You will design and automate scalable infrastructure using infrastructure-as-code while implementing monitoring and observability systems to meet SLIs and SLOs. Additionally, you will lead on-call incident responses and collaborate with development teams to optimize performance and capacity planning.
What they look for
Requirements
Candidates must have extensive experience with containerization, orchestration, and managing Linux systems in a cloud environment. Proficiency in Shell scripting or high-level programming languages like Python or Go is required, along with strong problem-solving skills.
Benefits
Full description
Ciklum is looking for a Senior Site Reliability Engineer to join our team full-time in Ukraine.
We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.
About the role:
As a Senior Site Reliability Engineer, become a part of the R&D team. In this role, you'll be a key contributor to our high-performance SaaS Cloud Platform.
As a Senior Site Reliability Engineer, you will design and automate scalable infrastructure with infrastructure-as-code, implement and tune monitoring and observability systems to meet defined SLIs and SLOs, lead on-call incident response and post-mortem reviews to continuously improve system resilience, collaborate with development teams on performance optimization and capacity planning, and drive the automation of routine operational tasks to minimize toil and maximize uptime.
Join our innovative, high-performing team, where the convergence of reliable cloud infrastructure and advanced data processes drives our success in a fast-paced, agile environment.
Responsibilities:
- You’ll develop, improve, and maintain Guardicore's Cyber Security SaaS cloud platform
- Lead problem-solving efforts for the entire technology stack in collaboration with other teams in the R&D
- Establish scalable, efficient, automated processes for large-scale data analyses
- Work closely with other R&D to develop a strategy for long-term data platform architecture
- Practice infrastructure as a code (IaC) and GitOps using technologies like Terraform and ArgoCD
- Collaborate with Guardicore's development and research groups to constantly improve our platform and infrastructure
- Participate in the on-call rotation supporting the applications and infrastructure
- Developed and evolved our tooling, logging, monitoring, and alerting mechanisms to increase observability and transparency
Requirements:
- Extensive experience with containerization and orchestration technologies (e.g., Docker, Kubernetes)
- Excellent problem-solving skills and ability to think critically about complex technical challenges and optimizing production systems
- Experience in managing and troubleshooting Linux systems
- experience with observability systems such as Datadog/Splunk/New Relic/Grafana, or similar
- Experience in Shell scripting and/or high-level Programming like Python and Go
- Experience working with cloud environments like GCP, Linode, AWS, and Azure
- Experience in a SaaS environment managing large-scale data sets - Advantage
- Excellent verbal and written English communication and presentation skills
What’s in it for you?
- Strong community: Work alongside top professionals in a friendly, open-door environment
- Growth focus: Take on large-scale projects with a global impact and expand your expertise
- Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications
- Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies
- Flexibility: Enjoy radical flexibility – work remotely or from an office, your choice
- Care: We’ve got you covered with company-paid medical insurance, mental health support, and financial & legal consultations
About us:
At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress.
As one of Ukraine’s largest IT companies and a top employer recognized by Forbes, we’ve spent over 20 years delivering meaningful tech solutions. We proudly support diverse talent and military veterans, recognizing their unique skills and perspectives they bring to shaping the future.
Explore, empower, engineer with Ciklum!
Interested already? We would love to get to know you! Submit your application. We can’t wait to see you at Ciklum.
#LI-NV1
Similar roles
-
Summer 2027 Site Reliability Internship
Tradeweb London, England, United Kingdom
-
Digital Site Reliability Engineer
Radisson Hotel Group Madrid, Community of Madrid, Spain
-
Site Reliability Engineer
WorldQuant Montevideo, Montevideo, Uruguay
-
Site Reliability Engineer II
Axon Boston, Massachusetts, United States · $116K–$165K/yr
-
Staff Site Reliability Engineer
Crunchyroll, LLC Los Angeles, California, United States · $210K–$263K/yr
-
Senior Site Reliability Engineer
Elevate Government Solutions Washington, District of Columbia, United States