About the role
The SRE will own release engineering, platform stability, and incident response within a high-throughput iGaming environment. They are responsible for automating operational tasks, maintaining CI/CD pipelines, and serving as an escalation point for production failures.
What they look for
Requirements
Candidates must have at least 2 years of experience in SRE or production operations with high-transaction web applications. Proficiency in AWS, Terraform, Jenkins, and observability tools like New Relic or Grafana is required.
Full description
Spinomenal is a dynamic force in the online casino sector, bursting onto the scene in 2014. Renowned as one of the fastest-growing content creators in the iGaming industry, Spinomenal thrives on fostering creativity and collaboration. By nurturing an environment where employees excel through teamwork and communication, Spinomenal maintains a rapid pace of innovation and development. Our dedication to collective effort and shared vision has propelled us to deliver captivating, cutting-edge gaming experiences to players worldwide.
About the position
Spinomenal is seeking a hands-on Production Manager / SRE to own release engineering, platform stability, and incident response in our high-throughput, low-latency iGaming environment. Serving as the gatekeeper to production across R&D, DevOps, QA, and Support, you will ensure high availability, secure CI/CD pipelines, and rapid anomaly diagnosis through automation.
About the team
null
Responsibilities
- Validate complex configurations, execute automated health checks, and own recovery and rollback runbooks for production deployments.
- Identify manual operational tasks (toil) and automate them using scripting and Infrastructure-as-Code principles to increase reliability.
- Maintain and optimize distributed tracing, dashboards, alerting thresholds, and log aggregation using New Relic, Grafana, and advanced SQL.
- Serve as a critical escalation point for complex, cross-layer production failures (Application, Network, Database, and Infrastructure).
- Drive deep-dive post-mortems and implement structural fixes to prevent incident recurrence.
- Govern infrastructure and configuration changes across environments in tight collaboration with DevOps to maintain environment parity and prevent drift.
- Partner with QA, Architects, DevOps, and Game Producers to champion SRE best practices, operational readiness standards, and resilient architecture.
Requirements
- At least 2 years of experience in a dedicated Site Reliability Engineering (SRE), Production Operations, or Release Engineering role handling high-transaction web applications.
- Deep experience engineering and maintaining Jenkins pipelines (Pipeline-as-Code / Jenkinsfiles) and familiarity with configuration tools (e.g., Ansible, Helm).
- Strong hands-on experience managing and troubleshooting cloud infrastructure (AWS: EC2, ECS, S3, Lambda, IAM policies, VPC routing) and Infrastructure as Code (IaC) using Terraform.
- Proficiency in Python or Bash for writing automation scripts, system utilities, and internal tooling.
- Advanced capability with observability stacks (New Relic, Prometheus/Grafana) and strong SQL skills for log parsing and database debugging.
- Mastery of web debugging (interpreting JSON payloads, analyzing API contracts, diagnosing HTTP status anomalies) and understanding of DNS, CDN/Redis caching, and proxies.
- Demonstrated capability in debugging distributed microservices architectures and defining Service Level Indicators/Objectives (SLIs/SLOs) and error budgets.
- Exceptional technical communication skills with a proactive mindset and the ability to stay calm under pressure during critical outages and on-call rotations.
Advantages
null
Similar roles
-
Staff Site Reliability Engineer
KEV Group Toronto, Ontario, Canada · $150K–$180K/yr
-
Site Reliability Engineer - Data Platform
IMC Amsterdam, North Holland, Netherlands
-
Staff Software Engineer, Site Reliability Engineering, Vertex AI
Google Warsaw, Masovian Voivodeship, Poland · PLN 480K–PLN 492K/yr
-
Site Reliability Engineering Manager
Conifers.ai Tel-Aviv, Tel-Aviv District, Israel
-
Senior Site Reliability Engineer
Precisely International Jobs Bielsko-Biała, Silesian Voivodeship, Poland
-
Network SRE
JPMorgan Chase & Co. Buenos Aires, Argentina