Spinomenal

SRE

Spinomenal Bnei Brak, Tel-Aviv District, Israel

Software Development · 51-200 employees

8 h ago
sre Mid (2-5 yrs) Full-time Israel
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The SRE will own release engineering, platform stability, and incident response within a high-throughput iGaming environment. They are responsible for automating operational tasks, maintaining CI/CD pipelines, and serving as an escalation point for production failures.

What they look for

Site Reliability Engineering Jenkins AWS Terraform Python Bash New Relic Prometheus Grafana SQL Ansible Helm Microservices CI/CD Infrastructure-as-Code Observability

Requirements

Candidates must have at least 2 years of experience in SRE or production operations with high-transaction web applications. Proficiency in AWS, Terraform, Jenkins, and observability tools like New Relic or Grafana is required.

Full description

Spinomenal is a dynamic force in the online casino sector, bursting onto the scene in 2014. Renowned as one of the fastest-growing content creators in the iGaming industry, Spinomenal thrives on fostering creativity and collaboration. By nurturing an environment where employees excel through teamwork and communication, Spinomenal maintains a rapid pace of innovation and development. Our dedication to collective effort and shared vision has propelled us to deliver captivating, cutting-edge gaming experiences to players worldwide.

About the position

Spinomenal is seeking a hands-on Production Manager / SRE to own release engineering, platform stability, and incident response in our high-throughput, low-latency iGaming environment. Serving as the gatekeeper to production across R&D, DevOps, QA, and Support, you will ensure high availability, secure CI/CD pipelines, and rapid anomaly diagnosis through automation.

About the team

null

Responsibilities

  • Validate complex configurations, execute automated health checks, and own recovery and rollback runbooks for production deployments.
  • Identify manual operational tasks (toil) and automate them using scripting and Infrastructure-as-Code principles to increase reliability.
  • Maintain and optimize distributed tracing, dashboards, alerting thresholds, and log aggregation using New Relic, Grafana, and advanced SQL.
  • Serve as a critical escalation point for complex, cross-layer production failures (Application, Network, Database, and Infrastructure).
  • Drive deep-dive post-mortems and implement structural fixes to prevent incident recurrence.
  • Govern infrastructure and configuration changes across environments in tight collaboration with DevOps to maintain environment parity and prevent drift.
  • Partner with QA, Architects, DevOps, and Game Producers to champion SRE best practices, operational readiness standards, and resilient architecture.

Requirements

  • At least 2 years of experience in a dedicated Site Reliability Engineering (SRE), Production Operations, or Release Engineering role handling high-transaction web applications.
  • Deep experience engineering and maintaining Jenkins pipelines (Pipeline-as-Code / Jenkinsfiles) and familiarity with configuration tools (e.g., Ansible, Helm).
  • Strong hands-on experience managing and troubleshooting cloud infrastructure (AWS: EC2, ECS, S3, Lambda, IAM policies, VPC routing) and Infrastructure as Code (IaC) using Terraform.
  • Proficiency in Python or Bash for writing automation scripts, system utilities, and internal tooling.
  • Advanced capability with observability stacks (New Relic, Prometheus/Grafana) and strong SQL skills for log parsing and database debugging.
  • Mastery of web debugging (interpreting JSON payloads, analyzing API contracts, diagnosing HTTP status anomalies) and understanding of DNS, CDN/Redis caching, and proxies.
  • Demonstrated capability in debugging distributed microservices architectures and defining Service Level Indicators/Objectives (SLIs/SLOs) and error budgets.
  • Exceptional technical communication skills with a proactive mindset and the ability to stay calm under pressure during critical outages and on-call rotations.

Advantages

null

Similar roles