Site Reliability Engineer - Senior
SOFTSWISS Kato Polemidia Municipality, Cyprus, Cyprus
Software Development · 1,001-5,000 employees
About the role
You will be responsible for investigating and resolving production issues while tuning system performance across high-load services. The role involves maintaining Kubernetes environments, improving observability, and developing internal tooling to reduce operational overhead.
What they look for
Requirements
Candidates must have solid experience with Go, strong debugging skills, and hands-on experience with Kubernetes and PostgreSQL. A deep understanding of distributed systems and experience with messaging systems like Kafka or RabbitMQ is required.
Benefits
Full description
Overview:
SOFTSWISS is looking for an experienced Site Reliability Engineer to join our Sportsbook team and work on a high-load, production-critical platform.
About Product:
SOFTSWISS Sportsbook
Sportsbook is a software platform designed for launching and managing a successful online sports betting business. The Sportsbook team’s goal is to develop truly innovative software that will make the process of creating and running a bookmaker website faster and easier for our clients. We are searching for driven individuals who are passionate about what they do and are ready to work on engaging tasks and exciting projects together.
Purpose of the role:
You will work on a high-load platform in a production-focused role that combines backend engineering and operational responsibility. The work is hands-on: debugging real issues, stabilizing systems and improving reliability rather than building features in isolation.
Key responsibilities:
- Investigate and resolve production issues across multiple services and integrations
- Perform deep debugging (logs, traces, metrics, DB queries) to identify root causes
- Tune system performance (PostgreSQL queries, caching, service behavior under load)
- Work directly with Kubernetes environments: deployments, configs, scaling, troubleshooting
- Improve observability: metrics, logging, tracing (Grafana, ELK stack)
- Support and maintain integrations between Sportsbook services and external client platforms
- Set up and configure new projects/environments (mainly Sportsbook-related)
- Write internal tooling and automation in Go to reduce manual work and operational overhead
- Work with messaging systems (Kafka, RabbitMQ) and diagnose related issues
- Collaborate with developers, QA, and DevOps to resolve incidents and improve system stability
Required Experience:
- Solid experience with Go (comfortable reading and writing production code)
- Strong debugging skills: ability to trace issues across services, logs, and data layers
- Experience working with PostgreSQL (query analysis, indexing, performance tuning)
- Hands-on experience with Kubernetes (kubectl, deployments, configs, troubleshooting)
- Familiarity with observability tools (Grafana, Kibana/ELK, logs/metrics/tracing)
- Experience with Redis (caching patterns, debugging)
- Experience with Kafka and/or RabbitMQ (consumer behavior, lag, retries, failures)
- Understanding of how distributed systems behave under load (timeouts, retries, race conditions)
- Comfortable working with production systems and handling incidents
- Ability to work independently: take an issue, investigate, and drive it to resolution
Nice to have:
- Experience with high-load or real-time systems
- Experience with Sportsbook or betting platforms
- Experience building internal tools or automation for engineering teams
Our Benefits:
- Private health insurance
- Sports benefits
- Comprehensive Mental Health Program
- Free English lessons (online)
- Local language courses
- Paid time off
- Maternity leave support
- Referral program rewards
- Upskilling, internal workshops, and participation in professional conferences and corporate events
Similar roles
-
Staff Site Reliability Engineer
NinjaTrader Chicago, Illinois, United States · $160K–$210K/yr
-
Senior Site Reliability Engineer – Network Observability
Blueprint Technologies $104K–$114K/yr
-
Senior Site Reliability Engineer (Cloud Platform)
Salve.Inno Consulting Denver, Colorado, United States
-
Amazon Connect SRE
Miratech Surat, Gujarat, India
-
AWS Site Reliability Engineer
Miratech Ahmedabad, Gujarat, India
-
Senior Data SRE / Cloud Platform Engineer
Flywire Valencia, Valencian Community, Spain · €49K–€62K/yr