Cloud & AI Platform Architect
LSports · Ashkelon, South District, Israel
IT Services and IT Consulting · 201-500 employees
About the role
You will own the architecture of the cloud platform, including GKE fleet design, CI/CD pipelines, and reliability standards. Additionally, you will design and govern the agentic AI ecosystem, focusing on orchestration, evaluation, and cost-aware production deployment.
What they look for
Requirements
The role requires over 8 years of experience in platform engineering or systems architecture with deep expertise in GCP, Kubernetes, and Infrastructure-as-Code. Candidates must also demonstrate 1-2 years of hands-on experience building production-grade LLM or agentic AI systems.
Benefits
Full description
LSports is a world-leading sports data provider, trusted by sportsbooks worldwide to deliver real-time data with unmatched accuracy and reliability. With technology that drives smarter trading and deeper engagement, we empower bookmakers to grow, innovate, and stay ahead of the game.
If you’re passionate about sports and technology and want to make your mark in a fast-moving industry, don't miss this chance! Step onto the field with us and help build the future of sports data – We are looking for a talented Cloud & AI Platform Architect.
About the Role
LSports runs a large-scale, real-time sports data platform on GCP (GKE fleets, Kafka streaming, GitOps delivery) and is building an agentic AI ecosystem on top of it: AI systems that plan, call tools, and act on real infrastructure, not just chat.
We are looking for one architect to own both layers. Your foundation is deep platform and DevOps architecture: reliability, scalability, security, and cost efficiency of the cloud platform every R&D team builds on. Your growth edge is the AI layer: designing how agents are orchestrated, how they access tools and data safely, how we evaluate whether they work, and how we govern their cost and behavior in production.
We know agentic AI is a young discipline. We are not looking for a decade of AI experience that does not exist. We are looking for a proven platform architect who has spent the last one to two years genuinely building LLM-based and agentic systems in production, and who wants to own where this field goes inside a company that takes it seriously.
What You’ll Do:
Platform & DevOps
- Cloud Architecture: Own the architecture of LSports' cloud platform, including GKE fleet design, VPC networking, IAM/Workload Identity, and multi-environment strategies across GCP (and AWS where relevant).
- IaC & GitOps: Lead Infrastructure-as-Code (Pulumi, Terraform) standards, ArgoCD deployment patterns, and secure CI/CD paved paths (GitHub Actions, GitLab CI) adopted by all R&D teams.
- Reliability & Observability: Own reliability architecture, including Disaster Recovery (DR) strategy and drills, P1 incident reduction, MTTR improvement, and observability platform architecture (Datadog, Prometheus, Grafana, OpenTelemetry) including usage and cost optimization.
- FinOps: Drive FinOps as an architectural discipline—rightsizing, idle-resource elimination, unallocated-spend attribution, and cost-aware design reviews using tools like Kubecost or GCP Cost Management.
Agentic AI
- Agent Architecture: Design LSports' agentic AI ecosystem, focusing on agent orchestration frameworks (e.g., LangChain, LangGraph, CrewAI, AutoGen), tool/MCP (Model Context Protocol) interfaces, vector databases/memory strategies (e.g., Pinecone, Qdrant, pgvector), and production AI deployment patterns.
- Evaluation & Guardrails: Build the evaluation and guardrail layer for AI in production using frameworks like Ragas, TruLens, or LangSmith. Define permission models, human-in-the-loop checkpoints, audit logging, and policy-as-code for AI usage.
- AI FinOps & Monitoring: Own AI cost observability, token spend monitoring, model routing strategies (optimizing for cost/quality/latency trade-offs), and real-time alerting on runaway model usage.
- Reference Architectures & Standards: Define reference architectures for teams building AI features, RAG patterns, prompt versioning/management, and model API standards (OpenAI API, Anthropic, open-source models via vLLM/Ollama), driving high-leverage AI automation across R&D.
Requirements
- Experience: 8+ years in platform/infrastructure engineering, DevOps, or systems architecture, with proven ownership of enterprise-scale platform decisions.
- Containers & Orchestration: Deep production expertise with Kubernetes (GKE strongly preferred), container runtime, and Service Mesh technologies.
- IaC & GitOps: Hands-on mastery of Infrastructure as Code using Terraform or Pulumi, along with GitOps practices via ArgoCD.
- Cloud Platform Depth: Strong GCP architecture depth—networking, IAM, Workload Identity, and cloud cost management.
- Production LLM / Agentic Systems: 1–2+ years of hands-on experience building LLM-based or agentic systems that have reached production—agent frameworks, tool/function calling, and shipped Retrieval-Augmented Generation (RAG) systems (not just tutorials or prototypes).
- Reliability Track Record: A proven track record of measurable reliability and cost outcomes (e.g., uptime improvements, MTTR reduction, and cloud spend optimization).
- Streaming Data: Solid experience with Apache Kafka or comparable streaming platforms at scale.
- Soft Skills: Ability to lead through influence, mentor engineering teams, and communicate complex architectural trade-offs clearly to both engineers and executives.
Bonus Points if you have:
- Experience with Model Context Protocol (MCP), multi-agent orchestration frameworks, or automated AI evaluation toolkits.
- FinOps certification (e.g., Certified FinOps Practitioner) or demonstrated large-scale cloud cost reduction wins.
- Experience with managed AI platforms like GCP Vertex AI or AWS Bedrock in production environments, as well as supporting GPU/inference workloads (e.g., Triton Inference Server, Ray Serve).
- Deep AWS architecture experience alongside GCP.
Why This Role
This is a rare seat: full architectural ownership of both the platform layer and the AI layer of a real-time data company, reporting directly to the DevEx Director. You will not inherit someone else's AI strategy — you will define it, on infrastructure you also own. If you are a platform architect who has been building with LLMs and wants that work to be your mandate rather than your side project, this is that role.
Life at LSports
We’re honest with each other and open to new ideas. We challenge ourselves to excel, knowing that good relationships never come at the expense of professionalism. We’re dynamic, always growing and improving to stay on top.
Come join a team that’s all about winning together, with integrity and heart.
All the Extras You Deserve
- Hybrid work model
- Work From Anywhere – up to 30 days a year
- Extended parental leave for both mothers and fathers
- Quarterly recharge days – company-wide shutdown to fully unplug
- Health insurance
- Birthday day off
And yes — we also care about the little things:
- Cibus
- Holiday eves off
- Team & company social events
- Dog-friendly office
- Kids summer camp
- Access to the Group Hug wellbeing platform