Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore
Palo Alto Networks Bengaluru, Karnataka, India
Computer and Network Security · 10,001+ employees
About the role
Drive reliability, uptime, and incident prevention across cloud services and high-volume threat inspection pipelines. Partner with developers to optimize infrastructure, automate operations, and ensure world-class availability through on-call rotations.
What they look for
Requirements
Requires 2-4 years of experience in SRE, DevOps, or Systems Engineering with proficiency in Linux, cloud platforms, and infrastructure automation. Candidates must have hands-on experience with observability tools, container orchestration, and scripting languages like Python or Bash.
Full description
Our Mission
At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.
Job Summary
The Team
Palo Alto Networks’ Cloud-Delivered Security Services (CDSS) is the intelligence engine of our Next-Generation Security platform. We provide a suite of AI-driven, subscription-based services—including Advanced Wildfire, Advanced Threat Prevention, DNS Security, URL Filtering, and IoT Security—integrated natively into our firewalls. Our infrastructure processes trillions of events daily, delivering real-time protection to over 85,000 global enterprises. Working in CDSS means building the backbone of global cybersecurity at a scale few companies in the world ever reach.
Job Summary
We are looking for an ambitious, technically sharp Staff Site Reliability Engineer to join our core WildFire platform team. In this role, you will take technical ownership of system resilience, deployment automation, and operational performance across WildFire public cloud infrastructure (AWS, GCP, Azure, OCI) and private cloud/appliance environments. You will work closely with feature developers to establish SLOs, optimize high-throughput threat inspection pipelines, automate platform operations, and participate in production on-call rotations to ensure sub-second malware detection telemetry and world-class availability.
Key Responsibilities
- Service Reliability & Ownership: Drive reliability, uptime, and incident prevention across WildFire’s cloud services, microservices, and high-volume malware analysis pipelines.
- Infrastructure Automation & IaC: Provision, configure, and maintain production cloud and hybrid environments using Infrastructure-as-Code (Terraform, Ansible) and GitOps (ArgoCD).
- SLO & Observability Execution: Implement and monitor SLIs/SLOs/SLAs across inspection pipelines. Build and manage observability dashboards and alerting systems using Prometheus, Grafana, and OpenTelemetry / Datadog / ELK.
- CI/CD & Release Engineering: Optimize and scale CI/CD pipelines (GitHub Actions, GitLab CI) to enable rapid, safe deployments of threat analysis engines and cloud microservices.
- On-Call & Incident Response: Participate actively in 24/7 production on-call rotations. Lead real-time incident troubleshooting, root-cause analysis (RCA), and blameless post-mortems for service disruptions.
- System Performance Tuning: Assist in capacity planning, resource optimization, and Linux system/container tuning for high-throughput, low-latency processing workloads.
- Engineering Collaboration: Partner with software and security developers during design phases to ensure new features are built with reliability, scalability, and observability in mind.
Qualifications
Required Qualifications
- Experience: 2 - 4 years of hands-on experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles.
- On-Call Availability: Demonstrated experience participating in production on-call rotations for high-availability systems.
- OpenTelemetry & Monitoring: Hands-on experience collecting logs, metrics, and traces using OpenTelemetry collector pipelines and routing telemetry data to modern observability backends.
- Cloud & Hybrid Infrastructure: Practical experience building and managing workloads on at least one major Cloud Platform (AWS, GCP, Azure, or OCI).
- Linux Systems Fundamentals: Strong hands-on proficiency with Linux/Unix administration, shell scripting, process management, and core networking concepts (TCP/IP, DNS, Load Balancing).
- Programming & Automation: Solid coding ability in Python or Bash to build operational scripts, tools, and platform automation.
- Containers & Orchestration: Hands-on experience with containerization (Docker) and orchestration (Kubernetes, Helm).
- Infrastructure as Code: Working experience using Terraform or Ansible to manage declarative infrastructure.
- Problem-Solving & Communication: Strong analytical skills to debug issues across complex microservice environments, coupled with clear technical communication skills.
Preferred Qualifications
- AI Tooling: Familiarity with leveraging AI/LLM development assistants to accelerate automation scripting and log analysis.
- Security & Hardening Awareness: Basic familiarity with Linux security concepts (SELinux, CIS benchmarks, or container security).
- Public Cloud Exposure: Direct experience operating or deploying infrastructure on GCP/AWS/Azure/OCI.
- Experience working with high-volume, low-latency data pipelines, message brokers (Kafka, RabbitMQ, Pub/Sub), or caching layers (Redis).
Our Commitment
We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.
We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.
Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.
All your information will be kept confidential according to EEO guidelines.
Is role eligible for Immigration Sponsorship? No. Please note that we will not sponsor applicants for work visas for this position.
Similar roles
-
Site Reliability Engineer, Enterprise Technology Services
Apple Singapore
-
Site Reliability Engineer II
Okta Bengaluru, Karnataka, India
-
Senior Site Reliability Engineer (GCP Cloud) - Advanced English - Chile
Thoughtworks Santiago, Santiago Metropolitan Region, Chile
-
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore
Goldman Sachs Singapore
-
Senior II Site Reliability Engineer
Akamai Krakow, Lesser Poland Voivodeship, Poland
-
Site Reliability Engineer (GCP DevOps)
66degrees Bengaluru, Karnataka, India