CLOUDSCOUTS SOFTWARE SOLUTIONS LLC

Senior Production Support Engineer

CLOUDSCOUTS SOFTWARE SOLUTIONS LLC Austin, Texas, United States

Information Technology & Services · 11-50 employees

7 h ago
support-engineer Principal (10+ yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The role involves managing AI-augmented incident triage, runbook execution, and root-cause analysis across core FinOps applications. Additionally, the engineer will build and maintain AI-powered anomaly detection and spend alerting systems to ensure high uptime.

What they look for

Production Support SRE Incident Management Root-cause Analysis AI-augmented Triage Runbook Execution Anomaly Detection Monitoring Alerting GCP Node.js Python Automation Scripting FinOps

Requirements

Candidates must have 8 to 13 years of experience in production support or SRE with a strong background in incident management and monitoring tools. Proficiency in GCP-native tooling and scripting languages like Node.js or Python is highly preferred.

Full description

Role: Senior Production Support Engineer Location: Austin, TX (Onsite)   Exp. Level: 8+ yrs   Role Purpose Owns AI-augmented incident triage, runbook execution, and root-cause analysis, and separately owns AI-powered anomaly detection and spend/usage alerting.   Key Responsibilities:

  • Own AI-augmented incident triage, runbook execution, and root-cause analysis across the 3 core FinOps applications and 50+ services.
  • Build and maintain AI-powered anomaly detection and spend/usage alerting.
  • Support the 99.99% uptime target during core PST business hours (8AM-5PM)
  • Operate within the support portion of delivery (L1/L2/L3 tiers) within the 35% run-and-support split.
  • Contribute to baselining incident response/resolution times post-transition.

  Must-Have:

  • 8–13 years in production support or SRE, including incident management.
  • Experience with monitoring/alerting tooling, ideally GCP-native.
  • Demonstrated root-cause analysis discipline at production scale.

  Nice-to-Have:

  • Experience with AI-based anomaly detection tools.
  • Node.js or Python for support automation/scripting.

Similar roles