Phizenix

PnP Pipeline On-call Support Engineer

Phizenix

IT Services and IT Consulting · 2-10 employees

3 h ago
Remote Mid (2-5 yrs) Full-time
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves monitoring PnP pipelines to ensure stability and performing initial triage on failures to reduce engineering overhead. You will also collaborate with cross-functional teams to resolve infrastructure issues and maintain measurement accuracy.

What they look for

Pipeline Monitoring Troubleshooting Data Integrity CI/CD Infrastructure Management Firmware Analysis Unified Test Framework Log Analysis AI Tooling Bug Fixing Performance Analysis Technical Support

Requirements

Candidates must have proven experience in monitoring alerts and dashboards with strong troubleshooting skills for failure logs. Familiarity with internal infrastructure tools and the ability to differentiate between infrastructure, firmware, and analyzer errors is preferred.

Full description

Role Summary

Your primary mission is to ensure the consistent health and stability of our PnP (Power and Performance) pipelines. You will serve as the first line of defense, monitoring pipelines and performing initial triaging to reduce manual overhead for our engineering teams. This role is critical to maintaining the velocity of our post-silicon PnP work as we scale.

Core Responsibilities

  • Pipeline Monitoring: Conduct regular, proactive health checks on functional and PnP-enabled flows. Ensure data integrity by verifying that test results are populating correctly on the RLS PnP dashboard.
  • Initial Triage: Investigate pipeline failures and categorize them to streamline resolution:
  • Infrastructure/Deployment Issues: CI pools, hargow issues, or platform environments (e.g., Artemis/Coleman HW/FW mismatches).
  • UTF / PnP Analyzer Issues: Troubleshoot problems within the Unified Test Framework (UTF) or artifact deployment.
  • Mature Pipeline Issues: Address issues specific to pipeline logic for stable flows that are no longer under active development.
  • Resolution & Escalation:
  • Perform bug fixes for straightforward issues (e.g., patching deployment or execution changes).
  • For infrastructure issues, post to the appropriate internal channels for escalation and follow-up.
  • Escalate complex or systemic issues to pipeline owners by providing necessary logs and context.
  • Collaborate with pipeline owners and XFN (Cross-Functional) DRIs to ensure proper placement of PnP trace tags and alignment of measurement windows.

Qualifications & Technical Requirements

  • Core Requirements:
  • Proven experience in monitoring alerts and dashboards.
  • Strong troubleshooting skills with the ability to identify patterns in failure logs and suggest improvements.
  • Ability to differentiate between infrastructure, firmware, and analyzer-specific errors.
  • Preferred Qualifications:
  • Familiarity with jest-e2e and the Unified Test Framework (UTF).
  • Experience working within internal infrastructure (e.g., buck, sandcastle, phabricator, internal AI tooling).
  • Ability to read and interpret CI pipeline logs effectively.
  • Ability to leverage AI tooling effectively (claude code, Metamate, etc.)