As the Pipeline SRE Lead, you will enhance the reliability of critical release pipelines by leading observability, incident management, and continuous improvement efforts.
Senior Site Reliability & Software Engineering Manager
📜 Description
- Build and lead senior engineering teams for fleet reliability in Palo Alto and Europe.
- Automate systems to achieve zero manual pages and ensure fleet availability.
- Integrate AI/ML techniques into SRE disciplines for predictive remediation.
- Drive software engineering for bare-metal provisioning and health monitoring.
- Establish observability metrics and customer-facing SLIs/SLOs.
🛠️ Requirements
- 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call.
- Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset.
- Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations.
- Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments.
- Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps.
- Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers).
- Custom Fleet Tooling: Hands-on experience building custom remediation controllers, event-driven fleet management software (Go, NetBox/DCIM), or OpenTelemetry/Prometheus pipelines.
- Facility Telemetry & Thermal Signals: Familiarity with high-density, liquid-cooled environments and integrating facility telemetry (power, thermal, flow) into compute health signals.
- Customer SLAs & Reporting: Proven experience constructing customer SLAs/SLOs, credit mechanics, and executive/customer-facing reliability reviews.
✨ Benefits
- 401(k) Plan with 4% company match (USA employees)
- Generous base, bonus, and additional incentive-based compensation.
- Dental
- Vision coverage for you and your dependents
- Company-paid life insurance and disability.
Full job description
Built to set the gold standard for integrated AI infrastructure
Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.
Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.
Crux AI is led by CEO, Ben Treynor Sloss, who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) discipline. At Crux AI, we treat operations fundamentally as a software engineering problem.
WHAT YOU'LL DO
We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA.
In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads.
This is a true hands-on, builder seat, not a supervisory position. In the early days, you will write the first remediation controllers, set reliability baselines, and take initial pages yourself to stay close to the work before expanding your team. You will lead an elite group of unusually senior software and reliability engineers: engineers with significantly greater technical and software depth than traditional operational SRE orgs. You must thrive on independence and revel in ambiguity, turning unknowns into concrete engineering priorities in a fast-paced, high-growth environment. While you will devise and participate in initial on-call rotations, your core mandate is to combine SRE disciplines with extensive AI/ML automation to drive operational pages down to zero.
In this role, you will:
Build & Lead Senior Engineering Teams: Recruit, lead, and mentor an initial team of senior software and reliability engineers across Palo Alto and Europe as fleet capacity ramps rapidly.
Automate Pages to Zero: Own fleet availability end-to-end; establish on-call rotations while relentlessly developing self-healing systems and predictive remediation to eliminate manual pages.
Embed AI/ML into SRE Disciplines: Apply agentic techniques, machine learning models, and automated diagnostic workflows extensively to telemetry collection, root-cause analysis, and predictive cluster recovery.
Own Bare-Metal & Fleet Lifecycle Software: Drive software engineering for bare-metal node provisioning, firmware deployment, thermal/stress burn-in validation, host/TPU health monitoring, and decommissioning.
Control Plane & Fabric Reliability: Own software reliability for cluster orchestration, scheduling, capacity allocation APIs, and high-performance TPU host/interconnect networks.
Define Observability & SLOs: Establish customer-facing SLIs/SLOs (job goodput, time-to-detect, node availability) and build the telemetry pipelines serving operators, executives, and customers.
SIGNALS OF SUCCESS
After 60 days in this role:
Baseline reliability framework defined (v1 SLOs, severity structure, change management); first AI/ML-driven automated remediation controller shipped to production; recruiting active for senior engineering hires in Palo Alto.
After 6 months:
First TPU cluster brought online under your team’s automated acceptance criteria; automated remediation pipeline running with pass rates tracked; observability v1 in daily production use; core senior team onboarded
After 1 year:
TPU fleet operating against published customer SLOs; >90% of node/fabric faults automatically quarantined and remediated without human paging; team scaled ahead of rapid capacity ramps.
EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH
Experiences
10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call.
Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset.
Hands-On Technical Depth (SRE + SWE): Track record shipping production code in Go, Python, or C++, with hands-on systems expertise across Linux OS kernels, bare-metal provisioning, firmware, and/or high-performance networking fabrics. Extensive experience with distributed systems and open source software.
Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations.
Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments.
Attributes
Possess a high tolerance for ambiguity. The first clusters will carry customer workloads while the SLOs are still being defined and the team is still being hired. You absorb that, translate unknowns into concrete near-term priorities, and never manufacture false certainty about reliability the data does not support.
Understands that the customer’s job is the unit of reliability. A node that is “up” while a training run stalls on a flapping link is down. You measure what customers experience — goodput, time-to-recover, lost progress — and hold the whole stack, and Google, to it.
Mindset
Builder Mindset & Ambiguity: A true "builder, not supervisory" orientation; comfortable operating with high autonomy, navigating ambiguity, and establishing structure amidst rapid growth.
AI-Agentic First: Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process.
Nice to have (Preferred, not required):
Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps.
Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers).
Custom Fleet Tooling: Hands-on experience building custom remediation controllers, event-driven fleet management software (Go, NetBox/DCIM), or OpenTelemetry/Prometheus pipelines.
Facility Telemetry & Thermal Signals: Familiarity with high-density, liquid-cooled environments and integrating facility telemetry (power, thermal, flow) into compute health signals.
Customer SLAs & Reporting: Proven experience constructing customer SLAs/SLOs, credit mechanics, and executive/customer-facing reliability reviews.
Salary Range Information
The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Crux
We offer generous base, bonus and additional incentive based compensation
Health, dental, and vision coverage for you and your dependents
Company-paid life insurance and disability
Full suite of other optional benefits
401(k) Plan with 4% company match (USA employees)
Similar jobs
Search more Site Reliability Engineer jobsStaff Site Reliability Engineer - Release Engineering
As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering, ensuring safe and efficient deployment systems.
As a Staff Site Reliability Engineer at Okta, you will lead the design and operation of scalable cloud infrastructure, ensuring high reliability and compliance with FedRAMP standards.
Staff Software Engineer - Databases SRE | Sweden | Remote
The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.
En tant qu'Ingénieur.e staff de fiabilité des sites, vous serez responsable de la fiabilité et de la scalabilité de la plateforme de données, tout en mentorant les équipes d'ingénierie.
