As a Staff Production Engineer, you will own high-risk technical domains, write production software, and collaborate closely with product teams to enhance reliability at scale.
Staff Software Engineer, Compute (Temporal Cloud)
📜 Description
- Create new managed compute primitives that feel first-class in Temporal Cloud: crisp abstractions, clean APIs, and an extension story across compute providers.
- Design self-optimizing autoscaling systems that scale worker fleets safely and predictably.
- Define the Open Source Server ↔ Cloud boundary for compute capabilities, ensuring a cohesive architecture.
- Architect, build, and operate services on the hot path of task execution where performance and correctness are customer-visible.
- Deliver real-world cloud integrations, including IAM boundaries and secure credentials handling.
- Make operability a feature through SLOs, tracing/metrics, and continuous hardening.
🛠️ Requirements
- Significant experience building distributed systems or multi-tenant platform services (design, implementation, and production operations).
- Strong systems fundamentals: concurrency, performance, reliability, and failure-mode thinking.
- A record of shipping platform primitives used by other engineers/customers (APIs, control planes, data planes).
- Comfort owning outcomes: SLOs, incident response, and improving on-call quality over time.
- Excellent written communication and crisp tradeoff thinking. (Go experience is a plus; judgment matters most.)
- Experience building cloud infrastructure platforms.
- Experience with IAM/security boundaries for cross-account execution models.
- Having built Kubernetes controllers / CRDs or heterogeneous worker fleet operations.
Full job description
Role Summary
Temporal powers durable execution for the world’s most demanding AI and enterprise systems. The Compute team at Temporal makes that durability feel effortless by building the managed compute primitives that run workers and other execution systems across Temporal Cloud.
We’re solving a hard platform problem: make compute transparent, elastic, safe-by-default, observable, and multi-cloud — without turning customers into infrastructure operators. We build platform primitives that span a variety of execution environments (e.g., serverless-style runtimes, containerized fleets, and customer-controlled deployments). Our work sits at the intersection of control plane + data plane: scaling decisions, isolation boundaries, multi-tenant safety, and production-grade observability. If you’ve built autoscaling systems, multi-tenant platforms, or cloud infrastructure, you’ll find a lot to love here.
See the "Compute team" link to read more --> Compute team
What You'll Do
Invent
Create new managed compute primitives that feel first-class in Temporal Cloud: crisp abstractions, clean APIs, and an extension story across compute providers.
Design self-optimizing autoscaling systems (signals, backstops, debouncing, guardrails) that scale worker fleets safely and predictably.
Define the Open Source Server ↔ Cloud boundary for compute capabilities, keeping the architecture cohesive and maintainable.
Own
Architect, build, and operate services on the hot path of task execution where performance and correctness are customer-visible.
Deliver real-world cloud integrations (e.g., customer-account execution): IAM boundaries, secure credentials handling, networking constraints, quotas, and failure modes.
Make operability a feature: SLOs, tracing/metrics, load and failure testing, incident reviews, and continuous hardening.
Ship end-to-end: API design, rollout strategy, backwards compatibility, and long-term maintenance.
Teach
Raise the bar through design docs, strong reviews, pragmatic technical leadership, and mentorship.
Lead through influence across teams (Server, SDKs, Security, Control Plane) to land coherent platform changes.
What You'll Bring
Significant experience building distributed systems or multi-tenant platform services (design, implementation, and production operations).
Strong systems fundamentals: concurrency, performance, reliability, and failure-mode thinking.
A record of shipping platform primitives used by other engineers/customers (APIs, control planes, data planes).
Comfort owning outcomes: SLOs, incident response, and improving on-call quality over time.
Excellent written communication and crisp tradeoff thinking. (Go experience is a plus; judgment matters most.)
Nice to Have
Experience building cloud infrastructure platforms.
Experience with IAM/security boundaries for cross-account execution models.
Having built Kubernetes controllers / CRDs or heterogeneous worker fleet operations.
Temporal Technologies is an Equal Opportunity Employer. Temporal Technologies does not discriminate on the basis of race, religion, color, sex, gender identity, sexual orientation, age, non-disqualifying physical or mental disability, national origin, veteran status, or any other basis covered by appropriate law. All employment is decided on the basis of qualifications, merit, and business need. We embrace and celebrate differences and diversity.
Temporal is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. If you need to request a reasonable accommodation, please let your Recruiter know so we can assist.
Similar jobs
Search more Software Engineer jobsSenior Software Engineer - SRE
As a Senior Software Engineer - SRE, you will design and implement solutions to enhance the availability and resiliency of Salesloft's services while collaborating on incident response and system optimization.
As a Senior Engineer, you will enhance a diverse software portfolio by improving UIs, optimizing workflows, fixing C++ bugs, and integrating tools with AI agents.
Senior Developer (C# / .Net)
In this senior engineering role, you will architect and develop robust.NET software solutions for electronic trading, ensuring reliability and scalability.
Senior Java Developer
As a Senior Java Developer, you will design, develop, and optimize enterprise-grade applications, ensuring high performance and reliability while collaborating with cross-functional teams.
