Remote Jobs RockRemote Jobs Rock

Engineering Manager, Site Reliability Engineering

๐Ÿ”ฅ 22 hours ago
Engineering ManagementObservabilityIncident ManagementLoad Testing

๐Ÿ“œ Description

  • Build and operate metrics, logs, traces, and alerting capabilities for observability.
  • Own incident tooling and practices, coordinating cross-team responses to improve recovery times.
  • Develop load and failure testing capabilities to validate critical paths and production readiness.
  • Lead performance engineering engagements to identify bottlenecks and deliver improvements.
  • Coach engineers and develop technical leaders while managing team performance.
  • Measure outcomes and track improvements in rollout safety and incident management.

๐Ÿ› ๏ธ Requirements

  • Demonstrated engineering management experience leading and developing engineers.
  • Depth in building and operating distributed systems or reliability platforms.
  • Experience with safe-change and performance judgment in production environments.
  • Ability to build capabilities that other teams adopt and lead hands-on engagements.

โœจ Benefits

  • ๐Ÿ’ฐ Competitive Salary & Equity
  • ๐Ÿ’น 401(k) Program with a 4% match ( US Only )
  • Dental
  • Vision and Life Insurance
  • ๐Ÿฉผ Short Term and Long Term Disability
  • ๐Ÿšผ Paid Parental
  • Medical
  • Caregiver Leave
  • ๐Ÿ Flexible Time Off (FTO) + Holidays
  • ๐Ÿš— Commuter Benefits ( In-Office & US Only )
Full job description

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the Role

Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.

This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure. You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries.

This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team.

What You'll Do

  • Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.

  • Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.

  • Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.

  • Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service ownersโ€”not just recommendations.

  • Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding toolsโ€”including Replitโ€”to prototype and automate. Apply rigorous review and verification to AI-generated changes.

  • Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.

  • Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.

What You'll Bring

  • Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a teamโ€”not only acted as its strongest individual contributor.

  • Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.

  • Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions.

  • Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.

Nice to Have

  • Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.

  • Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.

  • Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.

  • Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.

Full-Time Employee Benefits Include:

๐Ÿ’ฐ Competitive Salary & Equity

๐Ÿ’น 401(k) Program with a 4% match (US Only)

โš•๏ธ Health, Dental, Vision and Life Insurance

๐Ÿฉผ Short Term and Long Term Disability

๐Ÿšผ Paid Parental, Medical, Caregiver Leave

๐Ÿ Flexible Time Off (FTO) + Holidays

๐Ÿš— Commuter Benefits (In-Office & US Only)

๐Ÿ“ฑ Monthly Wellness Stipend

๐Ÿง‘โ€๐Ÿ’ป Autonomous Work Environment

๐Ÿ–ฅ In Office Set-Up Reimbursement (In-Office Only)

๐Ÿš€ Quarterly Team Gatherings

โ˜• In Office Amenities (In-Office Only)

Want to learn more about what we are up to?

Interviewing + Culture at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Vercel

Engineering Manager, Dashboard

๐Ÿ•’ 4 days ago
Vercel๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development๐Ÿค B2B

As an Engineering Manager on the Dashboard team, you will lead engineers to enhance the primary interface for developers, ensuring a seamless experience in building and managing applications.

Team LeadershipCoachingFrontend ArchitecturePerformance Optimization
Sentry

Engineering Manager, Events Analytics Platform

๐Ÿ•’ 4 days ago
Sentry๐Ÿ‘ฅ 201 - 500 employees๐Ÿข Computer Software

Lead a team of engineers to build and scale Sentry's Events Analytics Platform, driving architectural evolution and ensuring system stability while mentoring talent.

Software EngineeringPeople ManagementData PlatformsStorage Systems
Anthropic

Engineering Manager, Data Infrastructure

๐Ÿ•’ 4 days ago
Anthropic๐Ÿ‘ฅ 10,000+ employees๐Ÿข Research Services๐Ÿค B2B

Lead and scale the Data Warehouse & Streaming Infra team, owning the data platform that supports business decisions and enhances system safety at Anthropic.

Data WarehousingEvent StreamingChange Data CaptureData Ingestion
Cloudflare

Senior Engineering Manager, Observability

๐Ÿ•’ 10 days ago
Cloudflare๐Ÿ‘ฅ 10,000+ employees๐Ÿข Computer And Network Security๐Ÿค B2B

As the Engineering Manager for the Developer Observability Platform, you will lead a customer-focused team to enhance telemetry discoverability and drive impactful solutions.

ObservabilitySoftware EngineeringTeam ManagementCloud Infrastructure
OpenAI

Engineering Manager, Rosalind Workbench

๐Ÿ•’ yesterday
OpenAI๐Ÿ‘ฅ 10,000+ employees๐Ÿข Research Services

Lead the engineering team for Rosalind Workbench, focusing on hiring, technical direction, and delivering reliable software products that enhance scientific research.

Team LeadershipSoftware EngineeringTechnical PlanningArchitecture Design

Trusted by Remote Workers