Remote Jobs RockRemote Jobs Rock

Lead Site Reliability Engineer

📅 Jul 17
Infrastructure ManagementSite Reliability EngineeringAutomationKubernetes

📜 Description

  • Design and build systems to automate infrastructure management at scale.
  • Reduce operational toil by turning manual processes into reliable workflows.
  • Build internal tooling and platforms for safe self-service changes.
  • Improve the reliability and resilience of infrastructure components.
  • Implement systems for deploying and running applications in Kubernetes.
  • Contribute to architecture decisions across infrastructure, reliability, and security.

🛠️ Requirements

  • 10+ years of experience in infrastructure, SRE, or software engineering roles.
  • Strong software engineering skills—you build systems, not just scripts.
  • Experience managing production infrastructure at scale (cloud + containerized systems).
  • Experience with Infrastructure as Code (e.g., Terraform).
  • Experience running and troubleshooting distributed systems (Docker/Kubernetes).
  • Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.).
  • Proficiency in at least one programming language (Python, Go, JavaScript, etc.).
  • Strong communication and collaboration skills.

Benefits

  • Unlimited PTO and flexible work policy
  • Employee stock options
  • Medical
  • Dental
  • Vision plans with HSA (monthly employer contribution) and FSA options
  • 401k with 100% match up to 4% of annual employee compensation
  • Eligible new parents receive 16 weeks of paid parental leave
  • Home office stipend for new employees
  • Annual Learning & Development annual stipend
  • Well-being benefits include access to ClassPass, OneMedical, UrbanSitter, and Spring Health
🕒 16 days ago

As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering, ensuring safe and efficient deployment systems.

Site Reliability EngineeringBackend SystemsPlatform EngineeringSLO Frameworks
Grafanalabs

Staff Software Engineer - Databases SRE | Sweden | Remote

Grafanalabs👥 1001 - 5000 employees🏢 Computer Software
🕒 18 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

Site Reliability EngineeringKubernetesAWSGCP
Waymo

Site Reliability Engineer, Lead

Waymo👥 1001 - 5000 employees🏢 Internet
🕒 yesterday

As the Pipeline SRE Lead, you will enhance the reliability of critical release pipelines by leading observability, incident management, and continuous improvement efforts.

C++Large-scale Production Software SystemsMachine Learning SystemsIncident Management

Trusted by Remote Workers