Remote Jobs RockRemote Jobs Rock

Staff Site Reliability Engineer

📅 May 20
Site Reliability EngineeringDevOpsSystems EngineeringInfrastructure Engineering

📜 Description

  • Architect and implement comprehensive monitoring, logging, and tracing solutions.
  • Define and drive reliability standards by tracking Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
  • Lead incident management and response during high-impact incidents.
  • Drive automation and infrastructure as code to eliminate operational toil.
  • Optimize performance on Kubernetes and resolve performance bottlenecks.
  • Educate and mentor the engineering team to prioritize reliability.

🛠️ Requirements

  • 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).
  • Strong programming skills in languages like Python or Go. You write high-quality, well-tested code.
  • Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture.
  • Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.
  • Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing).
  • Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure.
  • Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.
  • Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices.
  • Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels.
  • A willingness to dive into understanding, debugging, and improving any layer of the stack.

Benefits

  • 💰 Competitive Salary & Equity
  • Dental
  • Vision and Life Insurance
  • 🚼 Paid Parental
  • Medical
  • Caregiver Leave
  • 🏝 Flexible Time Off (FTO) + Holidays
  • 🚗 Commuter
Ping Identity

Senior Staff Site Reliability Engineer

Ping Identity👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 11 days ago

As a Ping Identity Senior Staff Site Reliability Engineer, you will be involved in every facet of our Cloud-based services. You will establish solutions for building, deploying, and maintaining the infrastructure of one of the largest.

Cloud ServicesDevOpsCI/CD PipelinesInfrastructure Design
Zscaler

Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler👥 10,000+ employees🏢 Computer And Network Security🤝 B2B
🕒 4 days ago

As a Staff Site Reliability Engineer, you will architect, scale, and maintain cloud infrastructure, ensuring high availability and security for large-scale distributed systems.

Cloud ManagementContainer OrchestrationMonitoring SystemsLinux

Trusted by Remote Workers