Remote Jobs RockRemote Jobs Rock

Principal Site Reliability Engineer (Sovereign Cloud)

๐Ÿ•’ 5 days ago
Site Reliability EngineeringInfrastructure AutomationCloud Native ApplicationsKubernetes

๐Ÿ“œ Description

  • Support services running on large infrastructure for Sovereign Cloud.
  • Design, build, and operate reliable, secure cloud infrastructure.
  • Develop tools and automation frameworks for deployment.
  • Participate in on-call rotations for critical systems.
  • Collaborate with developers, researchers, and security experts.

๐Ÿ› ๏ธ Requirements

  • 7+ years as an engineer in Infrastructure, Operations, DevOps, or System Engineering.
  • 7+ years building high availability, scalable cloud native applications on AWS or GCP.
  • BS or MS in Computer Science or equivalent experience.
  • Expertise in configuration management with Ansible, Terraform, or Helm.
  • Experience in Site Reliability Engineering or Production Engineering.
  • Solid experience in Kubernetes and containers.
  • Proficiency with programming languages like Python, Java, and Golang.
Full job description

Join Palo Alto Networks as a Principal Site Reliability Engineer to support and enhance our Sovereign Cloud infrastructure through automation, security, and reliability.

Description

  • Support services running on large infrastructure for Sovereign Cloud.
  • Design, build, and operate reliable, secure cloud infrastructure.
  • Develop tools and automation frameworks for deployment.
  • Participate in on-call rotations for critical systems.
  • Collaborate with developers, researchers, and security experts.

Requirements

  • 7+ years as an engineer in Infrastructure, Operations, DevOps, or System Engineering.
  • 7+ years building high availability, scalable cloud native applications on AWS or GCP.
  • BS or MS in Computer Science or equivalent experience.
  • Expertise in configuration management with Ansible, Terraform, or Helm.
  • Experience in Site Reliability Engineering or Production Engineering.
  • Solid experience in Kubernetes and containers.
  • Proficiency with programming languages like Python, Java, and Golang.
Grafanalabs

Staff Software Engineer - Databases SRE | Sweden | Remote

Grafanalabs๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software
๐Ÿ•’ 21 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloudโ€™s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

Site Reliability EngineeringKubernetesAWSGCP
Grafanalabs

Staff Software Engineer - Databases SRE | UK | Remote

Grafanalabs๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software
๐Ÿ•’ 21 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloudโ€™s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

KubernetesAWSGCPAzure
Grafanalabs

Staff Software Engineer - Databases SRE | Germany | Remote

Grafanalabs๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software
๐Ÿ•’ 21 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloudโ€™s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

KubernetesAWSGCPAzure
Yuno

Site Reliability Engineer

Yuno๐Ÿ‘ฅ 201 - 500 employees๐Ÿข Financial Services
๐Ÿ•’ 29 days ago

As a Staff Site Reliability Engineer, you'll define and drive the reliability strategy for Yuno's AI agent infrastructure, ensuring it scales and remains robust across global operations.

๐Ÿ“ ๐Ÿ‡ช๐Ÿ‡บ Europe - Remote๐Ÿ’ผ Remote๐ŸŽบ Leadโญ Site Reliability Engineer๐Ÿ“ข ๐Ÿ‡ฌ๐Ÿ‡ง English Required๐Ÿ“ข ๐Ÿ‡ช๐Ÿ‡ธ Spanish Required
Event-driven ArchitectureMessaging SystemsAWSInfrastructure As Code
Plaid

Staff Site Reliability Engineer - Release Engineering

Plaid๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development
๐Ÿ•’ 18 days ago

As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering, ensuring safe and efficient deployment systems.

Site Reliability EngineeringBackend SystemsPlatform EngineeringSLO Frameworks

Trusted by Remote Workers