Remote Jobs RockRemote Jobs Rock

Site Reliability Engineer

🕒 28 days ago
AnsiblePuppetTerraformKubernetes

📜 Description

  • Be on an on-call rotation responding to production availability incidents and support service engineers with customer incidents.
  • Use your on-call shift to prevent incidents from ever happening.
  • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes.
  • Make monitoring and alerting focus on symptoms, not outages.
  • Document every action to turn findings into repeatable actions and automation.
  • Design, build, and maintain core infrastructure that scales to hundreds of thousands of concurrent users.

🛠️ Requirements

  • Strong programming skills in Python, Java, Golang, and Node.js.
  • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies.
  • Ability to collaborate and communicate asynchronously.
  • A proactive attitude towards fixing issues.
Full job description

Our client's Cloud Operations team is expanding its SRE function. Site Reliability Engineers keep all user-facing services and production systems running smoothly. SREs here are a blend of pragmatic operators and software craftspeople who apply sound engineering principles, operational discipline and mature automation to the environment and the codebase. The team specialises in systems — networking, the Linux kernel, and scaling, algorithms and distributed systems.

As an SRE you will

  • Be on an on-call rotation responding to production availability incidents, and support service engineers with customer incidents
  • Use your on-call shift to prevent incidents from ever happening
  • Run infrastructure with Ansible, Puppet, Terraform and Kubernetes
  • Make monitoring and alerting alert on symptoms, not outages
  • Document every action, so findings turn into repeatable actions — and then into automation
  • Improve the deployment process to make it as boring as possible
  • Design, build and maintain core infrastructure that scales to hundreds of thousands of concurrent users
  • Debug production issues across services and levels of the stack
  • Plan the growth of the infrastructure

You may be a fit if you

  • Think cloud-first, regardless of the flavour of public cloud
  • Think security-first
  • Think about systems — edge cases, failure modes, behaviours, specific implementations
  • Know your way around Linux and Windows
  • Know the use of config-management systems like Ansible or Puppet
  • Have strong programming skills — Python, Java, Golang, Node.js
  • Collaborate and communicate asynchronously, and document so nothing is learned twice
  • Have a go-for-it attitude: when you see something broken, you fix it
  • Have experience with Nginx, HAProxy, Docker, Kubernetes, Terraform or similar technologies

Projects you could work on

  • Coding infrastructure automation with Ansible and Terraform
  • Improving Prometheus monitoring or building new metrics
  • Helping release managers deploy and fix new versions of application software
  • Planning and executing the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS)
  • Developing a relationship with a product group and defining their SRE KPIs — the SRE practice here is early in its journey
Social-Discovery-Ventures

Site Reliability Engineer (SRE)

Social-Discovery-Ventures👥 501 - 1000 employees🏢 Computer Software
🔥 1 hour ago

As a Site Reliability Engineer (SRE), you will enhance infrastructure reliability, automate processes, and develop scalable production systems to support our social discovery platforms.

Site Reliability EngineeringInfrastructure ReliabilityAutomationKubernetes

Red Hat is seeking a Customer Site Reliability Engineer to ensure the reliability and performance of critical services in the OpenShift Managed Cloud Services team.

📍 🇦🇺 Australia - Remote💼 Full-Time🎹 Mid-level Site Reliability Engineer📢 🇬🇧 English Required📢 🇯🇵 Japanese Required
OpenshiftKubernetesLinuxAWS

Trusted by Remote Workers