Remote Jobs RockRemote Jobs Rock

Site Reliability Engineer

๐Ÿ“… Jul 9
Performance MonitoringCapacity PlanningScriptingAutomation

๐Ÿ“œ Description

  • Proactively monitor and analyse platform performance.
  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs.
  • Continually improve observability through monitoring and alerting.
  • Ensure the service is highly available and resilient.
  • Conduct assessments of capacity and plan for scaling.

๐Ÿ› ๏ธ Requirements

  • 7-12 years of experience
  • Experience in performance monitoring and analysis
  • Capacity planning experience
  • Scripting and automation skills, with experience in relevant technologies.
  • Experience with Infrastructure as Code, in particular, Terraform
  • Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora)
  • Experience with messaging and distributed asynchronous workloads
  • Experience with nginx or similar technologies
  • Familiarity with SRE processes.
  • Aware of DevOps principles like the 3 ways and 5 ideals.

โœจ Benefits

  • Hybrid work environment
  • Group Term Life Insurance paid out at 3x Annual CTC (Arbor India)
  • 32 days holiday (plus Arbor Holidays). This is made up of 25 days annual leave plus 7 extra companywide days given over Easter
  • Work time: 9.30 am to 6 pm (8.5 hours only)
  • Compensation - 100% fixed salary disbursement and no variable components
Full job description

We are looking for an enthusiastic and proactive Site Reliability Engineer to join our SRE team and help us ensure we provide world-class resilience and performance across the platform. The remit and focus of the role is to advise on all aspects of site reliability including availability, scalability, observability and capacity planning. Itโ€™s a broad and exciting role, so weโ€™re looking for someone up for a challenge - if youโ€™re an energetic and a collaborative Site Reliability Engineer, this is the role for you.

Key responsibilities

  • Proactively monitor and analyse platform performance.
  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs
  • Continually improve observability through monitoring and alerting, and dashboards, using tools such as DataDog or Prometheus for example.
  • Work with other teams to ensure it is effective and provides full coverage.
  • Ensure the service is highly available and resilient
  • Champion best practices in design for high availability
  • Devise runbooks and run game sessions to test our DR plan, H/A and backups
  • Conduct assessments of capacity and plan for scaling to meet current and future business needs.
  • Work closely with the Head of Platform Engineering and Head of SRE to strategize and implement scalable solutions.
  • Work closely with the Platform team, feature teams and, 2nd line support and other stakeholders to ensure a good level of service is provided for our customers and embed SRE practices.
  • Key player in the response and troubleshooting of incidents, ensuring rapid resolution and minimising downtime.
  • Participate in blameless postmortems to identify root cause and corrective actions
  • Develop and maintain playbooks and documentation

Requirements

  • 7-12 years of experience
  • Experience in performance monitoring and analysis
  • Capacity planning experience
  • Scripting and automation skills, with experience in relevant technologies.
  • Experience with Infrastructure as Code, in particular, Terraform
  • Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora)
  • Experience with messaging and distributed asynchronous workloads
  • Experience with nginx or similar technologies
  • Familiarity with SRE processes.
  • Aware of DevOps principles like the 3 ways and 5 ideals.

Desired Skills

  • Experience with other database technologies and cloud platforms.
  • Past experience with Enterprise solutions running at scale
  • Familiarity with Kanban and Agile development processes
  • Experience with containerisation, for example Docker
  • Familiarity with software best practices such as Refactoring, Clean Code, Domain-Driven Design and Test-Driven Development.

Benefits

The chance to work alongside a team of hard-working, passionate people in a role where youโ€™ll see the impact of your work everyday. We also offer:

  • Hybrid work environment
  • Group Term Life Insurance paid out at 3x Annual CTC (Arbor India)
  • 32 days holiday (plus Arbor Holidays). This is made up of 25 days annual leave plus 7 extra companywide days given over Easter, Summer & Christmas
  • Work time: 9.30 am to 6 pm (8.5 hours only)
  • Compensation - 100% fixed salary disbursement and no variable components
Okta

Senior Site Reliability Engineer

๐Ÿ•’ 11 days ago
Okta๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development

As a Senior Site Reliability Engineer, you will build and operate a Kubernetes-based platform, enhancing internal workflows and ensuring reliability for AI-driven processes.

KubernetesSite Reliability EngineeringPlatform EngineeringInfrastructure Engineering
Zscaler

Senior Staff Site Reliability Engineer- Golang

๐Ÿ•’ 12 days ago
Zscaler๐Ÿ‘ฅ 10,000+ employees๐Ÿข Computer And Network Security๐Ÿค B2B

As a Senior Staff Site Reliability Engineer, you will ensure the reliability and operational integrity of a global infrastructure, developing automation and fault-tolerant systems.

AnsiblePythonGoGCP
Funcional Health Tech

ESPECIALISTA INFRAESTRUTURA SRE

๐Ÿ•’ 20 days ago
Funcional Health Tech๐Ÿ‘ฅ 501 - 1000 employees๐Ÿข Information Technology And Services

As a Senior Infrastructure Analyst, you will enhance the reliability and resilience of environments and applications using SRE practices, automation, and infrastructure engineering.

SRE PracticesObservabilityAutomationInfrastructure Engineering
Crunchyroll

Senior Site Reliability Engineer

๐Ÿ•’ 22 days ago
Crunchyroll๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Entertainment

As a Senior Site Reliability Engineer, you will enhance the reliability, scalability, and security of Crunchyroll's data platforms while driving modern SRE practices and cross-functional collaboration.

Site Reliability EngineeringKubernetesGoogle Cloud PlatformInfrastructure As Code
GoDaddy

Senior Site Reliability Engineer

๐Ÿ•’ 28 days ago
GoDaddy๐Ÿ‘ฅ 5001 - 10,000 employees๐Ÿข Internet

As a Senior Site Reliability Engineer, you will lead initiatives to enhance reliability and operational excellence across GoDaddy's Commerce platform, collaborating with various engineering teams.

AWSLinuxKubernetesInfrastructure As Code

Trusted by Remote Workers