Remote Jobs RockRemote Jobs Rock

Site Reliability Engineer (SRE)

📅 Feb 4, 2025
Site Reliability EngineeringDevOpsELK StackTICK Stack

📜 Description

  • Ensure the reliability, availability, and performance of customer platforms and services.
  • Utilize ELK and TICK stacks for log and metrics monitoring.
  • Bridge the gap between development and operations.
  • Collaborate with teams to optimize system performance.

🛠️ Requirements

  • Minimum 3+ years of experience in Site Reliability Engineering, DevOps, or a related role.
  • Proficiency in the ELK stack (Elasticsearch, Logstash, Kibana) for log monitoring.
  • Experience with the TICK stack (Telegraf, InfluxDB, Chronograf, Kapacitor) for metrics monitoring.
  • Strong scripting skills in Python, Bash, or Ruby.
  • Understanding of Operating Systems: Ubuntu (OpenStack), Debian, and Redhat.
  • Familiarity with configuration management tools like Ansible, Puppet, or Chef.
  • Experience with containerization and orchestration tools like Docker and Kubernetes.
  • Bachelor's degree in computer science, Information Technology, or a related field.
Full job description

About VivaOps :

VivaOps is a leading DevSecOps platform company specializing in GitLab - The comprehensive DevOps platform, to transform and secure software development processes. We help organizations to streamline their DevSecOps journey by offering a complete range of GitLab services, from advisory, to implementation and managed services, to accelerate deployment, optimize security, and improve collaboration. Our deep expertise in GitLab enables clients to unlock the full potential of their development pipelines, driving efficiency, innovation, and competitive advantage.

Job Title: Site Reliability Engineer (SRE)

Location: Remote

Shift Timings: 5:30 PM to 3:00 AM IST to ensure support for global operations.

Job Description:

We are seeking a skilled Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have a strong background in both log and metrics monitoring stacks, specifically ELK (Elasticsearch, Logstash, Kibana) and TICK (Telegraf, InfluxDB, Chronograf, Kapacitor). As an SRE, you will be responsible for ensuring the reliability, availability, and performance of our customer’s platforms and services, bridging the gap between development and operations.

Required Skills and Qualifications:

● Minimum 3+ years of experience in Site Reliability Engineering, DevOps, or a related role.

● Proficiency in the ELK stack (Elasticsearch, Logstash, Kibana) for log monitoring.

● Experience with the TICK stack (Telegraf, InfluxDB, Chronograf, Kapacitor) for metrics monitoring.

● Strong scripting skills in languages such as Python, Bash, or Ruby.

● Understanding of Operating System:

Ubuntu(OpenStack) - Must have

Debian and Redhat etc.,

● DevOps Platforms

Gitlab - Good to have

Or similar

● Solid understanding of Grafana and Prometheus.

● Having worked with ServiceNow or something similar.

● Experience with configuration management tools like Ansible, Puppet, or Chef.

● Familiarity with containerization and orchestration tools like Docker and Kubernetes.

● Understanding of cloud platforms (Any of AWS, Azure, or GCP) and their services.

● Bachelor’s degree in computer science, Information Technology, or a related field.

● Excellent problem-solving skills and attention to detail.

● Strong communication and collaboration abilities.

About Inorg :

InOrg handles India Operations for VivaOps and is dedicated to empowering organizations to achieve global growth at scale. InOrg is establishing a Global Capability Center (GCC) for VivaOps.

Social-Discovery-Ventures

Site Reliability Engineer (SRE)

Social-Discovery-Ventures👥 501 - 1000 employees🏢 Computer Software
🔥 2 hours ago

As a Site Reliability Engineer (SRE), you will enhance infrastructure reliability, automate processes, and develop scalable production systems to support our social discovery platforms.

Site Reliability EngineeringInfrastructure ReliabilityAutomationKubernetes
SpaceX

Site Reliability Engineer (Manufacturing Infrastructure)

SpaceX👥 10,000+ employees🏢 Aviation And Aerospace Component Manufacturing🤝 B2B
🕒 5 days ago

As a Site Reliability Engineer for Manufacturing Infrastructure, you will enhance the reliability and scalability of systems supporting SpaceX's manufacturing processes.

Site Reliability EngineeringDevOpsInfrastructure As CodeLinux
Pagerduty

Site Reliability Engineer II

Pagerduty👥 1001 - 5000 employees🏢 Computer Software
🕒 6 days ago

As a Site Reliability Engineer II, you will enhance and maintain the foundational infrastructure that supports PagerDuty's real-time digital operations platform, ensuring reliability and scalability.

Site Reliability EngineeringDevOpsPlatform EngineeringLinux

Trusted by Remote Workers