Remote Jobs RockRemote Jobs Rock

Observability Engineer / Site Reliability Engineer

๐Ÿ“… Jul 17
Google Cloud Platform (gcp)GrafanaPrometheusTerraform

๐Ÿ“œ Description

  • Architect, optimize, and maintain observability frameworks across cloud environments, focusing on Google Cloud Platform (GCP) tools.
  • Design, deploy, and maintain observability stacks using Prometheus, Grafana, and cloud-native integrations.
  • Drive infrastructure-as-code (IaC) initiatives with Terraform and Ansible for automated deployments.
  • Build and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) using CI/CD pipelines.
  • Analyze Linux/Unix system architectures, optimizing performance across distributed environments.
  • Implement SRE best practices, establishing meaningful SLIs, SLOs, and Error Budgets.

๐Ÿ› ๏ธ Requirements

  • Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
  • OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
  • Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.
Full job description

About Ontrac Solutions

Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.

Role Overview

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.

Key Responsibilities

  • GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.

Required Skills & Qualifications

  • Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
  • OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
  • Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

__________________________________

Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are โ€” and gives you a portable verification credential you can share with any employer.

Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
Get verified โ†’ https://pinpointverify.com/ontrac

Aristanetworks

Site Reliability Engineer (SRE) - Engineering Productivity

Aristanetworks๐Ÿ‘ฅ 5001 - 10,000 employees๐Ÿข Computer Networking
๐Ÿ•’ 6 days ago

As a Site Reliability Engineer in the Engineering Productivity team, you will design, build, and operate scalable and reliable systems to enhance developer experience and support Arista's product development.

GoPythonShell ScriptingLinux
Cyberark1

Principal Site Reliability Engineer (Sovereign Cloud)

Cyberark1๐Ÿ‘ฅ 10,000+ employees๐Ÿข Computer & Network Security
๐Ÿ•’ 5 days ago

Join Palo Alto Networks as a Principal Site Reliability Engineer to support and enhance our Sovereign Cloud infrastructure through automation, security, and reliability.

Site Reliability EngineeringInfrastructure AutomationCloud Native ApplicationsKubernetes
Servicenow

Senior Reliability Engineer

Servicenow๐Ÿ‘ฅ 10,000+ employees๐Ÿข Computer Software
๐Ÿ•’ 11 days ago

Lead critical infrastructure initiatives as a Staff Site Reliability Engineer, architecting scalable solutions and driving innovation across the organization.

Site Reliability EngineeringDevOpsInfrastructure EngineeringAWS
Omnisend

Senior SRE Engineer

Omnisend๐Ÿ‘ฅ 201 - 500 employees๐Ÿข Internet
๐Ÿ•’ yesterday

As a Senior SRE Engineer, you will build a platform that empowers teams to manage their datastores and cloud resources while ensuring the stability and scalability of a Kubernetes cluster running hundreds of microservices.

MongodbPostgreSQLMySQLRedis
Anthropic

Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic๐Ÿ‘ฅ 10,000+ employees๐Ÿข Research Services๐Ÿค B2B
๐Ÿ•’ 4 days ago

You'll lead the deployment and verification of safety systems for AI model launches, ensuring safeguards are effectively configured and operational across various platforms.

Production Change ManagementDeploy PipelinesConfig Management SystemsCanary Analysis

Trusted by Remote Workers