Remote Jobs RockRemote Jobs Rock

Senior Site Reliability Engineer

📅 Sep 25, 2025
AWSKubernetesPythonGo

📜 Description

  • Own uptime, reliability, and performance of services running on AWS and Kubernetes.
  • Design and implement self-healing infrastructure using automation and AI agents.
  • Build LLM-powered operational tooling for intelligent alert triage and incident management.
  • Manage and scale Kubernetes workloads, ensuring cluster reliability and cost efficiency.
  • Build observability systems including metrics, logs, and tracing.
  • Lead incident response and continuous reliability improvements.

🛠️ Requirements

  • 5+ years in SRE / DevOps / Platform Engineering.
  • Strong hands-on experience with AWS infrastructure at scale.
  • Experience with production-grade Kubernetes clusters.
  • Proven ability to debug complex distributed systems under pressure.
  • Strong coding skills in Python or Go.
  • Experience implementing monitoring, alerting, and incident management systems.
  • Familiarity with LLM APIs such as the OpenAI API is a plus.
Full job description
Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com.

We’re a fast-moving AI Security Company building AI-native infrastructure and applications powered by LLMs and autonomous agents. Our stack is deeply integrated with AWS, Kubernetes, and OpenAI-based systems, and we’re rethinking reliability in a world where software can reason, adapt, and self-heal.

We’re hiring a Senior Site Reliability Engineer to own reliability across our cloud-native and AI-driven platform. You’ll work at the intersection of distributed systems, Kubernetes operations, and LLM-powered automation, building systems that don’t just scale—but think and fix themselves.

Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com. Saviynt is an amazing place to work. We are a high-growth, Platform as a Service company focused on Identity Authority to power and protect the world at work. You will experience tremendous growth and learning opportunities through challenging yet rewarding work which directly impacts our customers, all within a welcoming and positive work environment. If you're resilient and enjoy working in a dynamic environment you belong with us! Security & ComplianceThis role requires adherence to Saviynt’s information security and privacy policies and procedures, including annual security training. Saviynt is an equal opportunity employer and we welcome everyone to our team.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.

WHAT YOU BRING

  • 5+ years in SRE / DevOps / Platform Engineering.
  • Strong hands-on experience with:
    • AWS infrastructure at scale
    • Kubernetes (production-grade clusters)
    • Proven ability to debug complex distributed systems under pressure.
    • Strong coding skills (Python or Go)—you build internal platforms and tools.
    • Experience implementing monitoring, alerting, and incident management systems.
    • Bonus (AI / LLM Focus)

      • Experience working with LLM APIs such as the OpenAI API.
      • Familiarity with agent frameworks like:
        • LangChain
        • AutoGen
        • Built or experimented with:
          • AI agents for DevOps / SRE workflows
          • Retrieval-Augmented Generation (RAG) systems
          • Vector databases (Pinecone, Weaviate, etc.)
          • Exposure to AIOps or intelligent automation systems.

WHAT YOU WILL BE DOING

  • Own uptime, reliability, and performance of services running on AWS + Kubernetes (EKS).
  • Design and implement self-healing infrastructure using automation and AI agents.
  • Build LLM-powered operational tooling using APIs such as the OpenAI API for:
    • Intelligent alert triage
    • Incident summarization
    • Root cause analysis
    • Runbook automation
    • Manage and scale Kubernetes workloads:
      • Deployments, autoscaling, resource optimization
      • Cluster reliability and cost efficiency
      • Build and evolve observability systems:
        • Metrics (Prometheus), dashboards (Grafana)
        • Logs (ELK / OpenSearch)
        • Tracing (OpenTelemetry)
        • Define and enforce SLOs, SLAs, and error budgets tied to business metrics.
        • Automate infrastructure using Terraform and CI/CD pipelines.
        • Lead incident response, postmortems, and continuous reliability improvements.
        • Introduce chaos engineering practices to proactively test system resilience.
Earnin

Staff Site Reliability Engineer

Earnin👥 501 - 1000 employees🏢 Financial Services
🕒 4 days ago

Lead the evolution of EarnIn's reliability practices by implementing an AI-first operating model to enhance operational quality and incident response across critical services.

Site Reliability EngineeringAI OperationsIncident ManagementSoftware Engineering
Bitwarden

Senior Site Reliability Engineer - FedRAMP

Bitwarden👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 14 days ago

As a Senior Site Reliability Engineer - FedRAMP, you will manage and enhance the Bitwarden Gov cloud infrastructure, ensuring security, reliability, and compliance in a multi-cloud environment.

FedrampCloud NetworkingTraffic ManagementMulti-region Deployments
Servicetitan

Senior Site Reliability Engineer

Servicetitan👥 1001 - 5000 employees🏢 Software
🕒 6 days ago

Join our Site Reliability & Infrastructure Engineering team as a Senior Site Reliability Engineer, where you'll ensure the reliability and health of cloud applications while driving efficiency and innovation.

KubernetesSRE PrinciplesAWSAzure

Trusted by Remote Workers