Remote Jobs RockRemote Jobs Rock

Site Reliability Engineer

🕒 26 days ago
📍 🇪🇺 Europe - Remote🎺 Lead Site Reliability Engineer📢 🇬🇧 English Required📢 🇪🇸 Spanish Required💼 Remote
Event-driven ArchitectureMessaging SystemsAWSInfrastructure As Code

📜 Description

  • Set the technical direction for reliability across Yuno's infrastructure, focusing on AI agent provisioning and deployment.
  • Own the reliability strategy, driving architectural decisions and defining metrics for measuring reliability.
  • Design and implement event-driven communication and messaging layers for inter-service interactions.
  • Automate cloud infrastructure provisioning and ensure reliable scaling as transaction volumes increase.
  • Build observability tools to monitor platform health and lead incident management efforts.

🛠️ Requirements

  • Deep AWS — EC2, VPC, IAM, S3, and RDS — with strong networking fundamentals, since inter-service communication runs over the internal VPC
  • Infrastructure as Code — Terraform or Pulumi, reviewed in PRs rather than clicked in consoles
  • Kubernetes and Docker in production — container lifecycle, resource limits, health checks, and orchestration at scale
  • Observability and SLOs — Datadog fluency or equivalent (dashboards, monitors, APM, distributed tracing), and a track record defining and operating SLOs, SLIs, and error budgets across services
  • Chaos engineering and resilience testing — hands-on experience with fault injection, game days, or chaos experiments (Gremlin, Chaos Mesh, AWS FIS, or similar) to harden production systems
  • Distributed systems debugging — you've diagnosed async flows and cascading failures in production and can explain what broke and how you fixed it; comfortable coding for automation and tooling (Go, Python, or similar)
  • Databases — solid SQL (PostgreSQL) and NoSQL (MongoDB, Redis): when to use each, indexing, replication, and performance tuning
  • Proven technical leadership — you've set reliability standards, influenced architecture across teams, and mentored engineers, not just owned your own scope
  • English — advanced proficiency, written and spoken
  • AI / MLOps infrastructure — running AI workloads in production (model serving, LLM inference, GPU/resource management, and agent evaluation/observability tools like LangFuse, LangSmith, Braintrust, or MLflow)

Benefits

  • Home Office Bonus — a one-time allowance to set up your ideal home office
  • Work Equipment
  • Stock Options
  • Flexible Days Off
Grafanalabs

Staff Software Engineer - Databases SRE | Sweden | Remote

Grafanalabs👥 1001 - 5000 employees🏢 Computer Software
🕒 18 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

Site Reliability EngineeringKubernetesAWSGCP
Grafanalabs

Staff Software Engineer - Databases SRE | UK | Remote

Grafanalabs👥 1001 - 5000 employees🏢 Computer Software
🕒 18 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

KubernetesAWSGCPAzure
🕒 18 days ago

The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:.

KubernetesAWSGCPAzure
Filevine

Staff Site Reliability Engineer

Filevine👥 10,000+ employees🏢 Software Development🤝 B2B
📅 Aug 1

Filevine - Staff Site Reliability Engineer Staff Site Reliability Engineer Remote Engineering – Reliability / Full-time / Remote Submit your application • Resume/CV ATTACH RESUME/CV Couldn't auto-read resume.

Site Reliability EngineeringSystem ReliabilityPerformance OptimizationMonitoring Solutions

Trusted by Remote Workers