Remote Jobs RockRemote Jobs Rock

Senior Site Reliability Engineer

πŸ•’ 25 days ago
AWSPostgreSQLDevOpsSite Reliability Engineering

πŸ“œ Description

  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and other data stores.
  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
  • Proactively monitor production using CloudWatch dashboards and alerts, addressing operational issues before they impact users.
  • Lead production debugging and incident response, building and maintaining runbooks and participating in the on-call rotation.
  • Continuously refine infrastructure to ensure it is easily deployable and scalable, keeping infrastructure as code accurate.
  • Share knowledge of production operations with the team, fostering a culture of learning and growth.

πŸ› οΈ Requirements

  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.
  • 5+ years of experience in DevOps, site reliability, or platform operations with significant responsibility for production systems.
  • 3+ years of hands-on experience with AWS, emphasizing serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance.
  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation.
  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
  • Experience with production monitoring and alerting, incident response, and on-call ownership.
Full job description

As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle β€” from deployment to maintenance and updates β€” always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.

Responsibilities

  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.

  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.

  • Proactively monitor production β€” CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting β€” addressing operational issues before they impact users.

  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.

  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.

  • Share your knowledge of production operations with the team, fostering a culture of learning and growth.

Qualifications: Knowledge, Skills, & Abilities

  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.

  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.

  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).

  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.

  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.

  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).

  • Experience with production monitoring and alerting, incident response, and on-call ownership.

Mirantis

Senior Site Reliability Engineer (Golang / Kubernetes)

MirantisπŸ‘₯ 501 - 1000 employees🏒 Computer Software
πŸ•’ 5 days ago

As a Senior Site Reliability Engineer, you will define and measure reliability for a GPU-accelerated AI platform, owning service-level indicators and objectives while collaborating across teams.

Site Reliability EngineeringKubernetesAPI DevelopmentObservability
Bloomreach

Senior Site Reliability Engineer for Fuse team

BloomreachπŸ‘₯ 1001 - 5000 employees🏒 Internet
πŸ•’ 8 days ago

As a Senior Site Reliability Engineer on the Fuse team, you will lead the reliability and operability of complex product-data systems, ensuring dependable data management and observability across Bloomreach's platforms.

KubernetesGCPGoPython
Mirantis

Senior Site Reliability Engineer (Golang, Kubernetes)

MirantisπŸ‘₯ 501 - 1000 employees🏒 Computer Software
πŸ•’ 16 days ago

As a Senior Site Reliability Engineer, you will design, develop, and maintain cloud-based AI infrastructure, ensuring reliability and performance while mentoring team members.

DevOpsCloud TechnologiesKubernetesGolang
Mirantis

Senior Site Reliability Engineer (SRE)

MirantisπŸ‘₯ 501 - 1000 employees🏒 Computer Software
πŸ•’ 19 days ago

As a Senior Site Reliability Engineer, you will design, develop, and operate cloud-based AI solutions, ensuring the reliability and performance of container infrastructure while mentoring team members.

KubernetesOpenstackDevOpsCloud Infrastructure

Trusted by Remote Workers