Remote Jobs RockRemote Jobs Rock

Site Reliability Engineer

🕒 7 days ago
AWSCI/CDIncident ManagementInfrastructure As Code

📜 Description

  • Own the reliability and performance of production services running in AWS, including capacity planning, cost optimization, and architecture reviews
  • Design, build, and maintain fully automated CI/CD pipelines that take code from commit to production with minimal manual intervention
  • Lead incident management: serve in the on-call rotation, coordinate response during outages, run blameless postmortems, and drive remediation to completion
  • Define and track SLOs, SLIs, and error budgets in partnership with product and engineering teams
  • Build and improve observability through monitoring, logging, alerting, and distributed tracing
  • Manage infrastructure as code and eliminate toil through automation

🛠️ Requirements

  • 4+ years in SRE, DevOps, or cloud operations roles supporting production systems
  • Deep hands-on experience operating workloads in AWS (e.g., EC2, ECS/EKS, Lambda, RDS, S3, IAM, VPC, CloudWatch)
  • Proven experience with incident management: on-call ownership, incident command, root cause analysis, and postmortem processes
  • Demonstrated track record of building fully automated CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline, or similar)
  • Strong infrastructure-as-code skills with Terraform, CloudFormation, or CDK
  • Proficiency in at least one scripting or programming language (Python, Go, Bash)
  • Experience with containers and orchestration (Docker, Kubernetes)
  • Familiarity with observability tooling such as Datadog, Prometheus/Grafana, or the ELK stack
  • Clear communicator who stays calm under pressure and can explain complex issues to technical and non-technical audiences
  • AWS certifications (Solutions Architect, DevOps Engineer)

✨ Benefits

  • Competitive salary
  • Dental
  • Vision coverage
  • 401(k) with company match
  • Flexible PTO and hybrid work arrangement in Atlanta
Full job description

About SeekNow

Seek Now is transforming property inspections through technology, data, and human expertise. We deliver faster, smarter, more reliable insights to insurance carriers and single-family rental markets, and we’re just getting started. If you want to be part of a product-driven, tech-forward team building real-world impact at scale, you’re in the right place.

The Role

As a Site Reliability Engineer, you'll be responsible for the availability, scalability, and operational health of our AWS-hosted infrastructure. You'll lead incident response, build the automation that lets our engineering teams ship safely and often, and drive a culture of measurable reliability across the organization.

What You'll Do

  • Own the reliability and performance of production services running in AWS, including capacity planning, cost optimization, and architecture reviews
  • Design, build, and maintain fully automated CI/CD pipelines that take code from commit to production with minimal manual intervention
  • Lead incident management: serve in the on-call rotation, coordinate response during outages, run blameless postmortems, and drive remediation to completion
  • Define and track SLOs, SLIs, and error budgets in partnership with product and engineering teams
  • Build and improve observability through monitoring, logging, alerting, and distributed tracing
  • Manage infrastructure as code and eliminate toil through automation
  • Partner with development teams to embed reliability best practices into system design and release processes
  • Contribute to disaster recovery planning, security hardening, and compliance efforts

What We're Looking For

  • 4+ years in SRE, DevOps, or cloud operations roles supporting production systems
  • Deep hands-on experience operating workloads in AWS (e.g., EC2, ECS/EKS, Lambda, RDS, S3, IAM, VPC, CloudWatch)
  • Proven experience with incident management: on-call ownership, incident command, root cause analysis, and postmortem processes
  • Demonstrated track record of building fully automated CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline, or similar)
  • Strong infrastructure-as-code skills with Terraform, CloudFormation, or CDK
  • Proficiency in at least one scripting or programming language (Python, Go, Bash)
  • Experience with containers and orchestration (Docker, Kubernetes)
  • Familiarity with observability tooling such as Datadog, Prometheus/Grafana, or the ELK stack
  • Clear communicator who stays calm under pressure and can explain complex issues to technical and non-technical audiences

Nice to Have

  • AWS certifications (Solutions Architect, DevOps Engineer)
  • Experience with GitOps workflows (ArgoCD, Flux)
  • Background in security operations, compliance frameworks (SOC 2, ISO 27001)

Why You'll Love It Here

  • Tech-First Culture: We believe in building smart, scalable systems—and we invest in them.
  • Real-World Impact: Your work will touch thousands of users every day, improving workflows and outcomes.
  • Autonomy + Collaboration: Own your space while being part of a highly connected, supportive team.
  • Growth-Minded Environment: We prioritize learning, innovation, and pushing the limits of what’s possible.

What We Offer

  • Competitive salary
  • Comprehensive health, dental, and vision coverage
  • 401(k) with company match
  • Flexible PTO and hybrid work arrangement in Atlanta

Location

This role is based in Atlanta, GA, with a hybrid schedule.

SpaceX

Site Reliability Engineer (Manufacturing Infrastructure)

🕒 27 days ago
SpaceX👥 10,000+ employees🏢 Aviation And Aerospace Component Manufacturing🤝 B2B

As a Site Reliability Engineer for Manufacturing Infrastructure, you will enhance the reliability and scalability of systems supporting SpaceX's manufacturing processes.

Site Reliability EngineeringDevOpsInfrastructure As CodeLinux
Anyscale

Site Reliability Engineer, Platform Infrastructure (Foundations)

📅 Aug 26
Anyscale👥 201 - 500 employees🏢 Computer Software

As a Site Reliability Engineer, you will design and optimize critical infrastructure for distributed AI applications, ensuring high performance and reliability in cloud environments.

KubernetesCloud-native TechnologiesAWSAzure
Qualysoft

SRE Engineer - CyberSecurity

📅 Jul 28
Qualysoft👥 201 - 500 employees🏢 Information Technology And Services

As an SRE Engineer in CyberSecurity, you will implement and maintain software systems, automate processes, and ensure high availability while collaborating with development teams.

Linux/unix AdministrationAutomation ToolsContainerizationOrchestration Technologies

Trusted by Remote Workers