As a Staff Site Reliability Engineer, you will enhance the reliability and performance of Replit's infrastructure by implementing automation, leading incident management, and mentoring the engineering team.
Senior Site Reliability Engineer (SRE)
📜 Description
- Demonstrate proficiency in problem analysis.
- Collaborate with teams to offer pragmatic solutions and be accountable for their execution.
- Create, design, develop, and operate the shared infrastructure platform.
- Improve the developer experience by treating the platform as a product.
- Train and support software engineers on DevOps and SRE practices.
- Stay updated on new technologies and provide relevant solutions.
🛠️ Requirements
- Significant site reliability engineering experience.
- Value code simplicity, quality, and security.
- Ability to design pragmatic architectures to solve problems at scale.
- Desire to work in a fast, high-growth startup environment.
✨ Benefits
- Opportunity to join an international tech company.
- Collaborative work environment that values innovation and creativity.
- Competitive salary and benefits package.
- Professional development and career growth opportunities.
Full job description
Key Responsibilities:
✨It will be a perfect match if you:
⚙️ Our Tech stack
💡What’s in it for you ?
Similar jobs
Search more Site Reliability Engineer jobsSenior Site Reliability Engineer - FedRAMP
As a Senior Site Reliability Engineer, you will ensure the availability and performance of production SaaS services in Azure and AWS, focusing on reliability, incident response, and automation.
As a Senior Site Reliability Engineer, you will enhance the reliability and performance of Replit's infrastructure by implementing automation, monitoring solutions, and incident management practices.
Senior Site Reliability Engineer (Golang / Kubernetes)
As a Senior Site Reliability Engineer, you will define and measure reliability for a GPU-accelerated AI platform, owning service-level indicators and objectives while collaborating across teams.
As a Senior Site Reliability Engineer, you will build and operate monitoring systems, automate HPC cluster management, and troubleshoot complex infrastructure issues.
