As a Site Reliability Engineer, you will design and operate advanced retrieval pipelines, collaborate on innovative AI solutions, and enhance agent performance through real-world feedback.
Staff Engineer, Site Reliability Engineering
📜 Description
- Design and develop a scalable SRE ecosystem following SRE and DevSecOps best practices
- Develop reusable TypeScript scaffolding libraries for cloud-native components
- Build and enhance solutions using AWS, EKS, Kubernetes and Infrastructure as Code
- Drive automation, reliability, scalability and operational excellence across microservices
- Define and implement SRE and DevOps best practices across applications
- Collaborate with technology, product, operations and functional teams on SRE initiatives
🛠️ Requirements
- Strong experience in TypeScript development, coding and design patterns
- Hands-on experience with AWS cloud and cloud-native technologies
- Experience with AWS CDK, Terraform and Infrastructure as Code (IaC)
- Experience with Kubernetes, Docker and Amazon EKS
- Experience in SRE, DevOps, scalability, reliability and cloud automation
- Experience with CI/CD tools such as Jenkins and Git
- Knowledge of observability and monitoring tools such as CloudWatch, Splunk and Dynatrace
- Knowledge of service mesh technologies such as Istio
- Experience developing reusable scaffolding libraries and cloud-native components
- Ability to analyse application and infrastructure dependencies across microservices environments
Full job description
Company Description
👋🏼We're Nagarro.
We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale across all devices and digital mediums, and our people exist everywhere in the world (18,700 + experts across 39 countries, to be exact). Our work culture is dynamic and non-hierarchical. We're looking for great new colleagues. That's where you come in.
Job Description
REQUIREMENTS:
• Total experience 5.5+ years
• Strong experience in TypeScript development, coding and design patterns
• Hands-on experience with AWS cloud and cloud-native technologies
• Experience with AWS CDK, Terraform and Infrastructure as Code (IaC)
• Experience with Kubernetes, Docker and Amazon EKS
• Experience in SRE, DevOps, scalability, reliability and cloud automation
• Experience with CI/CD tools such as Jenkins and Git
• Knowledge of observability and monitoring tools such as CloudWatch, Splunk and Dynatrace
• Knowledge of service mesh technologies such as Istio
• Experience developing reusable scaffolding libraries and cloud-native components
• Ability to analyse application and infrastructure dependencies across microservices environments
• Experience working with distributed teams across multiple time zones
• Excellent communication, presentation and stakeholder collaboration skills
RESPONSIBILITIES:
• Design and develop a scalable SRE ecosystem following SRE and DevSecOps best practices
• Develop reusable TypeScript scaffolding libraries for cloud-native components
• Build and enhance solutions using AWS, EKS, Kubernetes and Infrastructure as Code
• Drive automation, reliability, scalability and operational excellence across microservices
• Define and implement SRE and DevOps best practices across applications
• Collaborate with technology, product, operations and functional teams on SRE initiatives
• Analyse business requirements and their impact across applications and cloud systems
• Establish standardized and automated onboarding paths for applications onto the SRE platform
• Evaluate emerging technologies and define strategies for cloud and SRE adoption
• Implement CI/CD, observability, monitoring and service mesh capabilities
• Identify and address reliability, scalability and operational challenges
• Provide technical guidance and support to distributed engineering teams
Qualifications
Bachelor’s or master’s degree in computer science, Information Technology, or a related field.
Additional Information
Videos To Watch
Similar jobs
Search more Site Reliability Engineer jobsAs a Senior Site Reliability Engineer, you will own the reliability of core production systems, ensuring they perform optimally under real traffic and contribute to a seamless user experience.
Contribute to infrastructure automation and operational resilience as a Staff Software Engineer - SRE & AIOps, implementing systems that enhance reliability and reduce operational toil.
Senior Site Reliability Engineer, AI Agents & Automation
Join our Site Reliability & Infrastructure Engineering team as a Senior Site Reliability Engineer, where you'll ensure the reliability and health of cloud applications while driving efficiency and innovation.
As Fingerprint's first dedicated Site Reliability Engineer, you will enhance platform reliability, implement measurable standards, and coach teams on operational excellence.
