The Principal Site Reliability Engineer will provide technical leadership to enhance the reliability, scalability, and security of cloud platforms while driving large-scale initiatives across multiple teams.
Principal Site Reliability Engineer, Platform
📜 Description
- Architect, scale, and own essential infrastructure.
- Build and maintain a Kubernetes-based platform supporting multiple teams and services.
- Build backend services (Golang) to support autonomous systems.
- Partner with product teams to launch new products on the platform.
- Grow high availability infrastructure while maintaining uptime metrics.
- Design and maintain observability infrastructure and dashboards.
🛠️ Requirements
- Min. of 8 years of deep experience in building and maintaining infrastructure for data-intensive, high-availability applications, including six years building and maintaining public cloud solutions.
- Deep understanding of cloud orchestration tools such as Kubernetes and Terraform.
- Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.
- Deep understanding of securing cloud infrastructure (preferably AWS and Kubernetes).
- Deep experience in one or more of the following languages: Golang (preference), Python, JavaScript, Rust.
- Deep experience in CI/CD tooling (GitHub Actions, ArgoCD, ArgoCD Image Updater, Artifactory).
- You are interested in robotic applications and developing software that assists robots.
- You are excited about robotics and the future of automation.
- You are a self-starter with infectious enthusiasm, energy, and problem-solving abilities.
Full job description
We’re Blue River, a team of innovators driven to create intelligent machinery that solves monumental problems for our customers. We empower our customers – farmers, construction crews, and foresters - to implement safer and more sustainable solutions, driving increased profitability with less reliance on scarce labor. We believe that focusing on the small stuff – pixel-by-pixel and task-by-task - leads to big gains.
Blue River Technology aligns with John Deere’s vision to “innovate on behalf of humanity” by quickly identifying and solving high-value, high-uncertainty challenges in AI, machine learning, computer vision, and robotics. BRT acts as a research and development flywheel, building not only new products but also new platforms that reliably create value for both Deere and its customers. From fully autonomous machines to highly precise farming equipment, BRT and Deere are partnering to create technical breakthroughs in industries like agriculture and construction.
Our people are at the heart of what we do. Through cross-disciplinary collaboration, this mission-driven team is eager to define the new frontier of robotics. We are always asking hard questions, rapidly iterating, and getting our boots in the field to figure it out. We won’t give up until we’ve made a tangible and positive impact on the planet!
Blue River Technology is based in Santa Clara, CA.
Summary
We are seeking a Principal Site Reliability Engineer to join the Platform organization, which accelerates company-wide adoption and scaling of automation and robotics. The Platform's product is a set of API services and infrastructure designed to overcome scaling hurdles, such as operational complexity and system exceptions, thereby enabling the rapid launch and scaling of new autonomy innovations and products at Blue River. In this role, you will join a fun, fast-moving engineering team to drive architectural decisions, mentor engineers across teams, and shape our platform's direction.
- Employment Type: Full-Time
- Work Location: Remote in the United States.
- Visa sponsorship is possible for this position.
Job Responsibilities
A combination, not necessarily all-inclusive, of the following:
- Architect, scale, and own essential infrastructure.
- Build and maintain a Kubernetes-based platform supporting multiple teams and services.
- Build backend services (Golang) to support autonomous systems.
- Partner with product teams to launch new products on the platform.
- Grow our high availability infrastructure while maintaining key metrics such as uptime.
- Build tooling to support our platform and development teams.
- Perform end-to-end performance analysis, identify areas for improvement, and implement robust solutions.
- Work with cloud vendors and external technical support for upgrades and rapid problem resolution.
- Participate in on-call rotation, triaging and resolving production incidents with thorough root cause analysis and postmortem documentation.
- Design and maintain observability infrastructure, dashboards, alerts, and log aggregation to ensure visibility into platform health and service performance.
- Collaborate with the security team to conduct regular risk assessments.
- Maintain the risk register and develop and implement mitigation plans.
- Assess intrusion detection alerts. Improve systems and services that digest threat feeds.
- In collaboration with IT and purchasing teams, establish and maintain the payment process for each SaaS service.
Required Experience and Skills
- Min. of 8 years of deep experience in building and maintaining infrastructure for data-intensive, high-availability applications, including six years building and maintaining public cloud solutions.
- Deep understanding of cloud orchestration tools such as Kubernetes and Terraform.
- Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.
- Deep understanding of securing cloud infrastructure (preferably AWS and Kubernetes).
- Deep experience in one or more of the following languages: Golang (preference), Python, JavaScript, Rust.
- Deep experience in CI/CD tooling (GitHub Actions, ArgoCD, ArgoCD Image Updater, Artifactory).
Preferred Experience and Skills
- You are interested in robotic applications and developing software that assists robots.
- You are excited about robotics and the future of automation.
- You are a self-starter with infectious enthusiasm, energy, and problem-solving abilities.
At Blue River, your base pay is one part of your total compensation package. For this position, the reasonably expected pay range is between $174,000 - $305,000/year for the level at which this job has been scoped. Your base pay will depend on several factors, including your experience, qualifications, education, location, and skills. This position is also eligible for an annual performance bonus and a competitive benefit package. During the recruitment process, we may identify an alternative role or level to which you are more suited. If your ideal role at Blue River differs from the advertised position, we will provide an updated pay range as soon as possible during the hiring process.
We’re passionate about creating an inclusive workplace that promotes and values diversity. While we have more work to do to advance diversity and inclusion, we’re investing in our programs, including recruiting, mentorship, career development, and learning & development, to ensure they support our Diversity, Equity, and Inclusion goals. We support each employee in living a full life, enabling a thriving career, and accomplishing a meaningful, challenging mission while collaborating with incredible people. We are dedicated to building a diverse and inclusive workplace, so if you’re excited about this role but your experience doesn’t align completely with the job description, we encourage you to apply anyway.
We are an equal-opportunity employer and do not discriminate based on race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive other benefits and privileges of employment. Please contact us to request an accommodation.
#LI-AN1
Similar jobs
Search more Site Reliability Engineer jobsStaff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)
As a Staff Site Reliability Engineer, you will architect, scale, and maintain cloud infrastructure, ensuring high availability and security for large-scale distributed systems.
As a Staff Site Reliability Engineer, you will enhance the Data Office's infrastructure, focusing on reliability, scalability, and developer experience across the Data Business Unit.
As a Staff Site Reliability Engineer, you will enhance the performance and reliability of Fivetran's infrastructure, ensuring robust deployment pipelines and effective incident response.
Staff Site Reliability Engineer
As a Site Reliability Engineer at Ping Identity, you will design, deploy, and maintain cloud infrastructure, ensuring high-quality software solutions through automated CI/CD pipelines.
