Remote Jobs RockRemote Jobs Rock

Senior Site Reliability Engineer

📅 Jul 23
ObservabilityOpentelemetryPrometheusGrafana

📜 Description

  • Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.
  • Contribute to and maintain Tulip's triage & remediation processes as a player/coach.
  • Perform incident response and debug production issues across the entire stack.
  • Design, build, and maintain the core infrastructure & tooling used by all of Tulip's engineering teams.

🛠️ Requirements

  • You’re excited about setting the technical agenda and coming up with novel, broad ideas
  • You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring
  • You know what a good SLA looks like, and can teach others how to spot one
  • You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication
  • 5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack)
  • Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale
  • Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy
  • Experience working with time-series data, ideally using promQL

Benefits

  • Direct impact on product and culture.
  • Company equity.
  • Dental
  • Vision
  • Life Insurance
  • AD&D Insurance
  • Flexible work schedule and unlimited vacation policy.
  • Virtual company events and happy hours.
  • Fitness subsidies.
Full job description

This role is located in Somerville, MA - We are a hybrid work environment and are in the office 3+ days/per week.

Tulip, the leader in AI-native frontline operations, is helping companies around the world equip their workforce with composable, connected apps, leading to higher quality work, improved efficiency, and end-to-end traceability across operations. Tulip’s cloud-native, no-code platform, powered by embedded AI, is driving the digital transformation of industrial environments through composable, human-centric solutions that go beyond disrupting the Manufacturing Execution System (MES) category.

A spinoff out of MIT, Tulip is headquartered in Somerville, MA, with offices in Germany, Hungary, Singapore, and Israel. Tulip has been recognized as a World Economic Forum Global Innovator, a 2024 Deloitte Technology Fast award winner, one of Energage’s Top Workplaces USA, and one of Built In Boston’s “Best Places to Work” and “Best Midsize Places to Work.”

About You:

  • You can reason about systems at scale: their edge cases, failure modes, and life cycles across
  • You’re excited about setting the technical agenda and coming up with novel, broad ideas
  • You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring
  • You know what a good SLA looks like, and can teach others how to spot one
  • You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication

What skills do I need?

  • 5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack)
  • Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale
  • Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy
  • Experience working with time-series data, ideally using promQL

Key Responsibilities:

  • Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.
  • Contributing to and maintaining Tulip's triage & remediation processes as a player / coach
  • Perform incident response and debug production issues across the entire stack
  • Design, build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering teams

Tech Stack:

  • TS and Go Services running on Kubernetes
  • MongoDB and PostGres DBs
  • Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry Observability tooling

Key Collaborators:

  • Engineering
  • Edge
  • DevOps
  • Hardware

Working At Tulip

We know even great candidates experience imposter syndrome. Even if you don’t match every requirement, applying gives you the opportunity to be considered.

We’re building a strong, diverse team that values hard work, families, and personal well-being. Benefits of working with us include:

  • Direct impact on product and culture
  • Company equity
  • Competitive benefits package including Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Commuter Benefits, Parental Leave, and 401(K)
  • Flexible work schedule and unlimited vacation policy
  • Virtual company events and happy hours
  • Fitness subsidies

We are an equal opportunity employer. At Tulip, we celebrate all. Qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. Help us build an inclusive community that will transform frontline operations.

The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range posted, actual compensation will be determined depending on multiple factors including job-related knowledge & skills, experience, business needs, geographical location, market compensation data, and internal equity. Expected compensation ranges for this role may change over time. The salary range for this position is $160,000 - $200,000 per year.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Please note that we may use AI-based tools to support parts of our hiring process. All data processing is carried out in compliance with local data protection laws, ensuring all personal candidate information is handled securely and ethically.

Crunchyroll

Senior Site Reliability Engineer

🕒 7 days ago
Crunchyroll👥 1001 - 5000 employees🏢 Entertainment

As a Senior Site Reliability Engineer, you will enhance the reliability, scalability, and security of Crunchyroll's data platforms while driving modern SRE practices and cross-functional collaboration.

Site Reliability EngineeringKubernetesGoogle Cloud PlatformInfrastructure As Code
Lambda

Senior Site Reliability Engineer - Fleet

🕒 13 days ago
Lambda👥 501 - 1000 employees🏢 Technology

As a Senior Site Reliability Engineer, you will build and operate monitoring systems, automate HPC cluster management, and troubleshoot complex infrastructure issues.

Site Reliability EngineeringHPC EngineeringDevOpsAI Infrastructure
Mirantis

Senior Site Reliability Engineer (Golang / Kubernetes)

🕒 14 days ago
Mirantis👥 501 - 1000 employees🏢 Computer Software

As a Senior Site Reliability Engineer, you will define and measure reliability for a GPU-accelerated AI platform, owning service-level indicators and objectives while collaborating across teams.

Site Reliability EngineeringKubernetesGoPython
Earnin

Staff Site Reliability Engineer

🕒 3 days ago
Earnin👥 501 - 1000 employees🏢 Financial Services

Lead EarnIn's reliability maturity by implementing an AI-first operating model to enhance incident response, operational quality, and reliability practices across engineering teams.

Site Reliability EngineeringAI-assisted WorkflowsIncident ResponseSlis

Trusted by Remote Workers