Remote Jobs RockRemote Jobs Rock

Senior Site Reliability Engineer

๐Ÿ•’ 8 days ago
AWSTerraformECSInfrastructure As Code

๐Ÿ“œ Description

  • Build and maintain cloud infrastructure to ensure resilience, security, and scalability as defaults.
  • Provide self-serve observability tools for monitoring, alerting, logging, and tracing.
  • Facilitate incident management with effective tooling and runbooks for quick debugging.
  • Set identity and least-privilege models for non-human actors to manage access safely.
  • Document infrastructure designs and operational procedures for both people and agents.
  • Coach engineers on reliability practices while balancing feature delivery with platform needs.

๐Ÿ› ๏ธ Requirements

  • Experience running incident command on production incidents and owning a platform during scaling.
  • Deep infrastructure skills: Infrastructure as Code (Terraform, Terragrunt), containers (ECS), AWS or GCP/Azure.
  • Strong observability skills with tools like Datadog and a grasp of incident management methodology.
  • Security-minded with knowledge of networking and cloud architecture fundamentals.
  • Experience working with coding agents and understanding their applications.
  • Ability to communicate technical complexity with a focus on business outcomes.
  • Collaborative team player with a supportive attitude.
  • Bonus: Passion for marketing, design, and user experience.
Full job description

Tracksuit exists to help marketers prove their brand building is working. We give teams the data they need to make smarter decisions, convince stakeholders, defend budgets, and track their progress. The brand tracking industry is dominated by 100-page reports, static data and big price-tags. We're doing things differently by being built for the modern marketer: always-on, accessible, and approachable.

We're now tracking more than 1,000 brands across 25 countries globally. With offices in Auckland, Sydney, London and New York City, we're scaling fast with a brilliant team of collaborative and ambitious humans. Our culture is defined by "high care, high performance". We strive to be the best and look after each other while we do it.

Are you our next Senior Site Reliability Engineer?

Weโ€™re on the lookout for a Senior Site Reliability Engineer to join Tracksuit, based in our Auckland office.

Youโ€™ll set how reliability, observability and security work at Tracksuit, from the AWS infrastructure underneath the platform through to the agents, MCP servers and model-backed features running on top. Teams build and run their own services, and the SRE team makes it straightforward for them to do that well, through the golden paths, guardrails, tooling and defaults that make the reliable way the easy way, plus the support to lean on when something does break.

Why this role is special

  • ๐Ÿงฑ Platform-wide impact: You set how reliability, observability and operational readiness work across the whole platform, and every team shipping on it feels the difference.
  • ๐Ÿค– Agentic systems in production: Youโ€™ll set the guardrails, identity and blast-radius controls for agents operating against real systems and data, and work out what could go wrong before it does.
  • ๐Ÿ›  Golden paths: Youโ€™ll make the platform legible to agents as well as people, through golden paths, MCP servers, skills and documentation that both can actually use.
  • ๐Ÿš€ Startup pace with scaleup reach: Five years in, weโ€™re working with over 1,000 brands across AU, NZ, USA and the UK, with plenty of scaling still ahead.
  • What you'll do

  • Build and maintain the cloud infrastructure and paved paths teams ship on, so resilience, security and scalability come as defaults rather than as decisions each team makes on its own
  • Give teams observability they can self-serve: monitoring, alerting, logging and tracing, plus automation for provisioning, deployments and the operational work nobody should be doing by hand
  • Make incidents easier to handle: the tooling, runbooks and practice that let whoever is closest to the problem debug it quickly, and post-incident reviews that turn into actual changes
  • Set the identity and least-privilege model for non-human actors, including agents, CI and MCP servers, so teams can give agents real access without real risk
  • Document infrastructure designs and operational procedures in a form both people and agents can act on, including machine-readable runbooks
  • Set the standard for what is safe to ship, and make it easy to meet through guardrails and checks in the pipeline rather than through gatekeeping
  • Give teams visibility and controls over cloud and inference spend, so the cost of what they run is something they can see and act on
  • Coach engineers in reliability practices, and work with Engineering and Product to balance feature delivery against platform needs
  • That's the role, so who are you?

  • Youโ€™ve run incident command on real production incidents, and youโ€™ve owned a platform through a meaningful scaling step
  • Deep infrastructure skills: Infrastructure as Code (Terraform, Terragrunt, CDK), containers (ECS), cloud platforms (AWS preferred, or GCP/Azure), CI/CD, and scripting in Python, TypeScript or Bash.
  • Strong on observability: Datadog or similar, distributed tracing and structured logging, and a solid grasp of incident management methodology.
  • Security-minded: Networking and cloud architecture fundamentals, plus an understanding of the security model for agentic systems, including credential handling, least privilege and prompt injection.
  • Works with agents: You use coding agents in your own work (Claude Code or equivalent) and have a point of view on where they help and where they donโ€™t.
  • Product-oriented: You navigate technical complexity with business outcomes in mind, and can clearly communicate system health, risks and trade-offs to technical and non-technical people.
  • Kind: This is a collaborative role, so we ultimately want a great, supportive team player.
  • Bonus points if youโ€™re passionate about marketing, design, and building exceptional user experiences.
  • Some of the tools we use: AWS, ECS, Terraform and Terragrunt, GitHub Actions, Datadog, Claude Code, Linear, Notion, Postgres, DynamoDB and Snowflake.

    Okta

    Senior Site Reliability Engineer (CI-CD/CTAP/Delivery team)

    ๐Ÿ•’ 21 days ago
    Okta๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development

    As a Senior Site Reliability Engineer, you will design and operate large-scale cloud infrastructure, enhance service reliability, and drive automation for Okta's Emerging Products Group.

    AWSGCPKubernetesTerraform
    Fingerprint

    Senior Site Reliability Engineer

    ๐Ÿ•’ 11 days ago
    Fingerprint๐Ÿ‘ฅ 51 - 200 employees๐Ÿข Computer Software

    As a Senior Site Reliability Engineer, you will own the reliability of core production systems, ensuring they perform optimally under real traffic and contribute to a seamless user experience.

    Site Reliability EngineeringCloud InfrastructureAWSTerraform
    Fingerprint

    Staff Site Reliability Engineer

    ๐Ÿ•’ 15 days ago
    Fingerprint๐Ÿ‘ฅ 51 - 200 employees๐Ÿข Computer Software

    As Fingerprint's first dedicated Site Reliability Engineer, you will enhance platform reliability, implement measurable standards, and coach teams on operational excellence.

    Site Reliability EngineeringSLI/SLO DesignIncident ManagementDistributed Systems
    Taboola

    Site Reliability Engineer, Traffic Infrastructure

    ๐Ÿ•’ 15 days ago
    Taboola๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software

    As a Site Reliability Engineer, you will manage the entire request path from user to service, ensuring optimal performance and reliability of traffic infrastructure at scale.

    Site Reliability EngineeringTraffic InfrastructureNetworkingLinux
    Anthropic

    Staff+ Site Reliability Engineer, Safeguards ML Infra

    ๐Ÿ•’ 25 days ago
    Anthropic๐Ÿ‘ฅ 10,000+ employees๐Ÿข Research Services๐Ÿค B2B

    You'll lead the deployment and verification of safety systems for AI model launches, ensuring safeguards are effectively configured and operational across various platforms.

    Production Change ManagementDeploy PipelinesConfig Management SystemsCanary Analysis

    Trusted by Remote Workers