Senior Site Reliability Engineer (CI-CD/CTAP/Delivery team)
As a Senior Site Reliability Engineer, you will design and operate large-scale cloud infrastructure, enhance service reliability, and drive automation for Okta's Emerging Products Group.
Tracksuit exists to help marketers prove their brand building is working. We give teams the data they need to make smarter decisions, convince stakeholders, defend budgets, and track their progress. The brand tracking industry is dominated by 100-page reports, static data and big price-tags. We're doing things differently by being built for the modern marketer: always-on, accessible, and approachable.
We're now tracking more than 1,000 brands across 25 countries globally. With offices in Auckland, Sydney, London and New York City, we're scaling fast with a brilliant team of collaborative and ambitious humans. Our culture is defined by "high care, high performance". We strive to be the best and look after each other while we do it.
Are you our next Senior Site Reliability Engineer?
Weโre on the lookout for a Senior Site Reliability Engineer to join Tracksuit, based in our Auckland office.
Youโll set how reliability, observability and security work at Tracksuit, from the AWS infrastructure underneath the platform through to the agents, MCP servers and model-backed features running on top. Teams build and run their own services, and the SRE team makes it straightforward for them to do that well, through the golden paths, guardrails, tooling and defaults that make the reliable way the easy way, plus the support to lean on when something does break.
Some of the tools we use: AWS, ECS, Terraform and Terragrunt, GitHub Actions, Datadog, Claude Code, Linear, Notion, Postgres, DynamoDB and Snowflake.
As a Senior Site Reliability Engineer, you will design and operate large-scale cloud infrastructure, enhance service reliability, and drive automation for Okta's Emerging Products Group.
As a Senior Site Reliability Engineer, you will own the reliability of core production systems, ensuring they perform optimally under real traffic and contribute to a seamless user experience.
As Fingerprint's first dedicated Site Reliability Engineer, you will enhance platform reliability, implement measurable standards, and coach teams on operational excellence.
As a Site Reliability Engineer, you will manage the entire request path from user to service, ensuring optimal performance and reliability of traffic infrastructure at scale.
You'll lead the deployment and verification of safety systems for AI model launches, ensuring safeguards are effectively configured and operational across various platforms.