As a Senior Site Reliability Engineer, you'll enhance reliability, observability, and security across Tracksuit's platform, enabling teams to deliver resilient services efficiently.
Site Reliability Engineer, Traffic Infrastructure
📜 Description
- Own the full path a request takes from user to service, including CDN configuration and load-balancing policy.
- Move proxy and CDN configuration into code with review, testing, and staged rollout.
- Lead cross-layer investigations to resolve complex issues.
- Plan capacity and drain procedures to maintain user experience during outages.
- Participate in on-call rotation and lead incident response efforts.
🛠️ Requirements
- 5+ years running production organization with internet-facing traffic infrastructure at large scale.
- Deep understanding of TCP/IP stack including BGP, anycast, and DNS-based traffic steering.
- Hands-on experience with L7 proxies in production: Nginx, HAProxy, or Envoy/Istio.
- Production experience with commercial CDNs such as Fastly or Cloudflare.
- Fluency in Linux networking internals and packet-level diagnosis.
- Strong coding skills beyond scripting for building tooling.
✨ Benefits
- (health, etc.), a fully stocked kitchen, and location-specific perks (gym partnerships, parking).
- Flexibility: We offer a hybrid work schedule (3 days in-office) with flexibility.
- Work with global leaders: Our publisher partners include Yahoo, Fox Sports, NBCU, ESPN, and CBS
- Ready to realize your potential?
Full job description
As a Site Reliability Engineer in the R&D Infrastructure department in our Tel Aviv Office, you'll own the path every request takes before it reaches a service: The CDN edge, the anycast VIPs behind it, a large fleet of open source load balancers and service mesh components. At peak that chain carries millions requests per second on the way to serving tens of billions of recommendations a day.
We need a networking specialist who doesn't stop at the wire. You bring in depth networking knowledge, native Linux fluency and modern coding practices with the diagnostic intuition to debug complex interactions between BGP enabled nodes, reverse proxies and application logs.
To thrive in this role, you'll need:
- 5+ years running production organization with internet-facing traffic infrastructure at very large scale, with deep understanding of TCP/IP stack (BGP, anycast, ECMP, DNS-based traffic steering, TCP, TLS, HTTP/2 and HTTP/3).
- Deep hands-on ownership of L7 proxies in production: Nginx, HAProxy or Envoy/Istio, covering configuration at scale, performance tuning and the class of bugs that only appear under real load.
- Production experience with a commercial CDN such as Fastly, Cloudflare, Akamai or CloudFront, including edge logic, cache policy and origin protection.
- Fluency in Linux networking internals and packet-level diagnosis: tcpdump, conntrack, nftables/iptables, socket and TCP tuning.
- Coding ability that goes past scripting. You will be building tooling and not only editing configuration files.
- Kubernetes networking depth covering CNI internals, kube-proxy, ingress and east-west traffic policy.
Bonus points if you have:
- A background in traffic observability, whether flow telemetry, real-user monitoring or synthetic probing across regions.
How you'll make an impact:
- Own the full path a request takes from user to service: CDN configuration, TLS termination, load-balancing policy and failover behavior.
- Move our proxy and CDN configuration into code with review, testing and staged rollout replacing hand edits in production.
- Lead the cross-layer investigations nobody else can close from an anycast withdrawal down to a single misbehaving upstream.
- Plan capacity and drain procedures for the traffic tier so that losing a POP or an entire data center stays invisible to users.
- Take part in the on-call rotation, lead incident response and turn post-mortems into permanent fixes in the traffic layer.
Why Taboola?
If you ask Taboolars what they love about working here, they’ll tell you that they’ve been empowered to realize their full potential while growing and learning with smart, talented people:
- Adam Singolda, Taboola Founder and CEO says: “You can copy anything from another business, but you can’t copy a company’s culture.”
- Well-being: Enjoy comprehensive benefits (health, etc.), a fully stocked kitchen, and location-specific perks (gym partnerships, parking).
- Flexibility: We offer a hybrid work schedule (3 days in-office) with flexibility.
- Work with global leaders: Our publisher partners include Yahoo, Fox Sports, NBCU, ESPN, and CBS, while our clients include major global brands.
Ready to realize your potential?
Taboola is an equal opportunity employer and we value diversity in all forms. We are committed to creating an inclusive environment for all employees and believe such an environment is critical for success. Employment is decided on the basis of qualifications, merit, and business need.-
Learn more about #TaboolaLife on LinkedIn, Facebook, Instagram, X, YouTube, & the Taboola Life Blog.
About Taboola
Taboola empowers businesses to grow through performance advertising technology that goes beyond search and social and delivers measurable outcomes at scale.
Taboola works with thousands of businesses who advertise directly on Realize, Taboola’s powerful ad platform, reaching approximately 600M daily active users across some of the best publishers in the world. Publishers like NBC News, Yahoo, and OEMs such as Samsung, Xiaomi and others use Taboola’s technology to grow audience and revenue, enabling Realize to offer unique data, specialized algorithms, and unmatched scale.
#LI-Hybrid
#LI-ID
Similar jobs
Search more Site Reliability Engineer jobsSenior Site Reliability Engineer (CI-CD/CTAP/Delivery team)
As a Senior Site Reliability Engineer, you will design and operate large-scale cloud infrastructure, enhance service reliability, and drive automation for Okta's Emerging Products Group.
Staff+ Site Reliability Engineer, Safeguards ML Infra
You'll lead the deployment and verification of safety systems for AI model launches, ensuring safeguards are effectively configured and operational across various platforms.
As the Senior Site Reliability Engineer for Project Volcano, you will define and drive the reliability posture, ensuring high availability and performance of the developer platform.
Seeking a senior Kubernetes-focused DevOps/SRE engineer to manage the developer platform and a customer-facing production region for enterprise GPU infrastructure.
