Senior Site Reliability Engineer (Golang / Kubernetes)
As a Senior Site Reliability Engineer, you will define and measure reliability for a GPU-accelerated AI platform, owning service-level indicators and objectives while collaborating across teams.
About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/
Job Description
NeoCloud is building an AI-operated GPU cloud — and because it is a customer-facing cloud service, reliability is the product. Tenants run mission-critical training, fine-tuning, and inference workloads on our GPU infrastructure and trust us with their SLAs. In this role you own the reliability of the customer-facing GPU cloud service end-to-end: from tenant onboarding and service provisioning, through workload execution, incident response, and post-incident recovery. You are the SRE who stands between raw infrastructure and the customer's experience — designing the observability, automation, and operational practices that make a 10,000-GPU cloud feel simple and dependable to the tenants who depend on it.
What You'll Own
Customer-Facing Ownership
Feed the AIOps Substrate
What Success Looks Like in Year 1
Requirements
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
As a Senior Site Reliability Engineer, you will define and measure reliability for a GPU-accelerated AI platform, owning service-level indicators and objectives while collaborating across teams.
As a Senior Site Reliability Engineer on the Fuse team, you will lead the reliability and operability of complex product-data systems, ensuring dependable data management and observability across Bloomreach's platforms.
Job Application for Senior Site Reliability Engineer for Fuse Team at Bloomreach.
Design, deploy, and operate the control plane for an AI-operated GPU cloud, ensuring automated remediation and efficient management of Kubernetes clusters optimized for GPU workloads.
As a Senior Site Reliability Engineer, you will enhance the reliability, scalability, and security of Crunchyroll's data platforms while driving modern SRE practices and cross-functional collaboration.