Remote Jobs RockRemote Jobs Rock

Staff AI Scheduling & Orchestration Engineer

๐Ÿ•’ 11 days ago
KubernetesAI Workload ExecutionDistributed Systems EngineeringBatch Scheduling Architectures

๐Ÿ“œ Description

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
  • Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
  • Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
  • Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

๐Ÿ› ๏ธ Requirements

  • Bachelorโ€™s or Masterโ€™s degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
  • Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
  • Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
  • Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
  • Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.
Full job description

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

We are seeking a Staff AI Scheduling & Orchestration Engineer to lead the workload placement logic that defines our AI-native NeoCloud platform. Standard Kubernetes scheduling is insufficient for the demands of large-scale AI; you will be responsible for eliminating "GPU stranding" and maximizing utilization across our expensive compute fleets. This role is pivotal in building a high-performance scheduling fabric that understands the physical realities of our hardwareโ€”from NVLink-connected GPU topologies to InfiniBand interconnects. You will work at the intersection of distributed systems and AI, driving the architectural decisions that enable our platform to handle massive-scale distributed training and inference jobs with industry-leading efficiency.

Key Responsibilities

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
  • Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
  • Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
  • Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
  • Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments.
  • Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer.

Qualifications

  • Bachelorโ€™s or Masterโ€™s degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
  • Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
  • Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
  • Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
  • Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

๐Ÿ•’ yesterday

You will design, develop, and maintain machine learning models and optimization algorithms to enhance workforce management processes and drive business success.

Machine LearningOptimization AlgorithmsPythonAWS
CoreWeave

Staff Software Engineer, Infrastructure - Marimo

CoreWeave๐Ÿ‘ฅ 10,000+ employees๐Ÿข IT Services And IT Consulting๐Ÿค B2B
๐Ÿ”ฅ 4 hours ago

As a Staff Engineer on Marimo's molab team, you will co-design and implement the backend architecture of molab, focusing on high availability, low latency, and system stability.

Software EngineeringBackend ArchitectureHigh Availability SystemsLow Latency Communication
JFrog

Senior Solutions Engineer

JFrog๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development๐Ÿค B2B
๐Ÿ”ฅ 9 hours ago

As a Senior Solutions Engineer, you will serve as a trusted technical advisor, helping enterprise customers modernize their software delivery practices and overcome complex challenges in DevOps and security.

๐Ÿ“ ๐Ÿ‡น๐Ÿ‡ผ Taiwan - Remote๐Ÿ’ผ Full-Time๐ŸŽธ Senior๐Ÿ’Ž Solutions Engineer๐Ÿ“ข ๐Ÿ‡จ๐Ÿ‡ณ Mandarin Required๐Ÿ“ข ๐Ÿ‡ฌ๐Ÿ‡ง English Required
Solutions EngineeringSales EngineeringDevOpsSoftware Engineering
๐Ÿ•’ yesterday

As a Senior Security Engineer at NVIDIA, you will architect and build foundational security systems for the DGX Cloud, ensuring the integrity of AI infrastructure.

Infrastructure SecuritySoftware EngineeringCloud-native ArchitectureKubernetes

Trusted by Remote Workers