Remote Jobs RockRemote Jobs Rock

Staff Slurm Cluster & HPC Engineer

🕒 10 days ago
📍 🇺🇸 San Jose - Remote💰 $180,000 – $260,000🎸 Senior🤖 AI Engineer📢 🇬🇧 English Required💼 Full-Time
SlurmHPCKubernetesGPU Scheduling

📜 Description

  • Design, deploy, and operate production Slurm clusters on bare metal and VMs with high availability.
  • Model physical fabric for topology-aware scheduling and validate placement quality for GPU fabrics.
  • Own multi-tenant scheduling policies, including account management and resource limits.
  • Lead implementation of Slinky slurm-operator for Kubernetes integration and shared resource management.
  • Utilize Slurm's cloud mechanisms to manage GPU node capacity between batch and Kubernetes workloads.
  • Build health-check systems for cluster reliability and automate cluster delivery using Terraform/Ansible.

🛠️ Requirements

  • 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
  • Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
  • Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
  • Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
  • Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
  • Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
  • Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
  • Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders
Reddit

Senior Staff Machine Learning Systems Engineer, Ads ML Platform

Reddit👥 1001 - 5000 employees🏢 Online Community/social Media
🔥 3 hours ago

Lead the technical strategy for the Ads ML engineer lifecycle, focusing on feature development and training iteration to enhance ML engineers' productivity.

Machine LearningInfrastructureDistributed SystemsML Platforms
Reddit

Senior Machine Learning Systems Engineer, Ads ML Experience Platform

Reddit👥 1001 - 5000 employees🏢 Online Community/social Media
🔥 3 hours ago

As a Senior Machine Learning Systems Engineer, you will design and build large-scale ML experimentation platforms and develop production-grade training orchestration frameworks to enhance ML lifecycle efficiency.

ML Experimentation PlatformsTraining Orchestration FrameworksInfrastructure EngineeringDistributed Systems
ZoomInfo

Senior Machine Learning Engineer

ZoomInfo👥 10,000+ employees🏢 Software Development🤝 B2B
🔥 8 hours ago

As a Senior Machine Learning Engineer, you will design and implement large-scale recommendation systems, optimize NLP models, and manage MLOps infrastructure to enhance AI-driven solutions.

Machine LearningNatural Language ProcessingRecommendation SystemsMlops
Zscaler

Staff Machine Learning Engineer - Data Lake, Anomaly Detection

Zscaler👥 10,000+ employees🏢 Computer And Network Security🤝 B2B
🔥 19 hours ago

As a Staff Machine Learning Engineer, you will enhance Zscaler's cloud security platform by integrating advanced AI technologies and designing scalable production systems.

AI/ML TechnologiesGenaiCloud PlatformsData Lakes

Trusted by Remote Workers