Remote Jobs RockRemote Jobs Rock

AI Platform Engineer, Training and Inference

📅 May 18
ML EngineeringMlopsRay EcosystemRay Train

📜 Description

  • Own the Ray ecosystem end-to-end, managing KubeRay on GKE and tuning Ray Core Task/Actor scheduling.
  • Operate distributed training with Ray Train, configuring TorchTrainer for multi-node H100 clusters and managing checkpoint lifecycle.
  • Build and operate the LLM inference mesh with Ray Serve, integrating vLLM, SGLang, and NVIDIA Triton.
  • Optimize inference performance through fractional GPU allocation and continuous batching.
  • Design and operate the model routing layer with capability-based and version-based routing.
  • Manage the full model promotion lifecycle from quality gate to canary rollout.

🛠️ Requirements

  • Experience in ML engineering with a focus on ML platforms or MLOps.
  • Proficient in the Ray ecosystem, including Ray Train, Serve, Core, and Data.
  • Hands-on experience with LLM serving engines like vLLM and NVIDIA Triton.
  • Knowledge of distributed training techniques such as DDP and NCCL.
  • Familiarity with reinforcement learning concepts and algorithms.
  • Strong programming skills in Python and PyTorch.

Benefits

  • You may also be eligible to participate in a Saviynt discretionary bonus plan, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

Trusted by Remote Workers