Remote Jobs RockRemote Jobs Rock

Senior AI Storage Infrastructure Engineer

๐Ÿ•’ 10 days ago
Container Storage Interface (csi)Gpudirect Storage (gds)NvmeKubernetes

๐Ÿ“œ Description

  • Design, deploy, and maintain robust Container Storage Interface (CSI) drivers for high-performance parallel file systems (e.g., Weka, Lustre, DAOS, VAST).
  • Architect and implement GPUDirect Storage (GDS) integrations for direct memory access (DMA) between NVMe drives and GPU memory.
  • Develop and manage local NVMe caching strategies for low-latency loading of massive model weights and datasets.
  • Optimize IOPS, throughput, and latency profiles across the containerized storage stack.
  • Collaborate with the GPU Systems & Fabric team to ensure storage optimization for RDMA and high-speed interconnects.
  • Implement automated monitoring and alerting for storage performance.

๐Ÿ› ๏ธ Requirements

  • Bachelorโ€™s or Masterโ€™s degree in Computer Science, Electrical Engineering, or a related field.
  • 5+ years of experience in distributed storage systems and high-performance file systems, with a deep understanding of POSIX compliance and file I/O semantics.
  • Deep expertise in the Kubernetes CSI paradigm, including building or extending volume plugins and storage operators.
  • Strong hands-on experience with block/file I/O at the Linux OS level and kernel-level performance tuning.
  • Familiarity with high-throughput networking protocols (RDMA, InfiniBand, RoCE) and how they interact with storage subsystems.
  • Proven track record of operating, debugging, and scaling large-scale storage environments in production or HPC settings.
  • Experience with infrastructure automation tools (e.g., Terraform, Ansible) and CI/CD pipelines.
  • Excellent technical communication skills, with the ability to influence cross-functional architectural decisions.
  • Experience working in high-velocity, high-growth engineering environments is strongly preferred.
Reddit

Senior Staff Machine Learning Systems Engineer, Ads ML Platform

Reddit๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Online Community/social Media
๐Ÿ”ฅ 10 hours ago

Lead the technical strategy for the Ads ML engineer lifecycle, focusing on feature development and training iteration to enhance ML engineers' productivity.

Machine LearningInfrastructureDistributed SystemsML Platforms
ZoomInfo

Senior Machine Learning Engineer

ZoomInfo๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development๐Ÿค B2B
๐Ÿ”ฅ 3 hours ago

As a Senior Machine Learning Engineer, you will design and implement large-scale recommendation systems, optimize NLP models, and manage MLOps infrastructure to enhance AI-driven solutions.

Machine LearningNatural Language ProcessingRecommendation SystemsMlops
Reddit

Senior Machine Learning Systems Engineer, Ads ML Experience Platform

Reddit๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Online Community/social Media
๐Ÿ”ฅ 10 hours ago

As a Senior Machine Learning Systems Engineer, you will design and build large-scale ML experimentation platforms and develop production-grade training orchestration frameworks to enhance ML lifecycle efficiency.

ML Experimentation PlatformsTraining Orchestration FrameworksInfrastructure EngineeringDistributed Systems
Databricks

Sr. Forward Deployed Engineer (FDE) - Retail

Databricks๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development๐Ÿค B2B
๐Ÿ”ฅ 15 hours ago

As a Forward Deployed Engineer (FDE), you will lead the design and implementation of data and AI solutions for customers using the Databricks platform, ensuring impactful project delivery.

Data EngineeringData PlatformsAnalyticsSoftware Engineering
Zscaler

Staff Machine Learning Engineer - Data Lake, Anomaly Detection

Zscaler๐Ÿ‘ฅ 10,000+ employees๐Ÿข Computer And Network Security๐Ÿค B2B
๐Ÿ”ฅ 14 hours ago

As a Staff Machine Learning Engineer, you will enhance Zscaler's cloud security platform by integrating advanced AI technologies and designing scalable production systems.

AI/ML TechnologiesGenaiCloud PlatformsData Lakes

Trusted by Remote Workers