Remote Jobs RockRemote Jobs Rock

Staff System Engineer – AI Infrastructure

🔥 1 hour ago
AI InfrastructureGPU EnvironmentsLinuxContainers

📜 Description

  • Investigate and design solutions for AI workloads across heterogeneous GPU environments.
  • Develop and optimize infrastructure solutions for production AI workloads, including inference and serving.
  • Diagnose issues across Linux, GPU runtimes, containers, networking, and storage.
  • Build diagnostic, benchmarking, and validation tooling to evaluate product performance.
  • Translate infrastructure findings into technical requirements and product improvements.

🛠️ Requirements

  • 8+ years of experience in systems software, distributed infrastructure, or related fields.
  • Hands-on experience with production GPU-based infrastructure supporting AI workloads.
  • Strong understanding of distributed systems, Linux, containers, and production infrastructure.
  • Demonstrated ability to diagnose and optimize bottlenecks involving GPU utilization and networking.
  • Experience with modern AI inference technologies and Kubernetes is a plus.

Benefits

  • Generous PTO Policy.
  • Flexible WFH Policy.
  • Mental & Physical Wellness programs.
  • Phone and Internet Reimbursement program.
  • Access to Continued Career Development.
Full job description

Join Cloudera as a Staff Systems Engineer – AI Infrastructure, where you'll drive hands-on engineering solutions for diverse AI workloads across various environments.

Description

  • Investigate and design solutions for AI workloads across heterogeneous GPU environments.
  • Develop and optimize infrastructure solutions for production AI workloads, including inference and serving.
  • Diagnose issues across Linux, GPU runtimes, containers, networking, and storage.
  • Build diagnostic, benchmarking, and validation tooling to evaluate product performance.
  • Translate infrastructure findings into technical requirements and product improvements.

Requirements

  • 8+ years of experience in systems software, distributed infrastructure, or related fields.
  • Hands-on experience with production GPU-based infrastructure supporting AI workloads.
  • Strong understanding of distributed systems, Linux, containers, and production infrastructure.
  • Demonstrated ability to diagnose and optimize bottlenecks involving GPU utilization and networking.
  • Experience with modern AI inference technologies and Kubernetes is a plus.

Benefits

  • Generous PTO Policy.
  • Flexible WFH Policy.
  • Mental & Physical Wellness programs.
  • Phone and Internet Reimbursement program.
  • Access to Continued Career Development.
Harvey

Staff Software Engineer, Model Infrastructure

Harvey👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 2 days ago

As a Staff Software Engineer on the Model Infrastructure team, you'll lead the design and development of systems that ensure high availability and operational excellence for AI inference at Harvey.

Software EngineeringDistributed SystemsCloud InfrastructureNetworking
Harvey

Senior Software Engineer, Model Infrastructure

Harvey👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 2 days ago

As a Staff Software Engineer on the Model Infrastructure team, you'll design and develop systems that ensure high availability and operational excellence for AI requests at Harvey.

Software EngineeringDistributed SystemsCloud InfrastructureNetworking
Cloudflare

Senior Software Engineer, Security Rules

Cloudflare👥 10,000+ employees🏢 Computer And Network Security🤝 B2B
🕒 yesterday

As a Senior Software Engineer on the Security Rules team, you will design and maintain software systems for Cloudflare's Application Security products, leading complex projects and mentoring fellow engineers.

RustGoC++Distributed Systems

Trusted by Remote Workers