Remote Jobs RockRemote Jobs Rock

Senior Software Engineer, AI Infrastructure

🕒 yesterday
KubernetesLLM Serving InfrastructureGPU SchedulingHelm

📜 Description

  • Design and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management.
  • Package the platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks.
  • Integrate the serving layer with the platform's API gateway, identity, and metering services.
  • Build the observability for operating GPU inference in production (serving metrics, GPU telemetry).
  • Contribute across a multi-service codebase and help set engineering direction through design docs and reviews.

🛠️ Requirements

  • 5+ years of software engineering experience in infrastructure, platform, or distributed systems.
  • Deep hands-on Kubernetes experience: building and operating production workloads and Helm charts, not just consuming managed clusters.
  • Experience with GPU workloads or LLM inference, or strong adjacent systems experience and a track record of learning fast.
  • Strong Go programming skills; solid CI/CD and infrastructure-as-code skills.
  • Fluency with AI-assisted development tools (Claude Code, OpenAI Codex) as part of your daily engineering workflow.
  • Comfortable with high autonomy on a small, remote-first, written-culture team.
  • Inference performance work (quantization, batching, caching) or distributed serving frameworks.
  • Enterprise deployment experience: air-gapped installs, SSO/OIDC, supply-chain security.
  • UI development experience (e.g. React/TypeScript), useful as the product's management surfaces grow.
  • Open-source contributions in the Kubernetes or ML-infrastructure ecosystems

Benefits

  • Professional development and training.
  • Attend conferences and working groups.
  • Company outings, happy hours, hackathons, and tech talks.
  • Receive a competitive compensation package with a strong benefits plan.
Full job description

Company Description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

Mirantis is building a new enterprise AI infrastructure product that lets organizations run and govern large language models on their own Kubernetes clusters. You will join a small senior team early, with broad ownership of the model-serving layer and its path to production.

What you'll do

  • Design and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management.

  • Package the platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks.

  • Integrate the serving layer with the platform's API gateway, identity, and metering services.

  • Build the observability for operating GPU inference in production (serving metrics, GPU telemetry).

  • Contribute across a multi-service codebase and help set engineering direction through design docs and reviews.

Qualifications

What we're looking for:

  • 5+ years of software engineering experience in infrastructure, platform, or distributed systems.

  • Deep hands-on Kubernetes experience: building and operating production workloads and Helm charts, not just consuming managed clusters.

  • Experience with GPU workloads or LLM inference, or strong adjacent systems experience and a track record of learning fast.

  • Strong Go programming skills; solid CI/CD and infrastructure-as-code skills.

  • Fluency with AI-assisted development tools (Claude Code, OpenAI Codex) as part of your daily engineering workflow.

  • Comfortable with high autonomy on a small, remote-first, written-culture team.

Nice to have:

  • Inference performance work (quantization, batching, caching) or distributed serving frameworks.

  • Enterprise deployment experience: air-gapped installs, SSO/OIDC, supply-chain security.

  • UI development experience (e.g. React/TypeScript), useful as the product's management surfaces grow.

  • Open-source contributions in the Kubernetes or ML-infrastructure ecosystems

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Kong

Forward Deployed Engineer

🔥 23 hours ago
Kong👥 501 - 1000 employees🏢 Computer Software

As a Forward Deployed Engineer, you will engage directly with enterprise customers to implement Kong's API and AI Connectivity platform, driving modernization and automation efforts.

API GatewaysAI ConnectivityCloud-native TechnologiesAutomation Frameworks
Reddit

Staff Software Engineer - Ingestion Platform

🕒 yesterday
Reddit👥 1001 - 5000 employees🏢 Online Community/social Media

Lead the development and evolution of Reddit's Ingestion Platform, focusing on reliable software for distributed data movement and enhancing user experience.

Distributed SystemsData InfrastructureSoftware DevelopmentData Pipelines

As a Senior Software Engineer II, you will lead the development of scalable back-end services and tackle distributed systems challenges while mentoring a talented engineering team.

KotlinJavaTypeScriptAWS

Trusted by Remote Workers