As a Software Engineer II, you will design and develop backend systems, ensuring accurate resolutions and timely communications for post-transaction interactions while collaborating with cross-functional teams.
Software Engineer, Infrastructure (Go) - remote in the EU
📜 Description
- Design, build, and maintain versioned REST and gRPC APIs for bare-metal server lifecycle and cluster operations.
- Develop asynchronous workflows for server enrollment, inspection, OS provisioning, and cluster bring-up.
- Design error-handling models and reconciliation loops for reliable long-running provisioning operations.
- Translate high-level API calls into infrastructure actions using Kubernetes primitives and controller patterns.
- Manage bare-metal provisioning flows including BMC/Redfish and PXE/iPXE.
🛠️ Requirements
- Strong experience designing RESTful APIs or gRPC services with knowledge of API versioning.
- Proficiency in Go, particularly within the backend/Kubernetes ecosystem.
- Deep understanding of Kubernetes primitives and controller/reconciler patterns.
- Hands-on experience with bare-metal provisioning flows and hardware inspection.
- Experience building workflow-driven or event-driven systems.
Full job description
Company Description
About Mirantis
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Job Description
We are looking for an experienced Full-Stack Software Engineer to design and implement the Infrastructure Services that power our GPU-as-a-Service platform. You will build the control plane that turns high-level API calls into real infrastructure actions — enrolling bare-metal servers, provisioning them, and assembling them into multi-tenant Kubernetes clusters on high-performance hardware.
You will own the full lifecycle of the infrastructure-level services—from the Server and MachineType APIs down to the provisioning workflows and reconciliation loops that keep the platform's view of hardware consistent with physical reality.
Key Responsibilities
• Infrastructure API Design: Design, build, and maintain the versioned REST and gRPC APIs for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations.
• Provisioning & Workflow Development: Develop the asynchronous workflows that drive server enrollment, inspection, OS provisioning, and cluster bring-up, exposing durable status to callers.
• System Reliability: Design the error-handling models, idempotency guarantees, and reconciliation loops necessary to manage long-running provisioning operations reliably.
Qualifications
• API Development: Strong experience designing RESTful APIs or gRPC services. You understand API versioning and gateway patterns.
• Proficiency in Go (preferred for backend/Kubernetes ecosystem)
• Kubernetes Knowledge: Deep understanding of Kubernetes primitives and controller/reconciler patterns. You will be interacting with systems like k0rdent, Metal3, and Cluster API to translate high-level API calls into infrastructure actions.
• Bare-Metal Provisioning: Hands-on experience with bare-metal provisioning flows — BMC/Redfish, PXE/iPXE, image management, and hardware inspection.
• Asynchronous Systems: Experience building workflow-driven or event-driven systems (e.g., Temporal) where operations are long-running and state must remain consistent across retries and failures.
Preferred Qualifications
• State Reconciliation: Experience building informers or reconciliation bridges that keep an external datastore consistent with Kubernetes resource state.
• Multi-Tenancy: Experience building platforms where strict data and network isolation between tenants is required.
• Infrastructure-as-Code: Familiarity with Terraform/OpenTofu and GitOps-driven configuration (ArgoCD or Flux).
• Hardware Domain: Familiarity with GPU server hardware, DPUs/NICs, and high-performance datacenter fabrics.
Additional Information
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to isamoylova@mirantis.com
By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.
We are a Leader for Container Management in G2 (#2 after AWS)!
Similar jobs
Search more Software Engineer jobsSoftware Engineer, Security Rules
As a Systems Engineer on the Security Rules team, you will design, implement, and maintain software systems for Cloudflare's Application Security offerings, ensuring reliability and scalability.
Software Engineer, Observability
As a member of the Observability team, you will maintain and enhance Lyft's logging and metrics infrastructure, ensuring operational health and performance across the platform.
Software Engineer, Network Performance & Reliability (Argo)
As a Software Engineer on the Argo team, you'll enhance network connectivity for Cloudflare's products, collaborating across engineering teams and utilizing cutting-edge technologies.
Systems Engineer - Global Resource Management (Data Residency)
As a Systems Engineer on the Global Resource Management team, you'll build and operate microservices and REST APIs to enhance resource management across Cloudflare's extensive platform.
