Senior Golang Software Engineer, Infrastructure - remote in the US
Design and implement Infrastructure Services for a GPU-as-a-Service platform, focusing on API development and provisioning workflows.
Kimchi is the AI platform inside CAST AI. We started by helping companies run LLMs on their own Kubernetes clusters and now we're providing a managed variant of those same capabilities.
Multi-model inference (MiniMax, Kimi, GLM-5, Nemotron, DeepSeek) with intelligent routing, an OpenAI-compatible API and deployment ranging from our GPUs to your own VPC. The inference layer is the foundation and the API is what sits in front of it as the primary channel for broadly and reliably distributing our AI services and powers our own Kimchi harness.
As a Senior Software Engineer, you will have the opportunity to work on different key features of our product. All of these are high-agency roles across multiple parts of the tech stack that minimize process friction that would otherwise prevent you from shipping.
In every team you will own features end-to-end: design, implementation, testing, production rollout. Most projects ship in 1-4 weeks. You'll work directly with product and other engineering teams on problems that don't have textbook solutions.
We are currently hiring Senior Software Engineers for the following teams:
Owns the inference API responsible for delivering our AI services across the world while making sure it's reliable and capable to scale in tandem with our company's growing ambitions, as well as our analytical platform, billing and role-based controls that enable our users to monitor and control their usage with ease - be it as a solo developer or a large enterprise. You'll own our infrastructure, datastores, analytics, observability and CI/CD pipeline.
OpenAI and Anthropic ship models. They also ship one harness each – the scaffolding that turns a raw model into something that can plan, execute, recover, and complete work. We ship a different kind of harness: one built for cost-conscious, long-horizon autonomy, running on inference infrastructure we control end-to-end.
A decent model with a great harness beats a great model with a bad harness. We've watched this play out. The gap between what today's models can do and what you see them doing is largely a harness gap – and that gap is where we operate.
The Agent Platform squad builds the infrastructure that lets teams run fleets of AI agents securely, whether in the cloud or self-hosted, so work can move confidently from a developer's laptop to production environments. The platform enforces least-privilege access to resources, requires human approval before agents can touch anything critical, and logs every action agents take for full auditability. The goal is to give engineering and operations teams the confidence to scale up agent usage without sacrificing security, control, or visibility into what their agents are actually doing.
As part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.
Please note that Cast AI does not provide any form of visa sponsorship/work permit.
#LI-Remote
Design and implement Infrastructure Services for a GPU-as-a-Service platform, focusing on API development and provisioning workflows.
Join the Release Engineering team to optimize build and deployment automation, create CI/CD solutions, and mentor fellow engineers while enhancing software delivery processes.
As a Staff Software Engineer, you will lead the design and evolution of secure payment infrastructure, focusing on building reliable APIs and ensuring system security and scalability.
Drive the development of control planes for AI workloads by designing and maintaining microservices in Rust and Go, while ensuring high-quality architecture and integration.
As a Senior Engineer on the Safety Experience team, you'll lead the charge in making Discord's platform safer for hundreds of millions of users by designing and implementing critical safety features.