Remote Jobs RockRemote Jobs Rock

Applied AI Engineer, Site Reliability Engineer - EMEA

📅 Jul 6
SREProduction EngineeringDevOpsKubernetes

📜 Description

  • Design for a fleet of Mistral platforms and apps, focusing on proactive reliability.
  • Operate Tier-1 customer environments, ensuring SLO compliance and managing incident responses.
  • Productize deployment, security, and scaling of Applied AI solutions.
  • Lead security operations for customer-side deployments, enforcing secure configurations.

🛠️ Requirements

  • 5+ years in SRE, Production Engineering, or DevOps with a record of shipping tooling.
  • Strong multi-tenant Kubernetes fluency and operations at scale.
  • Experience with observability stacks in production, including Prometheus and Grafana.
  • Proficient in Python and/or Golang for tooling and automation.
  • Strong written communication skills for runbooks and incident communications.

Benefits

  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal and transportation allowances
  • Other location-specific perks
Full job description

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

About the team

The Applied AI team is Mistral's customer-facing technical organization. We work directly with enterprise clients from pre-sales through implementation to deploy cutting-edge AI solutions that deliver measurable business impact.

Our team combines deep ML expertise with strong customer engagement skills, operating like startup CTOs who own end-to-end project execution. Our SRE team works transversally across customer engagements, enabling value creation through Mistral tech - at scale. By joining the team you will bridge the gap between cutting-edge AI research and real-world enterprise applications, ensuring our solutions are robust, scalable, and aligned with both customer needs and Mistral's technological vision.


About The Job

You will be one of the founding engineers of the Applied AI SRE sub-team.

Your mission, alongside the team, is to build and operate the framework to ensure Mistral’s solution delivery is reliable and sustainable - and applied uniformly across all our accounts, both Mistral-hosted and customer-hosted. You should already have a strong understanding of what operational excellence looks like, and you’re ready to scale your impact.

You will operate in four concurrent modes:

- BUILD - Design for a fleet of Mistral platforms and apps. Build proactivity to reduce reactivity.
Productize reliability, author runbooks, create SLO templates, implement observability.
- RUN - Operate the Tier-1 customer environments that Mistral are contracted to operate.
Ensure SLO compliance, own on-call and incident response, manage drift, partner with Technical Support as L3 escalation, champion high signal post-mortems.
- ENABLE - Productize how Mistral deploy, secure, and scale our Applied AI solutions.
Engineer on-demand provisioning, author security baseline packages, embed security guardrails, automate everything.
- SECURE - Own the security operations layer for our customer-side deployments.
Lead CVE response across the fleet, ship supply-chain integrity controls (SBOM, signed images, provenance), co-page with InfoSec on security incidents, enforce secure-config baselines.

This is a framework-first, fleet management role at heart. If you're excited by the difference between solving one customer's problem and structurally solving the class of problem for every customer, this is the role.


How We Work in Applied AI

• We care about people and outputs.
• What matters is what you ship, not the time you spend on it
• Bureaucracy is where urgency goes to vanish. You talk to whoever you need to talk to. The best idea wins, whether it comes from a principal engineer or someone in their first week.
• Always ask why. The best solutions come from deep understanding, not from copying what worked before
• We say what we mean. Feedback is direct, timely, and given because we care.
• No politics. Low ego, high standards.
• We embrace an unstructured environment and find joy in it.


About you

• Fluent in English.
• 5+ years in SRE, Production Engineering, or DevOps, with a record of shipping tooling.

• Strong multi-tenant Kubernetes fluency, namespace segmentation, network policy, RBAC, admission control, operations at scale.

• On-call discipline: incident response, blameless post-mortem culture, runbook-first mindset.

• Observability stack in production: Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Signoz.

• Infrastructure as code: Terraform, Ansible (or close equivalents).

• Proficient in Python and/or Golang for tooling and automation.

• Security mindset: you treat secure-SDLC, CVE response, and supply-chain integrity as reliability properties of the shipped artifact, not as someone else's job.

• Strong written communication skills: runbooks, post-mortems, and customer-facing incident comms are core deliverables of this role.

• Comfortable operating with high autonomy in an ambiguous, fast-paced environment — and disciplined enough to defend the team's scope when work tries to spill in.

• Solid Linux internals, networking debug, and distributed-systems fundamentals.


Strong plus

• Cloud or application security background (AppSec, K8s security, supply chain — SBOM, cosign, SLSA). At least one of our early hires must bring this; if it's you, flag it.

• Experience operating LLM / model-serving stacks in production

• Experience with multi-cloud or on-prem hybrid customer environments (AWS, GCP, Azure, sovereign clouds).

• Open-source contributions, particularly in SRE, observability, or security tooling.

 
 

By applying, you agree to our Applicant Privacy Policy.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Aristanetworks

Site Reliability Engineer (SRE) - Engineering Productivity

Aristanetworks👥 5001 - 10,000 employees🏢 Computer Networking
🕒 6 days ago

As a Site Reliability Engineer in the Engineering Productivity team, you will design, build, and operate scalable and reliable systems to enhance developer experience and support Arista's product development.

GoPythonShell ScriptingLinux
Earnin

Staff Site Reliability Engineer

Earnin👥 501 - 1000 employees🏢 Financial Services
🕒 3 days ago

Lead the evolution of EarnIn's reliability practices by implementing an AI-first operating model to enhance operational quality and incident response across critical services.

Site Reliability EngineeringAI OperationsIncident ManagementSoftware Engineering
Asana

Senior Software Engineer, Site Reliability

Asana👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 6 days ago

Asana’s rapid growth brings new challenges in keeping our systems fast, reliable, and resilient. As our product evolves, we’re making a major investment in reliability – and building a brand new SRE team in Warsaw is a key part of that.

AWSKubernetesDatadogMySQL

Trusted by Remote Workers