Remote Jobs RockRemote Jobs Rock

AI Researcher - Bolter

๐Ÿ•’ yesterday
AI ResearchML ResearchLLM Application DesignExperimental Design

๐Ÿ“œ Description

  • Design and run experiments to test agent behavior, focusing on reliability and context retention.
  • Build and own the evaluation layer to measure actual work completion of agents.
  • Research the latest advancements in LLM agents and tool use, translating insights into product development.
  • Produce clear, actionable recommendations for the engineering team based on research findings.
  • Prototype research ideas into working products and shape the research roadmap.

๐Ÿ› ๏ธ Requirements

  • Experience with LLM evals, hallucination mitigation, or production AI reliability at scale.
  • Experience building or evaluating agentic systems, tool use, or autonomous workflows.
  • A public track record - papers, open source, blog posts, side projects. We love researchers who ship outside work too.
  • Experience in product-led research - where the output is a shipped feature, not just a finding.
Full job description

Both short-term and permanent roles are possible.

About Bolter

Bolter is an AI agent platform built to help people get real work done, without needing to stitch together multiple tools or spend weeks setting things up.

You describe what you need in plain language and Bolter creates and runs agents that can carry out the work end to end. They retain context, remember how you work and can also create real, shareable apps such as trackers and dashboards.

Bolter is funded and currently at the validation-sprint stage, working with a small group of high-impact operators. The product already exists. We are now looking for a founding designer to lead a significant redesign and establish the design foundations for what comes next.

ย 

About the role

As Bolter's AI Researcher, your job is to make agents that do real work trustworthy, capable, and useful - and to figure out how before we build it at scale.

This is applied research with direct product impact. You design and run experiments, evaluate what works and what doesn't, and turn findings into decisions the engineering team ships. You'll work directly with the GM and the product engineer in a lean team, with specialised AI agents supporting implementation.

The research problem goes beyond improving a model. It is working out how an agent that retains context, carries out multi-step work and creates software itself can be relied on to do it correctly.

ย 

What youโ€™ll do

  • Design and run experiments to test how Bolter's agents behave - reliability, context retention, multi-step task completion, and where they fail.

  • Build and own the evaluation layer. Design evals that measure whether agents actually do the work, not just whether they sound plausible.

  • Research the frontier. Keep Bolter current on the state of the art in LLM agents, tool use, and reliability - and translate that into what we should build.

  • Turn findings into decisions. You produce clear, actionable recommendations the engineering team can ship, not just papers.

  • Prototype research into product. Take promising ideas from experiment to working prototype, and hand off what proves useful.

  • Shape the research roadmap as Bolter grows.

Why you're made for this

  • A track record of applied AI/ML research - ideally in LLM application design, agentic systems, evals, or production AI reliability.

  • Strong experimental design and analysis. You know how to test a hypothesis rigorously and read the result honestly.

  • Strong engineering fundamentals - you can prototype your own experiments, not just direct others to run them.

  • Deep familiarity with LLMs, tool use, context management, and the failure modes of agentic systems.

  • Strong written communication. Research that isn't understood and acted on doesn't help.

  • You're rigorous but pragmatic. You know when a finding is strong enough to act on and when it needs more evidence.

  • You're comfortable with ambiguity. Many of the problems we're solving don't have textbook answers.

  • You can leverage AI agents as a force multiplier - we run lean, with a small human core augmented by a fleet of specialised agents.

  • You care about craft, but you ship.

Bonus points

  • Experience with LLM evals, hallucination mitigation, or production AI reliability at scale.

  • Experience building or evaluating agentic systems, tool use, or autonomous workflows.

  • A public track record - papers, open source, blog posts, side projects. We love researchers who ship outside work too.

  • Experience in product-led research - where the output is a shipped feature, not just a finding.

While we think the above experience could be important, we're keen to hear from people who believe they have valuable experience to bring to the role. If you identify with the team and mission, but not all of our requirements, then please still apply!

Improbable Candidate Privacy Policy

Eversana1

Senior AI Application Engineer

Eversana1๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Pharmaceuticals
๐Ÿ•’ yesterday

As a Mid-Level AI Application Engineer, you will architect and implement scalable multi-agent systems, develop APIs, and optimize data pipelines to enhance the organization's AI capabilities.

AI Application DevelopmentAPI DevelopmentData Pipeline OptimizationMulti-agent Systems
Atlas-Technica

AI Engineer

Atlas-Technica๐Ÿ‘ฅ 201 - 500 employees๐Ÿข Information Technology And Services
๐Ÿ•’ 6 days ago

The AI Engineer will build, integrate, test, and deploy AI-enabled applications and enterprise solutions, translating architecture and security decisions into working software.

PythonAI-enabled ApplicationsMicrosoft AzureREST APIs
LangChain

Deployed Engineer, Professional Services

LangChain๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Technology, Information And Internet๐Ÿค B2B
๐Ÿ•’ 6 days ago

As a Deployed Engineer on the Professional Services team, you'll work with enterprise customers to design and build reliable production agents, translating workflows into actionable software specifications.

PythonTypeScriptJavaScriptLangchain

Trusted by Remote Workers