Remote Jobs RockRemote Jobs Rock

Engineering Manager, Model Infrastructure

📅 Jul 22
Software EngineeringTeam ManagementLarge-scale Distributed SystemsCloud Infrastructure

📜 Description

  • Lead and grow a high-performing team of software engineers responsible for the Model Infrastructure platform.
  • Define the technical roadmap for model reliability, scalability, and operational excellence.
  • Build systems for model provisioning, capacity management, failover, and incident response across AI providers.
  • Drive the evolution of the Unified Model Controller and Model Selector platform for intelligent traffic routing.
  • Improve observability through health dashboards, alerting, and end-to-end model telemetry.
  • Partner with Product Engineering to support new model launches and proactive production monitoring.

🛠️ Requirements

  • 8+ years of software engineering experience, including multiple years managing high-performing engineering teams.
  • Experience leading teams responsible for large-scale distributed systems or cloud infrastructure.
  • Strong technical background that enables you to guide architectural decisions and mentor senior engineers.
  • Experience operating highly available production services with strong reliability and operational excellence.
  • Experience building platforms that require scalability, observability, automation, and cost optimization.
  • Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors.
  • Excellent communication skills and the ability to influence technical strategy across organizations.
  • A passion for building teams and developing engineering talent.
  • Experience with AI infrastructure, LLM serving, or machine learning platforms.
  • Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems.
Full job description

Why Harvey

At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.

This is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched.

Our team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you.

At Harvey, the future of professional services is being written today — and we’re just getting started.

Role Overview

As the Engineering Manager for Model Infrastructure, you'll lead the team responsible for the platform powering every model request across Harvey. You'll partner closely with AI Research, Product Engineering, Infrastructure, and external AI providers to ensure our platform remains reliable, scalable, and cost-efficient as our business grows.

Model Infrastructure is one of Harvey's most strategic engineering organizations. Every product capability—from chat experiences and agents to document workflows and future reasoning systems—depends on this platform.

Over the next several years, the team will evolve beyond operating third-party models to building the infrastructure that enables Harvey to train, evaluate, deploy, and operate our own frontier AI models. This role offers the opportunity to shape the technical foundation of Harvey's AI platform and build an organization that will power the company's next phase of growth.

What You'll Do

  • Lead and grow a high-performing team of software engineers responsible for Harvey's Model Infrastructure platform.

  • Define the technical roadmap for model reliability, scalability, and operational excellence.

  • Build highly reliable systems for model provisioning, capacity management, failover, and incident response across multiple AI providers.

  • Own Harvey's multi-provider model platform, including provider integrations, SDK upgrades, API migrations, and onboarding new model providers.

  • Drive the evolution of our Unified Model Controller (UMC) and Model Selector platform to automatically detect degraded models and intelligently route traffic based on health, latency, quality, compliance, and cost.

  • Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end model telemetry.

  • Partner with Product Engineering to support new model launches, capacity planning, experimentation, and proactive production monitoring.

  • Lead initiatives to improve inference efficiency, reduce infrastructure costs, and increase model utilization across providers.

  • Build the infrastructure foundation for Harvey's future model training efforts, including data pipelines, model operations, training environments, and AI platform capabilities.

  • Partner with executive leadership on long-term AI infrastructure strategy and vendor relationships.

  • Recruit, mentor, and develop exceptional engineering talent while fostering a culture of technical excellence and operational ownership.

What You Have

  • 8+ years of software engineering experience, including multiple years managing high-performing engineering teams.

  • Experience leading teams responsible for large-scale distributed systems or cloud infrastructure.

  • Strong technical background that enables you to guide architectural decisions and mentor senior engineers.

  • Experience operating highly available production services with strong reliability and operational excellence.

  • Experience building platforms that require scalability, observability, automation, and cost optimization.

  • Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors.

  • Excellent communication skills and the ability to influence technical strategy across organizations.

  • A passion for building teams and developing engineering talent.

Nice to Have

  • Experience with AI infrastructure, LLM serving, or machine learning platforms.

  • Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems.

  • Experience building inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure.

  • Experience with Kubernetes, cloud infrastructure, distributed systems, and large-scale observability platforms.

  • Experience supporting GPU infrastructure, model training platforms, or ML infrastructure.

  • Familiarity with data platforms and technologies such as Spark, Kafka, Flink, Airflow, or Iceberg.

  • Experience leading organizations through periods of rapid growth and technical transformation.

Compensation

$260,000 - $340,000 USD

Depending on your location, an Applicant Privacy Notice may apply to you. You can find all of our Applicant Privacy Notices here.

#LI-AN2

Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made by emailing accommodations@harvey.ai

Fastly

Senior Engineering Manager - Containers at Edge

Fastly👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 10 days ago

As Senior Engineering Manager for Containers at the Edge, you will lead a high-performing team to architect and manage distributed container infrastructure, ensuring low latency and secure application delivery.

Software EngineeringSite Reliability EngineeringDistributed SystemsCDN
Cohere

Engineering Manager, FDE Agentic Platform

Cohere👥 10,000+ employees🏢 Software Development🤝 B2B
🕒 6 days ago

Lead and mentor a team of Forward Deployed Engineers to build agentic solutions, ensuring customer success and fostering a culture of innovation in a fast-paced environment.

Software EngineeringTeam ManagementCustomer SuccessAI Integration
OpenAI

Site Selection Lead

OpenAI👥 10,000+ employees🏢 Research Services
🕒 yesterday

The Site Selection Lead develops and implements strategies for site selection, ensuring that powered-land and greenfield/brownfield opportunities are attractive, executable, and scalable.

Site SelectionInfrastructure DevelopmentData Center DevelopmentEnergy Development
Cloudflare

Senior Engineering Manager - Cloudforce One

Cloudflare👥 10,000+ employees🏢 Computer And Network Security🤝 B2B
🕒 3 days ago

Lead the Threat Application Services team as a Senior Engineering Manager, focusing on building and scaling threat intelligence platforms and APIs to enhance enterprise security operations.

Software DevelopmentGoJavaScriptDistributed Systems

Trusted by Remote Workers