Remote Jobs RockRemote Jobs Rock

Principal Engineer, Data & Compute

📅 Jun 23
Large-scale Distributed SystemsGPU-based Cloud InfrastructureAI TrainingInference Workloads

📜 Description

  • Define and evolve the architecture for allocating and orchestrating training and inference workloads across thousands of GPUs.
  • Design systems for fast, reliable access to high-volume sensor and simulation data across geographies.
  • Build foundations for large-scale AI workloads to run seamlessly across hybrid and multi-cloud environments.
  • Act as a trusted partner to leadership in aligning compute investments with company strategy.
  • Uplift the engineering organization through architectural coaching and technical mentorship.

🛠️ Requirements

  • 10+ years designing and building large-scale distributed systems, with at least 4 years focused on GPU-based cloud infrastructure.
  • Proven experience enabling large-scale AI training, inference, or computer vision workloads in GPU clusters.
  • Deep understanding of petabyte-scale data architecture, including storage federation and high-throughput access.
  • Strong technical leadership with a track record of defining and communicating architectural strategy.
Full job description

About us

Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.

Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.

In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.

At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.

Make Wayve the experience that defines your career!

The Role

  • At Wayve, we are teaching machines to drive—not by coding rules, but by training end-to-end neural networks that learn from vast streams of real-world data. Achieving this requires unprecedented scale in both data infrastructure and compute orchestration. Our workloads span thousands of GPUs, petabytes of driving data, and geographically distributed training and inference clusters.
  • As Architect for AI Infrastructure, you will design and guide the evolution of the foundational compute and storage systems that fuel our model development lifecycle. Your leadership will directly accelerate AI research, enable rapid model deployment, and ensure our platform meets the demands of a company pushing the boundaries of autonomy.
  • You’ll sit at the strategic core of AI, systems, and cloud infrastructure—owning challenges that few companies have the ambition or scale to tackle.

Key Responsibilities

  • Global Compute Strategy – Define and evolve the architecture for how Wayve allocates and orchestrates training and inference workloads across thousands of GPUs and multiple data centers, ensuring optimal throughput, resiliency, and cost efficiency.
  • Petabyte-Scale Data Federation – Design systems that enable fast, reliable access to high-volume sensor and simulation data across geographies, ensuring the right data is always available for training, evaluation, and inference. Furthermore, preparing Wayve for being an exabyte-scale company.
  • Cross-Region GPU Job Execution – Build the foundations that enable large-scale AI workloads to run seamlessly across hybrid and multi-cloud environments.
  • Cloud Infrastructure Advisory – Act as a trusted partner to leadership in aligning compute investments and architecture with company strategy, growth plans, and performance goals.
  • Technical Leadership & Mentorship – Uplift the broader engineering org through architectural coaching, technical deep dives, and by cultivating a culture of operational and engineering excellence.

About You

In order to set you up for success at Wayve, we’re looking for the following skills and experience.

Essential

  • 10+ years designing and building large-scale distributed systems, with at least 4 years focused on GPU-based cloud infrastructure.
  • Proven experience enabling large-scale AI training, inference, or computer vision workloads in GPU clusters.
  • Deep understanding of petabyte-scale data architecture, including storage federation, high-throughput access, and data locality for AI workloads.
  • Strong technical leadership with a track record of defining and communicating architectural strategy, balancing long-term vision with delivery needs.
  • A natural mentor with a history of developing engineers and influencing technical direction across teams.
  • Advanced degree in Computer Science, Electrical Engineering, or a related field—or equivalent industry experience.

Desirable

  • Experience with multi-cloud orchestration, particularly in latency- or cost-sensitive training and inference pipelines.
  • Familiarity with systems like Ray, Kubernetes, Airflow, or Flyte, and deep fluency in AI/ML job scheduling, model lifecycle management, and infrastructure-as-code practices.
  • Background in supporting safety-critical or real-time inference use cases (e.g., robotics, autonomous vehicles, aerospace).
  • Passion for building infrastructure-as-a-product that delivers performance and simplicity to research and product teams alike.

This is a full-time role based in-office. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. This role is a full-time role based in Sunnyvale, CA (hybrid) and the reasonably estimated salary for this role ranges from $370,300 to $418,200, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.

#LI-HH1

Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.

We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply.

At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law.

For more information visit Careers at Wayve.

To learn more about what drives us, visit Values at Wayve

For US candidates only, please visit E-Verify Notice and Participation and Right to Work


DISCLAIMER: We will not ask about marriage or pregnancy, care responsibilities or disabilities in any of our job adverts or interviews. However, we do look to capture information about care responsibilities, and disabilities among other diversity information as part of an optional DEI Monitoring form to help us identify areas of improvement in our hiring process and ensure that the process is inclusive and non-discriminatory.

Plaid

Senior Machine Learning Engineer - Fraud

🕒 2 days ago
Plaid👥 10,000+ employees🏢 Software Development

As a Senior Machine Learning Engineer on the Fraud Data team, you will develop and optimize models to enhance fraud detection, leveraging insights from Plaid's extensive network data.

Machine LearningModel DevelopmentData ScienceFeature Engineering
OpenAI

Partner Applied AI Engineer

🕒 4 days ago
OpenAI👥 10,000+ employees🏢 Research Services

As a Partner AI Deployment Engineer, you'll lead technical engagements with systems integrators to drive the adoption of Generative AI solutions, ensuring partners deliver exceptional results for their customers.

Generative AITechnical ConsultingPythonJavaScript
Plaid

Senior Machine Learning Engineer - Embedded Insights

🕒 4 days ago
Plaid👥 10,000+ employees🏢 Software Development

As a Senior Machine Learning Engineer, you will build machine learning-powered products and features, supporting the Plaid App and driving product-market fit for a new business line.

Machine LearningSQLPythonData Visualization
Wayve

Staff / Senior Machine Learning Engineer, Reinforcement Learning

🕒 4 days ago
Wayve👥 501 - 1000 employees🏢 Computer Software

As a Senior / Staff Machine Learning Engineer, you will advance reinforcement learning methods for driving models, focusing on improving driving behavior through innovative techniques.

Reinforcement LearningBehavior CloningPythonPytorch
Babbel

Senior Machine Learning Engineer (all genders) - Babbel Labs

🕒 4 days ago
Babbel👥 501 - 1000 employees🏢 E-learning

As a Senior Machine Learning Engineer, you will enhance Babbel's learner-personalisation engine by owning subsystems, implementing features, and ensuring robust evaluation and monitoring.

Machine LearningML EngineeringProbabilistic ModelingBayesian Inference

Trusted by Remote Workers