Remote Jobs RockRemote Jobs Rock

Embedded AI Engineer, On-Device Models

📅 Jul 7
Embedded SystemsAI OptimizationC ProgrammingC++ Programming

📜 Description

  • Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware.
  • Define the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
  • Optimize models for constrained targets through quantization, pruning, distillation, and architecture-specific compilation.
  • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments.
  • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains.
  • Build the on-device runtime plumbing: model packaging, deployment pipelines, and lightweight telemetry.

🛠️ Requirements

  • Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
  • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
  • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
  • Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
  • A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.
  • Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.
  • Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

Trusted by Remote Workers