Remote Jobs RockRemote Jobs Rock

Technical Lead (Machine Learning)

πŸ“… Jul 27
Machine LearningLarge Language ModelsData PipelinesModel Optimization

πŸ“œ Description

  • Lead the end-to-end execution of machine learning systems, including data pipelines and production deployment.
  • Fine-tune and optimize models using techniques like LoRA, QLoRA, and model distillation.
  • Design and operate scalable inference systems focusing on latency and cost efficiency.
  • Develop and maintain data pipelines for training datasets.
  • Build evaluation frameworks to measure model performance and safety.
  • Partner with application engineering teams to integrate ML systems into products.

πŸ› οΈ Requirements

  • Proven experience building and deploying production-grade machine learning systems used by real users.
  • Strong expertise working with large language models and understanding model behavior, limitations, and failure modes.
  • Experience developing scalable ML infrastructure, training pipelines, and inference systems.
  • Strong software engineering skills with the ability to write maintainable, production-quality code.
  • Experience balancing real-world production constraints, including latency, reliability, scalability, cost, and safety.
  • Strong ownership mindset with the ability to independently drive technical initiatives from design through deployment.
  • Excellent communication and collaboration skills, with experience working in cross-functional, high-performing engineering teams.
  • PyTorch and/or JAX
  • GPU-based model training and inference systems
Full job description

About the Company

Our client is a stealth AI startup backed by one of Southeast Asia's leading technology companies and is currently building its global founding team.

The company is developing an AI-native communication platform designed to simplify everyday tasks by integrating AI directly into conversations. Instead of switching between multiple applications, users can plan, organize, compare, research, and complete tasks within a single intelligent assistant.

Serving a market of billions of users still relying on traditional productivity tools, the platform focuses on delivering reliable AI workflows, persistent context, multi-step reasoning, and seamless task execution. The mission is to create an AI assistant that significantly improves productivity while making everyday work simpler and more intuitive.

About the Role

Our client is seeking a Technical Lead, Machine Learning to lead the execution of its AI platform by translating research into scalable, production-ready machine learning systems. This role sits at the intersection of research, infrastructure, and product, with responsibility for ensuring models are trainable, deployable, observable, and optimized for real-world performance.

Working closely with research, engineering, and product teams, this position will drive the development of robust ML infrastructure while balancing performance, reliability, latency, and cost.

Key Responsibilities

  • Lead the end-to-end execution of machine learning systems, including data pipelines, training workflows, evaluation frameworks, inference architecture, and production deployment.
  • Fine-tune and optimize models using modern techniques such as LoRA, QLoRA, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and model distillation.
  • Design, build, and operate scalable inference systems with a focus on latency, cost efficiency, and reliability.
  • Develop and maintain data pipelines for both synthetic and real-world training datasets.
  • Build evaluation frameworks to measure model performance, robustness, safety, and bias in collaboration with research teams.
  • Optimize production deployments through GPU utilization, memory efficiency, inference optimization, and scaling strategies.
  • Partner closely with application engineering teams to integrate machine learning systems into backend, desktop, and mobile products.
  • Continuously improve production systems through rapid iteration, monitoring, and data-driven optimization.

Requirements

  • Proven experience building and deploying production-grade machine learning systems used by real users.
  • Strong expertise working with large language models and understanding model behavior, limitations, and failure modes.
  • Experience developing scalable ML infrastructure, training pipelines, and inference systems.
  • Strong software engineering skills with the ability to write maintainable, production-quality code.
  • Experience balancing real-world production constraints, including latency, reliability, scalability, cost, and safety.
  • Strong ownership mindset with the ability to independently drive technical initiatives from design through deployment.
  • Excellent communication and collaboration skills, with experience working in cross-functional, high-performing engineering teams.

Preferred Technical Skills

Experience with the following technologies is preferred:

  • Python
  • PyTorch and/or JAX
  • GPU-based model training and inference systems
Okta

Engineering Manager, Governance Intelligence

πŸ•’ 2 days ago
OktaπŸ‘₯ 10,000+ employees🏒 Software Development

As an Engineering Manager on the Governance Intelligence team, you will lead the development of AI-powered features for Okta's Identity Governance product, ensuring secure and scalable access for government customers.

JavaPythonTypeScriptAPI Development
Coder

Software Engineering Manager (Core Workspaces)

πŸ”₯ 10 hours ago
CoderπŸ‘₯ 51 - 200 employees🏒 Computer Software

Lead and grow the Core Workspaces engineering team at Coder, guiding technical direction and enhancing agent capabilities in development environments.

ReactTypeScriptGoLLMs
OpenRouter

Engineering Manager, Provider Ecosystem

πŸ•’ 2 days ago
OpenRouterπŸ‘₯ 51 - 200 employees🏒 Computer Software

Lead the Provider Ecosystem team at OpenRouter, managing engineering talent and driving the integration of diverse AI models and endpoints while ensuring high-quality performance.

Engineering ManagementDistributed SystemsAPI DesignLLM Inference
Xai

Expert Team Lead, Engineering

πŸ•’ 5 days ago
XaiπŸ‘₯ 1001 - 5000 employees🏒 Internet

Lead a team of AI Tutors to ensure high-quality training data and evaluations for SpaceXAI's models, while driving operational excellence and continuous improvement.

Team LeadershipData Quality MetricsAnnotation ProcessesGuideline-driven Workflows
NerdWallet

Engineering Manager, Consumer-Facing Web Product Team

πŸ•’ 5 days ago
NerdWalletπŸ‘₯ 501 - 1000 employees🏒 Consumer Services

Lead the Consumer Credit Cards Marketplaces team as a hands-on Engineering Manager, driving technical direction and mentoring engineers to deliver high-quality web shopping experiences.

Software EngineeringTechnical LeadershipTypeScriptReact

Trusted by Remote Workers