Remote Jobs RockRemote Jobs Rock

Researcher, Post Training

πŸ“… Jul 27
Post-training ExpertiseFine-tuningPreference OptimizationReinforcement Learning

πŸ“œ Description

  • Develop and test post-training methods including supervised fine-tuning and feedback-driven learning.
  • Create high-quality post-training datasets through curation and synthetic data generation.
  • Improve model behavior in instruction following, transcription, and multilingual performance.
  • Build benchmarks and failure taxonomies to assess model improvements.
  • Implement and scale training and evaluation pipelines with a focus on reproducibility.
  • Collaborate with research and product teams to adapt models for specific use cases.

πŸ› οΈ Requirements

  • Practical experience with fine-tuning, preference optimization, RLHF, or related methods.
  • Strong command of PyTorch and modern model-training workflows.
  • Experience designing evaluations and diagnosing complex model behavior.
  • Ability to write clean, production-quality Python and debug distributed training systems.
  • A record of publications, open-source work, or substantial independent research.
  • MSc, PhD, or equivalent practical experience in machine learning, speech processing, or NLP.
Full job description

About the role

As a Researcher in Post Training, you will shape how nyra labs models behave after pretraining.

You will develop methods that make speech models more accurate, controllable, robust, and aligned with what people actually need. This includes post-training recipes, feedback-driven learning, data curation, model evaluation, and the systems required to run reliable experiments at scale.

This is a research role with strong engineering ownership. You will take ideas from an initial hypothesis through experimentation, evaluation, and release.

Why we need you

Pretraining creates capability. Post training determines whether that capability becomes useful.

Speech models need to understand what should be preserved, how uncertainty should be handled, and how behavior should change across tasks, languages, speakers, and clinical contexts. Generic alignment methods rarely account for the details that matter in real speech: hesitations, repetitions, interruptions, atypical pronunciation, silence, and incomplete utterances.

nyra health gives us access to a uniquely large, therapist-labeled dataset of neurological speech, including millions of recordings from real clinical settings. You will help turn that asset into models that behave reliably for people underserved by existing speech technology.

About the company

At nyra health, we build software that supports clinics, therapists, and patients throughout neurorehabilitation. myReha delivers personalized therapy, while nyra insights helps clinical teams manage and understand patient progress.

nyra labs is the research arm of nyra health. We turn difficult problems encountered in practice into open models, datasets, benchmarks, and research that the wider community can build on.

If that resonates with you, we would love to hear from you.

What you’ll shape

  • Post-training methods: Develop and test approaches including supervised fine-tuning, preference optimization, distillation, feedback-driven learning, and reinforcement learning where useful.

  • Training data: Create high-quality post-training datasets through curation, annotation, synthetic data generation, and model-assisted data improvement.

  • Model behavior: Improve instruction following, verbatim transcription, uncertainty handling, long-form consistency, multilingual performance, and resistance to hallucinations.

  • Evaluation: Build benchmarks and failure taxonomies that reveal whether models are genuinely improving.

  • Experimental systems: Implement, debug, and scale training and evaluation pipelines with a strong focus on reproducibility.

  • Specialized models: Work with research and product teams to adapt foundation models to specific speech and clinical use cases.

  • Open releases: Contribute to publications, model releases, datasets, and technical reports.

Deepgram

Research Staff, LLMs

πŸ“… Aug 23
DeepgramπŸ‘₯ 10,000+ employees🏒 Software Development🀝 B2B

Join Deepgram as a Research Staff member to innovate and advance Large Language Models (LLMs) through experimental research and collaboration.

Large Language ModelsTransformer ArchitectureDeep LearningData Curation
Deepgram

Research Staff, Voice AI Foundations

πŸ“… Aug 23
DeepgramπŸ‘₯ 10,000+ employees🏒 Software Development🀝 B2B

As a Research Staff member, you will develop innovative Latent Space Models to tackle the challenges of building scalable and cost-effective voice AI solutions.

Latent Space ModelsNeural Audio CodecsGenerative ModelsEmbedding Systems
Snorkel AI

Research Scientist - Human-AI Systems

πŸ•’ 4 days ago
Snorkel AIπŸ‘₯ 1001 - 5000 employees🏒 Computer Software

As a Research Scientist, you'll advance frontier AI by creating high-quality data and environments, optimizing pipelines, and collaborating with experts to enhance model performance.

AIMachine LearningNLPLLMs
Wayve

Staff Research Scientist, Reinforce Learning

πŸ•’ 7 days ago
WayveπŸ‘₯ 501 - 1000 employees🏒 Computer Software

As a Research Scientist at Wayve Labs, you will develop cutting-edge AI systems for autonomous driving, focusing on machine learning, simulation, and robotics.

Machine LearningReinforcement LearningComputer VisionRobotics
Thinkingmachines

Research, Coding Agents

πŸ“… Aug 27
ThinkingmachinesπŸ‘₯ 51 - 200 employees🏒 Technology

Join a high-leverage team to enhance AI coding capabilities through research, design, and execution of RL training jobs and data generation.

PythonDeep Learning FrameworksReinforcement LearningData Generation

Trusted by Remote Workers