Remote Jobs RockRemote Jobs Rock

Research, Post-Training Evals

🕒 8 days ago
Computer ScienceMachine LearningPhysicsMathematics

📜 Description

  • Work closely with researchers and engineers across post-training and the broader research organization.
  • Create internal evaluations and research signals for capabilities and behaviors important to model research and post-training.
  • Develop usability evaluations that measure whether models are genuinely useful in real research and product workflows.
  • Improve evaluation robustness, including grader reliability and gaps between measured and intended behavior.
  • Build benchmark auditing methodologies to help researchers understand and use evaluation signals effectively.
  • Develop evaluations for personalized preferences, biases, and other nuanced dimensions of model behavior.

🛠️ Requirements

  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Experience designing, building, or analyzing evaluations, benchmarks, datasets, graders, or other measurement systems.
  • Strong written and verbal communication skills, with the ability to collaborate effectively across research and engineering teams.
  • Experience with LLMs, post-training, reinforcement learning, or agentic systems.
  • Experience with evaluation auditing, human evaluations, LLM-judges, or open-ended task evaluation.
  • Experience with agentic evaluation, harnesses, long-horizon tasks, or RL environments.
  • Experience evaluating preferences, personalization, biases, values, or other nuanced model behaviors.
  • Track record of developing new evaluation methodologies or research signals that meaningfully influenced model development.
  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX). Comfortable with debugging distributed training and writing code that scales.
  • Strong research judgment: clean ablations, honest baselines, and clear technical writing.
Snorkelai

Research Scientist - Frontier Benchmarks

Snorkelai👥 1001 - 5000 employees🏢 Computer Software
🕒 4 days ago

We're looking for a Research Scientist to collaborate with partners and lead the development of the next frontier benchmarks and datasets. This is a highly visible, customer-facing role at the intersection of research, company strategy,.

AIMachine LearningNLPExperimental Design
Anthropic

Research Scientist, Life Sciences

Anthropic👥 10,000+ employees🏢 Research Services🤝 B2B
🕒 14 days ago

Join the Life Sciences team as a Research Scientist to enhance AI capabilities in biological research, focusing on model training, evaluation, and real-world applications.

Machine LearningSoftware EngineeringComputational BiologyBioinformatics
🕒 yesterday

In this role, you will train and evaluate frontier models to enhance agent safety, develop scalable measurement systems, and collaborate on research-backed mitigations.

AI SafetyMachine Learning EngineeringQuantitative ResearchData Processing

Trusted by Remote Workers