Remote Jobs RockRemote Jobs Rock

Research Engineer, Post-Training (All Industry Levels)

📅 Jan 24, 2024
Alignment AlgorithmsLoss FunctionsData PipelinesQuality Signals

📜 Description

  • Develop alignment algorithms and loss functions to improve data sample efficiency.
  • Write data pipelines to process diverse web data into a format models can ingest.
  • Identify quality signals to understand our model's performance in the real world.
  • Design sampling algorithms to improve serving efficiency of large generative models.

🛠️ Requirements

  • "All Industry Levels": have at least PhD (or equivalent)
  • Write clear and clean production-facing and training code
  • Experience working with GPUs (training, serving, debugging)
  • Experience with data pipelines and data infrastructure
  • Strong understanding of modern machine learning techniques (reinforcement learning, transformers, etc)
  • Track-record of exceptional research or creative applied ML projects
  • Experience with product experimentation and A/B testing
  • Experience training large models in a distributed setting
  • Familiarity with ML deployment and orchestration (Kubernetes, Docker, cloud)
  • Publications in relevant academic journals or conferences in the field of machine learning
Full job description

About the role and team

Joining us as a Research Engineer on the Post-Training team, you'll be diving into the exciting world of fine-tuning AI models, optimizing their performance, and ensuring they meet the highest standards of quality and efficiency. Your work will directly contribute to our groundbreaking advancements in AI, helping shape an era where technology is not just a tool, but a companion in our daily lives. At Character.AI, your talent, creativity, and expertise will not just be valued—they will be the catalyst for change in an AI-driven future.

The Post-Training team is responsible for developing our powerful pretrained language models into intelligent, engaging, and aligned products.

As a Post-Training Researcher, you will work across teams and our technical stack to improve our model performance and training methods, including data, compute and algorithms. You will get to shape the conversational experience of millions of users per day.

What you'll do

  • Develop alignment algorithms and loss functions to improve data sample efficiency.

  • Write data pipelines to process diverse web data into a format models can ingest.

  • Identify quality signals to understand our model’s performance in the real world.

  • Design sampling algorithms to improve serving efficiency of large generative models.

Who you are

  • "All Industry Levels": have at least PhD (or equivalent)

  • Write clear and clean production-facing and training code

  • Experience working with GPUs (training, serving, debugging)

  • Experience with data pipelines and data infrastructure

  • Strong understanding of modern machine learning techniques (reinforcement learning, transformers, etc)

  • Track-record of exceptional research or creative applied ML projects

Nice to Have

  • Experience with product experimentation and A/B testing

  • Experience training large models in a distributed setting

  • Familiarity with ML deployment and orchestration (Kubernetes, Docker, cloud)

  • Publications in relevant academic journals or conferences in the field of machine learning

About Character.AI

Character.AI empowers people to connect, learn and tell stories through interactive entertainment. Over 20 million people visit Character.AI every month, using our technology to supercharge their creativity and imagination. Our platform lets users engage with tens of millions of characters, enjoy unlimited conversations, and embark on infinite adventures.


In just two years, we achieved unicorn status and were honored as Google Play's AI App of the Year—a testament to our innovative technology and visionary approach.


Join us and be a part of establishing this new entertainment paradigm while shaping the future of Consumer AI!

At Character, we value diversity and welcome applicants from all backgrounds. As an equal opportunity employer, we firmly uphold a non-discrimination policy based on race, religion, national origin, gender, sexual orientation, age, veteran status, or disability. Your unique perspectives are vital to our success.

Anthropic

Research Engineer, Cybersecurity RL (Reinforcement Learning)

Anthropic👥 10,000+ employees🏢 Research Services🤝 B2B
🕒 yesterday

As a Research Engineer in Cybersecurity, you'll advance AI capabilities in incident response and security analysis while developing novel approaches and implementing them in code.

Machine LearningCybersecuritySoftware EngineeringReinforcement Learning
Anthropic

Research Engineer, Machine Learning (RL Velocity)

Anthropic👥 10,000+ employees🏢 Research Services🤝 B2B
🕒 11 days ago

The RL Velocity team owns the efficiency and reliability of our RL Science stack - the infrastructure, tooling, and systems that let researchers iterate quickly on training runs.

Software EngineeringMachine LearningDistributed SystemsResearch Tooling
Anthropic

Research Engineer, Machine Learning (Reinforcement Learning)

Anthropic👥 10,000+ employees🏢 Research Services🤝 B2B
🕒 11 days ago

As a Research Engineer in Reinforcement Learning, you will advance the capabilities and safety of large language models through innovative research and engineering practices.

PythonAsync ProgrammingMachine Learning FrameworksPytorch

Trusted by Remote Workers