As a Deployment Engineer, you will onboard customers, provide technical support, and manage deployment challenges while enhancing the customer experience.
Software Engineer | AI Training Data & Evals Lab
📜 Description
- Build and maintain evaluation harnesses that measure the performance of AI models and agents on real-world tasks.
- Improve evaluation reliability, coverage, and signal quality through better rubrics and task design support.
- Develop tools that enable researchers and operators to run experiments efficiently without rebuilding workflows.
- Build and maintain APIs and backend services supporting human-in-the-loop workflows and quality-control processes.
- Improve data pipelines that transform expert work into structured training and evaluation datasets.
- Document technical decisions and system behavior clearly for effective collaboration.
🛠️ Requirements
- Strong software engineering fundamentals with professional experience in Node.js and TypeScript.
- Strong coding ability in Python and/or Go.
- Experience building, deploying, and owning production systems, including APIs and data pipelines.
- Solid understanding of distributed systems, scalability, and reliability.
- Experience with AWS or GCP and modern infrastructure technologies such as containers and Kubernetes.
- Strong written communication skills for effective collaboration in a distributed environment.
✨ Benefits
- Direct impact on the training data and evaluation systems used by leading AI labs.
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer | AI Training Data & Evals Lab based in United States.
This is a broad engineering role at the intersection of platform development, AI evaluation, and experimentation infrastructure.
You will build and operate production systems that researchers and operators rely on to develop and assess frontier AI capabilities.
Your work will include evaluation harnesses, backend services, data pipelines, training environments, and tools that make experimentation more reliable and repeatable.
You will have meaningful ownership from the start, with an expectation of delivering production improvements and becoming a trusted owner of core systems.
The environment is lean, async-first, and highly collaborative, with an emphasis on clear communication, sound technical judgment, and dependable execution.
You will work on practical engineering challenges closely connected to AI research while helping transform expert work into high-quality training and evaluation data.
This opportunity is well suited to a strong builder who enjoys autonomy, distributed systems, and working on infrastructure that directly influences how AI systems are evaluated and improved.
Accountabilities:
- Build and maintain evaluation harnesses that measure the performance of AI models and agents on real-world tasks.
- Improve evaluation reliability, coverage, and signal quality through better rubrics, task design support, and scoring approaches.
- Develop tools that enable researchers and operators to run experiments efficiently without repeatedly rebuilding the same workflows.
- Build and maintain APIs and backend services supporting human-in-the-loop workflows, task routing, and quality-control processes.
- Improve data pipelines that transform expert work into structured training and evaluation datasets.
- Strengthen system observability, scalability, and operational reliability through effective logging, metrics, monitoring, and debugging capabilities.
- Write clear, maintainable production code and actively participate in code reviews, architecture discussions, and technical design decisions.
- Document technical decisions and system behavior clearly so that other engineers and collaborators can build upon and operate the systems effectively.
- Take ownership of core systems from development through production operation, with an expectation of delivering meaningful improvements within the first 30–90 days.
- Strong software engineering fundamentals with professional experience in Node.js and TypeScript.
- Strong coding ability in Python and/or Go.
- Demonstrated experience building, deploying, and owning production systems, including APIs, backend services, and data pipelines.
- Solid understanding of distributed systems, scalability, reliability, and engineering trade-offs.
- Experience working with AWS or GCP and modern infrastructure technologies such as containers and Kubernetes.
- Proven track record of shipping and maintaining production systems that other people depend on, rather than working exclusively on prototypes.
- Strong written communication skills and the ability to collaborate effectively in an asynchronous, distributed environment.
- Comfortable taking ownership of ambiguous technical problems, making sound engineering decisions, and following projects through to production.
- Experience with evaluation frameworks, experimentation platforms, or machine-learning tooling is a plus.
- Experience with data pipelines, workflow orchestration, or internal platforms for research and operations teams is a plus.
- Experience working in early-stage environments or high-ownership B2B SaaS and platform teams is a plus.
- Full-time, fully remote position with a LATAM focus and meaningful overlap with U.S. time zones.
- Compensation of $7,000–$10,000 USD per month, based on experience.
- Significant ownership and opportunities to grow into larger systems, deeper technical leadership, and projects central to the organization’s growth.
- Lean, async-first working environment focused on clear writing, sound judgment, and strong follow-through.
- Opportunity to work on research-adjacent engineering challenges at the frontier of AI while building practical production platforms.
- Direct impact on the training data and evaluation systems used by leading AI labs.
- Structured hiring process including a practical take-home assignment, team review, technical screen, real-world work trial, and final offer stage.
Requirements:
Benefits:
Similar jobs
Search more Software Engineer jobsLMR Application Support & ROL Testing Engineer
Support and test the Radio Online (ROL) application by providing user support, conducting system testing, and collaborating within an Agile team to enhance application performance.
Develop and evolve a high-availability distributed PostgreSQL solution, focusing on architecture, software development, and performance optimization in a remote EMEA role.
Develop and evolve a high-availability distributed PostgreSQL solution, focusing on architecture, software development, and performance optimization within a remote, globally distributed team.
M365 Engineer — Purview Compliance & Messaging (f/m/d)
As an M365 Engineer, you will design and maintain compliance policies, lead stakeholder meetings, automate processes, and ensure the smooth operation of a hybrid Exchange environment.
