Draft hiring plan. The nine roles, location, pay, perks and hiring steps need owner confirmation. Applications record interest in these drafts.
All roles

Engineering

Applied AI Engineer, Evaluation

You run the benchmarks and checks that decide how Agent routes work between models. Every result goes public with its sample size and its caveats, so the method has to hold up.

  • Remote
  • Full time
  • Senior
  • Pay terms to confirm
Proof you can reproduce
  1. 1Fixture
  2. 2Measure
  3. 3Compare

Meet Sage“What is the interval?”

What success looks like

  • Routing choices rest on measured results, not on a feeling about a model.
  • Each public study states what it did not test.
  • A cheaper route ships only when the checks say it keeps the quality.

What you will do

  • Design tasks, graders and protocols, and declare them before the first run.
  • Run studies across models and harnesses, with intervals and honest tie calls.
  • Turn a finding into a product change, then measure the change.
  • Write the study in plain words for people who will check your work.

About you

  • You know the difference between a result and a story about a result.
  • You are comfortable with confidence intervals and with saying “we do not know yet”.
  • You write clearly and you do not oversell.

Nice to have

  • Experiment design
  • Cost and latency analysis
  • Public technical writing

A short note, please

With your resume, write a page or less on the work you are proudest of. Tell us what you chose to cut, and which of our values you would hold us to.

Start a conversation

Apply to the Applied AI Engineer, Evaluation draft

This role is a draft. The role, location, pay and hiring steps need owner confirmation. You can send your interest now. We will store it for the hiring team to review.

Your application

Explore the team.

Read the other role drafts and tell us what you would build first.