On demand

The Eval Stack Top AI Teams Are Building Right Now

Hosted by Encord

As models get more capable, automated evals stop telling you much. The signal that's left, whether the model is actually improving, comes from structured human judgment at scale. Most teams don't have the infrastructure to produce it; this session is about how those that do have built it. Covers where automated evals fall short, what separates a rigorous human eval pipeline from ad-hoc annotation, the failure modes teams keep hitting when they try to scale human feedback, and where this is all heading as models get more capable.

Watch the recording →
When
Watch any time
Format
On demand
For
AI teams
model evaluationhuman feedbackllm evaluationannotation