On demand Wed 19 Aug

Beyond Benchmarks: Evaluating Agents Against What They Are Actually Supposed to Do

Hosted by TestMu AI

Watch the recording →
When
Watch any time · recorded 19 Aug
Format
On demand
Speaker
Francesca Lazzeri
agent evaluationllm evaluationbenchmarksai agents