Viewpoint
Evals should closely reflect real production usage
Evals should closely reflect real production usage
Karanam says evaluation environments should resemble both production use and training environments, with evals drawn directly from how users actually interact with the product.
- Speaker
- Arjun Karanam
- Topic
- Production Evals
- Source timestamp
- 3:56
More from this interview
- AI progress needs experience alongside model intelligence
- Agent interactions should become signals for future improvement
- Continual learning can create systems that compound with use
- Feedback should update either models or their surrounding harness
- Full agent traces should include tools and sub-agents
- Corrective behavior provides richer feedback than simple ratings
- Agent tasks should be replayable for evaluation
- Harnesses should enable orchestration instead of rigid workflows
- Agents should access the same capabilities as users
- Tool responses should expose informative execution details
- Open weights enable companies to continually improve owned models
- Model routers can match intelligence to task requirements
- Continual learning should optimize the whole intelligence system
- Learning can avoid direct training on customer data
- Corrective feedback gives stronger reward signals than dissatisfaction
- Feedback scope determines where learned information should live
- Continual learning is most compelling near capability frontiers