Viewpoint
Agent improvement can be partially automated from traces
Agent improvement can be partially automated from traces
Chase presents LangSmith Engine as an attempt to automate trace curation, issue discovery, experimentation, and suggested changes to prompts, context, or harness code.
- Speaker
- Harrison Chase
- Source timestamp
- 17:30
More from this interview
- Agents combine a harness, model, and context
- A harness brings context to the model when needed
- Most agents share a simple model-and-tools loop
- Middleware can customize the core agent loop
- Start general and specialize as requirements become clearer
- Out-of-distribution tasks require more harness customization
- Custom harnesses should preserve model-native tool patterns
- Mission-critical agents should have task-specific benchmarks
- Agent benchmarks should measure more than accuracy
- Poor context often causes agent failures
- Detailed traces are necessary for debugging agents
- Production traces can drive continuous agent improvement
- Agent UX can generate useful implicit feedback
- Trace data can improve every major agent component
- Benchmark comparisons can reveal techniques worth adopting
- Harness customization exists on a spectrum
- Predictability can justify more controlled agent architectures