Agents combine a harness, model, and contextChase describes an agent as three main components: a harness that orchestrates a model and its context, and argues that owning an organization's intelligence involves control over all three.Harrison ChaseAgent HarnessesSource at 1:09Save viewpoint
A harness brings context to the model when neededChase says the harness's primary role is orchestrating fixed and dynamic context so the model receives the appropriate information at the appropriate point in its execution loop.Harrison ChaseAgent HarnessesSource at 2:02Save viewpoint
Most agents share a simple model-and-tools loopChase characterizes the core architecture of most agents as an LLM repeatedly generating responses, calling tools when needed, and receiving tool observations back into the loop.Harrison ChaseAgent HarnessesSource at 2:43Save viewpoint
Middleware can customize the core agent loopChase says hooks, plugins, and middleware can modify different stages of a basic agent loop to add capabilities such as file systems, sub-agents, memory, summarization, and context offloading.Harrison ChaseHarness CustomizationSource at 3:25Save viewpoint
Start general and specialize as requirements become clearerChase recommends starting with a general-purpose harness for faster time to value, then adding gates, checks, or more explicit cognitive architecture as the target use case becomes more specialized.Harrison ChaseOff-the-Shelf vs Custom HarnessesSource at 6:37Save viewpoint
Out-of-distribution tasks require more harness customizationChase argues that off-the-shelf harnesses work better when tasks resemble what models were trained on, while increasingly out-of-distribution tasks create a stronger need to customize the harness.Harrison ChaseOff-the-Shelf vs Custom HarnessesSource at 6:55Save viewpoint
Custom harnesses should preserve model-native tool patternsChase says a domain-specific harness can still retain implementations for subtasks that match a model's training, such as using the file-editing behavior best aligned with the particular model.Harrison ChaseModel-Specific ToolingSource at 7:46Save viewpoint
Mission-critical agents should have task-specific benchmarksChase expects companies building mission-critical agents to create benchmarks that define desired performance, catch regressions, and guide improvements to either the harness or the model.Harrison ChaseAgent EvaluationSource at 10:11Save viewpoint
Agent benchmarks should measure more than accuracyChase says agent evaluation should account for dimensions such as latency and token usage or cost in addition to task accuracy.Harrison ChaseAgent EvaluationSource at 12:59Save viewpoint
Poor context often causes agent failuresChase argues that agent failures frequently stem from inadequate context supplied to the model rather than insufficient model capability, making visibility into context construction important for debugging.Harrison ChaseAgent ObservabilitySource at 13:04Save viewpoint
Detailed traces are necessary for debugging agentsChase says message trajectories alone are insufficient for complete debugging and advocates detailed traces that expose tool calls, context accumulation, execution steps, and model interactions.Harrison ChaseAgent ObservabilitySource at 13:43Save viewpoint
Production traces can drive continuous agent improvementChase outlines a data flywheel in which teams run agents, collect and curate traces, experiment on the resulting data, and use those findings to improve the system.Harrison ChaseContinuous Agent ImprovementSource at 14:39Save viewpoint
Agent UX can generate useful implicit feedbackChase argues that thoughtful agent UX can capture behavioral feedback from users even when they do not explicitly provide thumbs-up or thumbs-down ratings.Harrison ChaseFeedback and CommunicationSource at 16:16Save viewpoint
Trace data can improve every major agent componentChase says feedback and trace data can be used to update the harness through harness engineering, the model through fine-tuning, and the context through memory mechanisms.Harrison ChaseContinuous Agent ImprovementSource at 17:16Save viewpoint
Agent improvement can be partially automated from tracesChase presents LangSmith Engine as an attempt to automate trace curation, issue discovery, experimentation, and suggested changes to prompts, context, or harness code.Harrison ChaseContinuous Agent ImprovementSource at 17:30Save viewpoint
Benchmark comparisons can reveal techniques worth adoptingChase says comparing different models and harnesses on the same benchmark can expose effective behaviors that can then be incorporated into an organization's own agent harness.Harrison ChaseAgent BenchmarksSource at 19:51Save viewpoint
Harness customization exists on a spectrumChase says organizations can range from using an off-the-shelf harness to adding middleware and hooks or building a highly controlled cognitive architecture, depending on task distribution and control requirements.Harrison ChaseOff-the-Shelf vs Custom HarnessesSource at 21:02Save viewpoint
Predictability can justify more controlled agent architecturesChase says some financial-services customers prefer more explicit cognitive architectures because they prioritize predictability and control over the flexibility of a general-purpose agent.Harrison ChaseOff-the-Shelf vs Custom HarnessesSource at 22:10Save viewpoint