Navigate Select ESC Close

The best AI agents are secretly teams

2026-07-02 Science & Technology
1.9k
69
2
LangChain
LangChain
194.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Ben Tannyhill is a product manager at LangChain, where he's building LangSmith Engine—an agent that finds and fixes your agent's failures. Engine continuously analyzes your production traces, clusters them into actionable issues, and opens pull requests to fix them. Engine's architecture is a lot like an org chart: a main model delegating to a team of cheaper, faster sub-agents. It launched in public beta at Interrupt 2026, and in this conversation, Ben unpacks why it uses a sandbox as a tool, how the team turned it into a self-improving agent that learns from its own traces, and the hard problem of testing a fix before it ships. We also discuss: • Why Engine is "the agent for agent engineers" • Making LangSmith agent-native with condensed trace views • Why the team keeps handing more control to the agent • Inside Engine's four sub-agents: the screener, verifier, and more • Giving Engine memory with an agent overview document • How to keep an always-on agent from blowing the inference budget • Where Insights, Polly, and Engine are converging Timestamps: 00:00 Introduction 01:25 LangSmith 101 02:22 Why Engine is "the agent for agent engineers" 03:49 Under the hood: Engine is a deep agent 06:08 Clustering millions of traces with condensed views 10:10 Why the team keeps handing more control to the agent 13:21 Why Engine uses a sandbox as a tool 14:11 Engine's four sub-agents and the org-chart analogy 16:51 Evals for Engine: IssueBench, Harbor, and synthetic environments 23:05 How Engine evolved: from noisy PRs to an issue inbox 25:56 Inside Engine's memory: the agent overview document 29:25 How to keep an always-on agent from blowing the inference budget 30:52 What models Engine uses 31:30 How Engine was rolled out: from Forge to public beta at Interrupt 34:18 Inside the two teams building Engine 35:53 Where Insights, Polly, and Engine are converging 40:06 The missing piece: testing a fix before it ships 42:22 Running a branched agent, and the write-access eval problem 46:35 Using Engine as long-term memory 47:39 Pointing Engine at coding-agent traces 48:49 Running Engine on Engine: the meta self-improvement loop References: • Anthropic: https://www.anthropic.com/ • Chat LangChain: https://chat.langchain.com/ • Claude Code: https://www.anthropic.com/claude-code • Claude Haiku: https://www.anthropic.com/claude/haiku • Claude Opus: https://www.anthropic.com/claude/opus • Codex: https://openai.com/codex/ • Context Hub: https://docs.langchain.com/langsmith/use-the-context-hub • Credit Genie: https://www.creditgenie.com/ • Deep Agents: https://docs.langchain.com/oss/python/deepagents/overview • Gemini: https://gemini.google.com/ • GPT-5.5: https://openai.com/index/introducing-gpt-5-5/ • Harbor: https://www.harborframework.com/ • Hex: https://hex.tech/ • Insights: https://docs.langchain.com/langsmith/insights • Interrupt: https://interrupt.langchain.com/ • LangGraph: https://www.langchain.com/langgraph • LangSmith: https://smith.langchain.com/ • LangSmith Chat (formerly Polly): https://docs.langchain.com/langsmith/chat • LangSmith Engine: https://www.langchain.com/langsmith/engine • LangSmith Observability: https://www.langchain.com/langsmith/observability • Mintlify: https://mintlify.com/ • OpenAI: https://openai.com/ • Palash Shah: https://www.linkedin.com/in/palash-sh/ • Terminal-Bench: https://www.tbench.ai/ • Unify: https://www.unifygtm.com/ Where to find Ben: • LinkedIn: https://www.linkedin.com/in/benjamintannyhill/ • Twitter/X: https://x.com/bentannyhill Where to find Harrison: • LinkedIn: https://www.linkedin.com/in/harrison-chase-961287118/ • Twitter/X: https://x.com/hwchase17 Where to find LangChain: • Website: https://langchain.com • Docs: https://docs.langchain.com/ Send feedback or questions to [email protected]

Top Comments (5)

@duythvn 2026-07-03

Those are valuable info and for some reason Im the first person to leave a comment. Thanks for sharing team

3
@LuminairPrime 2026-07-04

this is definitely the way to go. logging text is cheap. every word you and the model type needs to be kept and processed to make the next steps better.

2
@Dom-zy1qy 2026-07-04

The guy did a good job explaining what Engine is. Especially considering how abstract their whole system is. Have you guys considered training a binary classifier to identify problematic traces? That could be another preliminary filter before passing the summarized version through, perhaps.

1
@ayeoh47 2026-07-03

I took a shot every time he said the word agent and now I’m in the ER

0
@plpk 2026-07-04

One agent is enough. We're not all rich like y'all founder types.

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot