Navigate Select ESC Close

Harbor x LangChain: A Unified Stack for Evaluating Agents

2026-07-01 Science & Technology
2.6k
61
2
LangChain
LangChain
193.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

As agents increase in capabilities, evaluations have gotten more difficult. Agent harnesses like Claude Code, Pi, and Deep Agents now give agents access to entire computers to read files, execute scripts, run code, and more. Every agent now needs to run in its own clean, reproducible environment for a given task. Evaluating long-running, stateful agents requires a new eval runner. Harbor has emerged as the industry leader in this space. In this blog, we first explain why everyone running agent evals should know what Harbor is and then show how to integrate Deep Agents, LangSmith Sandboxes, and LangSmith Experiments into Harbor. We ultimately need to run agents in a real, reproducible, isolated environment, many times in parallel, with a deterministic check at the end. Harbor solves this problem and is now wired directly into Deep Agents, LangSmith Sandboxes, and LangSmith Observability. Join LangChain software engineer Nick for a look at: - Why traditional output-based evals fall short for complex, long-running AI agents. - Harbor, an open source evaluation framework that runs agents in isolated, reproducible sandboxes - The full workflow: building a Deep Agent, structuring an eval dataset, and tracking results in LangSmith 0:00 Why agent evals needed to change 0:32 From output strings to real environments 1:07 What makes a Deep Agent different 1:40 Building a deep research agent 2:47 Running the agent: noise pollution demo 3:55 Introducing Harbor 4:33 Datasets and tasks in Harbor 6:15 Turning a demo into a verifiable eval 7:57 Integrating Deep Agent with Harbor 8:32 Running Harbor via CLI 9:21 Viewing results in LangSmith 10:05 Recap and getting started https://www.harborframework.com/docs https://docs.langchain.com/oss/python/deepagents/overview https://docs.langchain.com/langsmith/harbor-integrations

Top Comments (2)

@GayathriG-h5h 2026-07-02

Pytest is not enough

0
@gordeyvasilev 2026-07-02

👍

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot