Navigate Select ESC Close

Building a Production Agent Eval Pipeline: Harbor + LangSmith + OpenAI SDK

2026-07-06 Science & Technology
1.5k
38
0
LangChain
LangChain
193.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Nick, a software engineer at LangChain, walks through running agent evaluations with Harbor and LangSmith. He builds a small OpenAI SDK agent that edits a sales CSV file, shows why testing purely on your local machine breaks down as tasks scale, then demonstrates how Harbor's isolated sandboxes, structured datasets, and LangSmith's observability platform let you run and score many agent tasks in parallel with full tracing. Chapters: 0:00 Why agent evaluations need to change 0:33 The three things every agent eval needs 1:02 Meet the demo agent: an OpenAI SDK data analyst 1:59 Running the agent locally and checking the results 2:41 Two problems with testing purely locally 3:05 Introducing Harbor: bring your agent, sandbox, and dataset 3:39 Anatomy of a Harbor dataset 4:32 How sandbox isolation actually works 5:23 Turning the demo into a real Harbor evaluation 6:01 Installing Harbor and setting up API keys 6:40 Running harbor run against two tasks 7:35 Viewing results in LangSmith datasets and experiments 8:09 Drilling into a single experiment and its trace 8:54 Wrap-up: bring your own agent and try it

Top Comments (1)

@stellarcoding 2026-07-08

well done! can you share the code used in the video ?

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot