Phase 7: AI Evaluation & Observability · 55 min · Python · LangSmith · Braintrust
Observability Platforms: LangSmith, Braintrust, HoneyHive, Phoenix
Trace everything. Evaluate continuously. Alert on regressions.
Hiring signal: FDE roles at AI companies test observability platform knowledge: candidates who can compare LangSmith (trace capture, LLM-as-judge), Braintrust (eval-first, Brainstore OLAP), HoneyHive (agent-specific), and Phoenix (open-source, self-hosted) pass. Candidates who use only printf debugging fail. Platform selection criteria (self-hosting requirements, agent evaluation depth, CI/CD integration) demonstrate production experience.
What you will learn
- Configure LangSmith: trace capture, LLM-as-judge evaluation, pattern detection, CI/CD integration
- Configure Braintrust: eval-first workflow, prompt versioning, Brainstore OLAP, production trace to test case
- Configure HoneyHive: agent-specific evaluation, session-level evaluations, agent session annotations
- Configure Phoenix: open-source, self-hosted, agent tracing, online evaluation, local deployment
- Select platforms based on criteria: self-hosting, agent evaluation depth, CI/CD integration, cost at scale
What You'll Learn
This lesson takes approximately 55 min. By the end, you will be able to:
- Configure LangSmith: trace capture, LLM-as-judge evaluation, pattern detection, CI/CD integration
- Configure Braintrust: eval-first workflow, prompt versioning, Brainstore OLAP, production trace to test case
- Configure HoneyHive: agent-specific evaluation, session-level evaluations, agent session annotations
- Configure Phoenix: open-source, self-hosted, agent tracing, online evaluation, local deployment
- Select platforms based on criteria: self-hosting, agent evaluation depth, CI/CD integration, cost at scale
The Problem
Trace everything. Evaluate continuously. Alert on regressions. LangSmith, Braintrust, HoneyHive, and Phoenix are the platforms FDEs use to get visibility into LLM applications: token-level traces, retrieval quality metrics, latency breakdowns, and cost tracking. This lesson covers platform selection, integration, and the dashboards that matter.
Unlock the full lesson
You've read the first 2 sections. The rest of this lesson covers LangSmith: Trace Capture and LLM-as-a-Judge, Braintrust: Eval-First Workflow, Phoenix: Open-Source Self-Hosted Observability, Platform Selection Criteria, Practical Application, What Hiring Managers Look For, Resources, Key Takeaways, Next Steps — plus a hands-on lab, quiz, and project artifact.
Create a free account to unlock Phase 0 and Phase 1 of every course — no credit card.
Browse all courses · View pricing · DeVenture Academy