LangSmith Helps Abridge Cut AI Release Cycles to Days
Healthcare developers Abridge and Included Health are using LangSmith to turn clinical reviews into automated evaluations, drastically accelerating their AI release cycles.

Medical AI companies are leveraging LangSmith to convert scarce clinical expertise into reusable evaluation infrastructure. By turning manual clinician feedback into durable datasets and automated judges, these teams are bypassing traditional development bottlenecks. For instance, clinical documentation startup Abridge has utilized LangSmith to shrink its software release cycle from one or two months down to just a few days. The company achieves this by building specialized, automated judges that evaluate notes for specific failure modes like style, completeness, and compliance.
Abridge combines reference-free judges, which score notes directly against source transcripts, with reference-based judges tailored to specific medical specialties. To validate these automated evaluators, the team uses LangSmith’s Align Evaluator to compare AI scores against human annotations. When deploying updates, Abridge runs offline evaluations and backtests before initiating a silent rollout to 10% to 15% of its customers. This gradual pipeline allows the team to gather real-world feedback on clinician edits before committing to a full release.
Similarly, Included Health uses LangSmith to monitor Dot, an AI healthcare guide built with LangGraph and Deep Agents. Following Dot's launch, Included Health saw a 75% increase in chat engagement, with the system successfully identifying more than 99% of high-risk situations. Clinician agreement with Dot's routing recommendations has remained above the company's 95% target. Furthermore, the team successfully migrated Dot's supergraph to Deep Agents in under two weeks without regressions by validating the change against their existing multi-turn simulation suite.
For AI practitioners, this shift transforms human evaluation from a recurring operational cost into a compounding asset. Instead of starting from scratch with each model update, developers can continuously test new iterations against historical clinician decisions stored in LangSmith annotation queues. Because these workflows involve protected health information, LangSmith supports managed cloud, self-hosted, and bring-your-own-cloud deployments, allowing teams to scrub sensitive data and maintain strict access controls.
This is our own summary of reporting by LangChain Blog



