Privacy and advertising choices

Git-Stars uses essential storage for site operation. Optional analytics and ad-measurement scripts stay disabled unless you accept them; partners such as Google may then use cookies or similar identifiers where required. Privacy Policy

LogoGit-Stars
Top StarsTrendingAI AgentsDaily PicksViral ReposInsights
LogoGit-Stars

Discover top GitHub projects with real rankings and AI insights

GitHub
Built withLogo of Git-StarsGit-Stars
Rankings
  • Top Stars
  • Trending
  • AI Agents
  • Daily Picks
  • Explore
Resources
  • Insights
  • Editorial Policy
About
  • About
  • Contact
Legal
  • Privacy Policy
  • Terms of Service
© 2026 Git-Stars. All Rights Reserved.
AI Agent Analysis
HL

harveyai/harvey-labs

A benchmark built to evaluate and improve agent capabilities for supporting legal work.

stars
1.2k
Language
Python
GitHub
Source and compliance noteLast synced: Aug 16, 2026

Git-Stars is independent and not affiliated with GitHub or this project. Analysis may be AI-assisted and based on public repository metadata plus short README-derived summaries. We do not mirror full README files, docs, issues, or social comments.

Original GitHub sourceMethodologyEditorial Policy

Overview

Harvey LAB is an open-source benchmark and execution harness for evaluating LLM agents on realistic legal work, providing a dataset of 1,671 tasks across 24+ practice areas with agent instructions, documents, and rubrics, plus a harness to run and score agents. Its core value proposition is to standardize and advance the measurement of AI agent capabilities in legal domains, enabling reproducible comparisons and improvements.

Installation

Clone the repo and follow the walkthrough in docs/tutorial.md; no single install command is provided, but the harness is Python-based and can be run via standard package management.

Problem solved

It addresses the lack of standardized, realistic benchmarks for legal AI agents, which often rely on generic QA datasets or proprietary evaluations. By providing a domain-specific, task-based benchmark with rubrics and an execution harness, it enables objective, reproducible assessment of agent performance on complex legal workflows, facilitating targeted improvements and fair comparisons across models and frameworks.

What you can build

Developers can build and evaluate AI agents that perform legal tasks such as M&A data-room analysis, contract review, legal research, and drafting, using the provided tasks and harness. The benchmark supports custom task creation, model adapters, and evaluation sweeps, allowing teams to measure agent performance, identify weaknesses, and iterate on agent designs. The ceiling includes achieving human-level or superhuman performance on specific legal workflows, with the potential to drive adoption of AI in legal practice by demonstrating reliability and quality.

Community sentiment

Positive

No community feedback yet.

Concerns

No concerns documented yet.

Bottom line

Harvey LAB is a valuable resource for AI researchers, legal tech developers, and law firms seeking to rigorously evaluate and improve AI agents for legal work. It is not suitable for those looking for a plug-and-play legal AI solution or for non-technical legal professionals without engineering support. The key trade-off is the significant effort required to set up and run the benchmark versus the benefit of obtaining domain-specific, actionable performance insights.

Analyzed by Git-Stars - 8/10/2026