
5 Open-Source LLM & Agent Evaluation Tools in 2026

An AI run can be tricky to evaluate whenever the failure isn't visible from the final answer — which is often the case. We're talking about RAG systems retrieving the wrong evidence, an agent choosing the wrong tool on its way to a plausible response, and a coding agent producing a convincing transcript without actually completing the task.
For this comparison, we selected five open-source projects that test different parts of an LLM or agent workflow: DeepEval, Ragas, Promptfoo, Inspect AI, and Opik. We'll compare them by what they evaluate, how their testing workflows differ, and how far they go beyond scoring a single model response.
Open-source LLM evaluation tools by evaluation layer
A compound quality score can hide a range of problems, so it helps to inspect the LLM or agent workflow layer by layer:
| Evaluation layer | What it asks | Example |
|---|---|---|
| Output | Did the model produce an acceptable response? | Correctness, relevance, safety, instruction following |
| Retrieval | Did the system retrieve the evidence needed to answer? | Context precision, recall, faithfulness |
| Trajectory | Did the agent take sensible steps and use tools correctly? | Tool selection, arguments, plan adherence, step efficiency |
| Task outcome | Did the agent actually complete the job? | File created, environment changed, test passed, goal reached |
AI memory benchmarks such as BEAM, LoCoMo, LongMemEval, HotPotQA, and AgentBench define tasks for measuring memory, retrieval, reasoning, and agent behavior, whereas evaluation tools provide the machinery for running those tasks and scoring the results.
The five tools we discuss below test different aspects of the LLM/agent workflow, and all are available under permissive open source licenses. Here's a quick overview:
| Tool | Main evaluation layer | Typical workflow | Agent evaluation | License |
|---|---|---|---|---|
| DeepEval | Outputs, conversations, and trajectories | Python tests, pytest, CLI, CI | Plans, tools, arguments, steps, task completion | Apache 2.0 |
| Ragas | Retrieval and grounded answers | Python metrics and evaluation datasets | Tool-call and goal metrics | Apache 2.0 |
| Promptfoo | Regression and security testing | YAML/CLI test matrices and CI | Tool and trajectory assertions | MIT |
| Inspect AI | Task and environment outcomes | Python evaluation tasks in controlled environments | Core focus | MIT |
| Opik | Experiments and trajectories | Datasets, experiments, metrics, UI | Full trajectory evaluation | Apache 2.0 |
1. DeepEval — pytest-style evals for LLM apps and agents
Run volume, quality trend, and per-metric scores in one shared dashboard.
DeepEval brings LLM evaluation into a software-testing workflow. Tests can run locally through Python or deepeval test run, use score thresholds for individual cases, and become CI checks that fail when model behavior drops below an accepted level. The project currently includes 50+ built-in metrics for RAG, chatbots, safety, multimodal systems, tool use, and agents.
Its agent evaluation includes four trajectory metrics: Task Completion, Step Efficiency, Plan Adherence, and Plan Quality. Tool Correctness and Argument Correctness target individual tool decisions. Together, these metrics can distinguish an efficient and effective agent from one that takes unnecessary steps, follows a weak plan, or chooses the correct tool with the wrong arguments.
DeepEval also supports custom evaluation criteria through G-Eval and DAG metrics, so application-specific requirements don't have to be forced into a generic relevance or correctness score. Many built-in metrics use an LLM judge; others add deterministic checks or use them exclusively. Tool Correctness, for example, can compare called tools directly against an expected set and only involve an LLM when judging optimality among available tools.
DeepEval itself is local-first, with shared dashboards, regression tracking, observability, and production monitoring provided through the separate Confident AI platform so the open-source framework can still run independently.
Solid choice for: Python applications that want repeatable LLM and agent tests in CI, including evaluations that need to inspect both final outputs and decisions made during an agent trajectory.
Probably not for: Engineers who mainly want a self-hosted evaluation workspace with built-in experiment management and collaborative dashboards. DeepEval handles the testing framework; those platform features belong to Confident AI.
2. Ragas — focused metrics for RAG and grounded generation

An MLflow trace of a Ragas run's retrieval and generation steps.
Ragas's angle is RAG + grounded generation evaluation, with a focus on metric design and dataset-based testing rather than environment execution or detailed trajectory inspection. Its metric library includes faithfulness, context precision and recall, factual correctness, answer relevance, semantic similarity, and noise sensitivity.
Giving retrieval and generation separate points of measurement helps pinpoint failures such as missing evidence, irrelevant context entering the prompt, or the model contradicting retrieved information. Depending on the test setup, Ragas supports both reference-based and reference-free evaluation.
Ragas also includes agent-facing metrics such as Tool Call Accuracy, Tool Call F1, Topic Adherence, and Agent Goal Accuracy. Their framework supports discrete, numeric, and ranking metrics for custom criteria as well.
Solid choice for: RAG systems that need separate measurements for retrieval quality, grounded generation, and factual correctness, plus applications that want custom dataset-level metrics without adopting a larger evaluation platform.
Probably not for: Agent evaluations where the main evidence is the complete execution path or the final state of an external environment. Ragas has agent metrics, but environment-based testing and detailed trajectory inspection aren't its primary focus.
3. Promptfoo — regression testing with a security edge

Pass rate, score distribution, and per-plugin results for two models in Promptfoo.
Promptfoo uses a configuration-first workflow for testing prompts, models, RAG systems, and agents. Test cases, providers, prompts, and assertions can be declared in YAML, run through the CLI, and added to CI/CD without building a separate Python evaluation suite. The resulting matrix compares the same inputs across model or prompt variants.
A Promptfoo test can check exact values, JSON structure, similarity, custom JavaScript or Python logic, latency, or an LLM rubric. Its documentation recommends using deterministic assertions wherever possible, which keeps straightforward requirements reproducible instead of routing every judgment through another model.
With tracing enabled, Promptfoo can reconstruct a trajectory from tool executions, shell commands, searches, guardrail decisions, and other recorded actions. Assertions can then check whether a tool was used, whether its arguments matched expectations, whether steps occurred in the required sequence, or whether the trajectory achieved its goal.
Security testing is another major aspect — red-team runs can generate adversarial cases for safety, security, and policy failures, and discovered failures can then become regression cases in the same test suite.
Promptfoo joined OpenAI in March 2026, but the project continues under the MIT license. Its CLI and library support models and applications from multiple providers and aren't restricted to OpenAI systems.
Solid choice for: Regression suites that compare prompts or models across many test cases, plus agent systems that need security testing and trajectory assertions in the same CI workflow.
Probably not for: Engineers who mainly want a Python-native metric library for statistical RAG analysis or a full self-hosted experiment platform. Promptfoo concentrates on declarative testing, comparisons, assertions, and adversarial evaluation.
4. Inspect AI — evaluations where agents have to act

Per-sample inputs, targets, answers, and scores in Inspect AI's log viewer.
Inspect AI, created by the UK AI Security Institute, is an MIT-licensed framework for evaluations where success depends on more than the final response. Its basic unit is a task, built from a dataset, a solver or agent, and a scorer. The same structure handles tool-using agents whose actions change files, services, or other resources during the test.
Inspect can run each evaluation sample inside a sandbox, with Docker built in and extensions available for Kubernetes, Modal, Daytona, EC2, Proxmox, Vagrant, and other environments. Files, commands, network access, and other resources can then be controlled at the task level instead of letting evaluated code execute directly on the host system.
The scorer can verify an outcome that exists outside the model's response. A coding-agent evaluation, for example, can check whether the expected file was created or whether tests pass after the agent edits a repository. Cybersecurity and tool-use evaluations can use the same pattern for terminal commands, services, or other environment state.
Inspect also supports external coding agents like Claude Code, Codex CLI, and Gemini CLI, which can run inside the same sandboxed tasks, allowing the complete agent system to be evaluated rather than only an individual model call.
Evaluation logs retain messages, tool calls, events, scores, and other execution details. The Inspect viewer can then be used to examine how the agent reached an outcome after the evaluation has finished.
Solid choice for: Coding agents, tool-using agents, cybersecurity tasks, and other evaluations where success needs to be verified through files, commands, tests, or environment state rather than response text alone.
Probably not for: Straightforward chatbot or RAG regression suites that only need answer-level metrics. Task environments and sandboxes add machinery that lighter frameworks such as DeepEval or Ragas may not require.
5. Opik — a self-hosted evaluation platform

A full tool-calling trajectory and agent graph traced in Opik.
Opik combines evaluation with the infrastructure needed to manage datasets, experiments, and repeated test runs. It provides a self-hostable workspace where evaluation results can be stored, compared, and inspected across model, prompt, or agent changes.
Its evaluation model has two main paths:
- Test Suites use natural-language assertions to produce pass/fail results for specific behaviors, with rules that can apply to an entire suite or individual cases.
- Datasets & Metrics run an application across a collection of test cases and assign quantitative scores using built-in or custom metrics. Each dataset evaluation creates an experiment that can be compared with previous runs.
Opik also keeps experiment history attached to those dataset runs — a dataset can be created manually, through the Python or TypeScript SDKs, or from earlier application traces, then reused to test a new prompt, model, retrieval configuration, or agent version.
For agents, Opik can score the trajectory rather than only the final answer. Evaluation metrics can access tool calls and intermediate operations recorded during execution, allowing tests to check whether the expected tools were selected, whether their order made sense, or whether the complete path reached the intended goal. Its dedicated TrajectoryAccuracy metric grades ReAct-style sequences of reasoning steps, actions, and observations against the task.
Opik also supports multi-turn evaluation using simulated users, which helps test agents whose behavior depends on several rounds of interaction rather than a single prompt-response pair.
A full Opik deployment adds more infrastructure than a lightweight library or CLI runner, which becomes justifiable when evaluation results need to persist across experiments and be reviewed through a shared platform.
Solid choice for: Engineers who want datasets, regression suites, quantitative metrics, experiment comparison, and agent trajectory evaluation inside one self-hostable open-source platform.
Probably not for: Projects that only need lightweight local evaluation or a few CI assertions. Running a complete platform adds operational overhead that a library or CLI-based framework can avoid.
Statefulness expands what needs to be evaluated
A production test suite often needs deterministic checks for known requirements, model-graded metrics for semantic quality, and outcome checks for actions taken outside the model. Persistent memory extends that evaluation across time, adding retention between sessions, updates to older information, later retrieval of relevant context, and the effect that retrieved memory has on the final response. No single score captures all of those behaviors.
In our own graph and memory evaluations, we used DeepEval Correctness alongside F1, Exact Match, and task-specific measurements. We found that an LLM judge can sometimes reward fluent, approximately correct answers more generously than token-level metrics do, producing sharply different DeepEval Correctness and F1 scores for the same system. That's why we believe multiple metrics make much more sense than a single judge score.
An evaluation stack can combine several of the tools we covered in this post. Ragas can measure retrieval and grounding, and recurring failures can become regression tests in DeepEval or Promptfoo. Inspect AI can verify outcomes inside an agent's environment, with Opik providing a place to compare repeated experiments over time. The appropriate tool depends on where failure can occur and what evidence is needed to verify it.


