How to Evaluate LLM Applications: Evals, Metrics and Tools That Work
Short answer
Evaluate an LLM application by building a test set of real inputs with expected outcomes, choosing metrics that reflect what matters to users (correctness, groundedness, format, safety, cost, latency), scoring with code checks where possible and LLM-as-judge calibrated against human ratings where not, and running the suite on every change. In production, add sampling, user feedback and monitoring so the test set keeps growing with real failures.
Key takeaways
- An evaluation set of real examples is the single most valuable asset in an LLM project.
- Prefer deterministic checks; use LLM-as-judge for subjective criteria, calibrated against humans.
- Evaluate components (retrieval, tool calls) and end-to-end results separately.
- Run evals on every prompt, model or code change — like unit tests.
Ask most teams how they know their LLM application works and you'll hear some version of: "We tried a bunch of questions and the answers looked good."
That's how prototypes get built. It's not how production systems stay reliable. LLM output varies, edge cases are endless, and every prompt tweak or model upgrade can fix one thing while silently breaking another.
Evaluation — "evals" — is how you replace "looks good" with evidence. Here's the playbook we use.
What is LLM evaluation?
LLM evaluation is the practice of measuring how well a language-model application performs on a defined set of inputs, using explicit criteria. A good evaluation setup lets you answer:
- Is this version better or worse than the last one?
- Is it good enough to launch?
- Which component — retrieval, prompt, model, tool — is causing failures?
- Did the new model release change behaviour?
It's the equivalent of a test suite in traditional software, adapted for outputs that aren't exactly predictable. It's also the most common missing piece in AI pilots that never reach production.
Step 1: Build an evaluation dataset
Your evaluation set is a collection of inputs with expected outcomes. It should be:
- Real: drawn from actual user questions, documents, tickets or transactions — not invented by engineers
- Representative: covering the common cases in roughly realistic proportions
- Edge-heavy: deliberately including hard cases, ambiguous inputs and known failure modes
- Labelled by domain experts: the people who know what "correct" looks like
What goes in each example?
| Field | Example |
|---|---|
| Input | The user question, document or task |
| Expected output or criteria | The correct answer, required fields, or a rubric |
| Reference sources | For RAG: the passages that support the answer |
| Tags | Case type, difficulty, language, customer segment |
How many? Start with 50–200. That's enough to catch most regressions and compare versions. Then grow it continuously: every production failure becomes a new test case.
Step 2: Choose metrics that match what users care about
Pick a small number of metrics tied to real quality. Common ones:
| Category | Metric | How to score |
|---|---|---|
| Correctness | Answer matches the expected answer | Exact match, rules, or LLM judge |
| Groundedness | Claims are supported by retrieved sources | LLM judge or NLI model |
| Completeness | All required points or fields present | Checklist, rules, or judge |
| Format | Valid JSON, required schema, length limits | Code |
| Retrieval quality | Right passages retrieved (recall@k, MRR) | Code, against labelled sources |
| Tool use | Correct tool and arguments | Code |
| Safety | No sensitive data, policy violations or harmful content | Rules, classifiers, judge |
| Cost & latency | Tokens, dollars and seconds per task | Code |
Resist the urge to track twenty metrics. Three to six well-chosen ones are easier to act on.
Step 3: Prefer deterministic checks
Whenever a criterion can be checked with code, do that:
- JSON parses and matches the schema
- Extracted invoice total equals the expected value
- The correct tool was called with the correct order ID
- The response cites at least one source
- No email addresses or card numbers appear in the output
Code checks are fast, free and perfectly consistent. Save model-based scoring for criteria that genuinely need judgment.
Step 4: Use LLM-as-a-judge carefully
For subjective criteria — is this answer helpful, accurate, well-grounded, appropriately toned? — a capable model can score outputs against a rubric. This is called LLM-as-a-judge, and it makes evaluating hundreds of open-ended answers practical.
It also has known weaknesses: judges can favour longer answers, prefer their own style, or be swayed by confident wording. To use it responsibly:
- Write specific rubrics. "Score 1 if the answer states the correct refund window from the source; 0 otherwise" beats "Rate helpfulness 1–10."
- Prefer binary or small scales over 1–10 scores.
- Ask for reasoning before the score.
- Give the judge the reference answer or sources.
- Calibrate against humans: have domain experts label a sample and check how often the judge agrees. Adjust the rubric until agreement is high.
- Re-check after changing the judge model.
Step 5: Evaluate components and the whole system
Complex applications fail in specific places. Evaluate each stage so you know where to fix things.
Evaluating RAG systems
- Retrieval: did the right passages come back? (recall@k, precision, MRR)
- Generation: given those passages, is the answer correct and grounded?
- End to end: is the final answer right?
If retrieval is poor, prompt changes won't help. Our article on why naive RAG fails covers the retrieval fixes.
Evaluating AI agents
Agents need three kinds of checks:
- Outcome: was the task completed correctly? (Was the right refund issued? Was the ticket resolved?)
- Trajectory: were the right tools called with the right arguments, in a reasonable number of steps, without loops?
- Safety: did the agent stay within its permissions and request approval when required? See AI agent guardrails.
Run agent evals against sandboxed tools or recorded responses so tests are repeatable and harmless.
Evaluating for security
Include adversarial cases: prompt injection attempts, requests for other users' data, jailbreaks. The OWASP Top 10 for LLMs is a good source of test ideas.
Step 6: Run evals on every change
Treat your eval suite like unit tests:
- Run it on every prompt change, model change, retrieval change and code change that touches the AI path
- Run it in CI and block merges that drop key metrics below a threshold
- Compare against the current production version, not just absolute scores
- Re-run when a provider releases a new model version, before switching
Tools such as promptfoo make CI integration straightforward; platforms like Langfuse, LangSmith and Braintrust store results and show comparisons over time.
Step 7: Keep evaluating in production
Offline evals can't anticipate everything. In production:
- Sample live traffic and score it with the same judges and checks
- Collect user feedback: thumbs up/down, edits to drafts, escalations
- Watch proxy signals: retry rates, "I don't know" rates, human override rates
- Feed failures back into the evaluation set
This closes the loop between LLM observability and evaluation: monitoring finds problems, evaluation prevents them from coming back.
Which LLM evaluation tools should you use?
| Tool | Good for |
|---|---|
| Ragas | RAG metrics: faithfulness, answer relevance, context precision and recall |
| DeepEval | Unit-test-style LLM evals with many built-in metrics |
| promptfoo | Prompt and model comparisons, red-teaming, CI integration |
| Langfuse | Open-source tracing, datasets, scoring and prompt management |
| LangSmith | Tracing and evaluation, especially with LangChain and LangGraph |
| Braintrust | Experiment tracking and evaluation workflows |
| Arize Phoenix | Open-source tracing and evaluation, OpenTelemetry-based |
The tool matters less than the habit. A spreadsheet of 100 labelled examples and a script that scores them beats an expensive platform nobody updates.
How we evaluate at Keyved
We build the evaluation set in the first week of every project, with your domain experts, from real examples. Every iteration after that is measured against it, and the results are part of what we hand over — so your team can keep testing after we're done.
That's how we can show you, rather than tell you, that a RAG system or AI agent is ready. See examples on our projects page, or get in touch if you'd like help building an evaluation suite for a system you already have.
Frequently asked questions
What are LLM evals?
LLM evals are tests that measure how well a language-model application performs on a defined set of inputs. They score outputs against expected results or criteria, such as correctness, faithfulness to sources, format and safety, so changes can be compared objectively.
How many examples do I need in an evaluation set?
Start with 50 to 200 real examples covering common cases and known edge cases. That is enough to catch most regressions. Grow it over time by adding production failures and new case types.
What is LLM-as-a-judge?
LLM-as-a-judge uses a language model to score another model's output against a rubric, such as whether an answer is correct, complete or supported by the sources. It scales evaluation of subjective criteria but must be validated against human ratings because judges have biases.
What tools are used for LLM evaluation?
Common tools include Ragas and DeepEval for RAG and general metrics, promptfoo for prompt testing in CI, Langfuse, LangSmith, Braintrust and Arize Phoenix for tracing and evaluation, and OpenAI Evals. Many teams also write custom checks in their existing test framework.
How do you evaluate an AI agent?
Evaluate the final outcome (was the task completed correctly?), the trajectory (were the right tools called with the right arguments, in a reasonable number of steps?) and safety (did it stay within permissions and ask for approval when required?). Use realistic scenarios with sandboxed tools.