Back to Blog
Engineering

How to Evaluate LLM Applications: Evals, Metrics and Tools That Work

By Keyved Engineering Team··6 min read

Short answer

Evaluate an LLM application by building a test set of real inputs with expected outcomes, choosing metrics that reflect what matters to users (correctness, groundedness, format, safety, cost, latency), scoring with code checks where possible and LLM-as-judge calibrated against human ratings where not, and running the suite on every change. In production, add sampling, user feedback and monitoring so the test set keeps growing with real failures.

Key takeaways

  • An evaluation set of real examples is the single most valuable asset in an LLM project.
  • Prefer deterministic checks; use LLM-as-judge for subjective criteria, calibrated against humans.
  • Evaluate components (retrieval, tool calls) and end-to-end results separately.
  • Run evals on every prompt, model or code change — like unit tests.

Ask most teams how they know their LLM application works and you'll hear some version of: "We tried a bunch of questions and the answers looked good."

That's how prototypes get built. It's not how production systems stay reliable. LLM output varies, edge cases are endless, and every prompt tweak or model upgrade can fix one thing while silently breaking another.

Evaluation — "evals" — is how you replace "looks good" with evidence. Here's the playbook we use.

What is LLM evaluation?

LLM evaluation is the practice of measuring how well a language-model application performs on a defined set of inputs, using explicit criteria. A good evaluation setup lets you answer:

  • Is this version better or worse than the last one?
  • Is it good enough to launch?
  • Which component — retrieval, prompt, model, tool — is causing failures?
  • Did the new model release change behaviour?

It's the equivalent of a test suite in traditional software, adapted for outputs that aren't exactly predictable. It's also the most common missing piece in AI pilots that never reach production.

Step 1: Build an evaluation dataset

Your evaluation set is a collection of inputs with expected outcomes. It should be:

  • Real: drawn from actual user questions, documents, tickets or transactions — not invented by engineers
  • Representative: covering the common cases in roughly realistic proportions
  • Edge-heavy: deliberately including hard cases, ambiguous inputs and known failure modes
  • Labelled by domain experts: the people who know what "correct" looks like

What goes in each example?

FieldExample
InputThe user question, document or task
Expected output or criteriaThe correct answer, required fields, or a rubric
Reference sourcesFor RAG: the passages that support the answer
TagsCase type, difficulty, language, customer segment

How many? Start with 50–200. That's enough to catch most regressions and compare versions. Then grow it continuously: every production failure becomes a new test case.

Step 2: Choose metrics that match what users care about

Pick a small number of metrics tied to real quality. Common ones:

CategoryMetricHow to score
CorrectnessAnswer matches the expected answerExact match, rules, or LLM judge
GroundednessClaims are supported by retrieved sourcesLLM judge or NLI model
CompletenessAll required points or fields presentChecklist, rules, or judge
FormatValid JSON, required schema, length limitsCode
Retrieval qualityRight passages retrieved (recall@k, MRR)Code, against labelled sources
Tool useCorrect tool and argumentsCode
SafetyNo sensitive data, policy violations or harmful contentRules, classifiers, judge
Cost & latencyTokens, dollars and seconds per taskCode

Resist the urge to track twenty metrics. Three to six well-chosen ones are easier to act on.

Step 3: Prefer deterministic checks

Whenever a criterion can be checked with code, do that:

  • JSON parses and matches the schema
  • Extracted invoice total equals the expected value
  • The correct tool was called with the correct order ID
  • The response cites at least one source
  • No email addresses or card numbers appear in the output

Code checks are fast, free and perfectly consistent. Save model-based scoring for criteria that genuinely need judgment.

Step 4: Use LLM-as-a-judge carefully

For subjective criteria — is this answer helpful, accurate, well-grounded, appropriately toned? — a capable model can score outputs against a rubric. This is called LLM-as-a-judge, and it makes evaluating hundreds of open-ended answers practical.

It also has known weaknesses: judges can favour longer answers, prefer their own style, or be swayed by confident wording. To use it responsibly:

  • Write specific rubrics. "Score 1 if the answer states the correct refund window from the source; 0 otherwise" beats "Rate helpfulness 1–10."
  • Prefer binary or small scales over 1–10 scores.
  • Ask for reasoning before the score.
  • Give the judge the reference answer or sources.
  • Calibrate against humans: have domain experts label a sample and check how often the judge agrees. Adjust the rubric until agreement is high.
  • Re-check after changing the judge model.

Step 5: Evaluate components and the whole system

Complex applications fail in specific places. Evaluate each stage so you know where to fix things.

Evaluating RAG systems

  • Retrieval: did the right passages come back? (recall@k, precision, MRR)
  • Generation: given those passages, is the answer correct and grounded?
  • End to end: is the final answer right?

If retrieval is poor, prompt changes won't help. Our article on why naive RAG fails covers the retrieval fixes.

Evaluating AI agents

Agents need three kinds of checks:

  • Outcome: was the task completed correctly? (Was the right refund issued? Was the ticket resolved?)
  • Trajectory: were the right tools called with the right arguments, in a reasonable number of steps, without loops?
  • Safety: did the agent stay within its permissions and request approval when required? See AI agent guardrails.

Run agent evals against sandboxed tools or recorded responses so tests are repeatable and harmless.

Evaluating for security

Include adversarial cases: prompt injection attempts, requests for other users' data, jailbreaks. The OWASP Top 10 for LLMs is a good source of test ideas.

Step 6: Run evals on every change

Treat your eval suite like unit tests:

  • Run it on every prompt change, model change, retrieval change and code change that touches the AI path
  • Run it in CI and block merges that drop key metrics below a threshold
  • Compare against the current production version, not just absolute scores
  • Re-run when a provider releases a new model version, before switching

Tools such as promptfoo make CI integration straightforward; platforms like Langfuse, LangSmith and Braintrust store results and show comparisons over time.

Step 7: Keep evaluating in production

Offline evals can't anticipate everything. In production:

  • Sample live traffic and score it with the same judges and checks
  • Collect user feedback: thumbs up/down, edits to drafts, escalations
  • Watch proxy signals: retry rates, "I don't know" rates, human override rates
  • Feed failures back into the evaluation set

This closes the loop between LLM observability and evaluation: monitoring finds problems, evaluation prevents them from coming back.

Which LLM evaluation tools should you use?

ToolGood for
RagasRAG metrics: faithfulness, answer relevance, context precision and recall
DeepEvalUnit-test-style LLM evals with many built-in metrics
promptfooPrompt and model comparisons, red-teaming, CI integration
LangfuseOpen-source tracing, datasets, scoring and prompt management
LangSmithTracing and evaluation, especially with LangChain and LangGraph
BraintrustExperiment tracking and evaluation workflows
Arize PhoenixOpen-source tracing and evaluation, OpenTelemetry-based

The tool matters less than the habit. A spreadsheet of 100 labelled examples and a script that scores them beats an expensive platform nobody updates.

How we evaluate at Keyved

We build the evaluation set in the first week of every project, with your domain experts, from real examples. Every iteration after that is measured against it, and the results are part of what we hand over — so your team can keep testing after we're done.

That's how we can show you, rather than tell you, that a RAG system or AI agent is ready. See examples on our projects page, or get in touch if you'd like help building an evaluation suite for a system you already have.

Frequently asked questions

What are LLM evals?

LLM evals are tests that measure how well a language-model application performs on a defined set of inputs. They score outputs against expected results or criteria, such as correctness, faithfulness to sources, format and safety, so changes can be compared objectively.

How many examples do I need in an evaluation set?

Start with 50 to 200 real examples covering common cases and known edge cases. That is enough to catch most regressions. Grow it over time by adding production failures and new case types.

What is LLM-as-a-judge?

LLM-as-a-judge uses a language model to score another model's output against a rubric, such as whether an answer is correct, complete or supported by the sources. It scales evaluation of subjective criteria but must be validated against human ratings because judges have biases.

What tools are used for LLM evaluation?

Common tools include Ragas and DeepEval for RAG and general metrics, promptfoo for prompt testing in CI, Langfuse, LangSmith, Braintrust and Arize Phoenix for tracing and evaluation, and OpenAI Evals. Many teams also write custom checks in their existing test framework.

How do you evaluate an AI agent?

Evaluate the final outcome (was the task completed correctly?), the trajectory (were the right tools called with the right arguments, in a reasonable number of steps?) and safety (did it stay within permissions and ask for approval when required?). Use realistic scenarios with sandboxed tools.

Want to see how we build these systems for clients?

Let's Talk

Keep reading