Back to Blog
Infrastructure

LLM Observability: How to Monitor AI Apps in Production

By Keyved Engineering Team··5 min read

Short answer

LLM observability is the ability to see what an AI application is doing in production: every prompt, retrieval step, tool call and model response, with its cost, latency and quality. Set it up by tracing each request end to end (OpenTelemetry's generative AI conventions or an LLM tracing tool), tracking token cost per user and feature, scoring quality on sampled traffic, collecting user feedback, and alerting on spikes in errors, latency, cost and quality drops.

Key takeaways

  • Uptime monitoring isn't enough: LLM apps fail silently through bad answers and runaway costs.
  • Trace every request end to end — prompts, retrieval, tool calls, model responses.
  • Track cost per request, per user and per feature, not just the monthly invoice.
  • Score quality on live traffic and alert on drops, not just on errors.

Your dashboards are green. Uptime is 100%. Error rates are flat. And yet customers are complaining that the AI assistant "got worse" last week, and finance wants to know why the OpenAI bill doubled.

This is the gap between traditional monitoring and LLM observability. AI applications can fail without throwing a single error: answers drift, retrieval misses, an agent loops through ten tool calls instead of two, or a single user's long prompts quietly multiply costs.

This guide covers what to observe, how to instrument it and what to alert on.

What is LLM observability?

LLM observability is the ability to understand what an AI application is doing in production — and why — from the data it emits. It extends standard observability (traces, metrics, logs) with AI-specific signals:

  • Prompts and responses, including system prompts and retrieved context
  • Intermediate steps: retrieval queries and results, tool calls, agent decisions
  • Token usage and cost per call, request, user and feature
  • Latency per step, including time to first token for streaming
  • Quality signals: automated scores, user feedback, human overrides
  • Model and prompt versions for every call

Why isn't traditional monitoring enough for AI apps?

ProblemTraditional APM sees it?LLM observability sees it?
Server down or API errorsYesYes
Slow responsesYesYes, with per-step breakdown
Answers getting worseNoYes, via quality scores and feedback
Retrieval returning irrelevant documentsNoYes, via traced retrieval results
Agent looping or calling wrong toolsNoYes, via step-level traces
Cost spike from one feature or userNoYes, via token cost attribution
Prompt injection attemptsNoYes, via logged inputs and detection

What should you trace in an LLM application?

Trace each user request end to end, with child spans for every meaningful step:

  1. Request: user (pseudonymised), session, feature, app version
  2. Input processing: validation, moderation, query rewriting
  3. Retrieval: query, filters, retrieved document IDs and scores, reranker results
  4. Model calls: provider, model, prompt version, parameters, input/output tokens, cost, latency, time to first token, finish reason
  5. Tool calls: tool name, arguments, result summary, duration, errors
  6. Guardrails: validation outcomes, blocked actions, approvals requested
  7. Response: final output, citations, and any user feedback attached later

With this structure, you can open any complaint ("the bot gave me the wrong delivery date") and see exactly what was retrieved, what the model saw and why it answered as it did.

Use OpenTelemetry where you can

OpenTelemetry has defined semantic conventions for generative AI — standard attribute names for model calls, token usage and related operations. Instrumenting with OpenTelemetry keeps you portable: the same traces can go to Grafana, Datadog, Honeycomb or an LLM-specific backend, and AI spans sit alongside your existing service traces.

Which metrics should you track?

Reliability

  • Error rate by type: provider errors, timeouts, rate limits, validation failures
  • Retry and fallback rates
  • Agent step counts and loop detections

Performance

  • End-to-end latency (p50, p95, p99)
  • Time to first token for streamed responses
  • Per-step latency: retrieval, model, tools

Cost

  • Tokens and dollars per request
  • Cost per user, per customer account and per feature
  • Cache hit rate (for prompt caching)
  • Cost per successful outcome — the number that matters to the business

See how to cut LLM API costs for what to do with these numbers.

Quality

  • Automated scores on sampled traffic (groundedness, correctness, format validity)
  • User feedback rate and sentiment
  • Human override and edit rates for drafts
  • "I don't know" and escalation rates
  • Retrieval hit rate (did sources get cited?)

Quality metrics come from the same checks you use in offline evaluation. Our guide to evaluating LLM applications explains how to build them.

How do you monitor quality in production?

Three complementary approaches:

  1. Online evaluation: score a sample of live requests (say 5–10%, or all high-value ones) with automated checks and LLM-as-judge rubrics.
  2. User feedback: thumbs up/down, corrections, escalations — attached to the trace so you can see what caused them.
  3. Drift detection: compare score distributions, input topics and output lengths week over week. Sudden changes often follow model updates, prompt changes or new types of user input.

Feed bad cases back into your evaluation set so the same failure can't recur unnoticed.

What should you alert on?

AlertExample trigger
Error spikeProvider error rate above 2% for 5 minutes
Rate limitingAny sustained 429 responses from a provider
Latencyp95 above your SLA for 10 minutes
Cost spikeHourly spend 2x the trailing average; any user above a daily cap
Quality dropGroundedness score down 10% day over day
FeedbackNegative feedback rate above threshold
Agent runawayTasks hitting step or spend limits
SecurityRepeated injection-pattern inputs from one user

Route alerts to the people who can act: engineering for errors and latency, the product owner for quality, finance for budget thresholds.

How should you handle sensitive data in logs?

Prompts and responses often contain personal or confidential information. Logging them is essential for debugging, so handle them deliberately:

  • Redact personal data (names, emails, card numbers, health details) before storage where possible
  • Restrict access to raw traces with role-based permissions
  • Encrypt at rest and in transit
  • Set retention limits that match your legal and contractual obligations
  • Keep data in-region if residency rules apply
  • Check the data terms of any third-party observability vendor

Which LLM observability tools should you consider?

ToolNotes
LangfuseOpen source, self-hostable; tracing, evals, prompt management
LangSmithTight integration with LangChain and LangGraph
Arize PhoenixOpen source; OpenTelemetry-native tracing and evaluation
HeliconeProxy-based logging and cost tracking with minimal code changes
BraintrustEvaluation-first, with tracing and experiments
Datadog LLM ObservabilityFits teams already on Datadog
Grafana + OpenTelemetryUnified with existing infrastructure dashboards

For regulated industries, self-hosted options like Langfuse or Phoenix keep sensitive traces inside your own infrastructure.

A practical rollout plan

  1. Week 1: instrument model calls with tokens, cost, latency and versions; log prompts and responses with redaction.
  2. Week 2: add spans for retrieval, tools and guardrails; link everything under one trace per request.
  3. Week 3: build dashboards for reliability, performance and cost; set alerts.
  4. Week 4: add online quality scoring and feedback capture; start the feedback-to-eval loop.

How we build observability at Keyved

Observability is built into our platform foundation, not bolted on per project. Every system we ship has OpenTelemetry tracing across services and model calls, cost attribution, dashboards in Grafana, error tracking with Sentry, and alerting from day one. You can read about how that foundation behaves under load in our 500 concurrent users deep dive.

If you have an AI system in production and limited visibility into it, we can instrument it without a rebuild. See our projects or get in touch.

Frequently asked questions

What is LLM observability?

LLM observability is the practice of collecting and analysing traces, metrics and logs from AI applications so you can understand their behaviour in production. It covers inputs and outputs, intermediate steps like retrieval and tool calls, token usage and cost, latency, errors and output quality.

How is LLM monitoring different from traditional APM?

Traditional application performance monitoring tracks uptime, errors and latency. LLM apps can be up, fast and error-free while giving wrong answers or spending too much. LLM observability adds prompt and response tracing, token cost tracking and quality evaluation.

What tools are used for LLM observability?

Popular options include Langfuse, LangSmith, Arize Phoenix, Helicone, Braintrust and Datadog LLM Observability. OpenTelemetry's generative AI semantic conventions allow traces to be sent to many backends, including Grafana, Honeycomb and Datadog.

Should you log full prompts and responses?

Usually yes, because they are essential for debugging, but handle them as sensitive data: redact personal information where possible, restrict access, encrypt them and set retention limits that match your privacy obligations.

What should you alert on in an LLM application?

Alert on error and timeout rates, provider rate-limit errors, latency percentiles, cost spikes per user or feature, drops in automated quality scores, rising negative feedback and unusual patterns such as repeated prompt injection attempts.

Want to see how we build these systems for clients?

Let's Talk

Keep reading