LLM Observability: How to Monitor AI Apps in Production
Short answer
LLM observability is the ability to see what an AI application is doing in production: every prompt, retrieval step, tool call and model response, with its cost, latency and quality. Set it up by tracing each request end to end (OpenTelemetry's generative AI conventions or an LLM tracing tool), tracking token cost per user and feature, scoring quality on sampled traffic, collecting user feedback, and alerting on spikes in errors, latency, cost and quality drops.
Key takeaways
- Uptime monitoring isn't enough: LLM apps fail silently through bad answers and runaway costs.
- Trace every request end to end — prompts, retrieval, tool calls, model responses.
- Track cost per request, per user and per feature, not just the monthly invoice.
- Score quality on live traffic and alert on drops, not just on errors.
Your dashboards are green. Uptime is 100%. Error rates are flat. And yet customers are complaining that the AI assistant "got worse" last week, and finance wants to know why the OpenAI bill doubled.
This is the gap between traditional monitoring and LLM observability. AI applications can fail without throwing a single error: answers drift, retrieval misses, an agent loops through ten tool calls instead of two, or a single user's long prompts quietly multiply costs.
This guide covers what to observe, how to instrument it and what to alert on.
What is LLM observability?
LLM observability is the ability to understand what an AI application is doing in production — and why — from the data it emits. It extends standard observability (traces, metrics, logs) with AI-specific signals:
- Prompts and responses, including system prompts and retrieved context
- Intermediate steps: retrieval queries and results, tool calls, agent decisions
- Token usage and cost per call, request, user and feature
- Latency per step, including time to first token for streaming
- Quality signals: automated scores, user feedback, human overrides
- Model and prompt versions for every call
Why isn't traditional monitoring enough for AI apps?
| Problem | Traditional APM sees it? | LLM observability sees it? |
|---|---|---|
| Server down or API errors | Yes | Yes |
| Slow responses | Yes | Yes, with per-step breakdown |
| Answers getting worse | No | Yes, via quality scores and feedback |
| Retrieval returning irrelevant documents | No | Yes, via traced retrieval results |
| Agent looping or calling wrong tools | No | Yes, via step-level traces |
| Cost spike from one feature or user | No | Yes, via token cost attribution |
| Prompt injection attempts | No | Yes, via logged inputs and detection |
What should you trace in an LLM application?
Trace each user request end to end, with child spans for every meaningful step:
- Request: user (pseudonymised), session, feature, app version
- Input processing: validation, moderation, query rewriting
- Retrieval: query, filters, retrieved document IDs and scores, reranker results
- Model calls: provider, model, prompt version, parameters, input/output tokens, cost, latency, time to first token, finish reason
- Tool calls: tool name, arguments, result summary, duration, errors
- Guardrails: validation outcomes, blocked actions, approvals requested
- Response: final output, citations, and any user feedback attached later
With this structure, you can open any complaint ("the bot gave me the wrong delivery date") and see exactly what was retrieved, what the model saw and why it answered as it did.
Use OpenTelemetry where you can
OpenTelemetry has defined semantic conventions for generative AI — standard attribute names for model calls, token usage and related operations. Instrumenting with OpenTelemetry keeps you portable: the same traces can go to Grafana, Datadog, Honeycomb or an LLM-specific backend, and AI spans sit alongside your existing service traces.
Which metrics should you track?
Reliability
- Error rate by type: provider errors, timeouts, rate limits, validation failures
- Retry and fallback rates
- Agent step counts and loop detections
Performance
- End-to-end latency (p50, p95, p99)
- Time to first token for streamed responses
- Per-step latency: retrieval, model, tools
Cost
- Tokens and dollars per request
- Cost per user, per customer account and per feature
- Cache hit rate (for prompt caching)
- Cost per successful outcome — the number that matters to the business
See how to cut LLM API costs for what to do with these numbers.
Quality
- Automated scores on sampled traffic (groundedness, correctness, format validity)
- User feedback rate and sentiment
- Human override and edit rates for drafts
- "I don't know" and escalation rates
- Retrieval hit rate (did sources get cited?)
Quality metrics come from the same checks you use in offline evaluation. Our guide to evaluating LLM applications explains how to build them.
How do you monitor quality in production?
Three complementary approaches:
- Online evaluation: score a sample of live requests (say 5–10%, or all high-value ones) with automated checks and LLM-as-judge rubrics.
- User feedback: thumbs up/down, corrections, escalations — attached to the trace so you can see what caused them.
- Drift detection: compare score distributions, input topics and output lengths week over week. Sudden changes often follow model updates, prompt changes or new types of user input.
Feed bad cases back into your evaluation set so the same failure can't recur unnoticed.
What should you alert on?
| Alert | Example trigger |
|---|---|
| Error spike | Provider error rate above 2% for 5 minutes |
| Rate limiting | Any sustained 429 responses from a provider |
| Latency | p95 above your SLA for 10 minutes |
| Cost spike | Hourly spend 2x the trailing average; any user above a daily cap |
| Quality drop | Groundedness score down 10% day over day |
| Feedback | Negative feedback rate above threshold |
| Agent runaway | Tasks hitting step or spend limits |
| Security | Repeated injection-pattern inputs from one user |
Route alerts to the people who can act: engineering for errors and latency, the product owner for quality, finance for budget thresholds.
How should you handle sensitive data in logs?
Prompts and responses often contain personal or confidential information. Logging them is essential for debugging, so handle them deliberately:
- Redact personal data (names, emails, card numbers, health details) before storage where possible
- Restrict access to raw traces with role-based permissions
- Encrypt at rest and in transit
- Set retention limits that match your legal and contractual obligations
- Keep data in-region if residency rules apply
- Check the data terms of any third-party observability vendor
Which LLM observability tools should you consider?
| Tool | Notes |
|---|---|
| Langfuse | Open source, self-hostable; tracing, evals, prompt management |
| LangSmith | Tight integration with LangChain and LangGraph |
| Arize Phoenix | Open source; OpenTelemetry-native tracing and evaluation |
| Helicone | Proxy-based logging and cost tracking with minimal code changes |
| Braintrust | Evaluation-first, with tracing and experiments |
| Datadog LLM Observability | Fits teams already on Datadog |
| Grafana + OpenTelemetry | Unified with existing infrastructure dashboards |
For regulated industries, self-hosted options like Langfuse or Phoenix keep sensitive traces inside your own infrastructure.
A practical rollout plan
- Week 1: instrument model calls with tokens, cost, latency and versions; log prompts and responses with redaction.
- Week 2: add spans for retrieval, tools and guardrails; link everything under one trace per request.
- Week 3: build dashboards for reliability, performance and cost; set alerts.
- Week 4: add online quality scoring and feedback capture; start the feedback-to-eval loop.
How we build observability at Keyved
Observability is built into our platform foundation, not bolted on per project. Every system we ship has OpenTelemetry tracing across services and model calls, cost attribution, dashboards in Grafana, error tracking with Sentry, and alerting from day one. You can read about how that foundation behaves under load in our 500 concurrent users deep dive.
If you have an AI system in production and limited visibility into it, we can instrument it without a rebuild. See our projects or get in touch.
Frequently asked questions
What is LLM observability?
LLM observability is the practice of collecting and analysing traces, metrics and logs from AI applications so you can understand their behaviour in production. It covers inputs and outputs, intermediate steps like retrieval and tool calls, token usage and cost, latency, errors and output quality.
How is LLM monitoring different from traditional APM?
Traditional application performance monitoring tracks uptime, errors and latency. LLM apps can be up, fast and error-free while giving wrong answers or spending too much. LLM observability adds prompt and response tracing, token cost tracking and quality evaluation.
What tools are used for LLM observability?
Popular options include Langfuse, LangSmith, Arize Phoenix, Helicone, Braintrust and Datadog LLM Observability. OpenTelemetry's generative AI semantic conventions allow traces to be sent to many backends, including Grafana, Honeycomb and Datadog.
Should you log full prompts and responses?
Usually yes, because they are essential for debugging, but handle them as sensitive data: redact personal information where possible, restrict access, encrypt them and set retention limits that match your privacy obligations.
What should you alert on in an LLM application?
Alert on error and timeout rates, provider rate-limit errors, latency percentiles, cost spikes per user or feature, drops in automated quality scores, rising negative feedback and unusual patterns such as repeated prompt injection attempts.