For AI agents: a documentation index is available at /docs/llms.txt. Append .md to any page URL for markdown, or send Accept: text/markdown.
Agent Analytics FAQ and troubleshooting
このページはまだあなたの言語に翻訳されていません。 現在取り組んでいますので、後でもう一度確認してください。
What is Agent Analytics?
What is Agent Analytics?
Agent Analytics is Amplitude's analytics product for AI agents. It captures every agent conversation and workflow (sessions, turns, traces, tool calls, token costs) and automatically scores each session on quality signals like task completion, response quality, user friction, and safety. Because everything lands as standard Amplitude events, you can connect agent quality directly to the product metrics you already track.
Can it connect agent quality to business outcomes like retention or conversion?
Yes, this is the core of the product. Evaluation results are emitted as standard Amplitude events, so you can build funnels, retention curves, and cohorts on them. For example, you can compare retention for users whose agent sessions completed successfully versus those who hit errors, or cohort users by the topics they ask about and see how each cohort converts. Most teams monitor their agents across several disconnected systems that a business user can't join together; Agent Analytics puts quality and outcomes in one place.
Is Agent Analytics only for chatbots?
No. It supports conversational agents, voice agents, background or autonomous "doer" agents, multimodal agents (including image inputs and outputs), and multi-agent systems. Multi-agent workflows are tracked with per-agent identifiers, and you can analyze cross-agent flows with Amplitude's journey charts.
How is it different from my observability and eval tools?
How does Agent Analytics fit alongside your existing tools (LangSmith, Langfuse, Braintrust, Datadog)?
You don't have to choose; most teams run all three layers, because they answer three different questions. Observability platforms tell you whether your agent is running: uptime, latency, errors. Offline eval tools tell you whether a change is safe to ship: synthetic test suites, regression testing before release. Agent Analytics tells you whether the agent is working: did it complete the user's task in production, was the user satisfied, and did the interaction drive the outcome you built the agent for, connected to the conversion and retention data you already have. Test changes with your eval stack, monitor infrastructure with your observability stack, and use Agent Analytics to learn what real sessions are doing to your business after release.
Why can't I just track agent events with regular product analytics?
Traditional engagement metrics break for agents: a long conversation can mean a delighted user or a deeply frustrated one, and identical usage patterns can represent opposite outcomes. Agent Analytics adds what product analytics can't see (conversation content, traces, and LLM-scored quality signals) and joins them with behavioral data.
Getting started and instrumentation
Do I have to use your SDK?
No. There are several paths, and they all produce the same event taxonomy:
- Amplitude AI SDK: lowest friction, full enrichment out of the box.
- HTTP API: send agent events directly from any stack.
- OpenTelemetry: already tracing with OTel? Route your GenAI spans to Amplitude as a second destination; no re-instrumentation of your app.
- Warehouse import: bring conversations from Snowflake, BigQuery, or Databricks with one mapping query. Refer to Import agent conversations from your data warehouse.
- Integrations: import conversations from LangSmith, Langfuse, Braintrust, Sierra, Decagon, Fin, or Agentforce. Refer to Import agent conversations from other tools.
If you use LangChain, the callback handler works with any model provider. For the full wire contract when emitting events yourself, refer to Send agent events without the AI SDK on the setup page.
Can I use the @amplitude/ai SDK in the browser? What about Lovable, Superblocks, or other AI app builders?
No. The AI SDK is Node.js and Python only; importing it in a browser bundle or an edge runtime (for example Cloudflare Workers) fails on Node-only modules. The supported path for browsers, edge runtimes, and AI app builders is to emit the [Agent] taxonomy directly with the standard @amplitude/analytics-browser SDK. The setup page carries the full contract and a paste-ready prompt for builder agents.
Should I build my own client-side idle timer to end sessions?
Almost never. There are two clocks, and only one is yours. Amplitude's clock always runs: if a session receives no events for 30 minutes (configurable, or -1 to disable), the server closes it and runs enrichment, and an explicit Session End from your app closes it immediately. If you also build your own short timer in the client (say, ending the session after two quiet minutes and rotating the session ID), you chop real conversations into fragments: a user who pauses to read a long answer comes back to a "new" session. Enrichment then judges half-conversations, task completion looks artificially low, and, since Agent Analytics is billed per session, one conversation bills as several. Close explicitly when the job genuinely ends, and let the 30-minute server clock catch abandonment. Only mirror an idle timer client-side if your product truly defines sessions that way, and keep it generous.
What is an agent session, and how is it different from an Amplitude session?
They are two different concepts, defined by two different properties:
- An agent session is defined by the
[Agent] Session IDevent property. It represents one job the user hands the agent, from start to finish: a conversation thread, a support ticket, a voice call, or a background run. You set it from an ID you already track (thread ID, ticket ID, call ID, job ID); everything sharing that ID, including sub-agent activity, rolls into the same agent session. - A regular Amplitude session is defined by the standard
Session IDproperty ($session_id). It's behavioral: the user's app or web visit, which powers Session Replay and product reports, and closes after inactivity.
The two are related but independent. A user might have three agent conversations in one product visit, or one agent workflow that runs long after they've left your app. Agent sessions don't interfere with your existing session definitions. To link them, forward the standard session ID to your backend and pass it on the agent session (browserSessionId), so agent events join the user's product session.
How do agent sessions end? Can I configure the timeout?
Sessions end one of two ways:
- Explicitly (recommended): call
trackSessionEnd()(Node) ortrack_session_end()(Python) when the job finishes, such as a closed ticket or a completed run. Quality-signal evaluation runs immediately. - By idle timeout: if you don't end the session, it closes automatically after 30 minutes of inactivity, measured from the last agent event received for the session. This is configurable per session with
idleTimeoutMinutes(Node) oridle_timeout_minutes(Python); today, also include anidle_timeout_minuteskey in the agent'scontextso the override reliably reaches the server (refer to where to set the override). Raise it for jobs with long natural gaps (for example,240for a support ticket worked over hours), or set it to-1to disable the idle close entirely and keep the session open until you end it explicitly (with a 90-day backstop).
When the same user returns with a new goal, start a new session with a new session ID rather than continuing the old one.
How do I link agent sessions to my existing Amplitude users?
Send the same User ID (or Device ID) you use in your product analytics; that's what joins agent activity to the rest of the user's journey. Session IDs don't participate in identity resolution. Anonymous or pre-signup users are merged later through Amplitude's standard identity resolution once they identify. Amplitude recommends deciding your ID strategy on day one of instrumentation, since it's the most common setup issue.
How do I instrument a thumbs up / thumbs down control?
Refer to Collect user feedback in Agent Analytics. The short version: send an [Agent] Score event with the name user-feedback, value 1 or 0, targeting the AI message being rated, with source user. A thumbs down sent this way counts as explicit dissatisfaction for Negative Feedback.
Can I import historical conversations?
Yes, through the HTTP import API, warehouse import, or an integration. Note that evaluators score sessions going forward; refer to the evaluators section below for how to score historical data.
Signals and evaluators
What gets scored automatically?
Amplitude scores every closed session on seven built-in signals, with no setup: Task Completed, Response Quality, Session Safety, User Intent, User Friction, Has Negative Feedback, and Data Quality Issues. Results land as properties on the session's [Agent] Session Record event, each with a rationale, so you can chart and segment on them right away. Signals are directional indicators, not ground truth. For the question each signal answers, its labels, and its properties, refer to Agent Analytics built-in signals.
What's the difference between signals and evaluators?
Signals are the fixed, built-in scores that run on every session automatically. Evaluators are custom checks you define for anything domain-specific ("did the agent quote the correct policy?", "did it stay in character?") using LLM-as-a-judge with your own criteria.
What is the recommended model to use for an evaluator?
Use a small, fast model and a binary (yes/no) question. In practice:
- Binary beats numeric. LLM judges are inconsistent on numeric ranges; a clear yes/no question calibrated against human labels is far more reliable. Built-in signals use binary judgments for the same reason.
- Small models are usually enough. For well-scoped binary checks, lightweight models (for example GPT-4o mini, Claude Haiku, Gemini Flash class) typically reach high agreement with human labels at a fraction of the cost, which matters when an evaluator runs on every production session. Reserve larger models for genuinely nuanced judgments, and validate with a dry run before activating either way.
- Don't debate the model choice; test it. Take a set of real sessions and have a human mark the expected result on each (the ground truth). Then dry-run the evaluator on those same sessions and count how often it agrees with the human labels. If agreement is around 90%, the evaluator is trustworthy enough to activate, whatever size the model is. If a small model reaches that bar, use the small model; if it falls short, improve the prompt or step up to a stronger model and measure again.
Custom evaluators run on your own model API key (bring your own key), so you control the model, the provider, and the spend. This applies to calibration runs as well, including dry runs and one-off runs over past sessions, not just live scoring.
Do evaluators run retroactively on past sessions?
Active evaluators score new sessions as they arrive. To score historical sessions, run the evaluator from the Runs page over specific sessions, a date range (sampling supported), or a dataset. Fully automatic retroactive scoring is on the roadmap.
How do I know I can trust the scores?
Calibrate before you activate: assemble a dataset of real sessions, dry-run the evaluator, label a sample with human ground truth, and compare. Activate when agreement is high; iterate on the prompt when it isn't. Evaluator versions are tracked so you can refine over time.
What kinds of evaluators can I create, and when should I use each?
- LLM evaluators judge each session against a prompt you write. Use them for criteria that need reading comprehension, such as policy compliance, groundedness, tone, or whether a refund decision was correct. They spend tokens per session, and their judgment isn't perfectly deterministic.
- Code evaluators apply deterministic pattern rules with no model calls: Regex for a True or False pattern match, and Keyword Classifier to label a session by the keywords it matches. Use them for exact, mechanical checks.
A binary output calibrated against human labels is usually the most reliable choice. Refer to Create and calibrate evaluators.
Users barely click thumbs up/down. Is that a problem?
No. Explicit feedback is famously sparse and unreliable. Agent Analytics doesn't depend on it: built-in signals score every session, and Amplitude captures implicit behavioral signals (copy, retry, abandonment, expressed frustration) automatically. Amplitude ingests explicit ratings too, when you have them. A thumbs down sent as a user-feedback Score counts as explicit dissatisfaction for Negative Feedback.
Pricing
How is Agent Analytics priced?
Agent Analytics is billed on agent sessions, not events, metered per month or per year depending on your contract. Every plan includes a free session allotment, with additional volume available as you scale. Contact your account team or refer to the pricing page for details.
Will agent events count against my Amplitude event volume?
Mostly no. Core agent events (user messages, AI responses, and tool calls) are metered separately as agent sessions and do not count against your existing event volume limits. The exception is custom eval-related events, which do count against your Amplitude event volume.
What about the cost of running evaluations?
The built-in quality signals are included in the base price. Custom LLM-as-a-judge evaluators run on your own model API key, so inference costs go to your provider at your rates, and you choose the model.
How does LLM cost tracking work?
Costs are computed from the token counts on your traces multiplied against a continuously maintained open-source model-price catalog (Pydantic's genai-prices project), so using standard model names matters. Heavily cached or batch-discounted workloads may diverge from your provider bill; treat provider billing as the financial source of truth and Agent Analytics cost data as the analytical view (per-agent, per-user, per-feature attribution your bill can't give you). Models the catalog can't price (nonstandard names, gateway aliases, fine-tuned ft: models) get no automatic cost; pass the cost explicitly for those.
Privacy and security
What happens to conversation content? Where does it go?
You control exactly what leaves your systems, at three levels. Full content is the default, and it's what built-in signals, topics, and LLM evaluators read.
- Full content (
full, the default): message content, with PII redaction on by default in the Node SDK and in the Python SDK from 1.17.0. The SDK scrubs emails, phone numbers, credit cards, SSNs, IP addresses, and base64-encoded image data client-side, before data leaves your environment. The phone and SSN patterns target US formats; add custom regex patterns or plug in your own redaction (for example, Presidio) for international locales or anything domain-specific. - Metadata only (
metadata_only): events and metrics, no message text at all. The gate covers every content-emitting channel in the SDK: message text ($llm_message.text), tool inputs and outputs, system prompts, exception messages, stack traces, framework-emitted content attributes (gen_ai.system_instructions,gen_ai.input.messages,gen_ai.output.messages), and score comments. - Enriched metadata (
customer_enriched): structured labels and evaluation results you compute, no raw content.
Your organization's standard Amplitude security controls govern stored content, and role-based permissions restrict access to it. Refer to Agent Analytics privacy modes for what each mode sends.
Does Amplitude train AI models on your data?
No. Your conversation data is used solely to deliver Agent Analytics to you, powering your sessions, signals, and evaluations. Amplitude does not use your data to train AI models.
What about GDPR / HIPAA / compliance?
Amplitude's standard compliance posture and DPAs apply to Agent Analytics data. With full content, PII redaction runs in your process before data leaves your environment. If message text can't leave your systems at all, the metadata-only and enriched-metadata modes still give you cost, latency, tokens, sessions, and any labels you compute yourself. Built-in signals, topics, and LLM evaluators need message text, so they don't run in those modes.
Can you limit who on your team sees conversation content?
Yes. Access to Agent Analytics, and to conversation content specifically, is controlled through Amplitude's Role-Based Access Control (RBAC) feature. Admins grant and manage these permissions at the role level.
What agents does it support?
Does it work with any model or framework?
Yes, the event schema is provider-agnostic. OpenAI, Anthropic, Google (including Vertex), Bedrock, open-source models, LangChain, Vercel AI SDK, custom stacks: if you can emit events, it's supported. Cost tracking works best with standard model identifiers.
What about third-party or hosted agents you don't control (support bots, platforms like Intercom Fin or Sierra)?
Yes. A scheduled job forwards the conversations the platform exposes, and you get session analytics, quality signals, and outcome attribution. Expect less depth than with first-party instrumentation, because hosted platforms typically don't expose token-level cost or internal tool calls. Amplitude publishes guides for Sierra, Decagon, Fin by Intercom, and Salesforce Agentforce. Refer to Import agent conversations from other tools.
Voice agents?
Yes. Instrument server-side and forward the transcript/analysis webhooks your voice platform provides.
Multi-agent systems?
Yes. Each agent gets its own identifier and version; sub-agents that share a session ID roll up into one session, and journeys across agents are analyzable in standard Amplitude charts.
Working with the rest of Amplitude
Can I chart agent data in regular Amplitude charts?
Yes. Built-in signals land on the [Agent] Session Record event and custom evaluator results on [Agent] Evaluator Result, so everything works in the standard chart builder (segmentation, funnels, retention, dashboards) right alongside your product events.
Should agent data live in the same project as my product analytics?
Yes, agent data can live in the same project as your product analytics; that's the simplest way to analyze them together. If you prefer to keep agent data in a separate project, Portfolio is also supported, so identity resolves across both. Your CSM or solutions engineer can help you choose the right setup.
Troubleshooting
For setup problems, such as events that don't arrive, missing Session Records, empty message content, missing costs, or uninformative signal results, refer to Troubleshoot Agent Analytics setup.
これは役に立ちましたか?