---
title: Import agent conversations from your data warehouse
description: "Import agent conversations stored in Snowflake, BigQuery, or Databricks into Agent Analytics with one SQL query and Amplitude warehouse import, or replay stored OpenTelemetry, OpenInference, and MLflow spans."
product: general
lang: en
token_estimate: 1559
---
# Import agent conversations from your data warehouse

> For AI agents: a documentation index is available at [/docs/llms.txt](/docs/llms.txt). Append `.md` to any page URL for markdown, or send `Accept: text/markdown`.

Warehouse import turns agent conversations already stored in Snowflake, BigQuery, or Databricks into `[Agent]` events, with no AI SDK and no forwarder job. You write one SQL query that maps your table to a standard shape, and Amplitude warehouse import runs it on the schedule you set. Conversations stored as OpenTelemetry, OpenInference, or MLflow spans take a separate path that replays them to Amplitude's OTLP endpoint.

Use warehouse import for conversations that already land in your warehouse, including history. For an agent you're still building, instrument it with the AI SDK or OpenTelemetry instead, which capture more detail as it happens.

## Set up with your AI coding agent (recommended)

Paste this prompt into your AI coding agent, such as Cursor, Claude Code, Windsurf, GitHub Copilot, or Codex. The agent fetches the guide, builds the job, and starts with a dry run.

```text
Connect the agent conversations in our warehouse to Amplitude Agent Analytics.

Fetch https://raw.githubusercontent.com/amplitude/Amplitude-AI-Node/main/docs/integrations/warehouses/README.md
and follow "Part 2: Coding agent procedure". Start with a dry run.
```

The agent stops and asks for values it can't infer, such as which field holds your user ID. Review its changes, then verify with the steps in [Get started](#get-started).

## What this means for you

- **You write**: One mapping query, called Stage 1, from your table to the standard columns. A shared second query, Stage 2, turns those columns into `[Agent]` events. Don't edit Stage 2.
- **Amplitude runs**: The import, on your schedule, usually hourly. You don't re-run the query by hand.
- **Your warehouse runs**: The query itself, on your warehouse compute. On Databricks, you also schedule a `MERGE` job that keeps an events table up to date.

## Supported source formats

Pick the guide that matches how your table stores conversations:

| Your table has | Guide | Path |
| --- | --- | --- |
| One row per message, tool call, or UI component | [Message rows](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/message-rows.md) | SQL warehouse import |
| One row per turn, with the user and agent side by side | [Turn rows](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/turn-rows.md) | SQL warehouse import |
| One row per conversation with an OpenAI-style `messages` array | [OpenAI messages](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/openai-messages.md) | SQL warehouse import |
| OpenTelemetry GenAI spans | [OTel GenAI spans](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/otel-genai.md) | Span replay |
| OpenInference spans, for example from Arize Phoenix | [OpenInference spans](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/openinference.md) | Span replay |
| MLflow Tracing traces | [MLflow traces](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/mlflow-tracing.md) | Span replay |

Start with the [warehouse import overview](https://github.com/amplitude/Amplitude-AI-Node/blob/main/docs/integrations/warehouses/README.md), which covers the shared columns, settings, and per-warehouse setup.

## Warehouse requirements

- **Snowflake**: A Snowflake source in Amplitude with the Event data type and timestamp-based sync on the `import_cursor` column. Store timestamps as UTC.
- **BigQuery**: A BigQuery source in Amplitude with the Event data type and time-based sync on the `import_cursor` column.
- **Databricks**: Amplitude's Databricks import reads a Delta table's change data feed and can't read views. Create an events table with change data feed turned on, then schedule a `MERGE` job that writes your query's output into it.

You also need a user ID column that matches the user ID in your product analytics. Amplitude drops rows without a conversation ID, a user or device ID, or an agent ID.

## What carries over

Warehouse import sends only what your table recorded. Model, tokens, cost, and tool calls arrive when your table has them, and stay empty when it doesn't. Amplitude doesn't estimate missing values. Each conversation gets one `[Agent] Session End`, so built-in signals run once it settles. To leave out or redact a field, do it in Stage 1, before data leaves your warehouse.

## Freshness and limits

- **Settle time**: A conversation imports after it has been idle for `settle_hours`, which defaults to 2. Set it longer than the longest real pause in a conversation.
- **Updates**: The import picks up new rows and rows with a later `updated_at`. A conversation that resumes after import keeps its first Session End.
- **Duplicates**: Amplitude deduplicates repeat events within 7 days. Re-importing rows older than that creates duplicates.
- **Backfill**: The first sync imports every settled conversation your query returns. Import a week first, check it, then widen how far back the query looks.

## Replay stored spans

For OpenTelemetry, OpenInference, or MLflow spans, export the spans from your warehouse and run `otlp-replay.mjs` from the guides repository. The script needs Node.js 18 or later and your Amplitude project API key, and posts the spans to Amplitude's OTLP endpoint, US or EU. Schedule the export and replay, for example hourly. Use `--dry-run` to preview without sending.

Without a conversation ID on the spans, each trace becomes its own session. Don't replay windows older than 7 days, because Amplitude only deduplicates within that window.

## Get started

1. Open the guide for your format from [Supported source formats](#supported-source-formats).
2. Write Stage 1 and check a sample with `check-agent-events.mjs`.
3. Create the warehouse source in Amplitude and run the first sync on one week of data.
4. Check Live Events and the session viewer, then run a second sync to confirm it adds no duplicates.

Guides last verified on October 6, 2026.

