# MCP Server Design: Lessons from Amplitude’s Tool Consolidation

Learn how Amplitude redesigned its MCP server to reduce context bloat, consolidate tools, and help AI agents complete tasks in fewer calls.

Source: https://amplitude.com/en-us/blog/lessons-from-consolidating-our-mcp-server

---

[Chanaka Perera](/blog/author/chanaka-perera)

[AI Engineer, Amplitude](/blog/author/chanaka-perera)

[](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Flessons-from-consolidating-our-mcp-server)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Flessons-from-consolidating-our-mcp-server)[](https://twitter.com/intent/tweet?url=https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Flessons-from-consolidating-our-mcp-server\&text=Lessons%20from%20consolidating%20our%20MCP%20server)[](mailto:?subject=Checkout%20this%20Amplitude%20Article\&body=Check%20this%20out%3A%20https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Flessons-from-consolidating-our-mcp-server)

When we first launched the [MCP server](https://amplitude.com/mcp-server), customer requests prompted teams at Amplitude to expose their functionality as tools. Usually that was one tool per API endpoint. The result was burdensome context bloat, overlapping tools, and a lot of steering to get the model to call the right one. By June, connecting to our server loaded 96 tools.

The problem was that every tool we added for one customer also created context for every customer. Niche requests rarely matched what other customers needed, so most users paid the context cost for tools they would never touch.

Our wake-up call came from a customer running Sonnet with a 200K context window. When their client listed our tools, just connecting to our serve&#x72;**&#x20;**&#x75;sed up **57%** of it before a single tool cal&#x6C;**.**

## **Why not just rely on tool search?**

Some clients, like Claude Code, support [**tool search**](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool), which loads tool definitions only when they're needed. For most other clients, turning on tool search is a decision made by the client’s owners, not the server. The default is t&#x6F;**&#x20;**&#x6C;oad every definition upfront..

From our instrumentation of Amplitude users connecting to our server, we concluded it's safest to **assume every tool costs tokens on every connection**, for three reasons:

- Most clients still load every tool definition upfront by default
- On Amazon Bedrock, tool search only works through the InvokeModel API, not the more common [Converse API](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool)
- Enterprise gateways often cache the tool list once and never refresh it

Token cost isn’t the only problem. Unnecessary tools also increase the chance of error. The more tools you expose, the more often the model picks the wrong one. Even with tool search, [Anthropic's own evals](https://www.anthropic.com/engineering/advanced-tool-use) show Opus 4.5 choosing correctly only 88% of the time, and older models do considerably worse. Anthropic also notes that wrong tool selection is one of the most common failures, especially when tools have similar names.

## **Build tools around outcomes, not endpoints**

We rebuilt our catalog around three rules.

****

**1. Don't mirror your API**

Our tools mirrored our API, so the model had to call several of them in the right order and pass IDs among them. Agents are bad at that. The easiest way to prevent overlap and cut description context was to combine tools based on how they were used.

We replaced groups of endpoint tools with one tool per object: cohort, chart, dashboard, flag. Each takes an action and routes to the existing handlers. Cohorts went from nine tools to one.

*One cohort tool, eleven actions, routed to existing handlers.*

Merging also shows you what's missing. There was no tool to list cohorts because there was no listing endpoint, so agents were calling the cohort lookup tool with invented IDs. list is a new route, not a rename.

Each action keeps its own permission level and audit event. This matters. Gate a merged tool at its riskiest action and read-only users lose the reads; gate it at the lowest, and you expose writes. We resolve access per action before the handler runs.

**2. Merge functionality around usage even if it crosses multiple product boundaries**

Tool boundaries should come from production usage, not from lines the API draws. We had one search tool per product: charts, dashboards, notebooks, cohorts. Instrumentation showed agents calling two or three of them in a row for the same request, because unclear directions like "find the retention analysis" doesn't specify which kind of thing it is.

Now there's one search tool that takes a list of entity types and searches all of them in one call. This allows the agent to do more dynamic search queries across multiple products rather than maintaining disjointed search calls for each entity.

****

**3. Merge tools if they need to be chained to achieve an outcome, but don't rely on the agent to chain calls**

Composability is important for certain use cases, but sometimes it can hurt the overall server experience. Tools designed like a REST API assume the caller will string them together in order.

Charts are a good example of this flow. Creating a chart in Amplitude takes three steps: create a chart edit, save it, add it to a dashboard. The dashboard rejects an unsaved edit, so order matters. We had tools that copied those steps: one to query the data and return a chart edit ID, one to render the chart from that ID, and one to save the edit.

Of 18,703 users who queried data, 20% went on to render a chart, and less than 5% saved it.

The rationale field explains why. Every tool on our server has one optional parameter called Tool Rationale which is a "Brief explanation of why you are calling this tool." Agents fill it in about three-quarters of the time without being told to. When query calls stop at that step, the rationales read like "save chart edit for dashboard analysis." The agent tried to save the chart, but never made it through the chain.

Now the agent sends a definition and gets data back in one call. Saving is also supported as a button on the rendered chart, covered in the MCP Apps section below. What used to be 16 chart tools became only four.**&#x20;**&#x4D;ulti-ste&#x70;**&#x20;**&#x74;ools like that were an important step to help us reduce context.

## **Two more places for context**

**Skills and CLI**

Tool descriptions load on every connection, so long instructions increase the cost for every user. We moved the descriptions into skills: markdown files the agent loads only when a task needs them. There are 37 today, covering things like building a chart, diagnosing errors, and running a weekly brief.

They live in a public, MIT-licensed repo ([amplitude/mcp-marketplace](https://github.com/amplitude/mcp-marketplace)) pulled into the server as a git submodule. Adding or fixing a skill is a pull request against that repo. The same files serve hosts that install the plugin directly and hosts that connect over MCP.

We serve skills as MCP resources, which is the spec's answer for this. In practice, only one in ten users is on a client that reads resources. The rest get them through a regular tool that fetches a skill by name, so we serve both and log which surface delivered the read.

For rarely used API operations, don't add tools. We put a CLI behind one tool so a new endpoint works the day it ships. Permissions come from the HTTP method, and unknown commands are treated as writes.

****

**MCP Apps**

MCP Apps let a person review a destructive or sensitive action before it runs, which is what makes those actions safe for public agents. A 'confirmed: true' parameter doesn't do this. The model will just set it.

We used the same flow for other actions that affect more than one person, like sharing entities and managing spaces.

Apps also let us add functionality without adding context. A tool with its visibility set to app is callable by the UI but never listed to the model, so it costs nothing on connection. Not every host supports MCP Apps yet, so keep a fallback.

## **Deployment and monitoring**

We shipped each change behind a feature flag, one per product area. Turning on a flag swaps the old tools for the new one in a single step, so no org ever sees both. We enabled the flags for internal orgs first, then rolled them out to customers in stages. Our CSMs knew what was changing and when, so they could give a heads-up to accounts with automations built on the old tools.

While each flag ramped, we watche&#x64;**&#x20;**&#x74;wo numbers:

- Error rate on tool calls
- Recovery rate: after a failed call, did the agent make a successful one within five minutes?

Recovery is the one we prioritized. A tool that fails with a clear message is fine. A tool that fails and leaves the agent guessing is not. Several of the merged tools added validation the old ones didn't have, so our error rate went up at first. That was the validation doing its job: a bad definition now fails at the schema with a message that tells the agent what to fix, instead of failing in the backend. Recovery stayed high through the rollout, so agents were using those messages to get to a successful call.

Once the change was live for all customers, we analyzed the MCP server like any other product, in Amplitude. We tracked retention across orgs and saw it grow. Over the last 12 weeks, among external customers, retention has been strong. 67% of users came back after one week, with 59% returning after four, and 50% after eight. Heavy users grew fastest. A smaller catalog brought in more people and let existing users do more.

## Start with how agents use your tools

We started with 96 tools because exposing each endpoint seemed like a reasonable way to give agents more capabilities. Production usage showed us where that approach broke down: overlapping tools, unfinished chains of calls, and context spent on functionality most users never needed.

If you're building an MCP server, look at where agents get stuck. Which tools do they call together? Where do they stop before completing the task? Use those patterns to decide what to combine, move into a skill, or handle in the UI. Then measure whether agents can finish the job—and recover when something goes wrong.

##### Ready to bring your data to every AI tool you use?

Connect Amplitude to your agents and IDE with the Amplitude MCP server.

[Set it up](/signup?source=blog-mcp\&topic=mcp\&siteLocation=blog-inline-cta)

About the author

Chanaka Perera

AI Engineer, Amplitude

[More from Chanaka](/blog/author/chanaka-perera)

Chanaka is an AI Engineer at Amplitude, where he’s building the MCP server that brings Amplitude’s behavioral context directly into your AI tools.

[More from Chanaka](/blog/author/chanaka-perera)

Topics

[Engineering](/blog/tag/engineering)

[MCP](/blog/tag/mcp)

#### Recommended Reading

[Read ](/blog/what-ai-engineers-do)

[Insights](/blog/what-ai-engineers-do)

###### [What AI engineers do in the age of agentic code](/blog/what-ai-engineers-do)

[Oct 1, 2026](/blog/what-ai-engineers-do)

[14 min read](/blog/what-ai-engineers-do)

[Read ](/blog/jev-analysis)

[Insights](/blog/jev-analysis)

###### [An analysis of Jev: Faster and cheaper, but watch the accuracy](/blog/jev-analysis)

[Sep 28, 2026](/blog/jev-analysis)

[13 min read](/blog/jev-analysis)

[Read ](/blog/ai-context-configuration)

[Product](/blog/ai-context-configuration)

###### [Your agent isn’t broken. It’s guessing.](/blog/ai-context-configuration)

[Sep 23, 2026](/blog/ai-context-configuration)

[6 min read](/blog/ai-context-configuration)

[Read ](/blog/headless-amplitude)

[Company](/blog/headless-amplitude)

###### [Bring headless Amplitude where you work](/blog/headless-amplitude)

[Sep 16, 2026](/blog/headless-amplitude)

[3 min read](/blog/headless-amplitude)
