PluginBench
MCP Server
Maintained
MIT

io.github.ref-tools/ref-tools-mcp MCP Server

io.github.ref-tools/ref-tools-mcp

Token-efficient documentation search for AI coding agents over public and private resources.

What is the io.github.ref-tools/ref-tools-mcp MCP server?

The Ref MCP server is a ModelContextProtocol server that gives AI coding agents like Claude access to fast, token-efficient search over technical documentation for APIs, libraries, and services. It minimizes context bloat by returning only the most relevant documentation sections, reducing both token usage and cost while improving agent accuracy.

Ref provides agentic search tools designed to find exactly the documentation context your coding agent needs while using minimal tokens. It filters repeated search results within a session, extracts only relevant sections from large documentation pages, and supports searching both public documentation and private resources like repositories and PDFs. This approach reduces context rot, lowers API costs, and keeps models performing at their best.

How to install io.github.ref-tools/ref-tools-mcp

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • REF_API_KEY
    required
    secret

    Your API key for Ref

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "ref-tools-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "ref-tools-mcp"
      ],
      "env": {
        "REF_API_KEY": "<YOUR_REF_API_KEY>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • ref_search_documentation — Search technical documentation across public and private sources. Accepts a full sentence or question query and returns relevant documentation results with URLs.
  • ref_read_url — Fetch and convert webpage content to markdown. Works in conjunction with search results to retrieve full documentation content from specific URLs.

Use cases

  • Search API documentation to find the exact endpoint or method your agent needs to call
  • Retrieve relevant code examples and best practices from library documentation without pulling in irrelevant sections
  • Access private repository documentation and PDFs alongside public documentation in a single search
  • Reduce token usage and API costs by fetching only the most relevant documentation sections for a coding task
  • Iteratively refine searches based on previous results to find the right context for complex multi-step coding problems

io.github.ref-tools/ref-tools-mcp MCP server FAQ

What is the Ref MCP server?

Ref is an MCP server that provides token-efficient search over technical documentation. It's designed for AI coding agents to find exactly the context they need while minimizing token usage and cost.

Is Ref free to use?

Ref requires an API key (sign up at ref.tools to get one). The README does not specify pricing details.

How do I install Ref in Cursor?

Use the Cursor deep link provided in the README, or manually add the HTTP configuration with your API key: {"type": "http", "url": "https://api.ref.tools/mcp?apiKey=YOUR_API_KEY"}. Alternatively, use the stdio method with npx and the REF_API_KEY environment variable.

Can Ref search private documentation?

Yes, Ref can search private resources including repositories and PDFs, not just public web documentation.

How does Ref reduce token usage?

Ref filters out repeated search results in a session, extracts only the most relevant sections from large documentation pages (targeting ~5k tokens), and uses search history to dropout less relevant content.

What authentication is required?

You need a REF_API_KEY obtained by signing up at ref.tools. This is passed either as a URL parameter (HTTP mode) or environment variable (stdio mode).

README (reference)

Source of truth, from the repository.

Documentation for your agent smithery badge Website License npm version

Ref MCP

A ModelContextProtocol server that gives your AI coding tool or agent access to documentation for APIs, services, libraries etc. It's your one-stop-shop to keep your agent up-to-date on documentation in a fast and token-efficient way.

For more see info ref.tools

Agentic search for exactly the right context

Ref's tools are designed to match how models search while using as little context as possible to reduce context rot. The goal is to find exactly the context your coding agent needs to be successful while using minimum tokens.

Depending on the complexity of the prompt, LLM coding agents like Claude Code will typically do one or more searches and then choose a few resources to read in more depth.

For a simple query about Figma's Comment REST API it will make a couple calls to get exactly what it needs:

SEARCH 'Figma API post comment endpoint documentation' (54 tokens)
READ https://www.figma.com/developers/api#post-comments-endpoint (385 tokens)

For more complex situations, the LLM will try to refine it's prompt as it reads results. For example:

SEARCH 'n8n merge node vs Code node multiple inputs best practices' (126)
READ https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.merge/#merge (4961)
READ https://docs.n8n.io/flow-logic/merging/#merge-data-from-multiple-node-executions (138)
SEARCH 'n8n Code node multiple inputs best practices when to use' (107)
READ https://docs.n8n.io/code/code-node/#usage (80)
SEARCH 'n8n Code node access multiple inputs from different nodes' (370)
SEARCH 'n8n Code node $input access multiple node inputs' (372)
READ https://docs.n8n.io/code/builtin/output-other-nodes/#output-of-other-nodes (2310)

Ref takes advantage of MCP sessions to track search trajectory and minimize context usage. There's a lot more ideas cooking but here's what we've implemented so far.

1. Filtering search results

For repeated similar searches in a session, Ref will never return repeated results. Traditionally, you dig farther in to search results by paging to the next result but this approach allows the agent to page AND adjust the prompt at the same time.

2. Fetching the part of the page that matters

When reading a page of documentation, Ref will use the agent's session search history to dropout less relevant sections and return the most relevant 5k tokens. This helps Ref avoid a big problem with standard fetch() web scraping which is when it hits a large documentation page you can easily end up pull in 20k+ tokens into context, most of which are irrelevant.

Why does minimizing tokens from documentation context matter?

1. More context makes models dumber

It's well documented that as of July 2025 that models get dumber as you put in more tokens. You might have heard about how models are great with long context now and that's kind of true but not the whole picture. For a quick primer on some research, checkout this video from the team at Chroma.

2. Tokens cost $$$

Imagine you are using Claude Opus as a background agent and you start by having the agent pull in documentation context and suppose it pulls in 10000 tokens of context with 4000 being relevant and 6000 being extra noise. At API pricing, that 6k tokens cost about $0.09 PER STEP. If one prompt ends up taking 11 steps with Opus, you've spent $1 for no reason.

Setup

There are two options for setting up Ref as an MCP server, either via the streamable-http server (recommended) or local stdio server (legacy).

This repo contains the legacy stdio server.

Streamable HTTP (recommended)

Install Ref MCP in Cursor

"Ref": {
  "type": "http",
  "url": "https://api.ref.tools/mcp?apiKey=YOUR_API_KEY"
}

stdio

Install Ref MCP in Cursor (stdio)

"Ref": {
  "command": "npx",
  "args": ["ref-tools-mcp@latest"],
  "env": {
    "REF_API_KEY": <sign up to get an api key>
  }
}

Tools

Ref MCP server provides all the documentation related tools for your agent needs.

ref_search_documentation

A powerful search tool to check technical documentation. Great for finding facts or code snippets. Can be used to search for public documentation on the web or github as well from private resources like repos and pdfs.

Parameters:

  • query (required): Query to search for relevant documentation. This should be a full sentence or question.

ref_read_url

A tool that fetches content from a URL and converts it to markdown for easy reading with Ref. This is powerful when used in conjunction with the ref_search_documentation tool that returns urls of relevant content.

Parameters:

  • url (required): The URL of the webpage to read.

OpenAI deep research support

Ref can be used as a source for deep research. OpenAI requires specific tool definitions so when used with an OpenAI client, Ref will provide the same tools with slightly different naming.

ref_search_documentation(query) -> search(query)
ref_read_url(url) -> fetch(id)

Development

npm install
npm run dev

Running with Inspector

For development and debugging purposes, you can use the MCP Inspector tool. The Inspector provides a visual interface for testing and monitoring MCP server interactions.

Visit the Inspector documentation for detailed setup instructions.

To test locally with Inspector:

npm run inspect

Or run both the watcher and inspector:

npm run dev

Local Development

  1. Clone the repository
  2. Install dependencies:
npm install
  1. Build the project:
npm run build
  1. For development with auto-rebuilding:
npm run watch

License

MIT

Related MCP servers

Search curated design styles, real product screens, and user flows for evidence-based design work.

Model Context Protocol server for Refgrow affiliate program management API

2
TypeScript
View repository →

AI agents pay APIs over Lightning: L402 payments, wallets, invoices, budgets, and agent-to-agent commerce.

9
C#
MIT
View repository →

Bridges AI chat interfaces with Claude Code CLI via shared markdown control files.

0
TypeScript
MIT
View repository →

Local multi-channel message emulator for E2E testing, with MCP tools for LLM agents.

SQSquire logo

Squire

Maintained

CLI-first remote runtimes for validation and offload tasks, exposed over MCP stdio.

0
Go
View repository →