io.github.TinySuiteHQ/tinysearch MCP Server
io.github.TinySuiteHQ/tinysearch
Fast, self-hosted web search and crawling for AI agents—returns ranked evidence, not full pages.
What is the io.github.TinySuiteHQ/tinysearch MCP server?
TinySearch is a self-hosted web-research tool for MCP agents that searches the web, crawls pages, and returns compact, source-grounded evidence chunks instead of full webpages. It uses local hybrid reranking (BM25 and ONNX embeddings) to minimize token usage and cost, with no required API keys or hosted accounts.
TinySearch reduces the cost of agent web research by moving page selection and passage ranking to the local machine before content reaches the model. It searches via DDGS or self-hosted SearXNG, crawls pages with Playwright, and returns only the highest-ranked evidence chunks with source URLs. This cuts input tokens by ~64% compared to naive search-and-fetch workflows, making it ideal for agents that need reliable, cost-efficient research without vendor lock-in.
How to install io.github.TinySuiteHQ/tinysearch
Copy-paste configuration for popular MCP clients.
MCP_TRANSPORTMCP transport to serve: stdio (default), sse, or streamable-http.
TINYSEARCH_CONFIG_PATHPath to a tinysearch_config.json inside the container, for overriding search/embedding defaults.
SEARXNG_URLSearXNG search endpoint URL. Only needed if not using the bundled Compose SearXNG service.
TINYSEARCH_EMBEDDING_BACKENDEmbedding implementation: onnx (default) or openai_compatible.
TINYSEARCH_EMBEDDING_MODELLocal ONNX preset (fast, balanced, quality) or a Hugging Face ONNX repository id.
OTEL_SERVICE_NAMEOpenTelemetry service.name resource attribute; defaults to tinysearch.
OTEL_EXPORTER_OTLP_ENDPOINTCommon OTLP endpoint for optional TinySearch traces and metrics.
OTEL_EXPORTER_OTLP_PROTOCOLOTLP transport: http/protobuf (default) or grpc.
OTEL_EXPORTER_OTLP_HEADERSsecretOptional authentication or routing headers for the OTLP exporter.
OTEL_RESOURCE_ATTRIBUTESAdditional standard OpenTelemetry resource attributes.
OTEL_TRACES_EXPORTERSet to otlp to enable traces explicitly or none to disable them.
OTEL_METRICS_EXPORTERSet to otlp to enable metrics explicitly or none to disable them.
OTEL_SDK_DISABLEDSet true to disable all OpenTelemetry signals.
Tools & capabilities
Tools this server exposes to the agent.
search— Fast backend-ordered web discovery without crawling or reranking; returns titles, URLs, previews, and dates.scrape_urls— Reads one to five known pages concurrently; returns selected Markdown chunks with optional hybrid reranking by query.get_current_datetime— Returns current UTC date and time for time-dependent queries.research— Legacy all-in-one search, crawl, and rerank pipeline; deprecated in favor of composing search with scrape_urls.
Use cases
- Research questions by searching the web and scraping the top results, receiving only ranked evidence passages instead of full pages
- Reduce token costs for agent workflows by filtering low-value content locally before it enters model context
- Build self-hosted web research without paying for metered search APIs or embedding services
- Cite sources accurately by retrieving original, unedited page text with source URLs attached to each evidence chunk
- Compose flexible research pipelines using search for discovery and scrape_urls for focused page reading
io.github.TinySuiteHQ/tinysearch MCP server FAQ
TinySearch searches the web, crawls pages, and returns compact, ranked evidence chunks with source URLs. It uses local BM25 and ONNX embedding reranking to filter low-value content before it reaches your model, reducing token usage by ~64% compared to naive search-and-fetch.
Yes. TinySearch uses DDGS (DuckDuckGo) for search and local ONNX embeddings by default—no API keys, hosted accounts, or per-request billing required. You can optionally add SearXNG, Brave Search, or other backends.
Add it to your MCP client config with `uvx --from "tinysuite-search[server]" tinysearch`, or use Docker. See the quick-start section in the README or the full installation guide at tinysuite.dev/docs/tinysearch/.
No authentication is required by default. DDGS search and local embeddings work out of the box. Optional paid backends (Brave Search, OpenAI embeddings) require their own API keys if you choose to use them.
Yes. TinySearch runs as a Python package, MCP server, Docker container, or FastAPI app. The Docker tier includes a bundled SearXNG instance for fully self-hosted search without external dependencies.
DDGS (default, no setup), self-hosted SearXNG, DuckDuckGo-only mode, and Brave Search API (with key). You can switch backends or add fallbacks without changing code.
README (reference)
Source of truth, from the repository.
TinySearch
<!-- mcp-name: io.github.TinySuiteHQ/tinysearch --> <p align="center"> <a href="https://tinysuite.dev"> <img src="assets/tinysearch-full-logo.png" alt="TinySearch" width="240" /> </a> </p> <p align="center"> <strong>Spend tokens on answers, not webpages.</strong> </p> <p align="center"> TinySearch searches, crawls, and reranks the web locally, then gives your agent only the evidence worth putting in its context. </p> <p align="center"> <a href="https://tinysuite.dev/docs/tinysearch/">Documentation</a> · <a href="#quick-start">Quick start</a> · <a href="#python-library">Python</a> · <a href="https://discord.gg/NG6u2zamR">Discord</a> </p>TinySearch is a self-hosted web-research tool for AI agents. It searches the web, reads the best pages, removes low-value content, and returns compact evidence with source URLs.
Your model receives the useful passages instead of paying to process entire webpages.
TinySearch is part of TinySuite, a suite of focused tools designed to make agentic operations cheaper by minimizing token usage through smart retrieval, selection, and context-management techniques.
<p align="center"> <img src="assets/tinysearch-readme.gif" alt="TinySearch returning source-grounded web evidence to an AI agent" width="780" /> </p>Choose a tier
| Tier | Use it when | Entry point | Search backend |
|---|---|---|---|
| 1. Python library | You are building with TinySuite or Python | pip install tinysuite-search | DDGS |
| 2. One-command MCP | An MCP client should launch TinySearch for you | uvx --from "tinysuite-search[server]" tinysearch | DDGS |
| 3. Docker + SearXNG | You want the full self-hosted stack and HTTP MCP | docker compose ... up -d | Bundled SearXNG |
Tiers 1 and 2 need no search service. Tier 3 adds a dedicated SearXNG service, persistent model storage, and a network MCP endpoint. See the installation guide for the Docker setup.
The expensive part of agent research is context
A search result is not yet useful evidence. Agents often have to open several pages, ingest navigation and boilerplate, and spend paid input tokens deciding which passages matter.
TinySearch moves that work in front of the model:
flowchart LR
A[Question] --> B[Search and crawl]
B --> C[Local hybrid reranking]
C --> D[Compact evidence<br/>with source URLs]
D --> E[Your agent]
That lowers cost in three ways:
- Smaller model context. Only the best-ranked evidence chunks are returned, within a controlled evidence budget.
- No metered search API required by default. TinySearch can search through DDGS without a paid search provider.
- Local retrieval by default. ONNX embeddings and hybrid reranking run on your machine instead of creating embedding API charges.
Search broadly. Read locally. Pay the model only for the evidence that matters.
This is retrieval, not summarization: TinySearch selects the passages worth keeping with local BM25 and embedding rerank, it doesn't run a model over the page to rewrite or condense it. Every returned chunk is the original page text, unedited, so what you cite is what the page actually said. That keeps the pipeline fast and free to run locally, at the cost of not compacting as aggressively as a dedicated reduction model could. A learned reduction step is a direction we may explore later; it isn't part of TinySearch today.
Actual savings depend on the pages, evidence limits, client model, and provider pricing. TinySearch reduces the web content sent to the model; it does not control what the client does with that evidence afterward.
<p align="center"> <img src="assets/token-savings-benchmark.svg" alt="Benchmark: TinySearch uses 64% fewer tokens than a naive search-and-fetch agent across 8 research queries, cutting modeled input cost per 1,000 queries from $55.08 to $20.03 at $3 per million tokens" width="900" /> </p>The cost panel uses an illustrative $3.00 per million input-token rate and excludes search, crawling, model output, and downstream agent use.
The naive baseline isn't a strawman product, it's the same pages TinySearch
crawled for each query, fed to the model unfiltered, the way a generic
"search, then fetch the page" tool (a plain web-search-plus-fetch loop, the
kind built into most coding agents) would. Measured against the current
recommended flow (search then scrape_urls, not the deprecated
all-in-one research tool) and counted on the actual MCP tool-result text,
TinySearch's primary interface. Reproduce or rerun it yourself:
python scripts/benchmark_token_savings.py --json-out report.json
Quick start
With uv installed, add TinySearch to any MCP
client:
{
"mcpServers": {
"tinysearch": {
"command": "uvx",
"args": [
"--python",
"3.12",
"--from",
"tinysuite-search[server]",
"tinysearch"
]
}
}
}
The client launches TinySearch over stdio when it needs it. No repository clone, hosted account, or paid search key is required.
Fast search starts without Chromium or an embedding model. The first scrape
initializes Chromium; focused scraping and the legacy research tool also
initialize the configured embedding model. Pre-warm both ahead of time if you
will use those workflows:
uvx --from "tinysuite-search[server]" tinysearch setup
<p align="center">
<img src="assets/demo_terminal_prompt.gif" alt="TinySearch CLI setup and first run in a terminal" width="780" />
</p>
Prefer Docker, a remote MCP endpoint, or a source checkout? Follow the installation guide.
Four MCP tools
| Tool | Use it when |
|---|---|
search(query) | You need fast, backend-ordered discovery without crawling or reranking |
scrape_urls(items) | You know one to five pages; each item may use * for its configured clean page-order token budget |
get_current_datetime() | A question depends on the current date or time |
research(query) | Legacy compatibility only; deprecated in favor of search followed by scraping |
TinySearch deliberately stays focused. It is a retrieval layer, not another agent, chat interface, hosted search product, or permanent web index.
See the complete MCP tool reference for parameters and response contracts.
What your agent gets
TinySearch does not spend another model call writing the final answer. The
recommended flow is search for lightweight discovery, then scrape_urls for
the pages worth reading.
Successful MCP tool-result text is XML. A search result looks like this:
<search_results>
<query>Python asyncio cancellation</query>
<results>
<result index="1">
<title>Coroutines and Tasks</title>
<url>https://docs.python.org/3/library/asyncio-task.html</url>
<search_preview>Tasks can be cancelled...</search_preview>
</result>
</results>
</search_results>
scrape_urls returns each page's selected Markdown chunks under one
<url_grounded_answers> batch root, and get_current_datetime returns
<current_datetime>. Dynamic values are escaped so retrieved content cannot
forge the XML boundaries around it.
MCP still uses its standard JSON-RPC transport envelope, including
protocol-level errors and optional structuredContent. Python and FastAPI keep
their structured JSON contracts for applications that need to store, inspect,
or transform the evidence.
How it works
searchreturns backend-ordered titles, URLs, previews, and upstream dates without starting Chromium or an embedding model.scrape_urlsreads one to five known pages concurrently. Omit an item's query or use"*"to keep clean Markdown in page order within the configured token budget.- Supply a focused item query when TinySearch should chunk and hybrid-rank that page before returning evidence.
The deprecated MCP research tool retains the older all-in-one search, crawl,
and rerank pipeline for compatibility. New MCP integrations should compose
search with scrape_urls instead.
Python library
TinySearch also works as a regular Python package:
pip install tinysuite-search
import asyncio
from tinysearch import scrape_urls, search
async def main():
results = await search("Python async tasks")
print(results["results"])
page_url = results["results"][0]["url"]
evidence = await scrape_urls([{
"url": page_url,
"query": "How does asyncio cancellation work?",
}])
print(evidence["results"])
asyncio.run(main())
The Python API returns stable, JSON-serializable results. search accepts a
per-call limit from 1 to 50. scrape_urls accepts a per-call max_tokens
budget (4,000 by default); omit an item's scrape query or use "*" for
page-order mode. Rendering structured evidence into an LLM prompt is explicit,
so applications can store, inspect, transform, or budget the result first.
The optional FastAPI app mirrors these surfaces. POST /search and
POST /research accept output_format (prompt or json) and always respond
with JSON; prompt mode places rendered text in the answer field.
POST /scrape accepts one to five { "url", "query" } items and always
returns structured per-item outcomes.
The app also exposes /health, /current_datetime, and read-only /config;
configuration writes require explicit environment opt-in.
Search backends
TinySearch selects a web-search backend from config, so you can start with no search service and add one later without changing code.
"ddgs"(native default): queries theddgspackage's automatic backend selection in-process. No SearXNG deployment required."searxng"(Docker default): queries a self-hosted SearXNG instance. Falls back toddgson backend failure unlesssearch_backend_fallbackis set tofalse."duckduckgo": skips SearXNG and queriesddgsin DuckDuckGo-only mode."auto": tries SearXNG, then falls back toddgson any backend failure.
Set the BRAVE_SEARCH_API_KEY environment variable to add Brave's official
Web Search API as a keyed fallback for the ddgs and duckduckgo backends.
Brave is only consulted when the primary call errors or returns no results.
Full key reference, SearXNG JSON-output setup, and Compose details live in the configuration reference.
External browser over CDP
TinySearch uses its bundled Playwright Chromium by default. To use a browser that you operate separately, set its Chrome DevTools Protocol endpoint in the config file:
{
"browser_cdp_url": "http://browser:9222"
}
Server processes also accept TINYSEARCH_BROWSER_CDP_URL. When either setting
is present, TinySearch connects through Crawl4AI instead of installing or
launching the bundled Chromium. The external browser owns its executable,
profile, proxy, and fingerprint configuration; TinySearch does not select or
install a particular browser backend.
Treat a CDP endpoint as privileged remote control of the browser. Keep it on a
private network or loopback interface, require authentication when it crosses
a host boundary, and do not expose port 9222 directly to the public internet.
When TinySearch itself runs in Docker, localhost refers to the TinySearch
container, so use an endpoint reachable from that container.
The CDP endpoint is operator-managed and cannot be changed through the HTTP
PUT /config endpoint, even when configuration writes are enabled. Set it in
the startup environment or the file selected by TINYSEARCH_CONFIG_PATH, then
restart TinySearch. HTTP clients can continue updating other settings by
omitting browser_cdp_url from their partial update.
Why TinySearch
- No vendor in the loop. No TinySearch account, no required API key, no per-request billing, no analytics service or hosted scraped-data cache. The infrastructure you'd otherwise pay a search API for runs on your machine.
- Source-grounded by construction. Every evidence chunk is the original page text, still attached to its originating URL, so a claim in your agent's answer traces back to one specific passage instead of stopping at "the vendor's model said this."
- Built around token efficiency. Page selection and passage selection happen locally, before content enters model context.
- Useful without paid infrastructure. DDGS search and local ONNX embeddings are the defaults.
- Bring your own stack when needed. SearXNG and OpenAI-compatible embedding providers remain optional.
- Works where agents already work. Use MCP over stdio, Streamable HTTP, Python, FastAPI, or Docker.
Part of TinySuite
TinySuite is a product suite built around one idea: agents should spend tokens on useful work, not operational overhead.
Each tool focuses on a different part of the agent workflow and uses targeted techniques to reduce unnecessary context before it reaches the model. TinySearch handles the web-research layer by turning pages into a small, ranked, source-grounded evidence packet.
Documentation
The README is the product overview. Detailed setup and operational material lives in the TinySuite documentation:
The repository also contains an annotated example configuration at
configs/tinysearch_config.json.
When not to use TinySearch
TinySearch is intentionally lightweight. Use a commercial search API, persistent crawler, or full search index when you need:
- guaranteed search coverage or an SLA
- large-scale or scheduled indexing
- long-term page storage and change history
- enterprise observability and access controls
Development
git clone https://github.com/TinySuiteHQ/TinySearch
cd TinySearch
python -m venv .venv
source .venv/bin/activate
pip install -e ".[server]"
python -m unittest discover tests
TinySearch supports Python 3.12 and newer. CI tests Python 3.12, 3.13, and 3.14 across Linux, macOS, and Windows.
Entrypoints
tinysearch.searchandtinysearch.scrape_urls: structured Python APItinysearch.research: legacy all-in-one structured Python research pipelinetinysearch.get_current_datetime: structured UTC date and timetinysearch.to_prompt: pure structured-evidence prompt renderertinysearch mcp: stdio MCP server (also the no-argument default)tinysearch serve: Streamable HTTP MCP servertinysearch.servers.fastapi_server:app: optional FastAPI application
Community
Questions, ideas, and bug reports are welcome:
Privacy and license
TinySearch reads public pages and returns selected excerpts to the calling client. Search, crawling, local embeddings, and reranking can run without sending page content to an embedding provider. If you choose an OpenAI-compatible embedding backend, that provider receives the text sent for vectorization.
TinySearch is available under the MIT License. Downloaded model weights remain subject to their respective model-card licenses. See NOTICE for third-party distribution details.
Related MCP servers
Web access for AI agents — search/scrape/extract/crawl via the auxiliar.ai gateway.
View repository →
io.github.Toby-Self/asymptotic-ethics
MCP server for an Asymptotic Ethics governance simulation with a verified compliance oracle.

io.github.Toby-Self/federated-ai-commons
MCP server for a federated AI-commons governance simulation with a verified compliance oracle.

Elasticsearch
Elasticsearch MCP Server with multi-version support (ES 5.x-9.x) and comprehensive API access

Kibana
Access and manage your Kibana instance via natural language with dynamic API discovery and Elastic Stack integration.

