PluginBench
MCP Server
Active
Apache-2.0

io.github.cyanheads/pubchem-mcp-server MCP Server

io.github.cyanheads/pubchem-mcp-server

Search PubChem chemical compounds, properties, safety data, bioactivity, and cross-references via MCP.

What is the io.github.cyanheads/pubchem-mcp-server MCP server?

The PubChem MCP server provides access to the PubChem chemical database through 10 tools and 6 resources. It enables searching compounds by identifier, formula, or structure; fetching physicochemical properties, safety data, bioactivity, interactions, and cross-references; and finding bioassays by biological target. The server runs as stdio, local HTTP, or via a public hosted endpoint.

This server integrates PubChem's PUG REST and PUG View APIs into your MCP client, letting you search chemical compounds, retrieve detailed properties and safety classifications, explore bioactivity profiles, and find related bioassays. It's useful for chemists, researchers, and AI agents needing rapid access to chemical data without API keys.

How to install io.github.cyanheads/pubchem-mcp-server

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • MCP_LOG_LEVEL

    Sets the minimum log level for output (e.g., 'debug', 'info', 'warn').

  • MCP_SESSION_MODE

    HTTP session mode: stateless, stateful, or auto. Stateless is sufficient for PubChem; auto resolves to stateful.

  • MCP_HTTP_HOST

    The hostname for the HTTP server.

  • MCP_HTTP_PORT

    The port to run the HTTP server on.

  • MCP_HTTP_ENDPOINT_PATH

    The endpoint path for the MCP server.

  • MCP_AUTH_MODE

    Authentication mode to use: 'none', 'jwt', or 'oauth'.

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "pubchem-mcp-server": {
      "command": "bun",
      "args": [
        "-y",
        "@cyanheads/pubchem-mcp-server",
        "run",
        "start:stdio"
      ],
      "env": {
        "MCP_LOG_LEVEL": "<YOUR_MCP_LOG_LEVEL>",
        "MCP_SESSION_MODE": "<YOUR_MCP_SESSION_MODE>",
        "MCP_HTTP_HOST": "<YOUR_MCP_HTTP_HOST>",
        "MCP_HTTP_PORT": "<YOUR_MCP_HTTP_PORT>",
        "MCP_HTTP_ENDPOINT_PATH": "<YOUR_MCP_HTTP_ENDPOINT_PATH>",
        "MCP_AUTH_MODE": "<YOUR_MCP_AUTH_MODE>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • pubchem_search_compounds — Search for compounds by name, SMILES, InChIKey, formula, substructure, superstructure, or 2D similarity; supports batching and optional property hydration.
  • pubchem_get_compound_details — Get physicochemical properties, descriptions, synonyms, drug-likeness assessment, and pharmacological classification for compounds by CID.
  • pubchem_get_compound_image — Fetch a 2D structure diagram (PNG) for a compound by CID in small (100x100) or large (300x300) size.
  • pubchem_get_compound_3d_structure — Fetch a 3D conformer with atomic coordinates and bonds for a compound by CID, as parsed JSON or raw SDF.
  • pubchem_get_compound_xrefs — Get external database cross-references (PubMed, patents, genes, proteins, CAS numbers, etc.) for a compound by CID.
  • pubchem_get_compound_safety — Get GHS hazard classification, signal words, pictograms, and precautionary statements for one or more compounds by CID.
  • pubchem_get_bioactivity — Get a compound's bioactivity profile including assay results, targets, and activity values; filter by outcome or molecular target.
  • pubchem_get_compound_interactions — Get drug-drug, drug-food, and chemical-target interactions for a compound by CID.
  • pubchem_search_assays — Find bioassays by biological target using gene symbol, protein name, Gene ID, or UniProt accession.
  • pubchem_get_summary — Get summaries for PubChem entities: assays, genes, proteins, and taxonomy records.

Use cases

  • Look up chemical compound properties, IUPAC names, and molecular formulas for research or documentation.
  • Check GHS hazard classifications and safety data for chemicals before handling or purchasing.
  • Explore bioactivity profiles to find which assays tested a compound and what biological targets it affects.
  • Search for compounds by structure (substructure, superstructure, or similarity) to discover related molecules.
  • Find external cross-references (patents, PubMed articles, genes) linked to a compound for literature review.

io.github.cyanheads/pubchem-mcp-server MCP server FAQ

What is the PubChem MCP server?

It's an MCP server that exposes PubChem's chemical database through tools for searching compounds, retrieving properties, safety data, bioactivity, and cross-references. No API key required—PubChem's API is freely accessible.

Is it free to use?

Yes. PubChem is a free public database, and this server has no subscription or authentication requirement. A public hosted instance is available at https://pubchem.caseyjhand.com/mcp.

How do I install it in Claude Desktop or Cursor?

Add the server to your MCP client config using stdio (npx or bunx) or Streamable HTTP. Direct install buttons are available in the GitHub repository for Claude Desktop and Cursor.

Do I need an API key?

No. PubChem's API is freely accessible without authentication. The server is read-only and idempotent.

What search strategies does it support?

Identifier (name, SMILES, InChIKey), formula (Hill notation), substructure/superstructure containment, and 2D Tanimoto similarity (threshold 70–100).

Can I batch requests?

Yes. Most tools support batching: compound search and details accept multiple CIDs, safety data accepts 1–25 CIDs per call, and summary supports up to 10 identifiers per call.

README (reference)

Source of truth, from the repository.

<div align="center"> <h1>@cyanheads/pubchem-mcp-server</h1> <p><b>Search the PubChem chemical database for compounds, properties, safety data, bioactivity, cross-references, and entity summaries via MCP. STDIO or Streamable HTTP.</b> <div>10 Tools • 6 Resources</div> </p> </div> <div align="center">

Version License Docker MCP SDK npm TypeScript Bun

</div> <div align="center">

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

</div> <div align="center">

Public Hosted Server: https://pubchem.caseyjhand.com/mcp

</div>

Overview

Chemical compound and bioassay data from PubChem's PUG REST and PUG View APIs. Search compounds by identifier, formula, or structure; fetch physicochemical properties, safety data, bioactivity, interactions, cross-references, and 3D structures; find bioassays by biological target. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

ToolDescription
pubchem_search_compoundsSearch for compounds by name, SMILES, InChIKey, formula, substructure, superstructure, or 2D similarity.
pubchem_get_compound_detailsGet physicochemical properties, descriptions, synonyms, drug-likeness, and classification for compounds by CID.
pubchem_get_compound_imageFetch a 2D structure diagram (PNG) for a compound by CID.
pubchem_get_compound_3d_structureFetch a 3D conformer (atomic coordinates and bonds) for a compound by CID, as parsed JSON or raw SDF.
pubchem_get_compound_xrefsGet external database cross-references (PubMed, patents, genes, proteins, etc.).
pubchem_get_compound_safetyGet GHS hazard classification and safety data for one or more compounds by CID (batch).
pubchem_get_bioactivityGet a compound's bioactivity profile: assay results, targets, and activity values; filter by outcome or molecular target.
pubchem_get_compound_interactionsGet drug-drug, drug-food, and chemical-target interactions for a compound by CID.
pubchem_search_assaysFind bioassays by biological target (gene symbol, protein, Gene ID, UniProt accession).
pubchem_get_summaryGet summaries for PubChem entities: assays, genes, proteins, taxonomy.

Resources

Compound and assay records are also exposed as URI-templated resources, backed by the same client methods as the tools; many MCP clients are tool-only and never surface resources.

ResourceDescription
pubchem://compound/{cid}Core physicochemical properties (JSON).
pubchem://compound/{cid}/safetyGHS hazard classification (JSON).
pubchem://compound/{cid}/image2D structure diagram (PNG).
pubchem://compound/{cid}/xrefsExternal cross-references (JSON).
pubchem://compound/{cid}/bioactivityBioassay activity profile (JSON).
pubchem://assay/{aid}BioAssay summary (JSON).

Capability reference

pubchem_search_compounds <sub>tool</sub>

  • Five search strategies: identifier (name/SMILES/InChIKey, batched 1-25), formula (Hill notation, optional allowOtherElements), substructure/superstructure containment, or 2D Tanimoto similarity (threshold 70-100, default 90)
  • Each strategy needs its own fields — identifier: identifierType + identifiers; formula: formula; substructure/superstructure/similarity: query + queryType — and a missing or blank one is rejected before the upstream call
  • Caps at 200 CIDs per page (default 20); offset pages to a ceiling of 10,000 — identifier lookups resolve every match up front so paging is free, while formula/structure/similarity searches cost more upstream per deep page
  • Optional properties hydration avoids a follow-up pubchem_get_compound_details call
  • Identifier mode reports unresolvedIdentifiers for inputs that resolved to no CID — no PubChem match, or a SMILES PubChem cannot interpret — while the rest of the batch still resolves, plus notices when multiple inputs collide on one CID
  • A query PubChem cannot search on (malformed SMILES or formula, a * wildcard atom, a CID with no record) fails fast with a search_query_rejected hint naming what to fix
  • Reports an exact totalFound when the full match set was observed, or a totalFoundAtLeast floor when a bounded upstream search saturated

pubchem_get_compound_details <sub>tool</sub>

  • Up to 100 CIDs per call; 27 available properties, defaulting to a core set of 14 (formula, weight, IUPAC name, SMILES forms, InChIKey, XLogP, TPSA, H-bond/rotatable-bond counts, heavy atom count, charge, complexity)
  • Optional textual descriptions, paged via descriptionOffset/maxDescriptions (default 3, up to 20) — fetched only for the first 10 CIDs in the batch, remaining CIDs listed in skippedCids
  • Optional synonyms for every found CID, paged via synonymOffset/maxSynonyms (default 20, up to 100)
  • Optional drug-likeness assessment (Lipinski Rule of Five + Veber rules), computed from the returned properties at no extra latency
  • Optional pharmacological classification (FDA classes/mechanisms, MeSH classes, ATC codes) — same 10-CID fan-out cap as descriptions
  • Per-CID found: false distinguishes a nonexistent CID from a real compound PubChem simply has no data for

pubchem_get_compound_image <sub>tool</sub>

  • Single CID; size is "small" (100x100) or "large" (300x300, default)
  • Returns base64-encoded PNG plus width/height
  • Typed cid_not_found error when PubChem has no record for the CID

pubchem_get_compound_3d_structure <sub>tool</sub>

  • Single CID; format="json" (default) returns parsed atoms (element + x/y/z) and bonds, format="sdf" returns the raw V2000 SDF text
  • maxAtoms/maxBonds cap the JSON preview (default 200 each); atomCount/bondCount always report the full totals, with any capping disclosed via enrichment
  • includeRawSdf bypasses the default 500-line cap on the raw SDF text
  • Optional includeAlternateConformerIds lists conformer IDs beyond the default
  • Typed no_3d_structure error when PubChem has no computed 3D coordinates (large molecules, mixtures, some salts)

pubchem_get_compound_xrefs <sub>tool</sub>

  • Single CID; one or more xrefTypes — string IDs (RegistryID, RN for CAS numbers, PatentID) and numeric IDs (PubMedID, GeneID, ProteinGI, TaxonomyID)
  • Paged per type: maxPerType up to 500 (default 50), with the same offset applied across every requested type
  • Each type reports its own totalAvailable and truncated flag
  • Empty-result notice distinguishes "this compound has none of the requested types" from a possibly-mistyped CID

pubchem_get_compound_safety <sub>tool</sub>

  • Batch of 1-25 CIDs
  • Returns GHS signal word, pictograms, hazard statements (H-codes), and precautionary statements (P-codes), with source attribution
  • Per-CID status: ok, no_ghs_data (compound exists, no deposited classification), or cid_not_found (no PubChem record at all) — kept distinct so a bad CID never reads as "no hazards on file"
  • Precautionary statements carry a decoded flag — false for codes needing label-specific fill text or outside the decoder table; the code itself is still authoritative

pubchem_get_bioactivity <sub>tool</sub>

  • Single CID; filter by outcomeFilter (active/inactive/all, default all) and/or targetGeneId/targetAccession
  • Caps at 100 results per page (default 20); offset reaches the rest
  • Reports totalAssays/activeCount/inactiveCount for the whole compound, plus filteredCount/returnedCount for the current page
  • Notices distinguish "no bioactivity data at all" from "the filter excluded everything" from "offset past the end"

pubchem_get_compound_interactions <sub>tool</sub>

  • Single CID; one or more kinds — drug-drug (DrugBank), drug-food, target (binding/activity from BindingDB, ChEMBL, and others); default ["drug-drug"]
  • maxEntries per kind per page (1-50, default 10); offset counts source records rather than returned entries, capped at 2,147,483,646
  • Each kind pages independently — paging[] reports per-kind totalRecords/nextOffset/truncated; the top-level nextOffset is populated only when exactly one requested kind still has records left
  • A kind that fails to retrieve is named in failedKinds without failing the kinds that succeeded

pubchem_search_assays <sub>tool</sub>

  • Search by targetType: genesymbol/proteinname (text), geneid (NCBI Gene ID), proteinaccession (UniProt)
  • Caps at 200 AIDs per page (default 50); offset pages to the total found
  • Rejects a blank targetQuery and a non-numeric geneid query before the upstream call
  • Reports totalFound across all pages and distinguishes "no match" from "offset past the end"

pubchem_get_summary <sub>tool</sub>

  • entityType: assay (AID), gene (NCBI Gene ID), protein (UniProt accession), or taxonomy (Tax ID); up to 10 identifiers per call
  • Per-identifier found flag; populated fields depend on entityType (taxonomy includes an ordered lineage, gene includes symbol/taxonomy)
  • Notice reports how many identifiers were not found and which ID type entityType expects

pubchem://compound/{cid} <sub>resource</sub>

  • Core physicochemical properties (the same default 14-property set as pubchem_get_compound_details), as application/json
  • Throws a typed not-found when the CID doesn't exist in PubChem
  • Use pubchem_get_compound_details to select specific properties or add descriptions, synonyms, drug-likeness, and classification

pubchem://compound/{cid}/safety <sub>resource</sub>

  • GHS hazard classification as application/json
  • status (ok/no_ghs_data/cid_not_found) is the only signal distinguishing a bad CID from a compound with no deposited classification — a resource read has no notice surface

pubchem://compound/{cid}/image <sub>resource</sub>

  • 2D structure diagram, 300x300 PNG, returned as a base64 blob
  • Use pubchem_get_compound_image for the 100x100 size option

pubchem://compound/{cid}/xrefs <sub>resource</sub>

  • Focused default set — RN (CAS), RegistryID, PubMedID — up to 25 IDs per type, as application/json
  • Use pubchem_get_compound_xrefs for the full set of xref types, a higher per-type cap, and offset paging

pubchem://compound/{cid}/bioactivity <sub>resource</sub>

  • Up to 25 assays as application/json, plus totalAssays/activeCount for the whole compound
  • Use pubchem_get_bioactivity to filter by outcome or target, raise the cap, or page with offset

pubchem://assay/{aid} <sub>resource</sub>

  • BioAssay summary as application/json — name, description, source, protocol, substance counts
  • Throws a typed not-found when the AID doesn't exist

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

PubChem-specific:

  • Covers both PUG REST (search, properties, cross-references, safety, bioactivity, interactions) and PUG View (textual descriptions, pharmacological classification) endpoints
  • Rate-limited client (5 req/s) with automatic request queuing, and retry with exponential backoff on 5xx errors and network failures
  • A cancelled tool call or resource read stops its PubChem work — queued requests, in-flight fetches, retry backoffs, and async-search polling — and fails with RequestCancelled
  • Hand-rolled V2000 SDF parser for 3D conformer atoms and bonds; drug-likeness (Lipinski/Veber) computed from already-fetched properties, adding no extra latency
  • All tools are read-only and idempotent — no API keys required, PubChem's API is freely accessible

Agent-friendly output:

  • Discriminated output contracts — per-CID status (ok / no_ghs_data / cid_not_found) and found flags let callers branch on data instead of matching an error string
  • Graceful partial failure — batch tools return per-item results alongside unresolvedIdentifiers, skippedCids, and failedKinds rather than failing the whole call
  • Response shaping — truncation disclosure (truncated, shown/cap, nextOffset) on every capped list, plus a totalFoundAtLeast floor in place of a count when an upstream search saturates
  • Typed error reasons — validation and not-found failures declare a reason (e.g. cid_not_found, missing_identifier_args, invalid_cid_query) with actionable recovery text, not generic messages

Getting started

Public Hosted Instance

A public instance is available at https://pubchem.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "pubchem-mcp-server": {
      "type": "streamable-http",
      "url": "https://pubchem.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file.

{
  "mcpServers": {
    "pubchem-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/pubchem-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "pubchem-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/pubchem-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "pubchem-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/pubchem-mcp-server:latest"]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • No API keys required — PubChem's API is freely accessible.

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/pubchem-mcp-server.git
  1. Navigate into the directory:
cd pubchem-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment (optional):
cp .env.example .env
# edit .env to override transport, session mode, storage, or logging defaults

Configuration

VariableDescriptionDefault
MCP_TRANSPORT_TYPETransport: stdio or http.stdio
MCP_HTTP_PORTPort for HTTP server.3010
MCP_HTTP_HOSTHost for HTTP server.127.0.0.1
MCP_SESSION_MODEstateless, stateful, or auto. PubChem needs no multi-round-trip input, so the server declares stateless; the example and Docker set it to match.stateless
MCP_AUTH_MODEAuth mode: none, jwt, or oauth.none
MCP_LOG_LEVELLog level (RFC 5424).info
STORAGE_PROVIDER_TYPEStorage backend.in-memory
OTEL_ENABLEDEnable OpenTelemetry.false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
    
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec
    

Docker

docker build -t pubchem-mcp-server .
docker run --rm -p 3010:3010 pubchem-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/pubchem-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point — registers tools/resources and inits the PubChem client.
src/mcp-server/tools/definitions/Tool definitions (*.tool.ts).
src/mcp-server/resources/definitions/Resource definitions (*.resource.ts).
src/services/pubchem/PubChem API client — rate limiting, retry, and response/SDF parsing.
scripts/Build, clean, devcheck, and tree generation scripts.
tests/Unit and integration tests.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for request-scoped logging
  • Wrap external API calls: validate the raw PubChem response → normalize to a domain type → return the output schema; never fabricate missing fields
  • Register new tools and resources in the index.ts barrel files

Contributing

Issues are welcome. Run checks before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP servers

Search PubMed, Europe PMC, and fetch full-text articles with citations and MeSH terms via MCP.

116
TypeScript
Apache-2.0
View repository →

Countries, timezones, elements, constants, HTTP status codes, unit conversion, and MIME type lookup.

1
TypeScript
Apache-2.0
View repository →

Search ReliefWeb humanitarian reports, disasters, jobs, training, and country profiles via MCP.

1
TypeScript
Apache-2.0
View repository →

Search Austrian federal and state law, court decisions, and the authentic Bundesgesetzblatt (RIS).

2
TypeScript
Apache-2.0
View repository →

Screen names against OFAC, EU, UK, UN sanctions lists; resolve entities via GLEIF. Screening aid.

1
TypeScript
Apache-2.0
View repository →

Query SEC EDGAR filings, XBRL financials, and company data through MCP. STDIO & Streamable HTTP.

6
TypeScript
Apache-2.0
View repository →