HKEx Filings MCP Server
io.github.simonplmak-cloud/hkex-filings
Access 25+ years of Hong Kong Stock Exchange regulatory filings via AI agents with live MCP gateway or local database pipeline.
What is the HKEx Filings MCP server?
The HKEx Filings MCP server scrapes Hong Kong Stock Exchange regulatory filings and exposes them to AI agents through a read-only MCP interface. It supports both a hosted live gateway (no setup required) and a local pipeline that ingests filings into nine database backends (PostgreSQL, MySQL, SQLite, MongoDB, Neo4j, ClickHouse, DuckDB, SurrealDB). The server extracts full text and tables from PDF/HTML/Excel documents and enables AI agents to search, browse, and retrieve filings with optional graph linking.
HKEx Filings lets AI agents query Hong Kong Stock Exchange regulatory filings spanning April 1999 to present. Use the hosted MCP gateway for instant live access with no API key, or run the local pipeline to store filings in your choice of nine databases. It speaks the undocumented HKEx JSON API directly for speed and resilience, extracts structured text and tables from documents, and integrates seamlessly with Claude, Cursor, ChatGPT, and other MCP-capable agents.
How to install HKEx Filings
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
get_server_info— Retrieve server metadata and capabilities.search_filings— Search filings within a 31-day window with optional filters by stock code, title, document type, category, and stock name.list_filing_facets— Browse available filing facets and metadata within a specified date window.get_filing— Download a single filing document and extract its text and tables.
Use cases
- Query live HKEx filings from an AI agent without setup using the hosted MCP gateway.
- Store 25+ years of HKEx regulatory filings in PostgreSQL, MySQL, SQLite, or other databases for offline analysis.
- Extract structured text and tables from filing PDFs/HTML/Excel documents for document processing pipelines.
- Search filings by stock code, company name, document type, and date range to support investment research.
- Build graph-linked filing databases with Neo4j or SurrealDB to track company relationships and cross-references.
HKEx Filings MCP server FAQ
It's a read-only MCP server that scrapes Hong Kong Stock Exchange regulatory filings (25+ years of history) and exposes them to AI agents. You can use a hosted live gateway with zero setup, or run a local pipeline to store filings in nine different databases.
Yes. The hosted MCP gateway is free and requires no API key. The local pipeline is open-source (MIT license) and free to run; optional database servers (PostgreSQL, MongoDB, etc.) are your choice.
Point your MCP client at the hosted gateway URL (https://hkex-listco-updates.ascent-partners.com/api/mcp) with type 'remote'. Ready-made configuration for Claude, Cursor, ChatGPT, VS Code/Copilot, and other agents is in the AI agent support docs.
No. The hosted MCP gateway requires no API key or authentication. The local pipeline needs only a DATABASE_TARGET environment variable and optional per-database credentials.
Nine: PostgreSQL, MySQL/MariaDB, SQLite, MongoDB, Neo4j, ClickHouse, DuckDB, and SurrealDB. You can configure multiple sinks simultaneously; reads are served from the first in your ordered list.
PDF, HTML, and Excel. Text and tables are extracted to Markdown and stored in your database for AI agents to query.
README (reference)
Source of truth, from the repository.
HKEx Filing Scraper

An open-source Python tool that scrapes 25+ years of Hong Kong Stock Exchange (HKEx) regulatory filings and ingests them into any combination of nine databases — with full-text and table extraction, chunk-level coverage, optional graph linking, and a read-only MCP server so AI agents can query the corpus or the live site.
<!-- mcp-name: io.github.simonplmak-cloud/hkex-filings -->It speaks the undocumented HKEx JSON API directly, which is faster and more resilient than driving a browser.
Vendors & integrations
Databases — nine first-class destinations, in documented popularity order (see the support matrix):
- PostgreSQL — production-grade open-source relational
- MySQL / MariaDB — GPL relational servers, one driver
- SQLite — zero-server file database, no install needed
- MongoDB — document database
- Neo4j — property-graph database
- ClickHouse — columnar analytics engine
- DuckDB — in-process analytical engine
- SurrealDB — multi-model graph + document database
AI clients — any MCP-capable agent; ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity.
Available on — PyPI · Glama · MCP Registry · hosted gateway.
Two ways to use it
| Hosted MCP gateway | Local pipeline | |
|---|---|---|
| What | A public endpoint you point an AI agent at | The hkex-scraper CLI |
| Setup | None — paste a URL | pip install + one environment variable |
| Data | Live from HKEx, nothing stored | Stored in your database(s) |
| Docs | Live MCP gateway · AI agent support | Getting started |
Use the hosted MCP gateway
POST, Streamable HTTP, no API key:
https://hkex-listco-updates.ascent-partners.com/api/mcp
Four read-only tools: get_server_info, search_filings (a window of at most 31 days, with
optional stock-code, title, document-type, category, and stock-name filters),
list_filing_facets (browse what a window contains), and get_filing (downloads one
document and extracts its text and tables).

Point a client at it — for example opencode:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"hkex-live": {
"type": "remote",
"url": "https://hkex-listco-updates.ascent-partners.com/api/mcp"
}
}
}
Then ask:
Use hkex-live to list the filings published between 2026-09-01 and 2026-09-18,
then summarise the interim report.
Ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode,
Manus, and Perplexity is in AI agent support — and for a stored corpus,
the stdio MCP server exposes a wider tool catalog and is published on
Glama. The gateway is listed
in the official MCP Registry as
io.github.simonplmak-cloud/hkex-filings.
Featured on Glama — the read-only stdio MCP server is also published on Glama, where Glama scans the built server and scores tool-definition quality (currently 4.7/5).
Quick start (local)
pip install hkex-filing-scraper # core; SQLite needs no server
pip install "hkex-filing-scraper[all]" # Excel + dotenv + every driver + the MCP server
cp .env.example .env # then set DATABASE_TARGET (below)
hkex-scraper --metadata-only --limit 100
Optional extras: excel, postgres, mysql, duckdb, mongodb, clickhouse, neo4j,
mcp, pdf, all, dev.
DATABASE_TARGET is an ordered, comma-separated list of sink ids; the order decides which
sink serves reads. To start with no server:
DATABASE_TARGET=sqlite
SQLITE_PATH=hkex.db
hkex-scraper runs the full pipeline (metadata + documents + graph); hkex-scraper --full-history covers everything since April 1999. The schema is created automatically.
Full install options and per-sink settings are in Getting started.
Database support
Every sink is a first-class destination; rows are in documented popularity order. The full matrix — licenses, capability differences, per-engine notes — is in Database sinks.
| Sink | Model | License | Extra | Idempotent upsert |
|---|---|---|---|---|
postgres | relational | PostgreSQL License | postgres | ON CONFLICT DO UPDATE |
mysql / mariadb | relational | GPLv2 | mysql | ON DUPLICATE KEY UPDATE |
sqlite | relational | Public domain | — | ON CONFLICT DO UPDATE |
mongodb | document | SSPL¹ | mongodb | update_one(upsert=True) |
neo4j | graph | GPLv3 (Community) | neo4j | MERGE |
clickhouse | columnar | Apache-2.0 | clickhouse | ReplacingMergeTree + read-merge |
duckdb | relational | MIT | duckdb | ON CONFLICT DO UPDATE |
surrealdb | graph + document | BSL 1.1¹ | — | UPSERT / RELATE |
¹ Source-available, not OSI-approved — labeled exceptions per ADR 0003.
Valid sink ids, in documented order: postgres, mysql, sqlite, mongodb, mariadb, neo4j, clickhouse, duckdb, surrealdb. Set one variable and the same run feeds every sink:
# Order sets read precedence.
DATABASE_TARGET=postgres,sqlite
POSTGRES_DSN=postgresql://user:password@localhost:5432/hkex
SQLITE_PATH=hkex.db
How it works
flowchart LR
A[HKEx JSON API] --> B[Phase 1: metadata]
B --> C[Canonical record]
C --> D{DATABASE_TARGET}
D --> E[(PostgreSQL)]
D --> F[(MySQL / MariaDB)]
D --> G[(SQLite)]
D --> H[(MongoDB)]
D --> I[(Neo4j)]
D --> J[(ClickHouse)]
D --> K[(DuckDB)]
D --> L[(SurrealDB)]
B --> M[Graph linking]
M --> D
B --> N[Phase 2: download and extract]
N --> C
- Phase 1 scrapes filing metadata through a JSF session, splitting the range into monthly
chunks and deduplicating on a 16-character MD5
filingId. - Phase 2 downloads each filing's PDF/HTML/Excel document, extracts text and tables to Markdown, and writes the payload.
- Graph linking (optional) writes
has_filingandreferences_filingedges whenCOMPANY_TABLEis set. - Failure isolation — a failure on one sink is logged and counted but never blocks another; the run exits non-zero if any configured sink failed.
Deeper detail: Architecture · ADR 0002.
Features
- Fast API scraping — direct HKEx JSON API; no browser or Selenium.
- Full history — every filing from April 1999 to today, with chunk-level coverage checks.
- Document processing — PDF/HTML/Excel text and structured tables, extracted to Markdown.
- Multi-sink — any ordered combination of nine databases, each with native idempotent upserts.
- AI-ready — a hosted live MCP gateway plus a local stdio MCP server.
- Resumable and observable — batching, parallel downloads, stalled-job detection, per-sink
counters, and
--coverage-report/--parity-report/--verify. - Optional dependencies — the core is
requests+beautifulsoup4; drivers and document extraction are extras with graceful fallbacks.
Documentation
- Getting started · Configuration · CLI
- Database sinks (matrix) — PostgreSQL, MySQL/MariaDB, SQLite, MongoDB, Neo4j, ClickHouse, DuckDB, SurrealDB
- Live MCP gateway · AI agent support · MCP server
- Architecture · Troubleshooting · Testing
- Roadmap · De-risking register · Upgrading
- What's new · Releasing · Legal & Terms of Use · Changelog
- Docs site: https://hkex-listco-updates.ascent-partners.com/ · Try it locally (
examples/)
Development
pip install -e ".[dev,all]"
ruff check # lint (py310, line-length 100)
ruff format --check # formatting
pytest # unit tests (no DB or network required)
Tests are pure unit tests; SQLite and DuckDB contract tests run in-process, and integration tests that need a server are skipped unless that sink is configured. See Testing.
Contributing
See CONTRIBUTING.md; report security issues per SECURITY.md. Ideas and questions are welcome in Discussions.
If this saves you time, a star helps others find it.
License
MIT — see LICENSE. That covers this project's code only; optional dependencies
carry their own licenses, notably the pdf extra (PyMuPDF / pymupdf4llm), which is
AGPL-3.0 and deliberately excluded from .[all]. See
docs/legal.md.
Data & Terms of Use: this is a research tool for the undocumented HKEx JSON API, and it is not affiliated with or endorsed by HKEx. Commercial redistribution of HKEx data may require a licensed HKEx feed; see docs/legal.md.
Related MCP servers
Intangible asset valuation: 14 tools, 124+ formulas for IP, technology, goodwill, PPA, impairment.
Startup valuation for AI agents: 14 tools, 80+ pre-revenue formulas.
View repository →Spec-driven development MCP server: 8 phases, bi-directional traceability, 7 quality gates.

SeekLink
Local semantic search for Markdown vaults with hybrid keyword + vector retrieval, no cloud required.

Temporal memory for AI with natural decay and reinforcement—remember preferences, decisions, and facts across conversations.

io.github.simplifier-ag/simplifier-mcp
MCP Server for Simplifier - manage Connectors, BusinessObjects, Datatypes and logins
