HKEx regulatory filings scraper — 25+ years of filings into nine databases, with full-text extraction, graph linking, and a hosted MCP server for AI agents
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add hkex-filing-scraper -- python -m hkex-filing-scraper{
"mcpServers": {
"hkex-filing-scraper": {
"command": "python",
"args": ["-m", "hkex-filing-scraper"]
}
}
}Resumen de MCP Servers
# HKEx Filing Scraper
<!-- mcp-name: io.github.simonmak-ascent/hkex-filing-scraper -->
> **An open-source scraper for 25+ years of HKEx regulatory filings** — into any of nine databases, with full-text extraction, graph linking, and a read-only MCP server for AI agents.

[](https://github.com/simonmak-ascent/hkex-filing-scraper/actions/workflows/ci.yml)
[](https://registry.modelcontextprotocol.io/v0/servers?search=io.github.simonmak-ascent/hkex-filings)
[](https://glama.ai/mcp/servers/simonmak-ascent/hkex-filing-scraper)
[](https://github.com/simonmak-ascent/hkex-filing-scraper/releases)
[](https://pypi.org/project/hkex-filing-scraper/)
[](https://www.python.org/downloads/)
[](https://hkex-listco-updates.ascent-partners.com/)
[](https://hkex-listco-updates.ascent-partners.com/ai-agents/)
[](https://github.com/astral-sh/ruff)
[](https://opensource.org/licenses/MIT)
## Overview
**The US has EDGAR full-text search. Japan has EDINET. Hong Kong has a search form that returns
one page at a time.** There is no bulk, machine-readable, full-text corpus of HKEx filings. This
builds one.
An open-source Python tool that scrapes 25+ years of Hong Kong Stock Exchange (HKEx)
regulatory filings and ingests them into **any combination of nine databases** — with
full-text and table extraction, chunk-level coverage, optional graph linking, and a
**read-only MCP server** so AI agents can query the corpus or the live site.
<!-- mcp-name: io.github.simonmak-ascent/hkex-filings -->
It speaks the undocumented HKEx JSON API directly, which is faster and more resilient than
driving a browser.
## Vendors & Integrations
**Databases** — nine first-class destinations, in documented popularity order (see the
[support matrix](docs/sinks/README.md)):
- [PostgreSQL](docs/sinks/postgresql.md) — production-grade open-source relational
- [MySQL](docs/sinks/mysql.md) / [MariaDB](docs/sinks/mysql.md) — GPL relational servers, one driver
- [SQLite](docs/sinks/sqlite.md) — zero-server file database, no install needed
- [MongoDB](docs/sinks/mongodb.md) — document database
- [Neo4j](docs/sinks/neo4j.md) — property-graph database
- [ClickHouse](docs/sinks/clickhouse.md) — columnar analytics engine
- [DuckDB](docs/sinks/duckdb.md) — in-process analytical engine
- [SurrealDB](docs/sinks/surrealdb.md) — multi-model graph + document database
**AI clients** — any MCP-capable agent; ready-made configuration for
[Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity](docs/ai-agents.md).
**Available on** — [PyPI](https://pypi.org/project/hkex-filing-scraper/) ·
[Glama](https://glama.ai/mcp/servers/simonmak-ascent/hkex-filing-scraper) ·
[MCP Registry](https://registry.modelcontextprotocol.io/) ·
[hosted gateway](https://hkex-listco-updates.ascent-partners.com/api/mcp).
## Two Ways to Use It
| | **Hosted MCP gateway** | **Local pipeline** |
| --- | --- | --- |
| What | A public endpoint you point an AI agent at | The `hkex-scraper` CLI |
| Setup | None — paste a URL | `pip install` + one environment variable |
| Data | Live from HKEx, nothing stored | Stored in your database(s) |
| Docs | [Live MCP gateway](docs/live-mcp.md) · [AI agent support](docs/ai-agents.md) | [Getting started](docs/getting-started.md) |

## Use the Hosted MCP Gateway
POST, Streamable HTTP, **no API key**:
```text
https://hkex-listco-updates.ascent-partners.com/api/mcp
```
Four read-only tools: `get_server_info`, `search_filings` (a window of at most 31 days, with
optional stock-code, title, document-type, category, and stock-name filters),
`list_filing_facets` (browse what a window contains), and `get_filing` (downloads one
document and extracts its text and tables).

Point a client at it — for example opencode:
```json
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"hkex-live": {
"type": "remote",
"url": "https://hkex-listco-updates.ascent-partners.com/api/mcp"
}
}
}
```
Then ask:
```text
Use hkex-live to list the filings published between 2026-09-01 and 2026-09-18,
then summarise the interim report.
```
Ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode,
Manus, and Perplexity is in [AI agent support](docs/ai-agents.md) — and for a stored corpus,
the [stdio MCP server](docs/mcp.md) exposes a wider tool catalog and is published on
[Glama](https://glama.ai/mcp/servers/simonmak-ascent/hkex-filing-scraper). The gateway is listed
in the [official MCP Registry](https://registry.modelcontextprotocol.io/) as
`io.github.simonmak-ascent/hkex-filings`.
> **Featured on Glama** — the read-only [stdio MCP server](docs/mcp.md) is also published on
> [Glama](https://glama.ai/mcp/servers/simonmak-ascent/hkex-filing-scraper), where Glama scans the
> built server and scores tool-definition quality (currently 4.7/5).
## Quick Start (Local)
```bash
pip install hkex-filing-scraper # core; SQLite needs no server
pip install "hkex-filing-scraper[all]" # Excel + dotenv + every driver + the MCP server
cp .env.example .env # then set DATABASE_TARGET (below)
hkex-scraper --metadata-only --limit 100
```
Optional extras: `excel`, `postgres`, `mysql`, `duckdb`, `mongodb`, `clickhouse`, `neo4j`,
`mcp`, `pdf`, `all`, `dev`.
`DATABASE_TARGET` is an ordered, comma-separated list of sink ids; the order decides which
sink serves reads. To start with no server:
```ini
DATABASE_TARGET=sqlite
SQLITE_PATH=hkex.db
```
`hkex-scraper` runs the full pipeline (metadata + documents + graph); `hkex-scraper
--full-history` covers everything since April 1999. The schema is created automatically.
Full install options and per-sink settings are in [Getting started](docs/getting-started.md).
## Database Support
Every sink is a first-class destination; rows are in documented popularity order. The full
matrix — licenses, capability differences, per-engine notes — is in
[Database sinks](docs/sinks/README.md).
| Sink | Model | License | Extra | Idempotent upsert |
| ---- | ----- | ------- | ----- | ----------------- |
| `postgres` | relational | PostgreSQL License | `postgres` | `ON CONFLICT DO UPDATE` |
| `mysql` / `mariadb` | relational | GPLv2 | `mysql` | `ON DUPLICATE KEY UPDATE` |
| `sqlite` | relational | Public domain | — | `ON CONFLICT DO UPDATE` |
| `mongodb` | document | SSPL¹ | `mongodb` | `update_one(upsert=True)` |
| `neo4j` | graph | GPLv3 (Community) | `neo4j` | `MERGE` |
| `clickhouse` | columnar | Apache-2.0 | `clickhouse` | `ReplacingMergeTree` + read-merge |
| `duckdb` | relational | MIT | `duckdb` | `ON CONFLICT DO UPDATE` |
| `surrealdb` | graph + document | BSL 1.1¹ | — | `UPSERT` / `RELATE` |
¹ Source-available, not OSI-approved — labeled exceptions per
[ADR 0003](docs/adr/0003-sink-support-policy.md).
Valid sink ids, in documented order: `postgres`, `mysql`, `sqlite`, `mongodb`, `mariadb`, `neo4j`, `clickhouse`, `duckdb`, `surrealdb`. Set one variable and the same run feeds every sink:
```ini
# Order sets read precedence.
DATABASE_TARGET=postgres,sqlite
POSTGRES_DSN=postgresql://user:password@localhost:5432/hkex
SQLITE_PATH=hkex.db
```
## How This Compares
Four ways to get HKEx filings, and what each one costs you.
| | This project | HKEXnews web search | Browser automation you write | Licensed HKEx feed |
| --- | --- | --- | --- | --- |
| Bulk export | Yes | No — page-at-a-time | Yes | Yes |
| History to April 1999 | Yes | Yes, manually | Depends on your code | Yes |
| Full text of documents | Extracted from PDF/HTML/Excel | No — you open each file | You build the extractor | Varies by contract |
| Structured tables | Extracted to Markdown | No | You build it | Varies |
| Coverage verification | Per-chunk, auditable | Not applicable | You build it | Vendor SLA |
| Lands in your engine | 9 engines, any combination | No | Whatever you wire up | Usually one format |
| Speed | JSON API, no browser | Manual | Slower — renders pages | Fast |
| Cost | Free, MIT | Free | Your time | Subscription |
| Commercial redistribution | See [docs/legal.md](docs/legal.md) | Restricted | Restricted | Licensed |
If you need licensed, redistributable, SLA-backed data, buy the feed. If you need a complete
local corpus for research, compliance, or RAG, this replaces the pipeline you would otherwise
write yourself.
## How It Works
```mermaid
flowchart LR
A[HKEx JSON API] --> B[Phase 1: metadata]
B --> C[Canonical record]
C --> D{DATABASE_TARGET}
D --> E[(PostgreSQL)]
D --> F[(MySQL / MariaDB)]
D --> G[(SQLite)]
D --> HLo que la gente pregunta sobre hkex-filing-scraper
¿Qué es simonmak-ascent/hkex-filing-scraper?
+
simonmak-ascent/hkex-filing-scraper es mcp servers para el ecosistema de Claude AI. HKEx regulatory filings scraper — 25+ years of filings into nine databases, with full-text extraction, graph linking, and a hosted MCP server for AI agents Tiene 16 estrellas en GitHub y su última actualización registrada es del 2026-10-03.
¿Cómo se instala hkex-filing-scraper?
+
Puedes instalar hkex-filing-scraper clonando el repositorio (https://github.com/simonmak-ascent/hkex-filing-scraper) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar simonmak-ascent/hkex-filing-scraper?
+
Nuestro agente de seguridad ha analizado simonmak-ascent/hkex-filing-scraper y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene simonmak-ascent/hkex-filing-scraper?
+
simonmak-ascent/hkex-filing-scraper es mantenido por simonmak-ascent. La última actividad registrada en GitHub es del 2026-10-03, con 0 issues abiertos.
¿Hay alternativas a hkex-filing-scraper?
+
Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.
Despliega hkex-filing-scraper en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/simonmak-ascent-hkex-filing-scraper)<a href="https://claudewave.com/repo/simonmak-ascent-hkex-filing-scraper"><img src="https://claudewave.com/api/badge/simonmak-ascent-hkex-filing-scraper" alt="Featured on ClaudeWave: simonmak-ascent/hkex-filing-scraper" width="320" height="64" /></a>Más MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.