PluginBench
MCP Server
Active
Apache-2.0

AISIX AI Gateway MCP Server

io.github.api7/aisix

Open-source Rust-native AI gateway that routes, governs, and observes LLM traffic through one OpenAI-compatible API.

What is the AISIX AI Gateway MCP server?

AISIX AI Gateway is a Rust-native gateway that puts a single OpenAI-compatible API in front of multiple LLM providers (OpenAI, Anthropic, Bedrock, Vertex, Azure OpenAI, DeepSeek). It runs as a static binary in your infrastructure and provides routing, rate limiting, guardrails, caching, and observability for LLM and AI-agent traffic.

AISIX is a production-grade gateway for managing LLM traffic across multiple providers. It unifies access to OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, and DeepSeek behind one OpenAI-compatible endpoint, with built-in routing, failover, rate limiting, content guardrails, response caching, and observability. Run it as a single static binary in your infrastructure with declarative YAML configuration or etcd-backed clustering.

How to install AISIX AI Gateway

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "aisix": {
      "command": "docker",
      "args": [
        "run",
        "-i",
        "--rm",
        "ghcr.io/api7/aisix:1.3.0"
      ]
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • OpenAI-compatible proxy — Chat completions, embeddings, image generation, audio, video, files, batches, fine-tuning, and realtime endpoints with native SSE streaming and tool calling support.
  • Routing & failover — Virtual/routing models with round-robin, consistent hashing, failover, least-cost, least-latency, and least-busy strategies; per-target priority tiers and retry budgets.
  • Ensemble models — Fan one request to multiple models concurrently and synthesize a single answer with a judge model.
  • Semantic routing — Dispatch requests by meaning: embed prompts, score against example utterances, and route to the best match.
  • Rate limiting & concurrency — RPS/RPM/RPH/RPD, TPM/TPD, and concurrency caps combined across caller keys, models, and policy scopes; per-process or Redis-backed.
  • Guardrails — Content-policy enforcement on input and output via keyword/regex, PII detection, Presidio, Lakera, OpenAI Moderation, AWS Bedrock Guardrails, Azure AI Content Safety, and Alibaba Cloud services.
  • Caching — Exact-match response cache with per-policy TTL and model/key scope matchers; memory and Redis backends; automatic prompt caching for Anthropic models.
  • MCP gateway — Front upstream MCP servers at /mcp with gateway-held credentials, per-server tool namespacing, and per-caller access control.
  • A2A agent gateway — Front Agent-to-Agent agents at /a2a/:agent with card serving and URL rewriting over JSON-RPC 2.0.
  • Inbound authentication — Caller API keys with SHA-256 hashing, model allowlists, and expiry; OIDC/JWT bearer tokens validated against Entra ID, Okta, Google Workspace, or any OIDC issuer.
  • Observability — Prometheus metrics, structured access logs, usage events, OTLP/GenAI span export, Datadog and Aliyun SLS exporters, and object-storage telemetry.
  • Declarative configuration — Single resources.yaml file with hot reloading via SIGHUP; supports provider keys, models, caller keys, guardrails, MCP servers, A2A agents, cache policies, and exporters.

Use cases

  • Route LLM requests from multiple applications to different providers (OpenAI, Anthropic, Bedrock) through one unified API endpoint.
  • Implement rate limiting, token budgets, and cost controls across all LLM traffic in your organization.
  • Add content guardrails and PII detection to all model requests without modifying application code.
  • Cache LLM responses to reduce costs and latency for repeated queries.
  • Monitor and observe all LLM traffic with Prometheus metrics, structured logs, and OTLP trace export.

AISIX AI Gateway MCP server FAQ

What is AISIX AI Gateway?

AISIX is an open-source, Rust-native gateway that provides a single OpenAI-compatible API in front of multiple LLM providers. It handles routing, rate limiting, guardrails, caching, and observability for all your LLM traffic.

Is AISIX free?

Yes, the open-source gateway is free and Apache-2.0 licensed. You run it in your own infrastructure. AISIX Cloud (a commercial control plane) is available separately for centralized management and team governance.

How do I install AISIX in Cursor or Claude?

AISIX is a gateway server, not an MCP server for Cursor/Claude. However, it does include an MCP gateway feature that can front upstream MCP servers. Install the gateway as a Docker container or static binary, configure it with resources.yaml, and point your applications to its OpenAI-compatible endpoint.

What authentication does AISIX support?

AISIX supports caller API keys (SHA-256 hashed with model allowlists and expiry) and OIDC/JWT bearer tokens validated against Entra ID, Okta, Google Workspace, or any OIDC issuer.

Which LLM providers does AISIX support?

AISIX supports OpenAI, Anthropic Claude, AWS Bedrock, Google Vertex AI (Gemini), Azure OpenAI, DeepSeek, and any OpenAI-compatible endpoint (Groq, Mistral, Together, Ollama, vLLM, etc.).

Does AISIX require a database or control plane?

The open-source gateway requires neither. It reads configuration from a declarative resources.yaml file with hot reloading, or from etcd for multi-replica clusters. AISIX Cloud adds an optional commercial control plane.

README (reference)

Source of truth, from the repository.

<div align="center">

AISIX AI Gateway

The open-source, Rust-native AI gateway for LLMs and AI agents

One OpenAI-compatible API in front of every model. Route, govern, secure, cache, and observe all your LLM and AI-agent traffic from a single control point — shipped as one static binary with low per-request overhead. Run it in your infrastructure for free, forever.

Built by the original creators of Apache APISIX.

License: Apache 2.0 Built with Rust Docs Discord Website

Start free · Documentation · Quickstart · AISIX Cloud · Roadmap

<br> <img src="assets/aisix-architecture.svg" alt="AISIX AI Gateway architecture — one OpenAI- or Anthropic-compatible API in front of OpenAI, Anthropic, Gemini/Vertex, Bedrock, Azure OpenAI, and DeepSeek, with API key auth, rate and token limits, guardrails, caching, routing and failover, and observability in between" width="100%"> </div>

AISIX AI Gateway is a Rust-native gateway that puts a single, OpenAI-compatible API in front of every LLM provider — OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, DeepSeek, and any OpenAI-compatible endpoint. It gives platform teams one place to route, govern, secure, and observe LLM traffic, with first-class SSE streaming and low gateway overhead.

It runs as a single static binary — low cold-start, lock-free config reads, and hot configuration reloads with no restarts: declare resources in one resources.yaml and reload on SIGHUP, or point the gateway at etcd for a multi-replica cluster. Run the open-source gateway in your infrastructure, or connect it to AISIX Cloud for centralized management with team governance, budgets, audit, and a dashboard.

AISIX AI Gateway (this repo) is the open-source product. It runs without a control plane using declarative configuration or etcd. When connected to AISIX Cloud, the same gateway serves as the data plane. AISIX Cloud adds a commercial control plane, either hosted by API7 (Hybrid Cloud) or hosted by you in your infrastructure (On-Premises). In both options, the gateway runs in your environment and calls providers directly; live AI traffic does not pass through the control plane or API7. The proxy API is identical throughout. Talk to us about AISIX Cloud →

⚡ Quickstart

One container. No control plane, no database, no configuration store — the gateway reads every dynamic resource from one declarative resources.yaml.

# config.yaml
resources_file: /etc/aisix/resources.yaml
proxy:
  addr: "0.0.0.0:3000"
admin:
  enabled: false          # a declarative gateway needs no admin listener
observability:
  metrics:
    prometheus:
      enabled: true
      addr: "0.0.0.0:9090"
# resources.yaml
_format_version: "1"

provider_keys:
  - display_name: openai-main
    provider: openai
    api_key: ${OPENAI_API_KEY}        # interpolated from the environment

models:
  - display_name: my-model
    provider: openai
    model_name: gpt-4o-mini
    provider_key: openai-main

api_keys:
  - display_name: local-dev
    key_env: CALLER_API_KEY           # hashed at load; the plaintext is never stored
    allowed_models: ["my-model"]
export OPENAI_API_KEY="YOUR_PROVIDER_KEY"
export CALLER_API_KEY="YOUR_CALLER_KEY"

docker run -d --name aisix \
  -v "$(pwd)/config.yaml:/etc/aisix/config.yaml:ro" \
  -v "$(pwd)/resources.yaml:/etc/aisix/resources.yaml:ro" \
  -e OPENAI_API_KEY -e CALLER_API_KEY \
  -p 3000:3000 -p 127.0.0.1:9090:9090 \
  ghcr.io/api7/aisix:latest        # proxy → :3000, metrics + status → :9090
#                                  ^ the metrics/status listener is unauthenticated;
#                                    keep it on loopback or a private network

Then call the gateway exactly like OpenAI:

curl http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer $CALLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"my-model","messages":[{"role":"user","content":"hello"}]}'

Edit resources.yaml and send SIGHUP (docker kill -s HUP aisix) to apply changes with no restart — an invalid file is rejected whole and the last good configuration keeps serving. Check a file before booting with aisix validate --resources resources.yaml.

Full walkthrough: the Gateway Quickstart · every field: the resources file reference. For a multi-replica cluster, point the gateway at etcd instead — resources_file and etcd are mutually exclusive.

✨ Why AISIX

  • One API, every model. Speak the OpenAI or Anthropic wire format in; the gateway translates to whichever provider each model points at. Point an OpenAI or Claude SDK at one base_url and switch models without changing code.
  • A real gateway, in Rust. Single static binary, low cold-start, lock-free config reads on the hot path, native streaming.
  • Open source, free forever. Apache-2.0 licensed and built to run in your infrastructure. Choose AISIX Cloud when you want centralized management through a control plane and dashboard.
  • Production controls built in. Routing & failover, rate limits, guardrails, caching, and observability ship in the box. (Budgets and spend caps are an AISIX Cloud feature — the gateway enforces the control plane's decisions.)

🧩 Features — available today

Covered by 183 end-to-end scenario files (496 cases) that run against real gateway processes.

  • OpenAI-compatible proxy (:3000) — chat/completions, completions, responses, embeddings, rerank, images/{generations,edits}, audio/{speech,transcriptions,translations}, videos (submit → poll → fetch), files, batches, fine_tuning/jobs, realtime, GET /v1/models, plus a root-level /passthrough/:provider/* escape hatch. Native SSE streaming, tool/function calling, JSON mode, vision/multimodal input, and reasoning-content support.
  • Anthropic Messages API — POST /v1/messages as a first-class route, working against any configured upstream: requests and responses (including streaming) are translated both ways when a model points at a non-Anthropic provider.
  • Routing & failover — virtual/routing models with six strategies: round_robin (smooth weighted round-robin), consistent_hash (session affinity keyed by header / cookie / API key / client IP), failover, plus metric-based least_cost, least_latency, and least_busy. Per-target priority tiers (active/backup pools), retry budgets, cooldowns, tag-conditional targets, and per-attempt timeouts.
  • Ensemble models — fan one request out to a panel of models concurrently, then have a judge model synthesize a single answer, with a minimum-successful-responses threshold.
  • Semantic routing — one virtual model that dispatches by the meaning of each request: it embeds the prompt, scores it against per-route example utterances, and routes to the best match (or a default). See the semantic routing docs.
  • Rate limiting & concurrency — RPS/RPM/RPH/RPD + TPM/TPD + concurrency caps, AND-combined across caller keys, models, and policy scopes (api_key / model / team / member / team_member). Counters are per-process by default, or shared across replicas with the Redis backend.
  • Guardrails — content-policy enforcement on input and output, in-process or through a provider: keyword/regex, built-in PII detection and redaction, Presidio, Lakera, OpenAI Moderation, AWS Bedrock Guardrails, Azure AI Content Safety (Prompt Shield + text moderation), and two Alibaba Cloud services. A block returns 422 content_filter; monitor mode records what would have happened without blocking.
  • Caching — exact-match response cache with per-policy TTL and model/key scope matchers; memory and Redis backends; cost-saved telemetry on every hit. Separately, automatic prompt caching can be enabled per direct Anthropic model to inject cache breakpoints, so callers get provider-side prompt discounts without changing their requests.
  • MCP gateway — front registered upstream MCP servers at /mcp with gateway-held credentials, per-server tool namespaces, and per-caller access. It serves every Streamable HTTP revision from 2025-03-26 through stateless 2026-07-28 without downstream sessions. Upstreams use initialize by default or server/discover with protocol_version: "2026-07-28". CI runs the official MCP suite's applicable tools-only protocol scenarios. Also exposes a REST API as MCP tools from its OpenAPI description.
  • A2A agent gateway — front A2A (Agent-to-Agent) agents at /a2a/:agent, serving each agent's card with URLs rewritten to the gateway, over JSON-RPC 2.0.
  • Inbound authentication — caller API keys (SHA-256 hashed, model allowlists, expiry, rotation), or OIDC/JWT bearer tokens validated against registered providers (Entra ID, Okta, Google Workspace, or any OIDC issuer) with JWKS caching.
  • Observability — Prometheus /metrics, structured per-request access logs, usage events, OTLP/GenAI span export (Langfuse, Honeycomb, Grafana Cloud, or any OTLP receiver), plus dedicated Datadog and Aliyun SLS log exporters and object-storage (S3/GCS/Azure Blob) telemetry.
  • Declarative configuration — one resources.yaml carries all ten resource collections (provider keys, models, caller keys, guardrails, MCP servers, A2A agents, cache policies, observability exporters, rate-limit policies, OIDC providers), validated against the same JSON Schemas the gateway uses at runtime. aisix validate checks a file offline; SIGHUP reloads it atomically.
  • Operational endpoints — /livez and /readyz on the proxy listener; /status/config, /status/ready, /status/models, and Prometheus /metrics on a dedicated metrics listener (:9090). The admin listener (:3001) additionally serves a read-only resource surface, OpenAPI 3 with a Scalar UI, and a playground. Resources are managed declaratively — through the resources_file (reloaded on SIGHUP) or direct etcd writes — not through the admin listener; its former write endpoints were removed.

🔌 Supported providers

AISIX dispatches through five native adapter families — distinct wire-protocol bridges, not one generic relabel. Whatever the upstream protocol, the client-facing API stays OpenAI-shaped.

Adapter familyReachesWire shape · auth
openaiOpenAI + any OpenAI-compatible vendor — DeepSeek, Groq, Mistral, Together, Fireworks, Perplexity, vLLM, Ollama, or self-hosted OpenAI-compatible endpointsOpenAI chat completions · Bearer
anthropicAnthropic ClaudeAnthropic Messages · x-api-key
bedrockAWS Bedrock — Anthropic, Meta Llama, Mistral, Cohere, Amazon Titan/Nova, AI21Bedrock Converse + /invoke · SigV4
vertexGoogle Vertex AI (Gemini)Vertex :generateContent · OAuth2
azure-openaiAzure OpenAIAzure deployments · api-key / Entra ID

Plus specialized handling for vendor quirks (e.g. DeepSeek reasoning content) and dedicated rerank / embeddings vendors (Cohere, Jina). Details in adapter protocol families.

☁️ Open source vs AISIX Cloud

Same gateway binary, same proxy API — in every form the gateway runs in your environment. AISIX Cloud adds a commercial control plane, either hosted by API7 (Hybrid Cloud) or hosted in your infrastructure (On-Premises).

<table> <tr> <td width="50%" valign="top"> <img src="assets/console-overview.png" alt="AISIX Cloud overview — requests, latency p50/p99, error rate and cost today, with a 7-day request-and-cost trend and data-plane health" width="100%"><br> <sub><b>Overview</b> — traffic, latency, error rate &amp; spend at a glance</sub> <br><br> <img src="assets/console-models.png" alt="AISIX Cloud models — alias an upstream LLM per provider (OpenAI, Anthropic, AWS Bedrock, DeepSeek) with model IDs and per-model rate limits" width="100%"><br> <sub><b>Models</b> — one alias per upstream: OpenAI, Anthropic, Bedrock, DeepSeek…</sub> <br><br> <img src="assets/console-guardrails.png" alt="AISIX Cloud guardrails — pre-input and post-output content policies (keyword blocklist, Azure Content Safety, AWS Bedrock) that block on violation" width="100%"><br> <sub><b>Guardrails</b> — pre-input &amp; post-output policies, block on violation</sub> </td> <td width="50%" valign="top"> <img src="assets/console-playground.png" alt="AISIX Cloud playground — pick a model, set system and user prompts, run, and read the response with live token and cost metering" width="100%"><br> <sub><b>Playground</b> — test any model with live token &amp; cost metering</sub> <br><br> <img src="assets/console-observability.png" alt="AISIX Cloud observability exporters — fan out chat-completion telemetry to OTLP, Datadog and object storage, with per-target delivery health" width="100%"><br> <sub><b>Observability</b> — fan out traces &amp; logs to OTLP, Datadog, object storage</sub> <br><br> <img src="assets/console-budgets.png" alt="AISIX Cloud budgets — organization and per-environment spend caps with progress bars, hard-stop versus warn-only, including an over-budget policy" width="100%"><br> <sub><b>Budgets</b> — hard-stop spend caps with warn-only tiers</sub> </td> </tr> </table> <p align="center"> <em>The AISIX Cloud dashboard — overview metrics, multi-provider models, guardrails, budgets (with hard-stop spend caps), and observability exporters, across all your gateways.</em> <br><br> <a href="https://aisix-demo.api7.ai/"><b>▶ Try the live dashboard demo — aisix-demo.api7.ai</b></a> </p>
Open-source gateway (this repo)AISIX Cloud (Hybrid Cloud or On-Premises)
PriceFree · Apache-2.0 · foreverCommercial — talk to us
ConfigurationDeclarative resources.yaml, or etcd for a clusterDashboard + Cloud Admin API, multi-environment
TenancySingle instance / namespaceOrg → Team → Member → Environment
Provider keysIn the resources file as ${VAR} env references, or in etcdEnvelope-encrypted at rest, write-only, in-place rotation
Inbound authCaller keys (SHA-256 hashed, model allowlists, expiry), or OIDC/JWT bearersSame, plus masked reveal, key ownership, and PATs
Budgets— (rate and token limits only)Per key / provider / env / org / team, hard-stop & alerts
RBACAdmin key = read-only resource surfaceOrg roles (owner / admin / member), invites
Audit log—Full org-scoped audit with diff viewer
Usage & costExport logs, metrics, and usage events yourselfManaged usage views, model pricing catalog, spend reporting
SurfaceStatus endpoints, OpenAPI read surface, playgroundFull dashboard + per-environment playground

→ Want the AISIX Cloud control plane, governance, budgets, and dashboard? Talk to API7 about Hybrid Cloud or On-Premises, or book a demo.

🏗️ Architecture

A single Cargo workspace; the aisix-server crate builds one binary named aisix that wires the crates together.

crates/
├── aisix-core           Config, snapshot, resource model, resources.yaml source, errors
├── aisix-etcd           Config provider + watch supervisor
├── aisix-gateway        Hub & bridge, SSE parser, provider trait
├── aisix-proxy          /v1/*, /mcp, /a2a handlers, routing, middleware
├── aisix-admin          Read-only resource surface + playground + OpenAPI
├── aisix-provider-*     openai · anthropic · azure-openai · bedrock · vertex
├── aisix-mcp            MCP gateway — server registry, tool ACL, transports
├── aisix-a2a            A2A agent gateway — agent cards, JSON-RPC bridge
├── aisix-ratelimit      fixed-window + token accounting + concurrency (local | redis)
├── aisix-cache          memory + redis backends
├── aisix-redis          shared Redis connection for cache + rate limits
├── aisix-guardrails     pre/post content-policy hooks
├── aisix-obs            tracing, metrics, access log, exporters
└── aisix-server         the `aisix` binary — bootstrap + CLI

🗺️ Roadmap

Highlights on the roadmap; tracked live in issues:

  • Semantic (embedding-similarity) response caching
  • More observability sinks — Langsmith, Helicone, Slack alerts
  • Prompt templates managed as gateway resources
  • Llama-Guard as a guardrail provider

Shipped since this list was last written: the MCP gateway, the A2A agent gateway, OIDC/JWT inbound auth, Redis-backed distributed rate limiting, and the Lakera, Presidio, PII, and OpenAI Moderation guardrails — see Features above.

🛠️ Development

Prerequisites: the Rust toolchain pinned in rust-toolchain.toml. Docker is only needed for the tests that exercise etcd, Redis, or provider emulators.

cargo check --workspace
cargo fmt --check
cargo clippy --workspace -- -D warnings
cargo test --workspace

# Coverage (matches the CI gate)
cargo llvm-cov --workspace --lcov --output-path lcov.info

# Run locally against a resources.yaml (no etcd needed). Copy the Quickstart's two files
# and change resources_file to the local path, e.g. resources_file: ./resources.yaml
cargo run -p aisix-server --bin aisix -- --config config.local.yaml

# Check a resources file without starting a listener
cargo run -p aisix-server --bin aisix -- validate --resources resources.yaml

💬 Community

If AISIX is useful to you, a ⭐ helps other engineers find it.

📄 License

Apache 2.0.

Related MCP servers

Control WeMo smart home devices - discover, toggle, adjust brightness, and manage HomeKit codes

0
Python
MIT
View repository →

Bulk-validates EU VAT numbers against the European Commission's official VIES service.

0
View repository →

Checks llms.txt, AI crawler access in robots.txt, and sitemap - with a 0-100 AI readiness score.

0
View repository →

Detect 7,500+ technologies on any website - CMS, ecommerce, analytics, frameworks. ~$5/1,000 sites.

0
View repository →

Discover and pay for APIs with USDC credits. No wallet, no gas, MCP-native marketplace.

0
View repository →
RYRyter Pro logo

Ryter Pro

Active

MCP server for humanizing and rewriting text with the Ryter Pro API.

1
Python
View repository →