PluginBench
MCP Server
Active
Apache-2.0

io.github.houtini-ai/gemini MCP Server

io.github.houtini-ai/gemini

Google Gemini image generation, video, and search-grounded chat inside Claude

What is the io.github.houtini-ai/gemini MCP server?

The Gemini MCP server integrates Google's Gemini models into Claude, providing tools for image generation with search grounding, video creation with Veo 3.1, SVG diagram generation, and deep research with Google Search. It includes 14 specialized tools built around Gemini 3 Pro Image, Nano Banana 2, Veo 3.1, and Gemini 3.1 Pro for chat and research.

This server brings Google Gemini's capabilities—particularly image generation, video synthesis, and search-grounded research—directly into Claude Desktop. Instead of switching between browser tabs, you get inline previews of generated images, videos, and SVGs, plus grounded chat and deep research that pulls live data from Google Search. It's designed for workflows where Gemini's strengths (text rendering in images, video with synchronized audio, data-driven image generation) complement Claude's reasoning.

How to install io.github.houtini-ai/gemini

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • GEMINI_API_KEY
    required
    secret

    Google AI Studio API key (get from https://aistudio.google.com/apikey)

~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "gemini": {
      "command": "npx",
      "args": [
        "-y",
        "@houtini/gemini-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "<YOUR_GEMINI_API_KEY>"
      }
    }
  }
}

Tools & capabilities

Tools this server exposes to the agent.

  • gemini_chat — Chat with Google Search grounding enabled by default; supports reasoning with configurable thinking levels
  • gemini_deep_research — Multi-pass grounded research with synthesis; runs search passes on fast models and synthesis on Pro with high thinking
  • generate_image — Image generation with optional search grounding; supports Nano Banana Pro (4K, text rendering) and Nano Banana 2 (fast)
  • edit_image — Conversational image editing on Nano Banana Pro; maintains context between edits via thought signatures
  • describe_image — Quick general image descriptions on Gemini 3.8 Flash
  • analyze_image — Structured image analysis and extraction on Gemini 3.1 Pro
  • load_image_from_path — Load local image files for use in other tools
  • generate_video — Video generation with Veo 3.1; creates 4–8 second clips up to 4K with synchronized audio
  • generate_svg — SVG diagram generation in four styles (technical, artistic, minimal, data-viz); outputs real vector markup
  • generate_landing_page — Self-contained HTML landing pages with inline CSS and vanilla JS; four style options
  • gemini_prompt_assistant — Access nine professional chart design systems (storytelling, financial, terminal, modernist, professional, editorial, scientific, minimal, dark)
  • gemini_help — Built-in documentation accessible without leaving Claude

Use cases

  • Generate data-driven infographics with live search data (weather forecasts, stock prices, sports results)
  • Create polished marketing landing pages from a brief and brand color
  • Produce 4–8 second cinematic videos with synchronized audio and optional first-frame animation
  • Design architecture diagrams and flowcharts as editable SVG code for repositories or web pages
  • Conduct multi-pass research on complex topics with synthesis and source attribution

io.github.houtini-ai/gemini MCP server FAQ

What is the Gemini MCP server?

It's an MCP server that exposes 14 Google Gemini tools—image generation, video, SVG diagrams, chat, and deep research—as Claude tools. Everything previews inline in Claude Desktop instead of requiring file exports.

Is it free?

Gemini 3 Flash is free-tier eligible, but Gemini 3.1 Pro (used for chat, deep research, and image analysis by default) is paid-only. You can override defaults to use free models, or check Google's current pricing page.

How do I install it in Claude Desktop?

Add the server to your claude_desktop_config.json with the npx command and your Gemini API key, then restart Claude. No separate npm install needed—npx pulls the package on first run.

Do I need a Google API key?

Yes. Get one free from Google AI Studio (aistudio.google.com/apikey) and set it as GEMINI_API_KEY in your config.

Can I use it in Claude Code (CLI)?

Yes, use `claude mcp add -e GEMINI_API_KEY=your-key -s user gemini -- npx -y @houtini/gemini-mcp`, optionally with GEMINI_IMAGE_OUTPUT_DIR for file storage.

What models does it use?

Gemini 3 Pro Image (Nano Banana Pro) for high-quality images with text, Gemini 3.1 Flash Image (Nano Banana 2) for speed, Veo 3.1 for video, and Gemini 3.1 Pro for chat and research. All are configurable per call or globally.

README (reference)

Source of truth, from the repository.

<div align="center"> <img src="https://raw.githubusercontent.com/houtini-ai/gemini-mcp/main/assets/logo.png" width="120" height="120" alt="Gemini MCP" /> </div>

Gemini MCP - Google Gemini image generation, video and search grounding inside Claude

npm version MCP Registry Known Vulnerabilities License: Apache 2.0

I've had this Gemini MCP server running in my Claude Desktop setup for the best part of a year now. It's one of the few I leave switched on permanently. Not because Gemini replaces Claude (it doesn't), but because grounded search, image generation, SVG diagrams and video are things Gemini happens to do well, and having them as tools inside Claude beats flipping between browser tabs.

Fourteen tools, built around the models people actually come looking for: Nano Banana Pro (gemini-3-pro-image) and Nano Banana 2 for images, Veo 3.1 for video with synchronised audio, and Gemini 3.1 Pro for chat and deep research with Google Search grounding. Everything previews inline in Claude Desktop through MCP Apps rather than landing as a file path you have to go and open.

One npx command. That's it.

<p align="center"> <a href="https://glama.ai/mcp/servers/@houtini-ai/gemini-mcp"> <img width="380" height="200" src="https://glama.ai/mcp/servers/@houtini-ai/gemini-mcp/badge" alt="Gemini MCP server" /> </a> </p>

Quick Navigation

What it makes | Get started | What it does | Image output | Configuration | Tools | Models | Requirements


What it makes

Everything below came out of the tools in this repo, unretouched, on the afternoon I wrote this. Prompts are in the captions so you can judge for yourself.

Image generation with search grounding. One generate_image call with use_search=true. Gemini looked up the actual Met Office forecast for the week, then drew it. The dates and temperatures are real (well, as real as a forecast gets).

London five-day weather infographic, generated by Nano Banana Pro from live search data

Prompt: "A polished editorial infographic poster: London weather, this week. Use real forecast data from search for the next 5 days..." - gemini-3-pro-image, 16:9, 2K.

Text that's actually spelled correctly. This is the thing Nano Banana Pro does that the previous generation of image models couldn't. Every label here is straight out of the model.

Four-step infographic explaining an MCP tool call, with correctly rendered text

Prompt: "A clean, magazine-quality infographic titled How an MCP tool call works, with four numbered steps..." - gemini-3-pro-image.

Generate, then edit. Left is generate_image. Right is edit_image on that file with one sentence of instructions: change the track to Monza in daylight, swap the rim lighting for window light, make the pedals red. Same cockpit, same camera angle.

generate_imageedit_image
Sim racing cockpit at night, rainy Spa on the monitorThe same cockpit edited to daytime Monza with red pedals

Image to video with Veo 3.1. The night-time cockpit above, passed to generate_video as firstFrameImage. Eight seconds at 1080p with generated audio (engine note, tyre hiss) - the GIF below is silent and squashed for GitHub, the real file is a proper MP4.

Veo 3.1 video: the cockpit comes to life, hands turning the wheel through Eau Rouge

SVG that you can actually use. Not a picture of a diagram - real vector markup you can drop into a page, edit by hand or commit to a repo. Both of these are the raw .svg files the tool wrote to disk.

generate_svg style=technicalgenerate_svg style=data-viz
Architecture diagram of this MCP server as an SVGGrouped bar chart of API latency by region as an SVG

A landing page from a paragraph. generate_landing_page with a brief, a company name and a brand colour. Self-contained HTML, inline CSS, the little chart in the hero is animated SVG. Screenshot of the file opened in Chrome, nothing else touched.

Generated SaaS landing page rendered in a browser

And the fast one. Nano Banana 2 (gemini-3.1-flash-image) for when you want volume rather than 4K. This took about ten seconds.

Flat-lay of a mechanical keyboard on a walnut desk, Nano Banana 2


Get started in two minutes

Step 1: Get a Gemini API key

Go to Google AI Studio and create one.

A word on the free tier, because the defaults lean on a paid model. Google's pricing page (as of 24 September 2026) gives Gemini 3 Flash Preview a free tier, but Gemini 3.1 Pro Preview is paid-only - and 3.1 Pro is the default for gemini_chat, gemini_deep_research (synthesis), analyze_image and generate_landing_page. On a free key those calls will fail unless you point them at a free model: set GEMINI_DEFAULT_MODEL=gemini-3-flash-preview (covers chat and landing pages), GEMINI_DEEP_RESEARCH_MODEL and GEMINI_IMAGE_ANALYSIS_MODEL likewise, or pass model per call. generate_svg already defaults to gemini-3-flash-preview. Check the pricing page for the other models before relying on them - Google changes these tiers.

Step 2: Add to your Claude Desktop config

Config file locations:

  • Windows: C:\Users\{username}\AppData\Roaming\Claude\claude_desktop_config.json
  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "gemini": {
      "command": "npx",
      "args": ["@houtini/gemini-mcp"],
      "env": {
        "GEMINI_API_KEY": "your-api-key-here"
      }
    }
  }
}

Step 3: Restart Claude Desktop

That's it. Tools show up automatically. npx pulls the package on first run, so there's no separate install.

Local build instead

For development, or if you'd rather not rely on npx:

git clone https://github.com/houtini-ai/gemini-mcp
cd gemini-mcp
npm install --include=dev
npm run build

Then point your config at the local build:

{
  "mcpServers": {
    "gemini": {
      "command": "node",
      "args": ["C:/path/to/gemini-mcp/dist/index.js"],
      "env": {
        "GEMINI_API_KEY": "your-api-key-here"
      }
    }
  }
}

Claude Code (CLI)

Claude Code doesn't read claude_desktop_config.json. Use claude mcp add instead:

claude mcp add -e GEMINI_API_KEY=your-api-key-here -s user gemini -- npx -y @houtini/gemini-mcp

With an output directory for images and video:

claude mcp add \
  -e GEMINI_API_KEY=your-api-key-here \
  -e GEMINI_IMAGE_OUTPUT_DIR=/path/to/output \
  -s user \
  gemini -- npx -y @houtini/gemini-mcp

Check with claude mcp get gemini - you want to see Status: Connected.


What it does

Chat with Google Search grounding

Use gemini:gemini_chat to ask: "What changed in the MCP spec in the last month?"

Grounding is on by default. Gemini searches Google before it answers, so you get this month's information rather than a training-cutoff guess, and the sources come back as markdown links. For questions where you want pure reasoning ("explain this code", that sort of thing) set grounding: false.

Runs on gemini-3.1-pro-preview unless you say otherwise. Pass model: "gemini-3.8-flash" if you'd rather have speed than depth. thinking_level works on every Gemini 3.x model: high for the hard stuff, low to keep it snappy.

Deep research

Use gemini:gemini_deep_research with:
  research_question="What are the current approaches to AI agent memory management?"
  max_iterations=5

Runs grounded search passes and then writes them up as one report. The passes run on gemini-3.8-flash with low thinking (they're gathering facts, not reasoning about them), and the synthesis at the end runs on gemini-3.1-pro-preview with high thinking. Two passes plus a synthesis is the default and lands in two to three minutes.

That split matters more than it sounds. Earlier versions ran everything on Pro with full thinking, and one pass on a broad question could run past Claude Desktop's four-minute timeout on its own. Keep max_iterations at 2 or 3 in Claude Desktop; in an IDE or an agent framework, 5 to 7 produces noticeably better synthesis. focus_areas takes an array if you want to steer each pass.

Image generation with search grounding

Use gemini:generate_image with:
  prompt="Stock price chart showing Apple (AAPL) closing prices for the last 5 trading days"
  use_search=true
  aspectRatio="16:9"

Default model is gemini-3-pro-image, which is Nano Banana Pro now that it's out of preview. It renders legible text, does 4K, and it's the only one that supports conversational editing. gemini-3.1-flash-image (Nano Banana 2) is the fast option - not far off in quality, a fraction of the time and the cost. There's a Lite variant too if you're doing hundreds.

With use_search=true, Gemini looks things up before it draws. Weather, prices, sports results, that kind of data-driven image works reliably. The full-resolution file always goes to disk; the inline preview is resized to fit the MCP transport cap but the original is untouched.

Video generation with Veo 3.1

Use gemini:generate_video with:
  prompt="A close-up shot of a futuristic coffee machine brewing a glowing blue espresso, steam rising dramatically. Cinematic lighting."
  resolution="1080p"
  durationSeconds=8

Google's Veo 3.1. Four to eight second clips at up to 4K with native, synchronised audio. It's asynchronous on Google's side and takes two to five minutes; the tool polls until it's ready so you don't have to.

Options worth knowing about:

  • aspectRatio - 16:9 landscape or 9:16 for vertical
  • generateAudio - on by default; dialogue and effects that match the prompt
  • firstFrameImage - animate from a still (that's how the cockpit clip above was made)
  • referenceImages - up to three, for character or style consistency
  • sampleCount - up to four variations in one call
  • seed - deterministic output across runs
  • generateThumbnail - pulls a frame out with ffmpeg, if you've got it in PATH
  • model - veo-3.1-lite-generate-preview if you want cheaper and faster

SVG generation

This is the one people underestimate. The output isn't a picture of a diagram, it's the SVG markup itself - drop it into a codebase, a slide, a web page, and it scales without a single raster artefact.

Use gemini:generate_svg with:
  prompt="Architecture diagram showing a microservices system with API gateway, three services, and a shared database"
  style="technical"
  width=1000
  height=600

Four styles:

StyleBest for
technicalArchitecture diagrams, flowcharts, system maps
artisticIllustrations, decorative graphics, icons
minimalClean data visualisations, simple charts
data-vizCharts, dashboards, infographics

You get real SVG code back. Edit it, animate it, embed it, commit it. No export step, no Figma.

Runs on gemini-3-flash-preview unless you pass model - it's quick, it handles SVG markup comfortably, and it has a free tier. Pass model: "gemini-3.1-pro-preview" for denser diagrams if you're on a paid key.

Image editing and analysis

Conversational editing. Nano Banana Pro keeps context between turns. Pass the thought signature from the previous call back in and it remembers what it was working on:

Use gemini:edit_image with:
  prompt="Change the colour scheme to blue and green"
  images=[{filePath: "C:/output/gemini-123.png", thoughtSignature: "fromPreviousCall"}]

Analysis, two tools for two jobs:

  • describe_image - quick general descriptions on gemini-3.8-flash
  • analyze_image - structured extraction and proper reasoning on gemini-3.1-pro-preview

Local files:

Use gemini:load_image_from_path with filePath="C:/screenshots/error.png"

Or just pass filePath straight into any image tool's images array and the server reads it itself - that skips the MCP transport limit entirely.

Media resolution control

Cut token usage by up to 75% where the task doesn't need the detail:

LevelTokensSavingsBest for
MEDIA_RESOLUTION_LOW28075%Simple tasks, bulk operations
MEDIA_RESOLUTION_MEDIUM56050%PDFs and documents (OCR saturates here)
MEDIA_RESOLUTION_HIGH1120defaultDetailed analysis
MEDIA_RESOLUTION_ULTRA_HIGH2000+per-image onlyMaximum detail

For PDF OCR, MEDIUM gives me identical text extraction to HIGH at half the tokens. I've not found a case where it didn't.

Landing page generation

Use gemini:generate_landing_page with:
  brief="A SaaS tool that helps developers monitor API latency"
  companyName="PingWatch"
  primaryColour="#6366F1"
  style="startup"
  sections=["hero", "features", "pricing", "cta"]

One self-contained HTML file: inline CSS, vanilla JS, no external dependencies. Styles are minimal, bold, corporate and startup. The PingWatch page in the gallery above is exactly this call.

Professional chart design systems

gemini_prompt_assistant carries nine chart design systems you can ask for by name:

SystemInspirationBest for
storytellingCole Nussbaumer KnaflicExecutive presentations
financialFinancial TimesEditorial journalism - FT pink, serif titles
terminalBloomberg / fintechHigh-density dark mode with neon
modernistW.E.B. Du BoisBold geometric blocks, stark contrasts
professionalIBM Carbon / TailwindEnterprise dashboards
editorialFiveThirtyEight / EconomistData journalism
scientificNature / ScienceAcademic rigour
minimalEdward TufteMaximum data-ink ratio
darkObservableModern dark mode

Help system

Use gemini:gemini_help with topic="overview"

The full documentation without leaving Claude. Topics: overview, image_generation, image_editing, image_analysis, chat, deep_research, grounding, media_resolution, models, all.


Image output and storage

By default, images come back as inline previews rendered directly in Claude, and the full-size file is written next to the package. Set GEMINI_IMAGE_OUTPUT_DIR if you'd rather they all landed somewhere sensible:

"env": {
  "GEMINI_API_KEY": "your-api-key-here",
  "GEMINI_IMAGE_OUTPUT_DIR": "C:/Users/username/Pictures/gemini-output"
}

Two files per image:

FileWhat it is
Full-resSaved to disk immediately, untouched
PreviewResized JPEG for inline transport, sized to fit under the cap

Gemini returns 2 to 5 MB images. The resize measures the non-image overhead in each response, works out the binary budget left, and steps the preview down (800, 600, 400, 300, 200px) until it fits under the 1 MB MCP transport limit. The full image is always there on disk.

Inline viewers in Claude Desktop

Image, SVG, video and landing-page results each open in an MCP App viewer with zoom, the saved path and a copy button. If you're on Claude Desktop and the viewer sat on "Waiting for image..." forever in an older version, that was Claude Desktop stripping the structured data the viewer reads (ext-apps#696). Since 2.7.0 the viewer fetches it back from the server itself, so it renders either way.


Configuration reference

VariableRequiredDefaultDescription
GEMINI_API_KEYYes-Google AI API key from AI Studio
GEMINI_DEFAULT_MODELNogemini-3.1-pro-previewModel for gemini_chat
GEMINI_DEEP_RESEARCH_MODELNogemini-3.1-pro-previewSynthesis model for gemini_deep_research
GEMINI_DEEP_RESEARCH_SEARCH_MODELNogemini-3.8-flashModel for the grounded search passes in gemini_deep_research
GEMINI_IMAGE_ANALYSIS_MODELNogemini-3.1-pro-previewModel for analyze_image
GEMINI_IMAGE_DESCRIBE_MODELNogemini-3.8-flashModel for describe_image
GEMINI_IMAGE_GENERATION_MODELNogemini-3-pro-imageModel for generate_image and edit_image
GEMINI_DEFAULT_GROUNDINGNotrueSet to false to turn Google Search grounding off by default
GEMINI_IMAGE_OUTPUT_DIRNo-Where generated images and videos are saved
GEMINI_ALLOW_EXPERIMENTALNofalseInclude experimental and preview models in auto-discovery
GEMINI_REQUEST_TIMEOUT_MSNo240000Per-request timeout for chat and analysis calls, in milliseconds
GEMINI_MCP_RETRY_ATTEMPTSNo3Total attempts per Gemini API request. Transient network failures (fetch failed, ECONNRESET, proxy or VPN drops) are retried with backoff. 1 disables it
GEMINI_MCP_LOG_FILENofalseWrite logs to ~/.gemini-mcp/logs/
DEBUG_MCPNofalseLog to stderr for debugging tool calls

Tools reference

ToolDescription
gemini_chatChat with Gemini 3.1 Pro. Google Search grounding on by default. Supports thinking_level
gemini_deep_researchGrounded search passes on Flash, synthesised into a report by 3.1 Pro. Default 2 passes
gemini_list_modelsLists the models your API key can see, live
gemini_helpDocumentation for every tool without leaving Claude
gemini_prompt_assistantExpert guidance for image generation with nine chart design systems
generate_imageImage generation with optional search grounding. Full-res saved to disk
edit_imageEdit images with natural-language instructions. Multi-turn continuity via thought signatures
describe_imageFast image descriptions on Gemini 3.8 Flash
analyze_imageStructured extraction and analysis on Gemini 3.1 Pro
load_image_from_pathRead a local image file and return base64 for any image tool
generate_videoVideo generation with Veo 3.1: 4 to 8 seconds at up to 4K with native audio
generate_svgProduction-ready SVG: diagrams, illustrations, icons, data visualisations
generate_landing_pageSelf-contained HTML landing pages with inline CSS and JS
gemini_viewer_payloadInternal. The inline viewers use it to fetch their display data; you'll never call it

Model reference

Checked against the live models API on 22 September 2026. gemini_list_models will tell you what your key can see today.

ModelUsed byNotes
gemini-3.1-pro-previewgemini_chat, gemini_deep_research (synthesis), analyze_image, generate_landing_pageDefault. Still the strongest reasoning model Google ships, preview label or not. Paid-only - no free tier
gemini-3.8-flashdescribe_image, gemini_deep_research (search passes)Default. Google's GA workhorse as of September 2026; a good gemini_chat choice when you want speed
gemini-3-flash-previewgenerate_svgDefault. Has a free tier (Google pricing page, 24 September 2026)
gemini-3.7-flash, 3.6, 3.5, 3.5-flash-lite, 3.1-flash-liteany text toolEarlier GA Flash releases, all accepted
gemini-3-pro-imagegenerate_image, edit_imageDefault. Nano Banana Pro, GA. 4K, real text, conversational editing
gemini-3.1-flash-imagegenerate_image, edit_imageNano Banana 2, GA. Near-Pro quality, Flash speed and price
gemini-3.1-flash-lite-imagegenerate_imageNano Banana 2 Lite. Fastest and cheapest
gemini-3-pro-image-preview, nano-banana-pro-preview, gemini-2.5-flash-imageimage toolsOlder IDs, still accepted so existing configs keep working
veo-3.1-generate-previewgenerate_videoDefault. Cinematic, native audio, up to 4K
veo-3.1-lite-generate-previewgenerate_videoCheaper and faster on the same API

Why the defaults are what they are. I went back and forth on this. Gemini 3.8 Flash is GA and newer, but 3.1 Pro is still what Google calls its strongest reasoning model, and reasoning is the point of gemini_chat and deep research - so Pro stays the default there, and Flash takes the lighter describe_image job. For images, Nano Banana Pro left preview and kept its quality lead, so it's the default; Nano Banana 2 is there when you'd rather trade a little quality for a lot of speed. Gemini's newer Omni video model uses a different API, so it's not wired in yet.

Gemini 3 notes: temperature is forced to 1.0 on every 3.x model (Google's requirement, lower values cause looping). thinking_level applies to gemini_chat.

Token budgets: max_tokens defaults to each model's full output ceiling as reported live by the models API (65,536 on current Gemini 3 text models; the 1M figure is input context). It's a cap, not consumption, so unused headroom costs nothing. Values below 4,096 are ignored because Gemini 3 thinking burns tiny budgets before you see any output, which looks exactly like a timeout, and values above the model's real limit are clamped.


Requirements

  • Node.js 18+
  • A Gemini API key from Google AI Studio
  • ffmpeg (optional, for video thumbnails)

Licence

Apache-2.0

Related MCP servers

Scores content on signals that get it cited by ChatGPT, Perplexity, and AI Overviews.

21
TypeScript
MIT
View repository →

Search Google's Knowledge Graph for structured entity data—people, places, organizations, concepts.

10
TypeScript
MIT
View repository →

Offload routine coding tasks from Claude to a local LLM, cutting token costs and API bills.

109
JavaScript
Apache-2.0
View repository →

Technical SEO in Claude: GSC + crawl + DataForSEO - audit, market sizing, content recon.

5
TypeScript
Apache-2.0
View repository →

Crawl websites for SEO errors and store results in SQLite (now superseded by SEO Audit Console).

16
TypeScript
Apache-2.0
View repository →

Use Grok as a peer code reviewer and second-opinion consultant inside Claude, Cursor, and Cline.

10
TypeScript
MIT
View repository →