PluginBench
MCP Server
Maintained
MIT

io.github.RLabs-Inc/gemini-mcp MCP Server

io.github.RLabs-Inc/gemini-mcp

Integrate Google's Gemini 3 with Claude via 30+ tools: image/video generation, research, code execution, TTS, and web search.

What is the io.github.RLabs-Inc/gemini-mcp MCP server?

The Gemini MCP server is a Model Context Protocol server that connects Google's Gemini 3 models with Claude, exposing 30+ tools for image generation, video creation, web research, code execution, text-to-speech, and document analysis. It enables powerful AI collaboration between Claude and Gemini for development, content creation, and research workflows.

This MCP server bridges Claude and Google's Gemini 3, giving Claude access to Gemini's capabilities including 4K image generation, video creation with Veo, real-time web search, deep research, code execution, text-to-speech with 30 voices, YouTube analysis, document processing, and more. Use it to extend Claude's abilities for creative generation, research, data extraction, and multimodal tasks.

How to install io.github.RLabs-Inc/gemini-mcp

Copy-paste configuration for popular MCP clients.

transport: stdio
Config generated by PluginBench — verify against the source before use.
Environment / auth
  • GEMINI_API_KEY
    required
    secret

    Your Google Gemini API key (get one at https://aistudio.google.com/apikey)

  • GEMINI_OUTPUT_DIR

    Directory for generated files (images, videos, audio)

Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "gemini-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "@rlabs-inc/gemini-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "<YOUR_GEMINI_API_KEY>",
        "GEMINI_OUTPUT_DIR": "<YOUR_GEMINI_OUTPUT_DIR>"
      }
    }
  }
}
Cursor
~/.cursor/mcp.json
{
  "mcpServers": {
    "gemini-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "@rlabs-inc/gemini-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "<YOUR_GEMINI_API_KEY>",
        "GEMINI_OUTPUT_DIR": "<YOUR_GEMINI_OUTPUT_DIR>"
      }
    }
  }
}
Windsurf
~/.codeium/windsurf/mcp_config.json
{
  "mcpServers": {
    "gemini-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "@rlabs-inc/gemini-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "<YOUR_GEMINI_API_KEY>",
        "GEMINI_OUTPUT_DIR": "<YOUR_GEMINI_OUTPUT_DIR>"
      }
    }
  }
}
VS Code
.vscode/mcp.json
{
  "servers": {
    "gemini-mcp": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y",
        "@rlabs-inc/gemini-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "<YOUR_GEMINI_API_KEY>",
        "GEMINI_OUTPUT_DIR": "<YOUR_GEMINI_OUTPUT_DIR>"
      }
    }
  }
}
Claude Code
claude mcp add gemini-mcp --env GEMINI_API_KEY=<YOUR_GEMINI_API_KEY> --env GEMINI_OUTPUT_DIR=<YOUR_GEMINI_OUTPUT_DIR> -- npx -y @rlabs-inc/gemini-mcp

Tools & capabilities

Tools this server exposes to the agent.

  • gemini-queryDirect queries to Gemini Pro or Flash models with configurable thinking levels (low/medium/high)
  • gemini-generate-imageGenerate images up to 4K resolution with style, aspect ratio, and seed control
  • gemini-start-image-editBegin a multi-turn image editing session for iterative refinement
  • gemini-continue-image-editRefine an image in an active editing session
  • gemini-end-image-editClose an image editing session
  • gemini-list-image-sessionsList all active image editing sessions
  • gemini-generate-videoGenerate videos using Veo 2.0 with async polling support
  • gemini-check-videoCheck video generation status and download when complete
  • gemini-analyze-codeAnalyze code for quality, security, performance, bugs, or general issues
  • gemini-analyze-textAnalyze text for sentiment, entities, key points, or summaries
  • gemini-brainstormCollaborative brainstorming between Claude and Gemini
  • gemini-summarizeSummarize content in brief, moderate, or detailed formats
  • gemini-run-codeExecute Python code written by Gemini with numpy, pandas, matplotlib, scipy, scikit-learn, tensorflow
  • gemini-searchReal-time web search with inline citations and source URLs
  • gemini-structuredGet JSON responses matching a provided schema
  • gemini-extractExtract entities, facts, keywords, sentiment, or custom fields from text
  • gemini-youtubeAnalyze YouTube videos by URL with optional timestamp clipping
  • gemini-youtube-summarySummarize YouTube videos in brief, detailed, bullet-point, or chapter formats
  • gemini-analyze-documentAnalyze PDFs and documents (DOCX, spreadsheets) with questions
  • gemini-summarize-pdfSummarize PDF documents in brief, detailed, outline, or key-points formats

Use cases

  • Generate 4K images and iteratively refine them through multi-turn editing conversations
  • Execute Python data analysis and visualization tasks with Gemini writing and running code
  • Conduct deep research on topics with autonomous web search and citation tracking
  • Analyze documents, PDFs, and YouTube videos to extract insights and summaries
  • Create videos, generate speech in 30 voices, and build Claude+Gemini collaborative workflows

io.github.RLabs-Inc/gemini-mcp MCP server FAQ

What is the Gemini MCP server?

It's an MCP server that connects Google's Gemini 3 models to Claude, exposing 30+ tools for image/video generation, web research, code execution, text-to-speech, document analysis, and more. It enables Claude to leverage Gemini's capabilities for creative, analytical, and research tasks.

Is it free to use?

The server itself is free (MIT licensed). You need a free Google Gemini API key from Google AI Studio (aistudio.google.com/apikey). Free tier has rate limits; paid plans available for higher usage.

How do I install it in Claude/Cursor?

Run: `claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY npx -y @rlabs-inc/gemini-mcp` (replace YOUR_KEY with your API key from Google AI Studio). For Cursor, use the same command or add to your MCP configuration.

What authentication is required?

You need a Google Gemini API key from Google AI Studio (free, takes seconds to generate). Set it via the GEMINI_API_KEY environment variable when installing the MCP server.

Can I use it as a CLI tool?

Yes. Install globally with `npm install -g @rlabs-inc/gemini-mcp`, set your API key with `gcli config set api-key YOUR_KEY`, then use commands like `gcli image`, `gcli search`, `gcli research`, `gcli speak`, etc.

How do I reduce context usage?

Use tool presets or explicit tool lists. Set `GEMINI_TOOL_PRESET=minimal` for just query/brainstorm, or `GEMINI_ENABLED_TOOLS=query,search,image-gen` to load only specific tools. Presets include: minimal, text, image, research, media, full.

README (reference)

Source of truth, from the repository.

MCP Server Gemini

A Model Context Protocol (MCP) server for integrating Google's Gemini 3 models with Claude Code, enabling powerful collaboration between both AI systems. Now with a beautiful CLI!

npm version MCP Registry

MCP Registry Support: Now discoverable in the official MCP ecosystem!

Features

FeatureDescription
Deep Research AgentAutonomous multi-step research with web search and citations
Token CountingCount tokens and estimate costs before API calls
Text-to-Speech30 unique voices, single speaker or two-speaker dialogues
URL AnalysisAnalyze, compare, and extract data from web pages
Context CachingCache large documents for efficient repeated queries
YouTube AnalysisAnalyze videos by URL with timestamp clipping
Document AnalysisPDFs, DOCX, spreadsheets with table extraction
4K Image GenerationGenerate images up to 4K with 10 aspect ratios
Multi-Turn Image EditingIteratively refine images through conversation
Video GenerationCreate videos with Veo 2.0 (async with polling)
Code ExecutionGemini writes and runs Python code (pandas, numpy, matplotlib)
Google SearchReal-time web information with inline citations
Structured OutputJSON responses with schema validation
Data ExtractionExtract entities, facts, sentiment from text
Thinking LevelsControl reasoning depth (minimal/low/medium/high)
Direct QuerySend prompts to Gemini 3 Pro/Flash models
BrainstormingClaude + Gemini collaborative problem-solving
Code AnalysisAnalyze code for quality, security, performance
SummarizationSummarize content at different detail levels

Quick Installation

MCP Server for Claude Code

# Using npm (Recommended)
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY npx -y @rlabs-inc/gemini-mcp

# Using bun
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY bunx @rlabs-inc/gemini-mcp

CLI (Global Install)

# Install globally
npm install -g @rlabs-inc/gemini-mcp

# Set your API key once (stored securely)
gcli config set api-key YOUR_KEY

# Now use any command!
gcli search "latest news"
glci image "sunset over mountains" --ratio 16:9

Get your API key: Visit Google AI Studio - it's free and takes seconds!

Installation Options

# With verbose logging
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY VERBOSE=true bunx -y @rlabs-inc/gemini-mcp

# With custom output directory for generated images/videos
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY GEMINI_OUTPUT_DIR=/path/to/output bunx -y @rlabs-inc/gemini-mcp

Available Tools

gemini-query

Direct queries to Gemini with thinking level control:

prompt: "Explain quantum entanglement"
model: "pro" or "flash"
thinkingLevel: "low" | "medium" | "high" (optional)
  • low: Fast responses, minimal reasoning
  • medium: Balanced (Flash only)
  • high: Deep reasoning for complex tasks (default)

gemini-generate-image

Generate images with Nano Banana Pro (Claude can SEE them!):

prompt: "a futuristic city at sunset"
style: "cyberpunk" (optional)
aspectRatio: "16:9" (1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9)
imageSize: "2K" (1K, 2K, 4K)
useGoogleSearch: false (ground in real-world info)
thinkingLevel: "high" (optional - minimal, low, medium, high)
personGeneration: "ALLOW_ALL" (optional - ALLOW_ALL, ALLOW_ADULT, ALLOW_NONE)
seed: 42 (optional - for reproducible results)

gemini-start-image-edit

Start a multi-turn image editing session:

prompt: "a cozy cabin in the mountains"
aspectRatio: "16:9"
imageSize: "2K"
useGoogleSearch: false
thinkingLevel: "high" (optional - minimal, low, medium, high)
personGeneration: "ALLOW_ALL" (optional - ALLOW_ALL, ALLOW_ADULT, ALLOW_NONE)
seed: 42 (optional - for reproducible results)

Returns a session ID for iterative editing.

gemini-continue-image-edit

Continue refining an image:

sessionId: "edit-123456789"
prompt: "add snow on the roof and make it nighttime"

gemini-end-image-edit

Close an editing session:

sessionId: "edit-123456789"

gemini-list-image-sessions

List all active editing sessions.

gemini-generate-video

Generate videos using Veo:

prompt: "a cat playing piano"
aspectRatio: "16:9" (optional)
negativePrompt: "blurry, text" (optional)

Video generation is async (takes 1-5 minutes). Use gemini-check-video to poll.

gemini-check-video

Check video generation status and download when complete:

operationId: "operations/xxx-xxx-xxx"

gemini-analyze-code

Analyze code for issues:

code: "function foo() { ... }"
language: "typescript" (optional)
focus: "quality" | "security" | "performance" | "bugs" | "general"

gemini-analyze-text

Analyze text content:

text: "Your text here..."
type: "sentiment" | "summary" | "entities" | "key-points" | "general"

gemini-brainstorm

Collaborative brainstorming:

prompt: "How could we implement real-time collaboration?"
claudeThoughts: "I think we should use WebSockets..."
maxRounds: 3 (optional)

gemini-summarize

Summarize content:

content: "Long text to summarize..."
length: "brief" | "moderate" | "detailed"
format: "paragraph" | "bullet-points" | "outline"

gemini-run-code

Let Gemini write and execute Python code:

prompt: "Calculate the first 50 prime numbers and plot them"
data: "optional CSV data to analyze" (optional)

Supports libraries: numpy, pandas, matplotlib, scipy, scikit-learn, tensorflow, and more. Generated charts are saved to the output directory and returned as images.

gemini-search

Real-time web search with citations:

query: "What happened in tech news this week?"
returnCitations: true (default)

Returns grounded responses with inline citations and source URLs.

gemini-structured

Get JSON responses matching a schema:

prompt: "Extract the meeting details from this email..."
schema: '{"type":"object","properties":{"date":{"type":"string"},"attendees":{"type":"array"}}}'
useGoogleSearch: false (optional)

gemini-extract

Convenience tool for common extraction patterns:

text: "Your text to analyze..."
extractType: "entities" | "facts" | "summary" | "keywords" | "sentiment" | "custom"
customFields: "name, date, amount" (for custom extraction)

gemini-youtube

Analyze YouTube videos directly:

url: "https://www.youtube.com/watch?v=..."
question: "What happens at 2:30?"
startTime: "1m30s" (optional, for clipping)
endTime: "5m00s" (optional, for clipping)

gemini-youtube-summary

Quick video summarization:

url: "https://www.youtube.com/watch?v=..."
style: "brief" | "detailed" | "bullet-points" | "chapters"

gemini-analyze-document

Analyze PDFs and documents:

filePath: "/path/to/document.pdf"
question: "Summarize the key findings"
mediaResolution: "low" | "medium" | "high"

gemini-summarize-pdf

Quick PDF summarization:

filePath: "/path/to/document.pdf"
style: "brief" | "detailed" | "outline" | "key-points"

gemini-extract-tables

Extract tables from documents:

filePath: "/path/to/document.pdf"
outputFormat: "markdown" | "csv" | "json"

Workflow: Claude + Gemini

The killer combination for development:

ClaudeGemini
Complex logicFrontend/UI
ArchitectureVisual components
Backend codeImage generation
IntegrationReact/CSS styling
ReasoningCreative generation

Example workflow:

  1. Ask Claude to design the backend API
  2. Use gemini-generate-image for UI mockups
  3. Ask Gemini to generate React components via gemini-query
  4. Use multi-turn editing to refine visuals
  5. Let Claude wire everything together

Environment Variables

VariableRequiredDefaultDescription
GEMINI_API_KEYYes-Your Google Gemini API key
GEMINI_OUTPUT_DIRNo./gemini-outputWhere to save generated files
GEMINI_MODELNo-Override model for init test
GEMINI_PRO_MODELNogemini-3-pro-previewPro model (Gemini 3)
GEMINI_FLASH_MODELNogemini-3-flash-previewFlash model (Gemini 3)
GEMINI_IMAGE_MODELNogemini-3-pro-image-previewImage model (Nano Banana Pro)
GEMINI_IMAGE_THINKING_LEVELNohighDefault thinking level for image generation (minimal, low, medium, high)
GEMINI_VIDEO_MODELNoveo-2.0-generate-001Video model
VERBOSENofalseEnable verbose logging
QUIETNofalseMinimize logging
GEMINI_ENABLED_TOOLSNo-Comma-separated list of tool groups to load (e.g., query,search,image-gen)
GEMINI_TOOL_PRESETNo-Preset profile: minimal, text, image, research, media, full

Tool Configuration

By default, all 37 tools are loaded. To reduce context usage, configure which tools to load:

Available Presets

PresetTool Groups
minimalquery, brainstorm
textquery, brainstorm, analyze, summarize, structured
imagequery, image-gen, image-edit, image-analyze
researchquery, search, deep-research, url-context, document
mediaquery, image-gen, image-edit, image-analyze, video-gen, youtube, speech
fullAll 18 tool groups (default)

Using Presets

# Minimal - query and brainstorm
GEMINI_TOOL_PRESET=minimal

# Text processing
GEMINI_TOOL_PRESET=text  # query, brainstorm, analyze, summarize, structured

# Image workflows
GEMINI_TOOL_PRESET=image  # query, image-gen, image-edit, image-analyze

# Research workflows
GEMINI_TOOL_PRESET=research  # query, search, deep-research, url-context, document

Using Explicit Tool Lists

# Only specific tools
GEMINI_ENABLED_TOOLS=query,search,image-gen

Combining Preset + Explicit

# Start with preset, add extras
GEMINI_TOOL_PRESET=minimal
GEMINI_ENABLED_TOOLS=search,image-gen  # Adds to minimal preset

Available Tool Groups

GroupTools
querygemini-query
brainstormgemini-brainstorm
analyzegemini-analyze-code, gemini-analyze-text
summarizegemini-summarize
image-gengemini-generate-image, gemini-image-prompt
image-editgemini-start-image-edit, gemini-continue-image-edit, gemini-end-image-edit, gemini-list-image-sessions
video-gengemini-generate-video, gemini-check-video
code-execgemini-run-code
searchgemini-search
structuredgemini-structured, gemini-extract
youtubegemini-youtube, gemini-youtube-summary
documentgemini-analyze-document, gemini-summarize-pdf, gemini-extract-tables
url-contextgemini-analyze-url, gemini-compare-urls, gemini-extract-from-url
cachegemini-create-cache, gemini-query-cache, gemini-list-caches, gemini-delete-cache
speechgemini-speak, gemini-dialogue, gemini-list-voices
token-countgemini-count-tokens
deep-researchgemini-deep-research, gemini-check-research, gemini-research-followup
image-analyzegemini-analyze-image

Manual Installation

Global Install

# Using npm
npm install -g @rlabs-inc/gemini-mcp

# Using bun
bun install -g @rlabs-inc/gemini-mcp

Claude Code Configuration

{
  "gemini": {
    "command": "npx",
    "args": ["-y", "@rlabs-inc/gemini-mcp"],
    "env": {
      "GEMINI_API_KEY": "your-api-key",
      "GEMINI_OUTPUT_DIR": "/path/to/save/files"
    }
  }
}

Troubleshooting

Rate Limits (429 Errors)

If you're hitting rate limits on the free tier:

  • Set GEMINI_MODEL=gemini-3-flash-preview to use Flash for init (higher limits)
  • Or upgrade to a paid plan

Connection Issues

  1. Verify your API key at Google AI Studio
  2. Check server status: claude mcp list
  3. Try with verbose logging: VERBOSE=true

Image/Video Issues

  • Ensure your API key has access to image/video generation
  • Check output directory permissions
  • Files save to GEMINI_OUTPUT_DIR (default: ./gemini-output)
  • For 4K images, generation takes longer

Previous Versions

0.7.2

Beautiful CLI with Themes! Use Gemini directly from your terminal:

# Install globally
npm install -g @rlabs-inc/gemini-mcp

# Set your API key once
gcli config set api-key YOUR_KEY

# Generate images, videos, search, research, and more!
gcli image "a cat astronaut" --size 4K
gcli search "latest AI news"
gcli research "quantum computing applications" --wait
gcli speak "Hello world" --voice Puck

5 Beautiful Themes: terminal, neon, ocean, forest, minimal

CLI Commands:

  • gcli query - Direct Gemini queries with thinking levels
  • gcli search - Real-time web search with citations
  • gcli research - Deep research agent
  • gcli image - Generate images (up to 4K)
  • gcli video - Generate videos with Veo
  • gcli speak - Text-to-speech with 30 voices
  • gcli tokens - Count tokens and estimate costs
  • gcli config - Manage settings

v0.6.x: Deep Research, Token Counting, TTS, URL analysis, Context Caching v0.5.x: 30+ tools, YouTube analysis, Document analysis v0.4.x: Code execution, Google Search v0.3.x: Thinking levels, Structured output, 4K images v0.2.x: Image/Video generation with Veo


Development

git clone https://github.com/rlabs-inc/gemini-mcp.git
cd gemini-mcp
bun install
bun run build
bun run dev -- --verbose

Scripts

CommandDescription
bun run buildBuild for production
bun run devDevelopment mode with watch
bun run typecheckType check without emitting
bun run formatFormat with Prettier
bun run lintLint with ESLint

License

MIT License


Made with Claude + Gemini working together

Related MCP servers

Give your AI agent stealth web scraping with Cloudflare bypass and CSS selection, powered by Scrapling.

67k
Python
BSD-3-Clause
View repository →

Give your AI coding agent full control of a live Chrome browser for automation, debugging, and performance analysis.

45k
TypeScript
Apache-2.0
View repository →

Let AI agents manage your Puter files, websites, and serverless workers over MCP.

43k
TypeScript
AGPL-3.0
View repository →

Browser automation for AI agents via MCP, powering ByteDance's Agent TARS hybrid GUI/DOM browser control.

37k
TypeScript
Apache-2.0
View repository →

Run arbitrary shell commands from an MCP-connected AI agent.

37k
TypeScript
Apache-2.0
View repository →

Filesystem access MCP server from ByteDance's UI-TARS/Agent TARS ecosystem.

37k
TypeScript
Apache-2.0
View repository →