PluginBench
Skill
Review
Audit score 70

fal-ai-media

affaan-m/ecc

Generate images, videos, and audio with fal.ai models via MCP.

What is fal-ai-media?

Unified media generation skill that covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound) via fal.ai MCP. Use when the user wants to generate or create media assets with AI.

  • Generate images from text prompts using Nano Banana 2 (fast) or Nano Banana Pro (high fidelity)
  • Create videos from text or images with Seedance, Kling, or Veo 3, with optional native audio
  • Generate speech and conversational audio with CSM-1B text-to-speech
  • Generate matching audio from video content using ThinkSound
  • Edit, inpaint, and style-transfer existing images
  • Estimate generation costs before running models

How to install fal-ai-media

npx skills add null --skill fal-ai-media
Prerequisites
  • fal.ai API key (obtain at fal.ai)
  • MCP server configured in ~/.claude.json with fal-ai command and FAL_KEY environment variable
Claude Code
Cursor
Windsurf
Cline

How to use fal-ai-media

  1. 1.Configure the fal.ai MCP server by adding the fal-ai entry to ~/.claude.json with your FAL_KEY
  2. 2.Use the search() or models() tool to discover available models for your task
  3. 3.Call generate() with the appropriate app_id and input_data (prompt, image_size, duration, etc.)
  4. 4.For image-to-video or image editing, first upload your source image using upload(), then reference the returned URL in generate()
  5. 5.Optionally use estimate_cost() before expensive generations to check pricing
  6. 6.Use the result() or status() tools to check async job completion

Use cases

Good for
  • Create product photos, thumbnails, and marketing images from text descriptions
  • Generate cinematic video content from text prompts or reference images for social media or presentations
  • Produce voiceovers and narration for videos or presentations using text-to-speech
  • Generate sound effects and ambient audio to match video content
  • Iterate on image designs quickly with Nano Banana 2 before finalizing with Pro
Who it's for
  • Content creators and social media managers
  • Marketing and design teams
  • Video producers and filmmakers
  • Product photographers and e-commerce teams
  • Developers building media generation features

fal-ai-media FAQ

Which model should I use for fast image generation?

Use fal-ai/nano-banana-2 for quick iterations and drafts. It's fast and cost-effective. Switch to fal-ai/nano-banana-pro for production images requiring higher fidelity and detail.

Can I generate video from an existing image?

Yes. Use image-to-video with Seedance or other video models by uploading your image first with upload(), then passing the image_url in the generate() call with a motion-focused prompt.

Do video models generate audio automatically?

Some models like Kling Video v3 Pro and Veo 3 support native audio generation. For other models, use ThinkSound to generate matching audio from the video afterward.

How do I make results reproducible?

Set the seed parameter to a fixed integer value in your generate() call. This ensures the same prompt and seed produce identical outputs.

What should I do if model IDs or parameters change?

This skill is drift-prone; fal.ai updates models and pricing frequently. Use search(), find(), or models() tools to fetch current model metadata before promising specific models or parameters.

Full instructions (SKILL.md)

Source of truth, from affaan-m/ecc.


name: fal-ai-media description: Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI. metadata: origin: ECC

fal.ai Media Generation

Drift-prone skill. fal.ai model IDs, pricing, inputs, and MCP tool names change quickly. Search or fetch the current model metadata before promising a specific model, parameter, output format, or cost.

Generate images, videos, and audio using fal.ai models via MCP.

When to Activate

  • User wants to generate images from text prompts
  • Creating videos from text or images
  • Generating speech, music, or sound effects
  • Any media generation task
  • User says "generate image", "create video", "text to speech", "make a thumbnail", or similar

MCP Requirement

fal.ai MCP server must be configured. Add to ~/.claude.json:

"fal-ai": {
  "command": "npx",
  "args": ["-y", "fal-ai-mcp-server"],
  "env": { "FAL_KEY": "YOUR_FAL_KEY_HERE" }
}

Get an API key at fal.ai.

MCP Tools

The fal.ai MCP provides these tools:

  • search — Find available models by keyword
  • find — Get model details and parameters
  • generate — Run a model with parameters
  • result — Check async generation status
  • status — Check job status
  • cancel — Cancel a running job
  • estimate_cost — Estimate generation cost
  • models — List popular models
  • upload — Upload files for use as inputs

Image Generation

Nano Banana 2 (Fast)

Best for: quick iterations, drafts, text-to-image, image editing.

generate(
  app_id: "fal-ai/nano-banana-2",
  input_data: {
    "prompt": "a futuristic cityscape at sunset, cyberpunk style",
    "image_size": "landscape_16_9",
    "num_images": 1,
    "seed": 42
  }
)

Nano Banana Pro (High Fidelity)

Best for: production images, realism, typography, detailed prompts.

generate(
  app_id: "fal-ai/nano-banana-pro",
  input_data: {
    "prompt": "professional product photo of wireless headphones on marble surface, studio lighting",
    "image_size": "square",
    "num_images": 1,
    "guidance_scale": 7.5
  }
)

Common Image Parameters

ParamTypeOptionsNotes
promptstringrequiredDescribe what you want
image_sizestringsquare, portrait_4_3, landscape_16_9, portrait_16_9, landscape_4_3Aspect ratio
num_imagesnumber1-4How many to generate
seednumberany integerReproducibility
guidance_scalenumber1-20How closely to follow the prompt (higher = more literal)

Image Editing

Use Nano Banana 2 with an input image for inpainting, outpainting, or style transfer:

# First upload the source image
upload(file_path: "/path/to/image.png")

# Then generate with image input
generate(
  app_id: "fal-ai/nano-banana-2",
  input_data: {
    "prompt": "same scene but in watercolor style",
    "image_url": "<uploaded_url>",
    "image_size": "landscape_16_9"
  }
)

Video Generation

Seedance 1.0 Pro (ByteDance)

Best for: text-to-video, image-to-video with high motion quality.

generate(
  app_id: "fal-ai/seedance-1-0-pro",
  input_data: {
    "prompt": "a drone flyover of a mountain lake at golden hour, cinematic",
    "duration": "5s",
    "aspect_ratio": "16:9",
    "seed": 42
  }
)

Kling Video v3 Pro

Best for: text/image-to-video with native audio generation.

generate(
  app_id: "fal-ai/kling-video/v3/pro",
  input_data: {
    "prompt": "ocean waves crashing on a rocky coast, dramatic clouds",
    "duration": "5s",
    "aspect_ratio": "16:9"
  }
)

Veo 3 (Google DeepMind)

Best for: video with generated sound, high visual quality.

generate(
  app_id: "fal-ai/veo-3",
  input_data: {
    "prompt": "a bustling Tokyo street market at night, neon signs, crowd noise",
    "aspect_ratio": "16:9"
  }
)

Image-to-Video

Start from an existing image:

generate(
  app_id: "fal-ai/seedance-1-0-pro",
  input_data: {
    "prompt": "camera slowly zooms out, gentle wind moves the trees",
    "image_url": "<uploaded_image_url>",
    "duration": "5s"
  }
)

Video Parameters

ParamTypeOptionsNotes
promptstringrequiredDescribe the video
durationstring"5s", "10s"Video length
aspect_ratiostring"16:9", "9:16", "1:1"Frame ratio
seednumberany integerReproducibility
image_urlstringURLSource image for image-to-video

Audio Generation

CSM-1B (Conversational Speech)

Text-to-speech with natural, conversational quality.

generate(
  app_id: "fal-ai/csm-1b",
  input_data: {
    "text": "Hello, welcome to the demo. Let me show you how this works.",
    "speaker_id": 0
  }
)

ThinkSound (Video-to-Audio)

Generate matching audio from video content.

generate(
  app_id: "fal-ai/thinksound",
  input_data: {
    "video_url": "<video_url>",
    "prompt": "ambient forest sounds with birds chirping"
  }
)

ElevenLabs (via API, no MCP)

For professional voice synthesis, use ElevenLabs directly:

import os
import requests

resp = requests.post(
    "https://api.elevenlabs.io/v1/text-to-speech/<voice_id>",
    headers={
        "xi-api-key": os.environ["ELEVENLABS_API_KEY"],
        "Content-Type": "application/json"
    },
    json={
        "text": "Your text here",
        "model_id": "eleven_turbo_v2_5",
        "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
    }
)
with open("output.mp3", "wb") as f:
    f.write(resp.content)

VideoDB Generative Audio

If VideoDB is configured, use its generative audio:

# Voice generation
audio = coll.generate_voice(text="Your narration here", voice="alloy")

# Music generation
music = coll.generate_music(prompt="upbeat electronic background music", duration=30)

# Sound effects
sfx = coll.generate_sound_effect(prompt="thunder crack followed by rain")

Cost Estimation

Before generating, check estimated cost:

estimate_cost(
  estimate_type: "unit_price",
  endpoints: {
    "fal-ai/nano-banana-pro": {
      "unit_quantity": 1
    }
  }
)

Model Discovery

Find models for specific tasks:

search(query: "text to video")
find(endpoint_ids: ["fal-ai/seedance-1-0-pro"])
models()

Tips

  • Use seed for reproducible results when iterating on prompts
  • Start with lower-cost models (Nano Banana 2) for prompt iteration, then switch to Pro for finals
  • For video, keep prompts descriptive but concise — focus on motion and scene
  • Image-to-video produces more controlled results than pure text-to-video
  • Check estimate_cost before running expensive video generations

Related Skills

  • videodb — Video processing, editing, and streaming
  • video-editing — AI-powered video editing workflows
  • content-engine — Content creation for social platforms