comfyui-workflow-builder
mckruz/comfyui-expert
Generate valid ComfyUI workflow JSON from natural language descriptions.
What is comfyui-workflow-builder?
Translates natural language requests into executable ComfyUI workflow JSON with correct node graphs, connections, and model settings. Use this to design txt2img, img2img, inpainting, ControlNet, LoRA, upscaling, and video generation pipelines without manually writing JSON.
- Parses natural language intent into workflow requirements (output type, source material, quality level)
- Validates against local inventory to select compatible checkpoints, identity models, and ControlNet nodes
- Generates valid ComfyUI workflow JSON with correct class_types, node connections, and output indices
- Handles specialized pipelines: txt2img, img2img, inpainting, identity preservation (InstantID/PuLID), LoRA stacking, video generation (Wan/AnimateDiff), upscaling, and face detailing
- Estimates VRAM requirements and optimizes settings for available hardware
- Queues workflows to ComfyUI API or saves to local project directory
How to install comfyui-workflow-builder
npx skills add https://github.com/mckruz/comfyui-expert --skill comfyui-workflow-builder- ComfyUI instance running and accessible via COMFYUI_URL environment variable
- curl or wget installed for API communication
- Local inventory.json file populated with available checkpoints, models, and custom nodes
How to use comfyui-workflow-builder
- 1.Set COMFYUI_URL environment variable to your ComfyUI instance (e.g., http://localhost:8188)
- 2.Describe your desired workflow in natural language (e.g., 'txt2img with FLUX, 1024x1024, identity-preserved character using InstantID')
- 3.The skill parses your request and checks local inventory for available models and nodes
- 4.Selects the appropriate pipeline pattern and generates valid workflow JSON
- 5.Either queues the workflow to ComfyUI API for immediate execution or saves it to projects/{project}/workflows/ for later use
Use cases
- Generate a text-to-image workflow for FLUX with a specific prompt and resolution
- Create an identity-preserved character image using InstantID + IP-Adapter with reference photos
- Build a video generation pipeline using Wan I2V from a reference image with motion control
- Design an inpainting workflow to edit specific regions of an existing image
- Stack multiple LoRAs for a trained character with upscaling and face detailing post-processing
- ComfyUI users who want to generate workflows programmatically without manual JSON editing
- AI image/video creators building complex multi-node pipelines with identity preservation or ControlNet
- Developers integrating ComfyUI workflow generation into larger automation systems
- Anyone iterating on workflow designs and needing quick JSON generation from descriptions
comfyui-workflow-builder FAQ
It supports any checkpoint, LoRA, ControlNet, or custom node that exists in your local inventory.json. Common models include FLUX, SDXL, SD1.5, InstantID, PuLID, IP-Adapter, AnimateDiff, Wan I2V, and FaceDetailer. Check your inventory to see what's available.
Yes. It handles Image-to-Video pipelines using Wan (high-quality) or AnimateDiff (fast with motion control), as well as Talking Head workflows that combine video generation with lip-sync.
The skill validates against inventory before generating. If a required model or node is missing, it will report what's unavailable. You must install the missing custom nodes or download the models to ComfyUI first.
No. It only generates workflows from natural language. Installation, custom node development, Python scripting, model training, and hardware advice are out of scope.
It uses approximate VRAM costs for each component (e.g., FLUX FP16 = 16GB, InstantID = +4GB) and sums them based on your pipeline. This is an estimate; actual usage depends on batch size, resolution, and hardware.
Full instructions (SKILL.md)
Source of truth, from mckruz/comfyui-expert.
name: comfyui-workflow-builder description: Generate, build, create, or design ComfyUI workflow JSON from natural language descriptions. Produces valid node graphs with correct class_types, connections, output indices, and model-appropriate settings. Handles txt2img, img2img, inpainting, ControlNet, LoRA stacking, upscaling, and face detailing pipelines. Does NOT cover ComfyUI installation, custom node development, Python scripting, model training, hardware advice, or architectural explanations. user-invocable: true metadata: {"openclaw":{"emoji":"🔧","os":["darwin","linux","win32"],"requires":{"anyBins":["curl","wget"]},"primaryEnv":"COMFYUI_URL"}}
ComfyUI Workflow Builder
Translates natural language requests into executable ComfyUI workflow JSON. Always validates against inventory before generating.
Workflow Generation Process
Step 1: Understand the Request
Parse the user's intent into:
- Output type: Image, video, or audio
- Source material: Text-only, reference image(s), existing video
- Identity method: None, zero-shot (InstantID/PuLID), LoRA, Kontext
- Quality level: Draft (fast iteration) vs production (maximum quality)
- Special requirements: ControlNet, inpainting, upscaling, lip-sync
Step 2: Check Inventory
Read state/inventory.json to determine:
- Available checkpoints → select best match for task
- Available identity models → determine which methods are possible
- Available ControlNet models → enable pose/depth control if available
- Custom nodes installed → verify all required nodes exist
- VRAM available → optimize settings accordingly
Step 3: Select Pipeline Pattern
Based on request + inventory, choose from:
| Pattern | When | Key Nodes |
|---|---|---|
| Text-to-Image | Simple generation | Checkpoint → CLIP → KSampler → VAE |
| Identity-Preserved Image | Character consistency | + InstantID/PuLID/IP-Adapter |
| LoRA Character | Trained character | + LoRA Loader |
| Image-to-Video (Wan) | High-quality video | Diffusion Model → Wan I2V → Video Combine |
| Image-to-Video (AnimateDiff) | Fast video, motion control | + AnimateDiff Loader + Motion LoRAs |
| Talking Head | Character speaks | Image → Video → Voice → Lip-Sync |
| Upscale | Enhance resolution | Image → UltimateSDUpscale → Save |
| Inpainting | Edit regions | Image + Mask → Inpaint Model → KSampler |
Step 4: Generate Workflow JSON
ComfyUI workflow format:
{
"{node_id}": {
"class_type": "{NodeClassName}",
"inputs": {
"{param_name}": "{value}",
"{connected_param}": ["{source_node_id}", {output_index}]
}
}
}
Rules:
- Node IDs are strings (typically "1", "2", "3"...)
- Connected inputs use array format:
["source_node_id", output_index] - Output index is 0-based integer
- Filenames must match exactly what's in inventory
- Seed values: use random large integer or fixed for reproducibility
Step 5: Validate
Before presenting to user:
- Every
class_typeexists in inventory's node list - Every model filename exists in inventory's model list
- All required connections are present (no dangling inputs)
- VRAM estimate doesn't exceed available VRAM
- Resolution is compatible with chosen model (512 for SD1.5, 1024 for SDXL/FLUX)
Step 6: Output
If online mode: Queue via comfyui-api skill
If offline mode: Save JSON to projects/{project}/workflows/ with descriptive name
Workflow Templates
Basic Text-to-Image (FLUX)
{
"1": {
"class_type": "LoadCheckpoint",
"inputs": {"ckpt_name": "flux1-dev.safetensors"}
},
"2": {
"class_type": "CLIPTextEncode",
"inputs": {"text": "{positive_prompt}", "clip": ["1", 1]}
},
"3": {
"class_type": "CLIPTextEncode",
"inputs": {"text": "{negative_prompt}", "clip": ["1", 1]}
},
"4": {
"class_type": "EmptyLatentImage",
"inputs": {"width": 1024, "height": 1024, "batch_size": 1}
},
"5": {
"class_type": "KSampler",
"inputs": {
"seed": 42,
"steps": 25,
"cfg": 3.5,
"sampler_name": "euler",
"scheduler": "normal",
"denoise": 1.0,
"model": ["1", 0],
"positive": ["2", 0],
"negative": ["3", 0],
"latent_image": ["4", 0]
}
},
"6": {
"class_type": "VAEDecode",
"inputs": {"samples": ["5", 0], "vae": ["1", 2]}
},
"7": {
"class_type": "SaveImage",
"inputs": {"filename_prefix": "output", "images": ["6", 0]}
}
}
With Identity Preservation (InstantID + IP-Adapter)
Extends basic template by adding:
- Load reference image node
- InstantID Model Loader + Apply InstantID
- IPAdapter Unified Loader + Apply IPAdapter
- FaceDetailer post-processing
See references/workflows.md for complete node settings.
Video Generation (Wan I2V)
Uses different loader chain:
- Load Diffusion Model (not LoadCheckpoint)
- Wan I2V Conditioning
- EmptySD3LatentImage (with frame count)
- Video Combine (VHS)
See references/workflows.md Workflow 4 for complete settings.
VRAM Estimation
| Component | Approximate VRAM |
|---|---|
| FLUX FP16 | 16GB |
| FLUX FP8 | 8GB |
| SDXL | 6GB |
| SD1.5 | 4GB |
| InstantID | +4GB |
| IP-Adapter | +2GB |
| ControlNet (each) | +1.5GB |
| Wan 14B | 20GB |
| Wan 1.3B | 5GB |
| AnimateDiff | +3GB |
| FaceDetailer | +2GB |
Common Mistakes to Avoid
- Wrong output index: CheckpointLoader outputs
[model, clip, vae]at indices[0, 1, 2] - CFG too high for InstantID: Use 4-5, not default 7-8
- Wrong resolution for model: FLUX/SDXL=1024, SD1.5=512
- Missing VAE: FLUX needs explicit VAE (
ae.safetensors) - Wrong model in wrong loader: Diffusion models need
LoadDiffusionModel, notLoadCheckpoint
Reference Files
references/workflows.md- Detailed node-by-node templatesreferences/models.md- Model files and pathsreferences/prompt-templates.md- Model-specific promptsstate/inventory.json- Current inventory cache
Related skills
More from mckruz/comfyui-expert and the wider catalog.

comfyui-api
Connect to ComfyUI, queue workflows, monitor execution, and retrieve results in online or offline mode.

comfyui-prompt-engineer
Craft model-specific prompts optimized for the target checkpoint and identity method. Handles FLUX, SDXL, SD1.5, and Wan video models with proper syntax, quality tags, and negative prompts. Use when generating or refining prompts for ComfyUI workflows.

comfyui-video-pipeline
Generate videos using ComfyUI with Wan 2.2, FramePack, or AnimateDiff. Handles image-to-video, text-to-video, talking heads, and motion-controlled animation. Use when creating any video content from character images or text descriptions.

documentation
Structure and write technical documentation using the Diátaxis framework—tutorials, how-to guides, reference, and explanations.

fastify-best-practices
Build fast, type-safe Node.js REST APIs with Fastify best practices and patterns.

init
Creates, updates, or optimizes an AGENTS.md file for a repository with minimal, high-signal instructions covering non-discoverable coding conventions, tooling quirks, workflow preferences, and project-specific rules that agents cannot infer from reading the codebase. Use when setting up agent instructions or Claude configuration for a new repository, when an existing AGENTS.md is too long, generic, or stale, when agents repeatedly make avoidable mistakes, or when repository workflows have changed and the agent configuration needs pruning. Applies a discoverability filter—omitting anything Claude can learn from README, code, config, or directory structure—and a quality gate to verify each line remains accurate and operationally significant.