Skip to content
FunCoding

Search

Search docs, Skills and MCP

flux-txt2img

Build Flux txt2img workflows with Flux.1 Dev (SRPO), Flux 2 Klein 9B, Turbo LoRAs, FluxGuidance, and DualCLIPLoader patterns

AI 与智能体797plugin/skills/flux-txt2img/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/artokun/comfyui-mcp/flux-txt2img/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Flux Text-to-Image Workflows

Overview

Flux is a guidance-distilled diffusion model family from Black Forest Labs. It uses a separate FluxGuidance node instead of KSampler CFG (which must always be 1.0). Three variants are available locally:

  1. Flux.1 Dev SRPO. Fine-tuned Flux.1 Dev with SRPO alignment. Uses DualCLIPLoader (T5XXL + CLIP-L). BF16 only.
  2. Flux 2 Klein 9B. Distilled Flux 2 variant. Uses single CLIPLoader (Qwen3-8B) + flux2-vae.safetensors. Fast 4-step generation.
  3. Flux 2 Turbo LoRA. Applied to Flux.1 Dev for 4-step generation.

Models

Flux.1 Dev SRPO

ComponentNodeModelNotes
UNETUNETLoaderflux.1-dev-SRPO-BFL-bf16.safetensors22.7GB, BF16 only — FP8 produces broken results
CLIPDualCLIPLoader (type=flux)clip_name1: t5xxl_fp8_e4m3fn.safetensors, clip_name2: clip_l.safetensorsT5XXL (4.7GB) + CLIP-L (235MB)
VAEVAELoaderae.safetensorsStandard Flux VAE (320MB). Z-Image uses the same VAE architecture but different weights — its VAE is a separate file (z-image-ae.safetensors), not this one

Flux 2 Klein 9B

ComponentNodeModelNotes
UNETUNETLoaderbigLove_klein1.safetensors17.3GB, Klein 9B variant
CLIPCLIPLoader (type=flux2)qwen_3_8b_fp8mixed.safetensorsQwen3-8B in text_encoders/ (8.3GB). Use flux2, NOT flux — both exist in the enum and flux fails at the sampler
VAEVAELoaderflux2-vae.safetensorsFlux 2 specific VAE (321MB)

Klein 9B vs Flux.1 Dev: Klein uses the Qwen3-8B text encoder (not T5XXL + CLIP-L). It has a different VAE (flux2-vae.safetensors). 9B distilled runs in 4 steps; 9B base needs ~50 steps at CFG 5.0. Fits in ~20GB VRAM with FP8.

Flux 2 Turbo LoRA (applied to Flux.1 Dev)

ComponentNodeModelNotes
LoRALoraLoaderModelOnlyflux2-turbo-lora.safetensors2.6GB, strength 1.0
Alt LoRALoraLoaderModelOnlyFlux2TurboComfyv2.safetensorsCommunity variant, same size

Conditioning

Provides separate prompt fields for each text encoder:

{
  "class_type": "CLIPTextEncodeFlux",
  "inputs": {
    "clip": ["<dual_clip>", 0],
    "clip_l": "short prompt for CLIP-L",
    "t5xxl": "detailed description for T5XXL",
    "guidance": 3.5
  }
}

clip_l captures key semantic features. t5xxl expands and refines descriptions. For simple use, put the same prompt in both fields. Guidance is built into this node, so no separate FluxGuidance is needed.

FluxGuidance (Alternative)

If using standard CLIPTextEncode instead of CLIPTextEncodeFlux, apply guidance separately:

{
  "class_type": "FluxGuidance",
  "inputs": {
    "conditioning": ["<clip_text_encode>", 0],
    "guidance": 3.5
  }
}

Guidance Values

ScenarioGuidanceNotes
Short prompts3.5–4.0Tighter prompt adherence
Long/complex prompts1.0–1.5More creative freedom
Realism2.5Less glossy skin, richer detail
Standard3.5Default for most use cases

Negative Conditioning

Flux does not support traditional negative prompts (guidance-distilled, CFG=1.0). Use ConditioningZeroOut:

{
  "class_type": "ConditioningZeroOut",
  "inputs": { "conditioning": ["<positive_cond>", 0] }
}

Or use an empty CLIPTextEncode for the negative input.

Sampler Settings

Flux.1 Dev SRPO

ParameterStandardNotes
steps20Range: 20–28
cfg1.0Always 1.0 — guidance is via FluxGuidance
sampler_nameipndmAuthor-recommended for SRPO
schedulerbetaAuthor-recommended for SRPO
guidance3.5Via CLIPTextEncodeFlux or FluxGuidance
denoise1.0

The SRPO author recommends the ipndm/beta combo. Standard Flux settings (euler/simple) also work, but ipndm/beta gives better results with this fine-tune.

Flux 2 Klein 9B (Distilled)

ParameterValueNotes
steps4Distilled model, 4 steps is optimal
cfg1.0Always 1.0
sampler_nameeuler
schedulersimple
denoise1.0

Flux 2 Klein 9B (Base/Undistilled)

ParameterValueNotes
steps50Full quality
cfg5.0Higher CFG for base model
sampler_nameeuler
schedulersimple

Flux.1 Dev + Turbo LoRA

ParameterValueNotes
steps4Turbo-distilled
cfg1.0
sampler_nameeuler
schedulersimple
lora_strength1.0

Resolutions

AspectResolutionMegapixels
Square1024x10241.0MP
Portrait 3:4896x11521.0MP
Landscape 4:31152x8961.0MP
Landscape 16:91344x7681.0MP
Portrait 9:16768x13441.0MP

Flux operates at ~1 megapixel natively. Dimensions should be multiples of 8.

Prompt Style

Natural language descriptions. No quality tags needed (unlike SDXL/Illustrious). Detailed, descriptive prompts work best.

Good: "A young woman with auburn hair sits at a sunlit cafe in Paris, wearing a cream linen blazer, soft bokeh background, shot on Sony A7III 85mm f/1.4"
Bad: "masterpiece, best quality, 1girl, cafe, paris"

Complete Workflow: Flux.1 Dev SRPO

{
  "1": { "class_type": "UNETLoader", "inputs": { "unet_name": "flux.1-dev-SRPO-BFL-bf16.safetensors", "weight_dtype": "default" }},
  "2": { "class_type": "DualCLIPLoader", "inputs": { "clip_name1": "t5xxl_fp8_e4m3fn.safetensors", "clip_name2": "clip_l.safetensors", "type": "flux" }},
  "3": { "class_type": "VAELoader", "inputs": { "vae_name": "ae.safetensors" }},
  "4": { "class_type": "CLIPTextEncodeFlux", "inputs": {
    "clip": ["2", 0],
    "clip_l": "<short prompt>",
    "t5xxl": "<detailed prompt>",
    "guidance": 3.5
  }},
  "5": { "class_type": "ConditioningZeroOut", "inputs": { "conditioning": ["4", 0] }},
  "6": { "class_type": "EmptyLatentImage", "inputs": { "width": 896, "height": 1152, "batch_size": 1 }},
  "7": { "class_type": "KSampler", "inputs": {
    "model": ["1", 0],
    "positive": ["4", 0],
    "negative": ["5", 0],
    "latent_image": ["6", 0],
    "seed": 42, "steps": 20, "cfg": 1, "sampler_name": "ipndm", "scheduler": "beta", "denoise": 1
  }},
  "8": { "class_type": "VAEDecode", "inputs": { "samples": ["7", 0], "vae": ["3", 0] }},
  "9": { "class_type": "SaveImage", "inputs": { "images": ["8", 0], "filename_prefix": "flux_srpo" }}
}

Complete Workflow: Flux 2 Klein 9B (Distilled, 4-Step)

{
  "1": { "class_type": "UNETLoader", "inputs": { "unet_name": "bigLove_klein1.safetensors", "weight_dtype": "default" }},
  "2": { "class_type": "CLIPLoader", "inputs": { "clip_name": "qwen_3_8b_fp8mixed.safetensors", "type": "flux2" }},
  "3": { "class_type": "VAELoader", "inputs": { "vae_name": "flux2-vae.safetensors" }},
  "4": { "class_type": "CLIPTextEncode", "inputs": { "clip": ["2", 0], "text": "<prompt>" }},
  "5": { "class_type": "ConditioningZeroOut", "inputs": { "conditioning": ["4", 0] }},
  "6": { "class_type": "EmptyFlux2LatentImage", "inputs": { "width": 1024, "height": 1024, "batch_size": 1 }},
  "7": { "class_type": "KSampler", "inputs": {
    "model": ["1", 0],
    "positive": ["4", 0],
    "negative": ["5", 0],
    "latent_image": ["6", 0],
    "seed": 42, "steps": 4, "cfg": 1, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
  }},
  "8": { "class_type": "VAEDecode", "inputs": { "samples": ["7", 0], "vae": ["3", 0] }},
  "9": { "class_type": "SaveImage", "inputs": { "images": ["8", 0], "filename_prefix": "flux_klein" }}
}

Klein uses a single CLIPLoader (not DualCLIPLoader) with type: "flux2" and the Qwen3-8B text encoder from text_encoders/. The CLIP loader path resolves from models/text_encoders/.

Flux-2-specific gotchas, all of which fail at the KSampler rather than the loader, so the error points at the wrong node:

  • type must be flux2, not flux. Both values exist in the CLIPLoader enum, so flux loads without complaint and then dies during sampling.
  • Use EmptyFlux2LatentImage, not EmptyLatentImage. Flux 2 uses a different latent channel count.
  • Klein 9B pairs with the Qwen3-8B encoder (qwen_3_8b* from Comfy-Org/vae-text-encorder-for-flux-klein-9b). The similarly-named qwen_3_4b ships in the klein-4b repo and is for the 4B model. Mismatching them raises mat1 and mat2 shapes cannot be multiplied (512x7680 and 12288x4096), where 7680 = 2560x3 (4B hidden size) and 12288 = 4096x3 (8B). It reads as a confusing CLIP error rather than a wrong-file error.

Complete Workflow: Flux.1 Dev + Turbo LoRA (4-Step)

{
  "1": { "class_type": "UNETLoader", "inputs": { "unet_name": "flux.1-dev-SRPO-BFL-bf16.safetensors", "weight_dtype": "default" }},
  "2": { "class_type": "LoraLoaderModelOnly", "inputs": { "model": ["1", 0], "lora_name": "flux2-turbo-lora.safetensors", "strength_model": 1.0 }},
  "3": { "class_type": "DualCLIPLoader", "inputs": { "clip_name1": "t5xxl_fp8_e4m3fn.safetensors", "clip_name2": "clip_l.safetensors", "type": "flux" }},
  "4": { "class_type": "VAELoader", "inputs": { "vae_name": "ae.safetensors" }},
  "5": { "class_type": "CLIPTextEncodeFlux", "inputs": {
    "clip": ["3", 0],
    "clip_l": "<short prompt>",
    "t5xxl": "<detailed prompt>",
    "guidance": 3.5
  }},
  "6": { "class_type": "ConditioningZeroOut", "inputs": { "conditioning": ["5", 0] }},
  "7": { "class_type": "EmptyLatentImage", "inputs": { "width": 1024, "height": 1024, "batch_size": 1 }},
  "8": { "class_type": "KSampler", "inputs": {
    "model": ["2", 0],
    "positive": ["5", 0],
    "negative": ["6", 0],
    "latent_image": ["7", 0],
    "seed": 42, "steps": 4, "cfg": 1, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
  }},
  "9": { "class_type": "VAEDecode", "inputs": { "samples": ["8", 0], "vae": ["4", 0] }},
  "10": { "class_type": "SaveImage", "inputs": { "images": ["9", 0], "filename_prefix": "flux_turbo" }}
}

LoRA Support

Custom LoRAs (jellyfish, etc.)

Apply Flux LoRAs with LoraLoaderModelOnly between UNET and KSampler:

{
  "class_type": "LoraLoaderModelOnly",
  "inputs": {
    "model": ["<unet_or_previous_lora>", 0],
    "lora_name": "<lora_file>.safetensors",
    "strength_model": 1.0
  }
}

Klein LoRAs

Klein 9B LoRAs go in the loras/Flux.2 Klein 9B/ subfolder:

  • klein_slider_detail.safetensors, a detail slider LoRA

VRAM Considerations

ModelVRAMNotes
SRPO BF16 + DualCLIP~24GBFills RTX 4090 exactly. Must use BF16 — FP8 is broken for SRPO
Klein 9B FP8 + Qwen3-8B~20GBFits comfortably on 4090
SRPO + Turbo LoRA~24GBSame as SRPO base
  • Always clear_vram before switching to Flux from another model family
  • T5XXL is the main VRAM consumer alongside the UNET; both stay loaded during sampling
  • CLIP-L is small (235MB) and negligible

Tips

  1. KSampler CFG must always be 1.0. All guidance is through CLIPTextEncodeFlux or FluxGuidance
  2. SRPO requires BF16. The FP8 quantization is known to produce broken results with this fine-tune
  3. For short prompts (1 to 2 sentences), increase guidance to 3.5 to 4.0. For long prompts (paragraph), decrease to 1.0 to 1.5
  4. Flux generates excellent text in images. Put text to render in quotes within your prompt
  5. Klein 9B is the fastest option at 4 steps. Use it for rapid iteration, then switch to SRPO for final quality

Sources

  • Official: none found.
  • Empirical: sampler values, wiring, and prompt notes from working graphs in packs/ and observed renders; not a vendor prompting guide.

Similar Skills

brand-guidelines
anthropics/skills180k

brand-guidelines

Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.

AI & agents

internal-comms
anthropics/skills180k

internal-comms

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

AI & agents

template-skill
anthropics/skills180k

template-skill

Replace with description of the skill and when Claude should use it.

AI & agents

mcp-builder
anthropics/skills180k

mcp-builder

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

AI & agents

algorithmic-art
anthropics/skills180k

algorithmic-art

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

AI & agents

academy-guide
anthropics/skills180k

academy-guide

Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy (academy.claude.com), Anthropic's learning hub. Trigger on: "how do I", "how can I", "getting started with", "what can Claude do", "teach me", "learn to use"; questions about artifacts, projects, skills, plugins, connectors, MCP; requests about rolling Claude out to a team, class, or organization; and any ask for training materials, onboarding content, or learning resources. Use it when the user is learning how to use a feature or product — not when they are mid-task and just want the task done. This skill composes with other skills: after consulting product documentation to answer how a Claude feature works, also check here for a matching course or tutorial — a docs-grounded answer and an Academy recommendation belong together. Only recommend on a strong match; never invent Academy content.

AI & agents