跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

generate2dmedia

The art-route layer of the Forge skills. route_media.py is the one command that generates an image or an image-to-video clip - through the paid API of any provider whose key is configured (OpenAI, Google Gemini, xAI, BytePlus ModelArk, fal.ai; the configured key is the owner's consent), else through the user's own signed-in Codex or Grok CLI - and prints one JSON line with the route, artifact, sha256 and cost estimate; with no route it answers no-route (codeart2d is then the last resort). Also the capability check (forge_doctor), API key configuration, the spend ledger, resume and Codex image adoption. Use when a Forge skill sends a generation here, for the capability check, or when the user asks about keys, routes or spend. Not for planning, processing or QA of sprites, video frames or maps (generate2dsprite, video2dsprite, generate2dmap) or for code-drawn art (codeart2d).

测试4.4kskills/generate2dmedia/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/0x0funky/agent-sprite-forge/generate2dmedia/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Generate 2D Media

The route layer. It makes raw media; the owning skill plans the asset before and checks it after.

Use when

  • Any Forge skill (or its pipeline script) needs a generated image, reference edit or image-to-video clip.
  • Once per session, to see which routes work here (forge_doctor.py).
  • The user asks to configure an API key, to use a particular route, or about spend.

Do not use when

Route order (owner decision 2026-10-06)

  1. API, every provider whose key is configured, in this order (the user's providers.order goes first). Images: OpenAI (gpt-image-2.5-sunburst), Google Gemini (gemini-3.1-flash-image; --tier hero takes gemini-3-pro-image for hero masters), xAI (grok-imagine-image-2.0), BytePlus ModelArk (Seedream 5.0), fal.ai (reference edits). Video: xAI (grok-imagine-video-1.5; --tier draft takes 1.5 Lite), BytePlus ModelArk (Seedance 2.0), fal.ai (Kling v3, Veo 3.1, Luma, MiniMax, Wan, Vidu, LTX). A configured key is the owner's consent: no per-call question.
  2. Local: your own image tool if you have one (Codex image_gen, Grok's native tools), otherwise the user's signed-in CLI: Codex (local CLI) image_gen with references attached, then Grok (local CLI) one-shot image or edit; video Grok (local CLI) in ACP mode. An installed CLI is used as is; its first successful run records the VERIFIED proof.
  3. codeart2d, only when the user explicitly asks for code-drawn art or route_media.py prints no-route (exit 3).

Always name the route that ran ("OpenAI API", "Google Gemini API", "xAI API", "BytePlus ModelArk API", "fal.ai API", "Codex (local CLI)", "Grok (local CLI)"). Details: route-media.md.

Capability check (once per session)

Run python "<skill-dir>/scripts/forge_doctor.py" --host-tools <tools> --save <output>/doctor.json, where <tools> lists the media tools in your own tool list (image_gen, image_edit, image_to_video) or none. Its ROUTES block shows which API keys are configured (yes or no, never a key), the local readiness of each CLI and the resolved order (routeOrder). If encoding.stdout fails, set PYTHONUTF8=1.

Generate (route_media.py)

NeedCommand
Which route would runpython "<skill-dir>/scripts/route_media.py" resolve --kind image (or video; --references N)
One image, still or editpython "<skill-dir>/scripts/route_media.py" image --prompt-file <txt> --reference <png> --size 1024x1024 --out-dir <new> (--reference repeatable, in the order the prompt names them)
Animate an approved stillpython "<skill-dir>/scripts/route_media.py" video --prompt-file <txt> --reference <still.png> --last-frame <still.png> --duration 6 --resolution 720p --out-dir <new>
A route or model of the user's choiceadd --route api, local, openai, gemini, xai, byteplus, fal, fal:<endpoint-id>, codex-cli, grok-cli or grok-acp; --model <id>, --tier draft or hero, --provider-option [PROVIDER:]KEY=VALUE
The plan and estimate firstadd --dry-run (no key is used, nothing is sent or written)

Success prints one ASCII JSON line and exits 0: {"status":"ok","route":"api:openai","artifact":"<out-dir>/generated.png","sha256":...,"estimateUsd":...} plus model, job and, for video, lastFrameUsed, durationRequested and durationUsed (a length is snapped to the nearest one the model renders; Grok (local CLI) renders 6 or 10 s and never pins a last frame) and crop when the frames were padded to 16:9 for a model without 1:1. Every clip is published without audio. No route prints {"status":"no-route","fallback":"codeart2d"} and exits 3. A failure prints one error: <route>: <code>: <message> line and exits 1. A route whose account cannot serve the request (no credit, no access, not signed in) passes it to the next route; a refusal such as moderation stops. Write prompts yourself into UTF-8 files; local routes take the size and every instruction from the prompt.

Keys

Each provider reads only its own key: OPENAI_API_KEY, GOOGLE_API_KEY or GEMINI_API_KEY (GOOGLE_API_KEY wins), XAI_API_KEY, ARK_API_KEY (BytePlus ModelArk) and FAL_KEY (fal.ai), from the environment or from the user config file, never from a project: %APPDATA%\agent-sprite-forge\config.json on Windows, ~/.config/agent-sprite-forge/config.json elsewhere ($XDG_CONFIG_HOME when set), as {"OPENAI_API_KEY": "...", "FAL_KEY": "..."} with optional "models" and "providers": {"order": [...]}. Never print, log, commit or paste a key, and never put one in a prompt. A Codex or Grok sign-in is never an API key.

Spend and the ledger

Every call, paid or local, is a line in <project>/.forge/ledger.jsonl with its estimate (estimateUsd is null when no verified price exists; say "price unknown"). There is no cap by default; --budget-usd, --max-calls, FORGE_MAX_PAID_REQUESTS and FORGE_SESSION_IMAGES / FORGE_SESSION_VIDEOS set one when the user asks. python "<skill-dir>/scripts/media_ledger.py" summary shows spend; media_ledger.py settle <id> --status ... closes an unknown outcome.

Lower-level tools

NeedCommand
One API request by hand (dry run until --execute)generate_media.py image or video with --provider, --model, --prompt-file, --out-dir (api-usage.md)
Poll or download a video or fal.ai job without paying againgenerate_media.py resume --job <out>/job.json
One local CLI run by handcli_media.py image --route codex-cli --prompt-file <txt> --output-dir <new> --execute (cli-routes.md)
A local run timed out or was interruptedcli_media.py resume --run <id>, then --adopt (never reruns the CLI)
Codex made an image that is not in the projectcli_media.py adopt --codex-thread <thread-id> --output-dir <new>
Many jobsgenerate_media.py batch or cli_media.py batch
Verify a local route ahead of time (optional)forge_doctor.py --verify-route <route> --execute (one quota call)

Host notes

  • Codex: image_gen is your own tool: with no configured key use it and adopt the result (cli_media.py adopt); scripts reach Codex (local CLI) through route_media.py. This skill is explicit-only in Codex.
  • Claude Code: no media tools; route_media.py takes the API or the local CLIs. <skill-dir> is ${CLAUDE_SKILL_DIR}; look at results with Read.
  • Grok: its native media tools are your own tools; the same grok executable is the "Grok (local CLI)" route for other hosts. Agent setups: agent-profiles/README.md.

Commands and outputs

Run each tool as one line from the user's project root: python "<skill-dir>/scripts/<tool>.py" .... Outputs go to a new --out-dir inside the project. Results are raw media: hand them to the owning skill for identity, size, alpha, motion and seam checks. Keep prompts, input hashes, job ids and result hashes; report the model requested and the model returned separately. Providers: provider-survey.md (last verified 2026-10-06, with the shutdown dates); capabilities: references/capabilities.json (python "<skill-dir>/scripts/media_providers.py" list); prices: references/prices.json; data contracts: references/schemas/media.schema.json.

相似的 Skill

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

测试

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

测试

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

测试

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

测试

ci-cd-and-automation
addyosmani/agent-skills103k

ci-cd-and-automation

Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.

测试

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

测试