跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

图像与视频 Skill

「图像与视频」分类共 451 个 Skill,按仓库 star 排序。分类自动生成,仅供参考。

backlog-capture
mvschwarz/openrig6.5k

backlog-capture

Quickly capture product ideas, feature requests, or insights from meetings and conversations. Rapid documentation with smart categorization and deduplication.

session-source-fork
mvschwarz/openrig6.5k

session-source-fork

Use when authoring a rig spec member or `rig expand` payload that needs to start a new managed seat from a prior runtime conversation source — `session_source: { mode: fork, ref: { kind, value } }`. v1 supports `mode: fork` with `ref.kind: native_id` for Claude and Codex. The new seat persists a NEW post-fork token; the parent token is NEVER written onto the new seat. NOT for restoring an existing seat or for artifact-backed mental-model rebuild.

algorithmic-art
ThinkInAIXYZ/deepchat6.4k

algorithmic-art

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

deepchat-cli
ThinkInAIXYZ/deepchat6.4k

deepchat-cli

Use DeepChat's bundled CLI control plane for model inference, image/video/speech generation, transcription, OCR, artifact inspection, public configuration, Skills, and MCP operations. Activate when a user asks to invoke DeepChat capabilities that are not already exposed as a more specific tool, compare models, run a benchmark, inspect DeepChat runtime state, or manage DeepChat through the CLI.

airbnb-listing-detail
browser-act/skills6.1k

airbnb-listing-detail

Fetches complete Airbnb listing details for a given numeric listing ID via the internal GraphQL API, returning title, room type, description, amenities, photos, coordinates, city, house rules, highlights, ratings, review count, bedroom configuration, and property overview. Use when user mentions Airbnb listing details, Airbnb property info, Airbnb room details, get Airbnb listing data, Airbnb amenities list, Airbnb house rules, Airbnb property description, Airbnb detail page scraper, Airbnb rooms detail, Airbnb property page data, Airbnb listing info, fetch Airbnb room details, pull Airbnb listing.

airbnb-search-listing
browser-act/skills6.1k

airbnb-search-listing

Extracts Airbnb accommodation search results from a destination query via SSR-embedded data, returning listing ID, URL, name, coordinates, rating, price, photos, and badge info for each result, plus pagination cursors for multi-page retrieval. Use when user mentions Airbnb search results, Airbnb listings, vacation rental search, short-term rental listings, scrape Airbnb, get Airbnb data, find rentals on Airbnb, Airbnb destination search, Airbnb property list, Airbnb stays search, Airbnb accommodation results, pull Airbnb listings, collect Airbnb search data, Airbnb scraper, Airbnb search page extraction, Airbnb search by destination.

sn-da-image-caption
OpenSenseNova/SenseNova-Skills5.7k

sn-da-image-caption

图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py 对图表/表格/截图/流程图进行 caption,(2) 将 caption 文本解析为结构化 DataFrame,(3) 基于提取数据重新生成可视化图表,(4) 导出为 Excel/CSV。**遇到以下任一情况就主动使用本 skill,不要自行猜测图片内容**:①用户出现触发词:图片分析 / 图表提取 / 表格识别 / OCR / 图片描述 / 截图分析 / 图表数据 / 提取图片中的数据 / 图片转表格 / 识别图片 / image caption / extract data from image / chart analysis / table OCR;②用户上传或指定了图片文件(.png / .jpg / .jpeg / .gif / .webp / .bmp)并要求理解、提取数据或分析内容;③任务需要从图表截图、表格截图、UI 截图、流程图中提取结构化信息;④用户要求将图片中的数据转为 Excel/CSV 或重新生成可视化图表。仅不用于:图片编辑(裁剪、滤镜、缩放)、图片生成、不含数据的风景/人物照片描述。

sn-image-base
OpenSenseNova/SenseNova-Skills5.7k

sn-image-base

Base-layer skill for the SenseNova-Skills project, providing low-level APIs for image generation, recognition (VLM), and text optimization (LLM). This skill does not preprocess inputs; it only calls backend services and returns results. This skill is not user-facing and is intended for upper-layer skills only.

sn-image-doctor
OpenSenseNova/SenseNova-Skills5.7k

sn-image-doctor

Environment diagnostic skill for SenseNova-Skills project. Checks that sn-image-base is properly installed and configured, validates dependencies and environment variables. Prompts user to configure missing required variables and saves them to .env file. After configuration, reloads environment and suggests agent restart if needed.

sn-image-imitate
OpenSenseNova/SenseNova-Skills5.7k

sn-image-imitate

Generates a new image that imitates the style of a reference image while updating content based on user intent. Uses a three-stage pipeline: image annotation (long caption), caption rewriting, and image generation. Use when user asks to "imitate style", "保持这个风格重画", "按这张图风格生成", or "style transfer with new content".

sn-image-resume
OpenSenseNova/SenseNova-Skills5.7k

sn-image-resume

Generates a designed portfolio-resume image from resume content provided in conversation text. Extracts optional style instructions, converts the resume into a fixed portfolio-resume layout prompt, and generates the final image through sn-image-base. Use when user asks to create "resume image", "portfolio resume", "简历图", "简历海报", or "个人简历视觉设计".

sn-ppt-tools
OpenSenseNova/SenseNova-Skills5.7k

sn-ppt-tools

Use when another PPT Skill needs web search, image search or download, or image generation and the host Agent's equivalent native capability is absent or has failed.

sn-search-image
OpenSenseNova/SenseNova-Skills5.7k

sn-search-image

USE FOR Google-backed image discovery via Serper.dev. Returns image URLs, page URLs, titles, and source domains.

get-prompt-from-image
wuyoscar/GPT-Image2-Skill5.7k

get-prompt-from-image

Analyze user-provided reference images and reverse-engineer high-fidelity AI image-generation prompts. Use when the user asks to recreate, imitate, reverse-engineer, or extract prompts from photographs, illustrations, 3D renders, products, characters, landscapes, typography, logos, posters, or other visual references. Do not use for requests that only require OCR or an ordinary image description.

gpt-image
wuyoscar/GPT-Image2-Skill5.7k

gpt-image

Generate or edit images with GPT Image 2 or 2.5 through the packaged CLI and Reference Gallery. Use for image requests including imprecise 'GPT 2.5' model names, posters, typography, reference edits, and inpainting; resolve the model choice before generation.

mirrord-quickstart
metalbear-co/mirrord5.4k

mirrord-quickstart

Guide users from zero to their first working mirrord session. Use when a user is new to mirrord, wants to install it, or needs help running their first session connecting to a Kubernetes cluster.

bb-methodology
elementalsouls/Claude-BugHunter4.8k

bb-methodology

Use at the START of any bug bounty hunting session, when switching targets, or when feeling lost about what to do next. Master orchestrator that combines the 5-phase non-linear hunting workflow with the critical thinking framework (developer psychology, anomaly detection, What-If experiments). Routes to all other skills based on current hunting phase. Also use when asking "what should I do next" or "where am I in the process."

dev-guide-generator
zebbern/claude-code-guide4.7k

dev-guide-generator

Generates complete technical tutorials from prerequisites and environment setup to core steps, troubleshooting, and a final cheatsheet. Trigger on requests to write a tutorial, create a setup guide, organize steps for beginners, or keywords like step-by-step, quickstart, or how-to guide.

Banana Pro Image Generation
xianyu110/awesome-openclaw-tutorial4.6k

Banana Pro Image Generation

使用 Gemini 3 Pro Image 生成图片,支持白板图、Logo设计、社交媒体配图等多种场景

consolidate-notes
inkeep/open-knowledge4.5k

consolidate-notes

Promote existing research into a stable-status canonical article under `articles/` in a Knowledge Base project (the `knowledge-base` starter pack). Read when a decision has actually been made and the team wants the source-of-truth written down, or when asked to consolidate, canonicalize, promote research, or supersede an older article. Carries the decision-confirmation gate, the `supersedes:` chain that keeps the evidence trail intact, and the canonical voice. Does not conduct new research — that is the sibling `research-with-sources` skill.

open-knowledge-discovery
inkeep/open-knowledge4.5k

open-knowledge-discovery

Read when the user asks what OpenKnowledge is, wants to install it on a repository, wants to open or preview a single markdown file that is not part of an OpenKnowledge project, wants to share an OpenKnowledge project with collaborators, asks whether OpenKnowledge supports a particular capability, or asks how `ok init` / `ok cowork` / OK Desktop set up a project. Do NOT load to perform OpenKnowledge reads/writes — the runtime guidance for editing markdown inside an initialized OK project ships as a separate project-local skill installed into each detected agent's skills dir (for example `.claude/skills/open-knowledge/`) whenever `ok init` runs.

generate2dmap
0x0funky/agent-sprite-forge4.4k

generate2dmap

Plan and build 2D game maps and scenes - top-down and side-scrolling levels, tilemaps, prop packs, parallax backgrounds and HD-2D battle or story plates - with collision, navigation and reachability checks, environment motion, a playable HTML preview and export to Tiled, Godot 4 or LDtk. Map art (terrain, tiles, props, plates, parallax layers) is image generation by default, through generate2dmedia route_media.py (the API when a key is configured, else the local Codex or Grok CLI); map data (layout, collision, exits, spawns) is written as data. Use for any map, level, room, scene, background, terrain or prop-kit request. Not for characters, creatures, items or FX (generate2dsprite), character animation (video2dsprite), code-drawn tiles or maps (codeart2d, only on request) or calling an image or video API by itself (generate2dmedia).

video2dsprite
0x0funky/agent-sprite-forge4.4k

video2dsprite

Turn one approved master still into a character's whole animated sprite set - one image-to-video clip per action (idle, walk, run, attack, jump, hurt, cast and so on), all from that still, generated through generate2dmedia route_media.py (the xAI API when a key is configured, else the local Grok CLI) - then gate each take, key it with a soft matte, register it to the master, pick loops or retime actions, finish (HD by default, pixel on request) and package fixed-canvas frames, alpha WebM, packed MP4 for iPhone and a PNG fallback with engine metadata and a JS runtime. Also processes a supplied clip. Use for character and creature animation, fluid motion and any video-to-sprite work. Not for FX, icon or prop sheets (generate2dsprite), maps or plate loops (generate2dmap), code-drawn animation (codeart2d, only on request) or calling a video API by itself (generate2dmedia).