Skip to content
FunCoding

Search

Search docs, Skills and MCP

scenario-grok-imagine-video

Use when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for consistent people, products, or clothing, prompt-based editing, extending a clip from its last frame, native audio with lip-synced dialogue and sound effects, or choosing between first-frame and reference conditioning. Keywords: Grok Imagine Video 1.5, R2V, Grok Edit Video, Grok Extend Video, xAI, T2V, I2V, V2V.

AI 与智能体923skills/scenario-grok-imagine-video/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/scenario-labs/skills/scenario-grok-imagine-video/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Scenario Grok Imagine Video

Overview

Grok Imagine, xAI's video family on Scenario, splits its modes across single-purpose members: text and first-frame generation, reference-to-video, prompt editing, and clip extension each live in their own model, so picking the member is picking the mode. Discover them with search and treat model_schema_get as the contract: members agree on prompt style and disagree on every cap. The Grok Imagine image models belong to the scenario-grok-imagine-image skill.

Connection and the core loop: see the scenario skill in this repo; model-agnostic video work: the scenario-video skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

Pick the member by intent (input names from the live schema):

IntentMemberInputs
Text-to-videoVideo 1.5prompt alone; aspectRatio auto lands on 16:9
Animate a stillVideo 1.5image + prompt; auto follows the image ratio
Identity, opening freeVideo 1.5 R2VreferenceImages + prompt tagging @image1
Restyle or swap in placeEdit Videovideo + a prompt naming only the changes
Continue a clipExtend Videovideo + a continuation prompt; duration is new seconds

Caps are per member, so read them off model_schema_get. At authoring time the generation members took 1 to 15 seconds, 1 to 4 numOutputs, and eight aspectRatio values including auto; 1080p existed only on Video 1.5's text and first-frame modes, while R2V stopped at 720p and defaulted to 480p. The earlier unversioned Grok Imagine Video takes the same inputs as Video 1.5, also caps at 720p, and runs cheaper per dry_run. R2V took up to 7 referenceImages; Extend generated 2 to 10 new seconds; Edit exposed no duration or resolution and preprocessed its source down to 8.7 seconds at 720p, so trim to the segment first (see scenario-video). No seed, negative prompt, or camera parameter exists anywhere in the family.

First frame or references

image locks frame one: the video opens on that exact composition and animates out of it. referenceImages (its own member, R2V) carries people, products, and clothing across the shot without deciding how it opens. Tags bind by array order: @image1 is referenceImages[0] (<IMAGE_1> also works), and one clean subject per reference keeps control; a busy reference dilutes it. If the opening must match a composition exactly, render that still and pass it as image; when only a face, product, or outfit must stay consistent, use R2V and tag each reference where it acts.

Sound is prompted, not switched

Generation members produce native audio and no parameter controls it: the prompt is the whole mixing desk. Left unaddressed, the track tends toward generic background music, so end every prompt with named sounds ("Audio: rain on glass, distant traffic") or "no music". Dialogue written in quotes with a delivery verb lip-syncs: She says calmly: "We're live." Structure the rest like a director: scene, camera, style, motion, audio, in present tense, one continuous shot, one dominant camera move, physical actions instead of named emotions. 80 to 150 words touching at least three of those layers beat a one-line scene; editing prompts run shorter, naming only what changes. Extension prompts describe the next beat, not a new scene: open with a bridge ("the shot continues"), keep the established light and cast, and restate the audio.

Worked example: dialogue over an animated hero still

  1. search with target="models", query="grok imagine video", public=true. Match the member by name: first-frame animation wants e.g. model_xai-grok-imagine-video-1-5 (a live hit at authoring time: re-discover each session).
  2. model_schema_get with that id: fields, caps, and defaults before anything else.
  3. upload_asset the hero still (see the scenario skill) to get its asset id.
  4. model_run with that model_id, dry_run=true, and the exact parameters={"prompt": "The frame shows a knight on a cliff at dusk. She lowers her sword, turns to the camera, and says quietly: 'It ends tonight.' Slow dolly-in, wind lifting her cloak. Audio: wind, distant thunder, her line clear. No music.", "image": "asset_abc", "duration": 8, "resolution": "1080p"} for the cost estimate; re-estimate after any change to duration, resolution, or numOutputs.
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second model_run.
  6. asset_display the output and watch it with sound: the audio direction is judged by ear.

Common mistakes

  • Leaving audio to the default: generic music appears; direct the sound or state "no music" in every prompt.
  • Prompting an exact opening frame in R2V: references guide identity, not frame one; pass the still as image on Video 1.5.
  • Expecting 1080p everywhere: at authoring time it existed only on Video 1.5 text and first-frame runs, and R2V silently defaults to 480p.
  • Feeding Edit Video a long clip and expecting full-length output: the source is preprocessed down (8.7 seconds at authoring time); trim to the segment first.
  • Reading Extend's duration as total length: it counts only new footage.
  • Stacking camera moves or contradictory directions ("zoom in as the camera pulls back"): one dominant move per shot.

Similar Skills

brand-guidelines
anthropics/skills180k

brand-guidelines

Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.

AI & agents

internal-comms
anthropics/skills180k

internal-comms

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

AI & agents

template-skill
anthropics/skills180k

template-skill

Replace with description of the skill and when Claude should use it.

AI & agents

mcp-builder
anthropics/skills180k

mcp-builder

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

AI & agents

algorithmic-art
anthropics/skills180k

algorithmic-art

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

AI & agents

academy-guide
anthropics/skills180k

academy-guide

Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy (academy.claude.com), Anthropic's learning hub. Trigger on: "how do I", "how can I", "getting started with", "what can Claude do", "teach me", "learn to use"; questions about artifacts, projects, skills, plugins, connectors, MCP; requests about rolling Claude out to a team, class, or organization; and any ask for training materials, onboarding content, or learning resources. Use it when the user is learning how to use a feature or product — not when they are mid-task and just want the task done. This skill composes with other skills: after consulting product documentation to answer how a Claude feature works, also check here for a matching course or tutorial — a docs-grounded answer and an Academy recommendation belong together. Only recommend on a strong match; never invent Academy content.

AI & agents