设计与多媒体

视频生成 Video Generate

根据提示词自动挑选视频模型,支持文生视频、图生视频、首尾帧、数字人与对口型。

它能做什么

作为 dLazy 视频 CLI 的客户端运行。根据提示词从约 17 个视频模型(即梦、可灵、Veo、Seedance、万相、PixVerse、Vidu、Happy Horse,以及数字人和对口型模型)中自动挑选合适的一个,并提交生成任务。覆盖的生成方式包括文生视频、图生视频、首尾帧视频、参考图视频、数字人视频、口型对齐。使用前需要 dLazy API key,可通过 `dlazy login` 设备码流程一次性登录,或手动执行 `dlazy auth set`。生成结果与上传素材均托管在 dLazy 云存储上。

什么时候用它

  • 用一段文字描述生成短视频
  • 把静态图片变成视频
  • 用首尾帧生成转场视频
  • 把新音频对到已有视频的口型上

技能文档

视频生成 Video Generate

English · 中文

Video generation skill. Automatically selects the best dlazy CLI video model based on the prompt.

中文说明 / Chinese Overview

视频生成技能:根据提示词自动选择最佳的视频生成模型并产出视频。支持文生视频、图生视频、图片转视频、首尾帧生成视频、参考图生成视频、数字人视频、视频对口型。

中文触发关键词:视频生成、AI 视频生成、文字生成视频、文生视频、图生视频、图片变视频、让图片动起来、一键生成视频、短视频生成、视频模型、即梦视频、可灵视频、Veo 视频、Seedance 视频、通义万相视频、PixVerse 视频、数字人视频、对口型、唇形同步。

Trigger Keywords

  • generate video
  • text to video
  • animate image

Authentication

All requests require a dLazy API key. The recommended way to authenticate is dlazy login:

dlazy login

This runs a device-code flow (also works in remote shells) and automatically saves your API key to the local CLI config — no manual copy/paste required.

Alternative: Set the Key Manually

If you already have an API key, you can save it directly:

dlazy auth set YOUR_API_KEY

The CLI saves the key in your user config directory (~/.dlazy/config.json on macOS/Linux, %USERPROFILE%\.dlazy\config.json on Windows), with file permissions restricted to your OS user account. You can also supply the key per-invocation via the DLAZY_API_KEY environment variable.

Getting Your API Key Manually

  1. Sign in or create an account at dlazy.com
  2. Go to dlazy.com/dashboard/organization/api-key
  3. Copy the key shown in the API Key section

Each key is scoped to your dLazy organization and can be rotated or revoked at any time from the same dashboard.

About & Provenance

You can install on demand without persisting a global binary by running:

npx @dlazy/[email protected] 

Or, if you prefer a global install, the skill's metadata.clawdbot.install field declares the exact pinned version (npm install -g @dlazy/[email protected]). Review the GitHub source before installing.

How It Works

This skill is a thin client over the dLazy hosted API. When you invoke it:

  • Prompts and parameters you provide are sent to the dLazy API endpoint (api.dlazy.com) for inference.
  • Any local file paths you pass to image / video / audio fields are uploaded to dLazy's media storage (files.dlazy.com) so the model can read them — the same flow as any cloud-based generation API.
  • Generated output URLs returned by the API are hosted on files.dlazy.com.

This is the standard SaaS pattern; the skill itself does not access network or filesystem resources beyond what the dLazy CLI already handles. See dlazy.com for the full service terms.

Piping Between Commands

Every dlazy invocation prints a JSON envelope on stdout. Any flag value can be a pipe reference that pulls from the upstream command's envelope, so you can chain steps without copying URLs by hand.

ReferenceResolves to
-Upstream's natural value for this field (scalar or array)
@NThe N-th output's primary value (e.g. @0 = first output url)
@N.Drill into the N-th output (@0.url, @1.meta.fps)
@*All outputs' primary values as an array
@stdinThe whole upstream JSON envelope
@stdin:Jsonpath into the whole envelope (@stdin:result.outputs[0].url)

Examples

# Generate an image and feed its url straight into image-to-video
dlazy seedream-4.5 --prompt "a red fox in snow" \
  | dlazy kling-v3 --image - --prompt "fox starts running"

# Generate an image, then add TTS narration over a still
dlazy seedream-4.5 --prompt "lighthouse at dawn" \
  | dlazy keling-tts --text "Welcome to the coast." --image @0.url

# Fan-out: pass every upstream output url into a batch step
dlazy seedream-4.5 --prompt "city skyline" --n 4 \
  | dlazy superres --images @*

Required flags can be entirely sourced from the pipe — --field - satisfies the requirement when an upstream value exists. If stdin is empty, the CLI fails with code: "no_stdin".

Usage

This skill handles all video generation requests by selecting the best dlazy video model.

Available Video Models

  • dlazy happyhorse-1.0: Happy Horse 1.0 video model — one model covers text-to-video (t2v), first-frame-to-video (i2v), reference-to-video (r2v), and video editing (edit). The selected mode is automatically routed to the matching sub-model.
  • dlazy heygen-lipsync-speed: HeyGen Lipsync Speed: Fast lip-sync model, ideal for scenarios requiring rapid generation.
  • dlazy jimeng-dream-actor: Jimeng character/action-driven video model, supports reference image and video input, suitable for character acting, action transfer, and style-consistent generation.
  • dlazy jimeng-i2v-first: Jimeng first-frame-to-video model, uses first frame + text to generate video. Suitable for single-shot scenes that naturally animate static images.
  • dlazy jimeng-i2v-first-tail: Jimeng first/last-frame video model; constrains shot start/end frames. Good for transitions and clearly resolved action.
  • dlazy jimeng-omnihuman-1.5: Jimeng digital human model: combines any-ratio character/subject image with audio to generate high-quality digital human videos.
  • dlazy kling-v3: Kling V3 general video model, supports text + up to 4 reference images, suitable for stable short video clips and daily creative workflows.
  • dlazy kling-v3-omni: Kling Omni video model, supports multiple reference images, duration, mode (std/pro), and optional audio. Suitable for highly controlled video synthesis tasks.
  • dlazy pixverse-c1: PixVerse C1 video model (strong on action, VFX, and high-motion scenes) — one model covers text-to-video, image-to-video, first/last-frame-to-video, and reference-to-video: t2v when no images, i2v with first frame only, kf2v with first+last frames, r2v with reference images.
  • dlazy seedance-2.0: ByteDance's latest video generation model. Supports multi-modal reference (images, video, audio) to generate videos, as well as first/last frame and text-to-video modes.
  • dlazy seedance-2.0-fast: Fast version of ByteDance's Seedance 2.0. Generates videos faster with support for multi-modal references, first/last frame, and text-to-video.
  • dlazy sync-lipsync-3: fal.ai sync-lipsync v3 — given an input video and audio, generate a new video where the speaker's lip movement matches the audio. Good for dubbing, localization, and re-syncing virtual presenters.
  • dlazy veo-3.1: High-quality video generation model, supports text-to-video and single-image-driven video. Suitable for ad shorts and cinematic sequences (slower speed, higher quality).
  • dlazy veo-3.1-fast: Fast video generation model, supports text-to-video and single/multi-image/first-last frame driven. Suitable for time-sensitive previews and rapid iterations.
  • dlazy video-replicate: Video replicate tool: extracts the first frame and audio from the source video, runs video understanding for a prompt, and returns a Seedance 2.0 replicate bundle (first frame + audio + video).
  • dlazy videoretalk: Tongyi VideoRetalk lip sync / lip-sync (mouth sync, dubbing) video model — takes a talking-person video plus a voice audio track and regenerates the video so the speaker's mouth/lips match the new audio. Use this for lip syncing a person video to new speech. Optionally provide a reference face image to pick the target person when the video contains multiple faces.
  • dlazy videoseg: Video human segmentation tool: invokes Aliyun's async SegmentVideoBody and returns a same-length black/white mask video, suitable for downstream compositing or matting.
  • dlazy viduq2-i2v: Vidu image-to-video model, supports reference image-driven video, duration/resolution/ratio, and audio settings, suitable for image animation and short clips.
  • dlazy wan2.7: Tongyi Wanxiang 2.7 video model — one model covers text-to-video, first/last-frame-to-video, and reference-to-video: uses text-to-video when no images are provided, first/last-frame-to-video when frames are provided, and reference-to-video when reference images are supplied.

CRITICAL INSTRUCTION FOR AGENT:

  1. Select the most appropriate video model.
  2. Run dlazy -h to check parameters.
  3. Execute the command.

Error Handling

CodeError TypeExample Message
401Unauthorized (No API Key)ok: false, code: "unauthorized", message: "API key is missing or invalid"
501Missing required parametererror: required option '--prompt ' not specified
502Local file read errorError: Image file/Video file not found: C:\path\to\your\file
503API request failed (no balance)ok: false, code: "insufficient_balance"
503API request failed (server error)HTTP status code error (500 server crash)
504Asynchronous task execution failed=== Generation Failed === / {Specific error reason returned by backend, for example "Prompt violates safety policy"}

AGENT CRITICAL INSTRUCTION:

  1. If the execution result returns code: "insufficient_balance", you MUST explicitly inform the user that their credits are insufficient and guide them to recharge: https://dlazy.com/dashboard/organization/settings?tab=credits
  2. If the execution result returns code: "unauthorized" or indicates missing API key, you MUST inform the user to get their API key from https://dlazy.com/dashboard/organization/api-key and save it using dlazy auth set and resume the task.

Tips

Visit https://dlazy.com for more information.

相关技能

按意图把图像、视频、音频请求分发到合适的 dlazy CLI 模型。

作者 dlazy68 次安装1 星标

按提示词自动选用合适的 dlazy CLI 音频模型,完成语音、音乐、音效与声音克隆。

作者 dlazy65 次安装1 星标

Powerful video generation with Kling v3. Supports high-quality text-to-video and image-to-video. 使用可灵 (Kling) v3 强大的视频生成模型,支持高质量的文生视频与图生视频。

65 次安装

Generate high-quality cinematic effects videos with Google Veo 3.1. 使用 Google Veo 3.1 模型,生成高质量的电影级特效视频,支持文生视频与图生视频。

63 次安装