Design & media

Dlazy Audio音频生成

Try it

Route prompts to the right dlazy CLI audio model for TTS, music, sound effects, and voice cloning.

What it does

Thin client over the dLazy hosted API that selects the matching `dlazy` audio model for the request. Supports text-to-speech, original music, short sound effects, and voice cloning across providers such as ElevenLabs, Suno, Gemini, Doubao, Kling, Qwen, and Vidu. Authenticates with a dLazy API key (via `dlazy login` or `dlazy auth set`), sends prompts to `api.dlazy.com`, and returns output URLs hosted on `files.dlazy.com`.

When to use it

  • Generate multilingual narration or dubbing for a video
  • Compose 10–300s background music from a natural-language prompt
  • Create 1–22s sound effects for foley, alerts, or game SFX
  • Clone a voice from a clean sample for custom TTS

The skill document

音频生成 Audio Generate

English · 中文

Audio generation skill. Automatically selects the best dlazy CLI audio/TTS model based on the prompt. 音频生成技能。根据提示词自动选择最佳的 dlazy CLI 音频/TTS 模型。

Trigger Keywords / 触发关键词

  • generate audio
  • text to speech, TTS
  • generate music, sound effect

Authentication

All requests require a dLazy API key. The recommended way to authenticate is dlazy login:

dlazy login

This runs a device-code flow (also works in remote shells) and automatically saves your API key to the local CLI config — no manual copy/paste required.

Alternative: Set the Key Manually

If you already have an API key, you can save it directly:

dlazy auth set YOUR_API_KEY

The CLI saves the key in your user config directory (~/.dlazy/config.json on macOS/Linux, %USERPROFILE%\.dlazy\config.json on Windows), with file permissions restricted to your OS user account. You can also supply the key per-invocation via the DLAZY_API_KEY environment variable.

Getting Your API Key Manually

  1. Sign in or create an account at dlazy.com
  2. Go to dlazy.com/dashboard/organization/api-key
  3. Copy the key shown in the API Key section

Each key is scoped to your dLazy organization and can be rotated or revoked at any time from the same dashboard.

About & Provenance

You can install on demand without persisting a global binary by running:

npx @dlazy/cli@1.2.3 

Or, if you prefer a global install, the skill's metadata.clawdbot.install field declares the exact pinned version (npm install -g @dlazy/cli@1.2.3). Review the GitHub source before installing.

How It Works

This skill is a thin client over the dLazy hosted API. When you invoke it:

  • Prompts and parameters you provide are sent to the dLazy API endpoint (api.dlazy.com) for inference.
  • Any local file paths you pass to image / video / audio fields are uploaded to dLazy's media storage (files.dlazy.com) so the model can read them — the same flow as any cloud-based generation API.
  • Generated output URLs returned by the API are hosted on files.dlazy.com.

This is the standard SaaS pattern; the skill itself does not access network or filesystem resources beyond what the dLazy CLI already handles. See dlazy.com for the full service terms.

Piping Between Commands

Every dlazy invocation prints a JSON envelope on stdout. Any flag value can be a pipe reference that pulls from the upstream command's envelope, so you can chain steps without copying URLs by hand.

ReferenceResolves to
-Upstream's natural value for this field (scalar or array)
@NThe N-th output's primary value (e.g. @0 = first output url)
@N.Drill into the N-th output (@0.url, @1.meta.fps)
@*All outputs' primary values as an array
@stdinThe whole upstream JSON envelope
@stdin:Jsonpath into the whole envelope (@stdin:result.outputs[0].url)

Examples

# Generate an image and feed its url straight into image-to-video
dlazy seedream-4.5 --prompt "a red fox in snow" \
  | dlazy kling-v3 --image - --prompt "fox starts running"

# Generate an image, then add TTS narration over a still
dlazy seedream-4.5 --prompt "lighthouse at dawn" \
  | dlazy keling-tts --text "Welcome to the coast." --image @0.url

# Fan-out: pass every upstream output url into a batch step
dlazy seedream-4.5 --prompt "city skyline" --n 4 \
  | dlazy superres --images @*

Required flags can be entirely sourced from the pipe — --field - satisfies the requirement when an upstream value exists. If stdin is empty, the CLI fails with code: "no_stdin".

Usage

This skill handles all audio generation requests by selecting the best dlazy audio model.

Available Audio Models

  • dlazy doubao-tts: ByteDance Doubao speech synthesis model. Supports multiple languages, voices, and highly natural streaming audio output, suitable for news broadcasts and audiobooks.
  • dlazy elevenlabs-dialogue: ElevenLabs eleven_v3 multi-voice dialogue: assign a different voice per line (up to 10) and render the whole conversation in one shot. Supports audio tags like [giggling], [whispers] — great for character dialogue, podcasts, and short skits. Before picking a voice, you can search for the right one via elevenlabs-search.
  • dlazy elevenlabs-music: ElevenLabs music_v1 model — generates 10–300s original music from a natural-language prompt. Good for BGM, ads, and short-video soundtracks.
  • dlazy elevenlabs-search: Search the ElevenLabs voice library by keyword, source, and category. Returns a playable preview for each matched voice so you can pick the right one before running TTS. Search with ONE short English descriptor (deep / young / narrative); extra words narrow the match and full sentences usually return nothing.
  • dlazy elevenlabs-sfx: ElevenLabs text-to-sound model — generates 1–22s short sound effects from a description. Suitable for foley, ambience, alerts, and game SFX.
  • dlazy elevenlabs-tts: ElevenLabs eleven_v3 text-to-speech with 12 curated multilingual voices and stability/similarity/style controls. Great for dubbing, audiobooks, and character dialog. Before picking a voice, you can search for the right one via elevenlabs-search.
  • dlazy elevenlabs-voice-clone: ElevenLabs Instant Voice Cloning (IVC). Upload a clean voice sample to clone a custom voice usable with ElevenLabs TTS.
  • dlazy gemini-2.5-tts: Gemini-powered high-quality text-to-speech. Supports bilingual (EN/CN) and various emotional voices.
  • dlazy keling-sfx: Sound effect generation model: supports text-to-SFX and matching SFX/BGM for reference videos. Suitable for foley, ambient sounds, and short video audio completion.
  • dlazy keling-tts: Text-to-speech model (TTS), supports language, voice, speed, and output format settings. Suitable for dubbing, audiobooks, and voice broadcasts.
  • dlazy kling-audio-clone: Custom voice (Kling), cloned voice used for dubbing or binding to subjects.
  • dlazy qwen-audio-clone: Alibaba Bailian qwen3-tts voice cloning. Upload a clean voice sample to clone a custom voice usable in subsequent TTS calls.
  • dlazy qwen-tts: Alibaba Bailian qwen3-tts text-to-speech. Choose from curated system voices (including dialects) or design a custom voice from a natural-language description.
  • dlazy suno-music: Suno V5.5 music generation. Inspiration mode (auto lyrics) or custom mode (manual style/title/lyrics). Generates music with or without vocals, with fine-grained controls over style weight, weirdness and audio weight.
  • dlazy vidu-audio-clone: Clone a real human voice and use it to read the specified text.

CRITICAL INSTRUCTION FOR AGENT:

  1. Select the most appropriate audio model.
  2. Run dlazy -h to check parameters.
  3. Execute the command.

Error Handling

CodeError TypeExample Message
401Unauthorized (No API Key)ok: false, code: "unauthorized", message: "API key is missing or invalid"
501Missing required parametererror: required option '--prompt ' not specified
502Local file read errorError: Image file/Video file not found: C:\path\to\your\file
503API request failed (no balance)ok: false, code: "insufficient_balance"
503API request failed (server error)HTTP status code error (500 server crash)
504Asynchronous task execution failed=== Generation Failed === / {Specific error reason returned by backend, for example "Prompt violates safety policy"}

AGENT CRITICAL INSTRUCTION:

  1. If the execution result returns code: "insufficient_balance", you MUST explicitly inform the user that their credits are insufficient and guide them to recharge: https://dlazy.com/dashboard/organization/settings?tab=credits
  2. If the execution result returns code: "unauthorized" or indicates missing API key, you MUST inform the user to get their API key from https://dlazy.com/dashboard/organization/api-key and save it using dlazy auth set and resume the task.

Tips

Visit https://dlazy.com for more information.

Questions people ask

Do I need an API key?
Yes. Run `dlazy login` for a device-code flow, or `dlazy auth set YOUR_API_KEY` to save one manually. You can also pass `DLAZY_API_KEY` per invocation. Keys are scoped to your dLazy organization and can be rotated from the dashboard.
Which audio models are available?
The skill can route to doubao-tts, elevenlabs-tts, elevenlabs-dialogue, elevenlabs-music, elevenlabs-sfx, elevenlabs-search, elevenlabs-voice-clone, gemini-2.5-tts, keling-tts, keling-sfx, kling-audio-clone, qwen-tts, qwen-audio-clone, suno-music, and vidu-audio-clone.
Can outputs be chained with image or video commands?
Yes. Every `dlazy` invocation prints a JSON envelope; use pipe references such as `-`, `@0.url`, `@*`, or `@stdin:result.outputs[0].url` to pass upstream values to the next step without copying URLs.

Related skills

轻量级文本转语音工具,支持多语言TTS与基础音效生成,适合个人内容创作。Use when 需要文本翻译、多语言转换、本地化处理时使用。不适用于专业医学法律翻译认证。适用于独立开发者、企业团队和自动化工作流场景。支持中文交互,无需复杂配置即开即用。输出结果可直接使用,减少二次加工成本。提供结构化输出和错误处理机制。

1 installs

轻量级AI图片生成工具,支持文生图与基础图片编辑,适合个人创意原型制作。Use when 需要AI模型调用、智能对话、Agent编排、LLM应用时使用。不适用于需要100%确定性的关键决策。适用于独立开发者、企业团队和自动化工作流场景。支持中文交互,无需复杂配置即开即用。输出结果可直接使用,减少二次加工成本。

1 installs

Generate customized speech that highly restores the timbre by uploading reference audio using Kling Audio Clone. 使用可灵 (Kling) 声音克隆模型,通过上传参考音频,生成高度还原该音色的定制语音。

64 installs

Generate multilingual, highly natural audio using Gemini 2.5 text-to-speech. 使用 Gemini 2.5 强大的文本转语音能力,生成多语言、高自然度的音频。

61 installs

Synthesize text into natural and fluent speech using Doubao TTS. 使用豆包 (Doubao) TTS 文本转语音模型,将文字合成为自然流畅的语音播报。

70 installs

ElevenLabs text-to-sound model — generates 1–22s short sound effects from a description. Suitable for foley, ambience, alerts, and game SFX. ElevenLabs 文本生音效模型,根据描述生成 1–22 秒短音效。适合拟音、环境声、提示音与游戏音效。

8 installs

More from dlazy

Browse all skills

Route image, video, and audio requests to the right dlazy CLI model based on intent.

by dlazy