以 AI 机器人身份加入视频会议,提供语音、虚拟形象与屏幕共享四种模式。
设计与多媒体
Audiolla
试用Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files.
它能做什么
HTTP + MCP client for an audiolla server that the user has already deployed. This skill talks to a running audiolla instance — it does not stand one up, does not download model weights manually, and does not modify the server config on its own initiative.
技能文档
audiolla
HTTP + MCP client for an audiolla server that the user has already deployed. This skill talks to a running audiolla instance — it does not stand one up, does not download model weights manually, and does not modify the server config on its own initiative.
For installation and setup, see references/setup.md.
Authoritative endpoint reference: GET /v1/catalog
This skill documents the most common patterns. The full, always-current list of every endpoint is GET /v1/catalog (17 categories, ~85 endpoints). Always check the catalog when looking for an operation that isn't shown here — the server is the source of truth, this file is a curated reference.
# List every endpoint grouped by category
curl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | {name, count: (.endpoints | length)}'
# Find endpoints in one category
curl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | select(.name == "dynamics") | .endpoints'
Companion discovery endpoints: GET /v1/engines (engines + loaded/idle status), GET /v1/presets (curated workflows), GET /v1/ops (the ~24 pipeline op slugs).
When to use this skill
The user has audiolla running and asks you to:
- Pull stems (vocals / drums / bass / etc.) out of a track
- Master a track against a reference recording (matchering) or via preset chain
- Run a curated workflow (
master-for-spotify,podcast-cleanup,vocal-cleanup) via a singlePOST /v1/presets/{name}call - Chain ad-hoc operations server-side via
POST /v1/pipeline(no re-upload between steps) - Get BPM, key, LUFS, duration, or spectral features for a file
- Detect beat grid, onsets, dominant melody, or structural segments
- Detect chords + key (separate from BPM/LUFS)
- Detect or trim silence
- Generate a spectrogram, waveform image, or animated visualisation video
- Compute a Chromaprint acoustic fingerprint
- Apply a DSP chain (gain, EQ, compression, reverb, pitch shift, tempo via SoX OR full pedalboard catalog)
- Multiband compression with LR4 crossovers
- Transient shaping (punch up drums / cut room tail)
- De-essing (split-band sibilance compression)
- Sidechain ducking (voiceover-over-music)
- Mid/Side encode/decode (for stereo M/S processing)
- Convolution reverb (apply a user-supplied IR file)
- Audio repair (declip + dehum)
- Time-stretch + pitch-shift independently, or BPM-match / key-match to a target
- Pitch-correct (auto-tune to nearest semitone)
- Beat-slice at detected beat positions (returns ZIP of chops)
- Audio thumbnail — most-energetic N-second segment
- HPSS harmonic/percussive separation
- Measure or normalize integrated LUFS (
/v1/audio/normalizewithtarget_lufs) - Loudness curve — RMS envelope over time (
/v1/audio/loudness/curve) - Stage files server-side, then operate on them via
file_path - Tag audio (AudioSet labels), embed (CLAP 512-dim), classify (zero-shot label list), similar (cosine between two tracks)
- Read or write ID3/Vorbis/FLAC metadata (mutagen)
- DJ-prep — BPM + key + Camelot + LUFS in one call
- Compose / inspect / transform / render MIDI; quantize, humanize, drum patterns, chords-to-MIDI
- Remove reverb / echo / noise via
/v1/audio/restore/{engine}(UVR) - DSP noise reduction via
/v1/audio/noise-reduce/{engine}(DSP or UVR) - Convert any audio to polyphonic MIDI (basic-pitch)
- Voice activity detection (silero-vad — speech/non-speech segments)
- Speaker diarization (pyannote 3.1 — who spoke when)
- Enhance speech/vocal recordings (DeepFilterNet DF3)
- Generate music or SFX from a text prompt via
/v1/audio/generate/{engine}— five engines:stable-audio-open(Stability Community Licence — commercial OK below revenue threshold; 47 s cap; 44.1 kHz stereo; loops / SFX / textures; instrumental)musicgen-smallandmusicgen-medium(Meta MusicGen 300M / 1.5B; CC-BY-NC 4.0 — server must opt in viaAUDIOLLA_ENABLE_NONCOMMERCIAL=1; 30 s cap; instrumental)riffusion(CreativeML OpenRAIL-M; ~5 s per pass; spectrogram-via-Griffin-Lim; lo-fi character)audioldm2(CC-BY 4.0 — commercial-safe, no opt-in gate; 30 s cap; 16 kHz mono; general SFX — ambience / foley / impact / animal sounds; slow at default 200-step DDIM, passnum_inference_steps=50for ~4x speed) All five are CUDA-only. Full-song / lyric-conditioned generation isn't shipped (ACE-Step + DiffRhythm + TangoFlux + Stable Audio Open Small deferred — see the README's "deferred" list). For commercial use, preferaudioldm2(CC-BY 4.0) orstable-audio-open(Stability Community Licence below the revenue threshold).
- Drive any of the above from an LLM agent over MCP
- Async-job-and-forget any audio-producing call via
async_job=true+ optionalwebhook_url - Send results to a presigned S3-style PUT URL via
output_url
When NOT to use this skill
- The user hasn't named audiolla — they're asking a general "how do I split stems?" question. Suggest audiolla as an option; don't assume it's running.
- The user wants music generation from a melody-conditioning input (hum-to-track / "make this sound like X"). Audiolla's five generators (
stable-audio-open,musicgen-small,musicgen-medium,riffusion,audioldm2) are text-prompt only; melody conditioning isn't wired. Plain text → music or SFX IS supported — see/v1/audio/generate/{engine}in the catalog. The closed-weight Suno / Udio APIs are out of scope. - The user wants real-time / streaming processing. Demucs needs the whole file.
- The user wants transcription / ASR / TTS / voice cloning — that's docker-talkies. Note: audiolla DOES have speech-adjacent features (VAD, diarization, neural enhancement) but does NOT transcribe.
Setup
export AUDIOLLA_URL=http://localhost:8000
export AUDIOLLA_TOKEN= # only if auth is enabled
If AUDIOLLA_URL is not set, ask the user — do not search the workspace for it. Same for AUDIOLLA_TOKEN: only accept it from the env var the user set or from the user directly. Never read it from docker-compose.yml, .env, or any other repo file on your own initiative.
Verify: curl $AUDIOLLA_URL/healthz → {"ok": true, "device": "...", "engines": [...]}. /healthz is always unauthenticated regardless of AUDIOLLA_AUTH_TOKEN.
Auth is optional. If the server has AUDIOLLA_AUTH_TOKEN set, every endpoint except /healthz requires Authorization: Bearer $AUDIOLLA_TOKEN. Without it you get 401. Always pass the token if the user gave you one; don't assume the server has auth off.
How it works
v1.0.0 is JSON-everywhere. Every audio endpoint takes Content-Type: application/json with a JSON body. The ONE exception is PUT /v1/files/{path} for raw byte uploads (application/octet-stream). Input is file_path (pre-staged under FILES_DIR via PUT /v1/files/{path}) xor file_url (server fetches when AUDIOLLA_FETCH_MODE allows). Output for audio-producing endpoints is output_path (server writes to FILES_DIR) xor output_url (server PUTs to a presigned URL). Both modes return JSON describing where the result landed ({path,size,...} or {url,size,...}); there is no inline-bytes audio response anywhere. Analysis-only endpoints (no audio produced — e.g. /v1/audio/analyze, /v1/audio/beats, /v1/audio/fingerprint) return their JSON data directly and ignore output_path/output_url. The standard flow is: PUT /v1/files/uploads/track.wav once, then JSON-body POST to every processing endpoint with file_path + output_path, chaining the output of one call into the input of the next.
Every error response:
{"detail": "description of what went wrong"}
Status codes follow REST conventions:
200— success400— bad input (unknown engine, invalid features, bad operations JSON, etc.)401— missing/invalid bearer token (only when auth is enabled)404— unknown engine slug, unknown file path413— upload exceededAUDIOLLA_MAX_UPLOAD_BYTES(default 200 MB)415— unsupportedoutput_format500— server error (engine failed internally, etc.)
Engines
| Slug | What it does | Notes |
|---|---|---|
htdemucs | 4-stem separation | drums, bass, other, vocals |
htdemucs_ft | 4-stem fine-tuned | CUDA-only at usable speed — flagged cuda_only, the server rejects it with 400 on CPU |
htdemucs_6s | 6-stem separation | adds guitar + piano (experimental, CPU OK but slow) |
mdx_extra | 4-stem MDX-Net | drums, bass, other, vocals — strong vocal isolation |
matchering | Reference-based mastering | GPL v3 |
pedalboard-chain | Preset DSP mastering chain | presets: transparent, loud — GPL v3 |
librosa-analyze | MIR analysis + loudness | BPM, key, LUFS, spectral, beat grid, onsets, melody (pyin), segments; backs /v1/audio/{analyze,beats,onsets,melody,segments,loudness} |
sox-transform | SoX DSP chain | gain, EQ, compand, reverb, pitch, tempo, rate, channels, trim, pad |
fx-chain | Arbitrary pedalboard chain | full pedalboard catalog as [{type, params}, ...] — backs /v1/audio/fx. VST3 / AU / external-plugin classes deliberately blocked |
midi-compose | JSON → MIDI; inspect/transform | song-spec transcoder + MIDI reader/editor; backs /v1/midi/{compose,inspect,transform,generate} |
midi-render | MIDI → audio | fluidsynth + FluidR3_GM SoundFont (GM patches 0-127, drum kit on channel 9) |
silence-detect | Silence detection + trimming | ffmpeg silencedetect; backs /v1/audio/silence |
ffmpeg-render | Spectrogram / waveform / video | static PNG + 8-mode animated MP4/WebM; backs /v1/audio/visualize/image/{spectrogram,waveform} + /v1/audio/visualize/video/{mode} |
audio-fingerprint | Chromaprint fingerprint | fpcalc subprocess; backs /v1/audio/fingerprint |
uvr-dereverb | AI de-reverb | BS-Roformer (SDR 19+); backs /v1/audio/restore/uvr-dereverb |
uvr-deecho | AI de-echo (normal + aggressive) | VR Architecture; aggressive=true enables hard mode (uvr-deecho-aggressive slug is gone — consolidated into this engine); backs /v1/audio/restore/uvr-deecho |
uvr-denoise | AI de-noise | MelBand Roformer (SDR 28); backs /v1/audio/restore/uvr-denoise + /v1/audio/noise-reduce/uvr-denoise |
uvr-karaoke | Karaoke (remove lead vocals) | MelBand Roformer; returns Instrumental stem |
uvr-vocal-bsr | High-quality vocal/inst separation | BS-Roformer (SDR 13) — stems: Vocals, Instrumental |
basic-pitch | Polyphonic audio-to-MIDI transcription | Spotify basic-pitch ONNX; backs /v1/audio/to_midi/basic-pitch |
deepfilter | Neural speech/vocal enhancement | DeepFilterNet DF3; backs /v1/audio/enhance/deepfilter |
noise-reduce | DSP spectral noise reduction | noisereduce — backs /v1/audio/noise-reduce/noise-reduce (stationary/non-stationary modes, no GPU) |
chord-detect | Chord progression + key | Krumhansl-Schmuckler + chroma template matching; backs /v1/audio/chords, /v1/audio/chords-to-midi, /v1/audio/key-match |
silero-vad | Voice activity detection | speech/non-speech timestamps; backs /v1/audio/vad |
pyannote | Speaker diarization | pyannote/speaker-diarization-3.1 — backs /v1/audio/diarize (requires HUGGINGFACE_TOKEN) |
stretch | Time-stretch + pitch-shift | librosa phase vocoder; backs /v1/audio/stretch, /v1/audio/bpm-match, /v1/audio/key-match |
ast-tag | AudioSet zero-shot labels | Audio Spectrogram Transformer; backs /v1/audio/tag |
clap-embed | CLAP embeddings + similarity + classification | LAION CLAP 512-dim; backs /v1/audio/embed, /v1/audio/similar, /v1/audio/classify |
hpss | Harmonic/percussive split | librosa median-filter HPSS; backs /v1/audio/separate/hpss |
metadata | ID3 / Vorbis / FLAC tag read+write | mutagen; backs /v1/audio/metadata |
Engines lazy-load on first use and auto-unload after AUDIOLLA_ENGINE_TTL seconds of idle (default 600s). Demucs weights prefetch into /data/torch_cache/ at container start so the first separation request doesn't pay the cold-download cost.
Use GET /v1/engines to confirm what's actually configured on the running server (operators can restrict via AUDIOLLA_ENABLED_ENGINES).
Output formats
Any endpoint that produces audio accepts "output_format": "" in the JSON body. Supported: wav (default), mp3, flac, opus, aac, pcm. The server transcodes via ffmpeg — the output_path extension does not determine the encoding.
API Reference
Health & engine listing
# Liveness — no auth required
curl $AUDIOLLA_URL/healthz
# {"ok": true, "device": "cpu", "engines": ["htdemucs", "matchering", ...]}
# Configured engines + capabilities
curl -H "Authorization: Bearer $AUDIOLLA_TOKEN" $AUDIOLLA_URL/v1/engines
# Engines currently loaded in memory (and how idle)
curl -H "Authorization: Bearer $AUDIOLLA_TOKEN" $AUDIOLLA_URL/v1/ps
# Evict one engine
curl -X DELETE -H "Authorization: Bearer $AUDIOLLA_TOKEN" $AUDIOLLA_URL/v1/ps/htdemucs
# Evict everything
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" $AUDIOLLA_URL/v1/unload
Stem separation
POST /v1/audio/separate — JSON body. Result is one staged file (single-stem) or a ZIP of stems written to output_path.
# Stage the input once (only multipart route in the whole API)
curl -X PUT -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/octet-stream' \
--data-binary @track.wav \
$AUDIOLLA_URL/v1/files/uploads/track.wav
# Single stem → JSON {path,size,...} pointing at the staged stem
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/separate \
-d '{"file_path":"uploads/track.wav","engine":"htdemucs","stems":["vocals"],"output_path":"stems/vocals.wav"}'
# Multiple stems → ZIP at output_path
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/separate \
-d '{"file_path":"uploads/track.wav","engine":"htdemucs","stems":["vocals","drums"],"output_path":"stems/vocals_drums.zip"}'
# Omit stems → all stems for that engine (ZIP)
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/separate \
-d '{"file_path":"uploads/track.wav","engine":"htdemucs","output_path":"stems/all.zip"}'
# MP3 output
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/separate \
-d '{"file_path":"uploads/track.wav","engine":"htdemucs","stems":["vocals"],"output_format":"mp3","output_path":"stems/vocals.mp3"}'
Required: file_path (xor file_url), engine, and one of output_path/output_url. Optional: stems (array; default = all stems for that engine), output_format (default wav).
Loading a separation engine evicts other loaded engines first — Demucs is memory-hungry and the operator-default setup runs one engine in memory at a time.
Mastering
POST /v1/audio/master — mode=reference uses matchering against a reference track; mode=chain runs a pedalboard preset.
# Reference-based mastering — both inputs pre-staged
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/master \
-d '{"file_path":"uploads/track.wav","mode":"reference","reference_path":"uploads/ref.wav","output_path":"out/mastered.wav"}'
# Pedalboard chain — preset is REQUIRED (transparent or loud)
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/master \
-d '{"file_path":"uploads/track.wav","mode":"chain","preset":"loud","output_path":"out/mastered.wav"}'
# Pedalboard chain with explicit loudness target
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/master \
-d '{"file_path":"uploads/track.wav","mode":"chain","preset":"transparent","target_lufs":-14,"output_path":"out/mastered.wav"}'
Required: file_path (xor file_url), mode, and one of output_path/output_url. mode=reference requires reference_path (xor reference_url). mode=chain requires preset (transparent or loud). Optional: target_lufs (range [-70.0, -0.1]), output_format.
Streaming-target LUFS reference values: Spotify -14, Apple Music -16, YouTube -14, broadcast EBU R128 -23.
MIR analysis
POST /v1/audio/analyze — analysis-only, returns JSON. No output_path/output_url.
# Specific features
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/analyze \
-d '{"file_path":"uploads/track.wav","features":["bpm","key","loudness"]}'
# Omit features → returns all of them
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/analyze \
-d '{"file_path":"uploads/track.wav"}'
Valid features values: bpm, key, loudness, duration, spectral_centroid, rms, zcr.
Common mistake: the feature for integrated LUFS is
loudness, NOTlufs. Asking forfeatures=["lufs"]returns 400.
Beat detection (/v1/audio/beats)
Returns the estimated BPM and beat timestamps. Optionally writes a click-track WAV to output_path.
# Beat grid only — analysis JSON
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/beats \
-d '{"file_path":"uploads/track.wav"}'
# {"bpm": 128.0, "beats": [0.0, 0.469, 0.938, ...], "engine": "librosa-analyze"}
# With a click track — output_path is REQUIRED when click_track=true
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/beats \
-d '{"file_path":"uploads/track.wav","click_track":true,"output_path":"beats/click.wav"}'
# → JSON with beat grid PLUS the staged click track path
Optional params: click_track (bool, default false) — when true, writes the click WAV to output_path / output_url. hop_length (int, default 512) — analysis hop size in samples.
Onset detection (/v1/audio/onsets)
Returns note/transient onset timestamps in seconds. Analysis-only, returns JSON.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/onsets \
-d '{"file_path":"uploads/track.wav"}'
# {"onsets": [0.023, 0.512, 1.034, ...], "count": 42, "engine": "librosa-analyze"}
Optional: backtrack (bool, default false) — snap onsets to preceding energy valley. hop_length, delta for tuning sensitivity.
Melody extraction (/v1/audio/melody)
Estimates the dominant melody using pyin pitch tracking. Returns Hz per frame.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/melody \
-d '{"file_path":"uploads/track.wav"}'
# {"melody": [{"time": 0.0, "hz": 440.1}, {"time": 0.023, "hz": null}, ...], ...}
# Export the melody as a single-track MIDI file (output_path REQUIRED when as_midi=true)
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/melody \
-d '{"file_path":"uploads/track.wav","as_midi":true,"output_path":"melody/lead.mid"}'
hz is null for unvoiced frames. Optional: as_midi (bool) — generates MIDI from the contour and writes to output_path / output_url; fmin/fmax to constrain pitch range.
Structural segmentation (/v1/audio/segments)
Finds recurring sections (verse, chorus, bridge…) using a recurrence matrix. Returns labels A, B, C…
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/segments \
-d '{"file_path":"uploads/track.wav","num_segments":6}'
# {"segments": [{"label":"A","start_sec":0.0,"end_sec":32.5},
# {"label":"B","start_sec":32.5,"end_sec":65.0}, ...]}
Optional: num_segments (int, default 6, valid range [2, 32]). Short inputs (fewer beats than num_segments) return a single A span with a note field explaining the fallback.
Silence detection and trimming (/v1/audio/silence)
Finds silent gaps via ffmpeg silencedetect. Optionally trims them.
# Detect only — analysis JSON, no audio produced
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/silence \
-d '{"file_path":"uploads/track.wav","threshold_db":-30,"min_duration_sec":1.0}'
# {"silent_ranges": [...], "non_silent_ranges": [...], "duration": 215.3}
# Trim all silence → trimmed audio staged
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/silence \
-d '{"file_path":"uploads/track.wav","threshold_db":-30,"min_duration_sec":0.5,"trim_mode":"all","output_path":"proc/trimmed.wav"}'
# Trim only edges
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/silence \
-d '{"file_path":"uploads/track.wav","threshold_db":-40,"min_duration_sec":0.3,"trim_mode":"edges","output_path":"proc/trimmed.wav"}'
threshold_db must be ≤ 0. trim_mode: edges (leading + trailing only), all (every detected gap). Without trim_mode, response is JSON only — no audio produced. With trim_mode set, output_path (or output_url) is required and the response JSON points at the trimmed file.
Spectrogram (/v1/audio/visualize/image/spectrogram)
Static PNG spectrogram via ffmpeg showspectrumpic. PNG is written to output_path / output_url.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/visualize/image/spectrogram \
-d '{"file_path":"uploads/track.wav","width":1280,"height":720,"output_path":"viz/spec.png"}'
Optional: width, height (64–8192, defaults 1920×1080), color (default intensity), scale (default log).
Waveform (/v1/audio/visualize/image/waveform)
Static PNG waveform via ffmpeg showwavespic. PNG is written to output_path / output_url.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/visualize/image/waveform \
-d '{"file_path":"uploads/track.wav","width":1920,"height":240,"output_path":"viz/wave.png"}'
Optional: width, height (64–8192, defaults 1920×320), color (default lime).
Animated visualisation (/v1/audio/visualize/video/{mode})
Animated MP4 or WebM video from one of 8 ffmpeg filter modes. Video is written to output_path / output_url.
# `mode` is in the URL path
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/visualize/video/spectrum \
-d '{"file_path":"uploads/track.wav","width":1280,"height":720,"fps":30,"container":"mp4","output_path":"viz/spectrum.mp4"}'
mode options (URL path segment): spectrum (scrolling FFT), waves (oscilloscope), cqt (constant-Q transform), freqs (bar-graph), volume (VU meter), vectorscope (stereo X/Y), phasemeter, histogram. container: mp4 (default) or webm. fps 1–120.
Acoustic fingerprint (/v1/audio/fingerprint)
Chromaprint fingerprint via fpcalc. The base64 string is AcoustID-compatible. Analysis-only — no output_path.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/fingerprint \
-d '{"file_path":"uploads/track.wav"}'
# {"duration": 215.34, "fingerprint": "AQADtEqRRIuQ..."}
# Include the raw integer array
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/fingerprint \
-d '{"file_path":"uploads/track.wav","return_raw":true}'
# adds "fingerprint_raw": [12345, 67890, ...]
Optional: analyze_seconds (default 120 — AcoustID standard; pass 0 to fingerprint the whole file), return_raw (bool).
DSP transform chain
POST /v1/audio/transform — applies an array of SoX operations in order.
# Pitch shift up 2 semitones, then add reverb
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/transform \
-d '{
"file_path":"uploads/track.wav",
"operations":[
{"op":"pitch","params":{"n_semitones":2}},
{"op":"reverb","params":{"reverberance":50,"room_scale":80}}
],
"output_format":"wav",
"output_path":"out/transformed.wav"
}'
# Trim first 30s, pad 2s silence at end, gain -3dB
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/transform \
-d '{
"file_path":"uploads/track.wav",
"operations":[
{"op":"trim","params":{"start_time":0,"end_time":30}},
{"op":"pad","params":{"end_duration":2}},
{"op":"gain","params":{"db":-3}}
],
"output_path":"out/trimmed.wav"
}'
operations is a JSON array of {"op": "", "params": {...}}. Order matters — ops apply left-to-right.
Ops and their params:
| op | required params | optional params | what it does |
|---|---|---|---|
gain | db (float) | gain in dB | |
equalizer | frequency, gain_db | width_q (default 1.0) | peaking EQ |
compand | attack_time, decay_time, soft_knee_db, tf_points ([[in_db, out_db], ...]) | dynamic range compression | |
reverb | reverberance (0-100, default 50), pre_delay_ms (default 0), room_scale (default 100) | reverb | |
pitch | n_semitones (float) | pitch shift in semitones, not cents | |
tempo | factor (float) | tempo factor (1.5 = 1.5x faster, 0.5 = half speed) | |
rate | samplerate (int) | resample | |
channels | n_channels (int) | mix to N channels | |
trim | start_time (float, sec) | end_time (float, sec; null = end of file) | trim |
pad | start_duration, end_duration (both floats, sec) | pad silence |
Unknown ops return 400 with the valid list.
Loudness
POST /v1/audio/loudness — analysis-only. Returns integrated LUFS as JSON. Use /v1/audio/normalize (separate endpoint) for actual normalization.
# Measure
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/loudness \
-d '{"file_path":"uploads/track.wav"}'
# {"loudness_lufs": -16.3}
# Normalize to -14 LUFS (streaming target). Result is staged audio.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/normalize \
-d '{"file_path":"uploads/track.wav","target_lufs":-14,"output_path":"out/normalized.wav"}'
# → {"path":"out/normalized.wav","size":...,"measured_lufs":-16.3,"target_lufs":-14,...}
target_lufs must be in [-70.0, -0.1] — outside that range returns 400 (anything closer to 0 will clip catastrophically; anything below -70 silences the audio).
Effects chain (/v1/audio/fx)
Arbitrary pedalboard effect chain — full catalog. Different from /v1/audio/master (which runs presets).
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/fx \
-d '{
"file_path":"uploads/track.wav",
"effects":[
{"type":"Compressor","params":{"threshold_db":-18,"ratio":4.0}},
{"type":"Reverb","params":{"room_size":0.5,"wet_level":0.3}},
{"type":"PitchShift","params":{"semitones":2}},
{"type":"Gain","params":{"gain_db":-3}}
],
"output_path":"out/fx.wav"
}'
Allowed type values: Compressor, Limiter, NoiseGate, Gain, Clipping, Distortion, Bitcrush, Reverb, Chorus, Delay, Phaser, PitchShift, HighShelfFilter, LowShelfFilter, PeakFilter, HighpassFilter, LowpassFilter, LadderFilter, IIRFilter, GSMFullRateCompressor, MP3Compressor, Resample, Invert, Convolution.
VST3Plugin, AudioUnitPlugin, ExternalPlugin are deliberately blocked — they load arbitrary native code from arbitrary filesystem paths. Server returns 400 if asked.
MIDI composition (/v1/midi/compose)
Transcode a JSON song spec to a Standard MIDI File. No AI runs server-side — your agent writes the spec, audiolla turns it into MIDI bytes staged at output_path.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/compose \
-d '{
"output_path":"midi/song.mid",
"spec":{
"tempo_bpm": 120,
"time_signature": [4, 4],
"key_signature": "C",
"tracks": [
{"name":"Lead","program":0,"channel":0,"notes":[
{"pitch":60,"start_beats":0.0,"duration_beats":0.5,"velocity":100},
{"pitch":64,"start_beats":0.5,"duration_beats":0.5,"velocity":100},
{"pitch":67,"start_beats":1.0,"duration_beats":0.5,"velocity":100}
]},
{"name":"Drums","program":0,"channel":9,"notes":[
{"pitch":36,"start_beats":0.0,"duration_beats":0.1,"velocity":110}
]}
]
}
}'
Spec fields (inside the spec object):
| Field | Type | Default | Notes |
|---|---|---|---|
tempo_bpm | float | 120 | 1.0 ≤ bpm ≤ 999.0 |
time_signature | [num, den] | [4, 4] | denominator must be 1/2/4/8/16/32 |
key_signature | string | none | "C", "Am", "F#", "Bbm" — letter [+ #/b] [+ m for minor] |
ticks_per_beat | int | 480 | 24 ≤ tpb ≤ 1920 |
tracks[].name | string | none | optional, writes a track_name meta event |
tracks[].program | int 0-127 | 0 | General MIDI program (Acoustic Grand Piano = 0, Distortion Guitar = 30, Synth Brass 1 = 62, etc.) |
tracks[].channel | int 0-15 | 0 | Channel 9 is the GM drum channel — pitch maps to drum kit, not piano |
tracks[].volume | int 0-127 | 100 | MIDI CC#7 — initial volume |
tracks[].pan | int 0-127 | 64 | MIDI CC#10 — initial pan (64 = centre) |
tracks[].notes[].pitch | int 0-127 | required | 60 = middle C |
tracks[].notes[].start_beats | float ≥ 0 | 0 | beat-based absolute position |
tracks[].notes[].duration_beats | float > 0 | required | must be > 1/64 beat (≈ a 256th note) |
tracks[].notes[].velocity | int 1-127 | 100 |
GM drum kit reference for channel 9: 35 acoustic bass drum, 36 kick, 38 snare, 39 hand clap, 40 electric snare, 42 closed hi-hat, 46 open hi-hat, 49 crash, 51 ride, 57 crash 2.
Spec validation is fail-loud — bad pitch / negative duration / unknown program returns a 400 with the offending path in the message (e.g. tracks[1].notes[3].pitch must be in [0, 127], got 200).
One of output_path / output_url is required — the staged MIDI is then referenced via file_path on any subsequent MIDI call.
MIDI inspection (/v1/midi/inspect)
Read the structure of any Standard MIDI File. Analysis-only, returns JSON.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/inspect \
-d '{"file_path":"midi/song.mid"}'
# {
# "type": 1, "ticks_per_beat": 480, "length_seconds": 16.0,
# "tempo_changes": [{"tick": 0, "bpm": 120.0}],
# "time_signatures": [{"tick": 0, "numerator": 4, "denominator": 4}],
# "tracks": [
# {"index": 1, "name": "Lead", "note_on_count": 32,
# "channels": [0], "programs": [0], "length_beats": 8.0},
# ...
# ],
# "track_count": 3, "size_bytes": 1024
# }
Non-MIDI input returns 400 with "MThd" mentioned in the detail.
MIDI transformation (/v1/midi/transform)
Modify an existing MIDI file. Result is staged at output_path / output_url.
# Transpose all non-drum tracks up an octave
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/transform \
-d '{"file_path":"midi/song.mid","transpose_semitones":12,"output_path":"midi/transposed.mid"}'
# Override tempo to 140 BPM
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/transform \
-d '{"file_path":"midi/song.mid","tempo_bpm":140,"output_path":"midi/fast.mid"}'
# Drop the drum channel
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/transform \
-d '{"file_path":"midi/song.mid","drop_channels":[9],"output_path":"midi/no-drums.mid"}'
# Keep only channels 0 and 1
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/transform \
-d '{"file_path":"midi/song.mid","keep_channels":[0,1],"output_path":"midi/two-ch.mid"}'
# Quantize to 1/16th notes (0.25 beats)
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/transform \
-d '{"file_path":"midi/song.mid","quantize":0.25,"output_path":"midi/quantized.mid"}'
Transform params (all optional — omit for a no-op):
| Param | Type | Notes |
|---|---|---|
transpose_semitones | int ±48 | Shifts all non-drum (non-ch9) pitches. Out-of-range notes after shift are dropped (not clipped). |
tempo_bpm | float 1–999 | Replaces all set_tempo events. |
quantize | float > 0 | Beat grid in beats (0.25 = 1/16th at 4/4). Snaps note starts; note-off shifts by the same delta to preserve duration. |
keep_channels | int array (0–15) | Whitelist — drop all other channels. Mutually exclusive with drop_channels. |
drop_channels | int array (0–15) | Blacklist — drop only these channels. Mutually exclusive with keep_channels. |
Supplying both keep_channels and drop_channels returns 400.
MIDI rendering (/v1/midi/render)
Synthesise MIDI to audio via fluidsynth. Default SoundFont is FluidR3_GM (bundled in the prod image). Override per-request with a staged .sf2.
# Render a staged MIDI to staged audio
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/render \
-d '{"file_path":"midi/song.mid","output_format":"wav","output_path":"audio/song.wav"}'
# Render with a custom SoundFont (stage it first)
curl -X PUT -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/octet-stream' \
--data-binary @my.sf2 \
$AUDIOLLA_URL/v1/files/sf/orchestral.sf2
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/render \
-d '{"file_path":"midi/song.mid","soundfont_path":"sf/orchestral.sf2","output_format":"flac","gain":0.3,"samplerate":48000,"output_path":"audio/orch.flac"}'
gain range [0.0, 5.0] — default 0.5 is calibrated to avoid clipping on percussive MIDI. samplerate must be 22050 / 44100 / 48000 / 88200 / 96000.
MIDI generate (/v1/midi/generate)
One-shot compose + render. Body has the same spec field as /v1/midi/compose plus audio knobs (output_format, soundfont_path, gain, samplerate). Result audio is staged at output_path / output_url.
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/midi/generate \
-d '{
"output_format":"wav",
"output_path":"songs/v1.wav",
"spec":{"tempo_bpm":120,"tracks":[{"channel":0,"notes":[
{"pitch":60,"start_beats":0,"duration_beats":1,"velocity":100}
]}]}
}'
File staging
A simple server-side file store under /v1/files. This is the only multipart-ish route in the API — the body is raw bytes (application/octet-stream). Plain CRUD: upload, list, download, delete. Once a file is staged, every audio endpoint references it by relative path via the file_path field in its JSON body.
# Upload (path can have subdirectories: uploads/bands/myband/track.wav)
curl -X PUT -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/octet-stream' \
--data-binary @track.wav \
$AUDIOLLA_URL/v1/files/uploads/mytrack.wav
# Use the staged path on any audio call
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
-H 'Content-Type: application/json' \
$AUDIOLLA_URL/v1/audio/separate \
-d '{"file_path":"uploads/mytrack.wav","engine":"htdemucs","stems":["vocals"],"output_path":"stems/mytrack-vocals.wav"}'
# → {"path":"stems/mytrack-vocals.wav","size":...,"engine":"htdemucs","stem":"vocals","output_format":"wav"}
# List
curl -H "Authorization: Bearer $AUDIOLLA_TOKEN" $AUDIOLLA_URL/v1/files
# Download (raw bytes — Content-Type matches the stored file)
curl -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
$AUDIOLLA_URL/v1/files/uploads/mytrack.wav -o copy.wav
# Delete
curl -X DELETE -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
$AUDIOLLA_URL/v1/files/uploads/mytrack.wav
Path traversal (.., leading /, etc.) is rejected with 400. Symlinks are not followed. Size cap is AUDIOLLA_MAX_UPLOAD_BYTES.
Input and output modes (every audio endpoint)
Every audio endpoint accepts exactly one of two input forms — supplying zero or both returns 400:
file_path— relative path under FILES_DIR (pre-staged viaPUT /v1/files/{path})file_url— remote URL the server fetches (subject to theAUDIOLLA_FETCH_MODEpolicy — see below)
Audio-producing endpoints (separate, master, transform, normalize, fx, restore, enhance, visualize, midi compose/transform/render/generate, melody-as-midi, beats-with-click-track, etc.) require exactly one of:
output_path— server writes the result toFILES_DIR /; response is JSON{path, size, ...}output_url— server PUTs the result to a presigned URL; response is JSON{url, size, ...}
output_path and output_url are mutually exclusive — supplying both is 400. Supplying neither is 400 too (no inline-bytes audio response exists in v1.0.0) — except when async_job=true, which auto-stages to jobs/{job_id}.{ext} if neither is set.
Analysis-only endpoints (/v1/audio/analyze, /v1/audio/onsets, /v1/audio/fingerprint, /v1/audio/loudness, beats without click_track, silence without trim_mode, etc.) ignore output_path / output_url — they return their JSON data directly.
The master endpoint additionally accepts reference_path xor reference_url for the reference track in mode=reference — same exactly-one-of rule.
Remote URLs (file_url / output_url)
The server-side URL fetch is disabled by default. To enable it, the operator sets:
AUDIOLLA_FETCH_MODE = disabled | allowlist | denylist (default: disabled)
AUDIOLLA_FETCH_HOSTS = comma-separated host patterns (required when mode=allowlist)
AUDIOLLA_FETCH_SCHEMES = https,http (default: https only)
AUDIOLLA_FETCH_TIMEOUT = 30s (per fetch/upload)
AUDIOLLA_FETCH_ALLOW_PRIVATE = false (allow private/loopback IPs)
AUDIOLLA_FETCH_MAX_REDIRECTS = 5
Host patterns are exact match (bucket.s3.amazonaws.com) or single-wildcard subdomain (*.s3.amazonaws.com, matches any .s3.amazonaws.com but NOT s3.amazonaws.com itself).
Always-on protections regardless of mode:
- DNS-resolved private / loopback / link-local / metadata-service IPs (
169.254.169.254) rejected unlessAUDIOLLA_FETCH_ALLOW_PRIVATE=true - Only schemes in
AUDIOLLA_FETCH_SCHEMESaccepted;file://,gopher://, etc. always rejected - Each redirect's
Locationre-validated through the
相关技能
把自然语言描述转为结构化 JSON,并由 mcp-diagram-generator MCP 服务生成 Draw.io、Mermaid 或 Excalidraw 图表文件。
通过一次 REST API 调用,向 10 个社交平台发布视频、图片、文字与文档。
从 AdMapix API 拉取广告创意、应用、榜单和收入预估等数据,原样返回结构化 JSON。
在本地磁盘以分类纯 Markdown 文件保存需要长期留存的事实,与智能体内置记忆并存。
按用户明确指令,在得到大脑(Get笔记)中保存、搜索并管理笔记与知识库。
psyb0t 的更多技能
浏览全部技能对接用户自部署的 mt5-httpapi MetaTrader 5 网关,每次涉及真实资金的写操作都必须逐笔确认后再执行。
面向反爬检测栈 QA 与授权测试场景的 Docker 浏览器自动化工具。
自托管、OpenAI 兼容的语音服务,一个容器搞定转写、翻译与合成。
在固定白名单的 SSH 沙箱里跑 ffmpeg、sox、ImageMagick 处理音视频和图片。
通过 SSH 调用 Qwen3-TTS 生成语音,支持预设音色、声音克隆与声音设计。
一个端点统一管控多个 IMAP/SMTP 邮箱,跨账号并行完成读取、检索、发送、标记与删除。