Volcengine Speech Synthesis MCP Server
The massive 'TikTok Voice' TTS API — generate natural speech with ByteDance's iconic voice models.
Ask AI about this MCP Server
Vinkius AI Gateway supports streamable HTTP and SSE.

Works with every AI agent you already use
…and any MCP-compatible client


















What is the Volcengine Speech Synthesis MCP Server?
The Volcengine Speech MCP Server gives AI agents like Claude, ChatGPT, and Cursor direct access to Volcengine Speech. The massive 'TikTok Voice' TTS API — generate natural speech with ByteDance's iconic voice models. Powered by the Vinkius AI Gateway — no API keys, no infrastructure, connect in under 2 minutes.
Volcengine Speech MCP Server: see your AI Agent in action
Built-in capabilities (7)
create_custom_voice
Requires 10-50 high-quality audio recordings of a single speaker. Training takes 1-3 days. Once complete, use the custom voice_type in synthesize_speech. Create a custom voice model from training audio samples
get_audio_formats
Use MP3 for web delivery, WAV for editing, OGG Opus for efficient streaming, or PCM for raw processing. List supported audio output formats
get_task_status
Returns whether processing, completed, or failed. Check status of an async TTS task
list_voices
Essential for choosing the right voice before synthesis. Includes the famous TikTok voice styles. List all available TTS voice models
synthesize_long_text
Ideal for articles, audiobooks, and lengthy documentation. Use this when your text exceeds the standard 1024 character limit. Synthesize speech from long text (over 1024 characters)
synthesize_speech
Supports multiple languages (Chinese, English, Japanese), various voice styles (female, male, child, trendy, news), and adjustable speed/volume. Returns audio data or URL. Ideal for narration, accessibility, multi-language content, and the iconic TikTok voice effects. Convert text to speech using Volcengine TTS
synthesize_ssml
Use SSML tags like <break>, <emphasis>, <prosody> for natural-sounding output with precise timing and intonation control. Convert SSML (Speech Synthesis Markup Language) to speech
What this connector unlocks
Connect Volcengine Speech Synthesis (ByteDance's TTS platform) to any AI agent and generate stunning natural speech — including the iconic TikTok voices — through natural conversation.
What you can do
- Text-to-Speech — Convert any text to natural-sounding speech
- TikTok Voices — Use the exact voice models behind TikTok's viral TTS effects
- Multi-Language — Synthesize in Chinese, English, Japanese, and more
- SSML Support — Fine-grained control with pauses, emphasis, and prosody
- Long-Form Audio — Synthesize articles, audiobooks, and lengthy documents
- Custom Voices — Train personalized voice models from audio samples
- Speed/Volume Control — Adjust speech rate and volume dynamically
How it works
1. Subscribe to this server 2. Enter your Volcengine Access Key and Secret Key 3. Start generating speech from Claude, Cursor, or any MCP clientWho is this for?
- Content Creators — Generate voiceovers for videos, reels, and TikToks
- Accessibility Teams — Add speech output to apps and websites
- Audiobook Producers — Convert long-form text to natural narration
- Developers — Integrate TikTok-quality TTS into applications
Frequently asked questions
Give your AI agents the power of Volcengine Speech
Access Volcengine Speech and 2,500+ MCP servers — ready for your agents to use, right now. No glue code. No custom integrations. Just plug Vinkius AI Gateway and let your agents work.
