ClaudeChatGPTPerplexityGeminiMicrosoft CopilotRaycastMeta AIGrokZ.aiQwenKimi
DeepSeekMistralCursorVS CodeWindsurfJetBrainsClineLovableVercel AI SDKLangChain

Use Volcengine Speech with your AI.

Connect your account once and let the AI you already use work with it, without building another integration. The massive 'TikTok Voice' TTS API. generate natural speech with ByteDance's iconic voice models.

Included with plan

Ask AI about this Connector

Developed, maintained, and hosted by Vinkius.

MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED

Waiting for input…

Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.

ChatGPTClaudeCursorPerplexityGeminiMicrosoft CopilotRaycastMeta AI

Complete set · 5 capabilities

The complete Volcengine Speech capability set.

These are the exact actions your AI can choose when you ask it to work with Volcengine Speech.

Capability set01 / 02

01-03

3 capabilities in this set.

Part of 5 available through Volcengine Speech.

  1. 01

    Get audio formats

    Use MP3 for web delivery, WAV for editing, OGG Opus for efficient streaming, or PCM for raw processing. List supported audio output formats

  2. 02

    List voices

    Essential for choosing the right voice before synthesis. Includes the famous TikTok voice styles. List all available TTS voice models

  3. 03

    Synthesize long text

    Ideal for articles, audiobooks, and lengthy documentation. Use this when your text exceeds the standard 1024 character limit. Synthesize speech from long text (over 1024 characters)

Capability set02 / 02

04-05

2 capabilities in this set.

Part of 5 available through Volcengine Speech.

  1. 04

    Synthesize speech

    Supports multiple languages (Chinese, English, Japanese), various voice styles (female, male, child, trendy, news), and adjustable speed/volume. Returns audio data or URL. Ideal for narration, accessibility, multi-language content, and the iconic TikTok voice effects. Convert text to speech using Volcengine TTS

  2. 05

    Synthesize ssml

    Use SSML tags like <break>, <emphasis>, <prosody> for natural-sounding output with precise timing and intonation control. Convert SSML (Speech Synthesis Markup Language) to speech

Observed, not estimated

777ms average. Fast in production.

Volcengine Speech is checked daily against the live service.

Daily averagePeak 998ms
Aug 30Today
Fastest day
741ms
Slowest day
998ms
14-day trend
Stable+1%

Connect your client

One URL. Every client.

Activate the Connector, copy your link, and paste it into the client you already use. 5 capabilities arrive ready to run.

Preview access · not provider authentication

The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Volcengine Speech, so you can see the experience inside your AI.

It does not authenticate your account with Volcengine Speech. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.

Volcengine Speech Connector

You're all set. Choose your MCP client and follow the setup instructions.

Connector linkhttps://edge.vinkius.com/vk_preview_w3fkAFQ6qN01Upq1pce75bdrQJBDCI1PUIsyEF2C/mcp

Claude Desktop

Follow the steps below to connect in seconds.

  1. 1In Claude Desktop, open Settings → Connectors.
  2. 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
  3. 3Click Add and start a new chat — Volcengine Speech capabilities are ready to use.
Configuration · claude_desktop_config.jsonCopy
{
  "mcpServers": {
    "volcengine-speech-synthesis-mcp": {
      "url": "https://edge.vinkius.com/vk_preview_w3fkAFQ6qN01Upq1pce75bdrQJBDCI1PUIsyEF2C/mcp"
    }
  }
}
  • Claude
  • ChatGPT
  • Cursor
  • VS Code
  • Windsurf
  • Claude Code
  • JetBrains
  • Cline

Step-by-step instructions for each client are in the guide. How to connect

FAQ

Questions Volcengine Speech owners ask.

  • 01

    What makes Volcengine TTS different from other TTS services?

    Volcengine powers the iconic TikTok TTS effects used in billions of videos. It offers industry-leading Chinese speech quality, trendy social media voices, and ByteDance's proprietary neural voice technology.

  • 02

    Which languages are supported?

    Chinese (Mandarin), English, Japanese, and more. Use language parameter: 'zh' for Chinese, 'en' for English, 'ja' for Japanese. Each language has multiple voice styles.

  • 03

    What's the max text length?

    Standard synthesis supports up to 1024 characters per request. For longer texts, use the synthesize_long_text capability which automatically handles chunking and combining results for articles and audiobooks.