Use Cartesia with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. Generate lifelike AI voices, clone speech, and transcribe audio with Cartesia's state-of-the-art Sonic models directly from your AI agent.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 20 capabilities
The complete Cartesia capability set.
These are the exact actions your AI can choose when you ask it to work with Cartesia.
01-04
4 capabilities in this set.
Part of 20 available through Cartesia.
- 01
Get usage credits
Get credit usage statistics
- 02
Create pronunciation dict
Create a new pronunciation dictionary
- 03
Delete pronunciation dict
Delete a pronunciation dictionary
- 04
Delete voice
Delete a voice
05-08
4 capabilities in this set.
Part of 20 available through Cartesia.
- 05
Get agent
Get details for a specific voice agent
- 06
List agent calls
List calls and transcripts for a specific agent
- 07
List agents
List all voice agents
- 08
List pronunciation dicts
List pronunciation dictionaries
09-12
4 capabilities in this set.
Part of 20 available through Cartesia.
- 09
List voices
List available voices
- 10
Clone voice
Clone a voice from a 5s audio clip
- 11
Generate access token
Generate a short-lived access token for client-side requests
- 12
Get voice
Get details for a specific voice
13-16
4 capabilities in this set.
Part of 20 available through Cartesia.
- 13
Infill bytes
Generate audio to smoothly connect two existing segments
- 14
Localize voice
Adapt a voice to a new language/dialect
- 15
Update voice
Update voice metadata
- 16
Update pronunciation dict
Update a pronunciation dictionary
17-20
4 capabilities in this set.
Part of 20 available through Cartesia.
- 17
Voice changer bytes
Change voice of an audio clip while preserving intonation
- 18
Stt batch
Transcribe audio file to text (Batch STT)
- 19
Tts bytes
Generate text-to-speech audio bytes
- 20
Tts sse
Generate text-to-speech via Server-Sent Events
Observed, not estimated
1060ms average. Fast in production.
Cartesia is checked daily against the live service.
- Fastest day
- 840ms
- Slowest day
- 1291ms
- 14-day trend
- Slowing+31%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 20 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Cartesia, so you can see the experience inside your AI.
It does not authenticate your account with Cartesia. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
Cartesia Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_fqUkC4o6TEWq9kAU8DlY8MAb7pfJxo3R952S85p7/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — Cartesia capabilities are ready to use.
{
"mcpServers": {
"cartesia-voice-ai-mcp": {
"url": "https://edge.vinkius.com/vk_preview_fqUkC4o6TEWq9kAU8DlY8MAb7pfJxo3R952S85p7/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions Cartesia owners ask.
- 01
Can I generate audio in different formats like MP3 or WAV?
Yes. Using the tts_bytes capability, you can specify the output_format_container as 'mp3', 'wav', or 'raw', and configure the sample rate and encoding to match your needs.
- 02
How do I transcribe an existing audio file to text?
Use the stt_batch capability. Provide the base64 encoded audio file, specify the model (e.g., 'ink-whisper'), and the language code to receive a full transcription.
- 03
Is it possible to clone a voice using this integration?
Absolutely. The clone_voice capability allows you to create a new voice model by uploading a short (approx. 5s) base64 encoded audio clip.
Explore
More in AI Frontier
ElevenLabs AI Connector
Generate lifelike speech, clone voices, and create sound effects using ElevenLabs' industry-leading AI audio t
ViewNVIDIA Audio AI Connector
Transcribe speech, generate voices, translate audio, and clone voices via NVIDIA Audio APIs.
ViewCAMB.AI AI Connector
Translate and dub audio content into dozens of languages using AI voices that sound natural and preserve speak
ViewElevenLabs AI Connector
Generate lifelike speech from text with neural voice synthesis that clones voices and supports dozens of langu
View
Suggestions
iFLYTEK Open Platform / 讯飞开放平台 AI Connector
China's leading voice and NLP platform — convert speech to text, synthesize voice, and analyze text via AI.
ViewPlay.ht (Voice Cloning) AI Connector
Generate ultra-realistic speech and clone voices instantly using Play.ht's advanced AI voice engines directly
ViewMaestra AI Connector
Automate transcription, translation, and AI voiceovers via the Maestra.ai REST API.
ViewSpiritme AI Connector
Create AI-generated videos with digital human presenters that deliver personalized messages in multiple langua
View
