Use Baseten with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. Manage your Baseten AI models. orchestrate deployments, list secrets, and run serverless inference predictions autonomously.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 6 capabilities
The complete Baseten capability set.
These are the exact actions your AI can choose when you ask it to work with Baseten.
01-03
3 capabilities in this set.
Part of 6 available through Baseten.
- 01
Get model
Get a specific Baseten model
- 02
List models
List Baseten managed models
- 03
List deployments
List active inferences bounds matching a specific model
04-06
3 capabilities in this set.
Part of 6 available through Baseten.
- 04
Predict
Formulate the explicit tensor shapes or dictionaries strictly matching the deployed instance. Invoke a serverless model inference prediction
- 05
Get deployment
Get explicit details of a running deployment
- 06
List secrets
List securely managed workspace secrets without showing values
Observed, not estimated
849ms average. Fast in production.
Baseten is checked daily against the live service.
- Fastest day
- 660ms
- Slowest day
- 1104ms
- 14-day trend
- Stable+2%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 6 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Baseten, so you can see the experience inside your AI.
It does not authenticate your account with Baseten. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
Baseten Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_Iok4x74580t8DfXgX0kQJnXTan6vttePrqZ5XRGU/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — Baseten capabilities are ready to use.
{
"mcpServers": {
"baseten-mcp": {
"url": "https://edge.vinkius.com/vk_preview_Iok4x74580t8DfXgX0kQJnXTan6vttePrqZ5XRGU/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions Baseten owners ask.
- 01
Can the AI agent run a prediction directly against my hosted model?
Yes. By pushing a correctly formatted JSON payload to the 'predict' capability, the agent securely triggers inference on the GPU instances, returning the exact calculated response data transparently to your editor context.
- 02
Is my workspace and environmental secret data kept safe?
Baseten secret fetching natively obscures variable values. When you use 'list_secrets', the agent simply evaluates the key names and identifiers existing across your environment to verify configurations without exposing plaintext passwords.
- 03
How do I check auto-scaling configurations for an explicitly deployed model?
You can examine exactly how instances are managed by using 'get_deployment'. Tell the agent to target an active deployment ID and it maps the scaling limits, replica status, and container bounds out-of-the-box.
Explore
More in AI Frontier
TrueFoundry AI Connector
Universal LLM Gateway & ML deployment hub: invoke 1000+ proxy models and manage MCP service instances natively
ViewReplicate AI Connector
Run ML models via Replicate — generate images, text, audio and video from community models, track predictions
ViewReplicate AI Connector
Automate machine learning workflows via Replicate — run models, manage predictions, and search for AI assets d
ViewModelbit (ML Model Deployments) AI Connector
Deploy and call machine learning models directly from your AI agent using Modelbit's inference endpoints.
View
Suggestions
Metatext AI Connector
No-code NLP and AI model management via Metatext — run inference and manage datasets.
ViewReplicate AI Connector
Equip your AI to dynamically search, run, and monitor thousands of open-source machine learning models hosted
ViewCerebras Inference AI Connector
Access lightning-fast AI inference via Cerebras Wafer-Scale Engine — generate chat completions, manage models,
ViewDeep Talk AI Connector
Equip your AI agent to analyze conversation datasets, extract topics, and monitor sentiment via the Deep Talk
View
