ClaudeChatGPTPerplexityGeminiMicrosoft CopilotRaycastMeta AIGrokZ.aiQwenKimi
DeepSeekMistralCursorVS CodeWindsurfJetBrainsClineLovableVercel AI SDKLangChain

Use Baseten with your AI.

Connect your account once and let the AI you already use work with it, without building another integration. Manage your Baseten AI models. orchestrate deployments, list secrets, and run serverless inference predictions autonomously.

Included with plan

Ask AI about this Connector

Developed, maintained, and hosted by Vinkius.

MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED

Waiting for input…

Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.

ChatGPTClaudeCursorPerplexityGeminiMicrosoft CopilotRaycastMeta AI

Complete set · 6 capabilities

The complete Baseten capability set.

These are the exact actions your AI can choose when you ask it to work with Baseten.

Capability set01 / 02

01-03

3 capabilities in this set.

Part of 6 available through Baseten.

  1. 01

    Get model

    Get a specific Baseten model

  2. 02

    List models

    List Baseten managed models

  3. 03

    List deployments

    List active inferences bounds matching a specific model

Capability set02 / 02

04-06

3 capabilities in this set.

Part of 6 available through Baseten.

  1. 04

    Predict

    Formulate the explicit tensor shapes or dictionaries strictly matching the deployed instance. Invoke a serverless model inference prediction

  2. 05

    Get deployment

    Get explicit details of a running deployment

  3. 06

    List secrets

    List securely managed workspace secrets without showing values

Observed, not estimated

849ms average. Fast in production.

Baseten is checked daily against the live service.

Daily averagePeak 1104ms
Aug 20Today
Fastest day
660ms
Slowest day
1104ms
14-day trend
Stable+2%

Connect your client

One URL. Every client.

Activate the Connector, copy your link, and paste it into the client you already use. 6 capabilities arrive ready to run.

Preview access · not provider authentication

The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Baseten, so you can see the experience inside your AI.

It does not authenticate your account with Baseten. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.

Baseten Connector

You're all set. Choose your MCP client and follow the setup instructions.

Connector linkhttps://edge.vinkius.com/vk_preview_Iok4x74580t8DfXgX0kQJnXTan6vttePrqZ5XRGU/mcp

Claude Desktop

Follow the steps below to connect in seconds.

  1. 1In Claude Desktop, open Settings → Connectors.
  2. 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
  3. 3Click Add and start a new chat — Baseten capabilities are ready to use.
Configuration · claude_desktop_config.jsonCopy
{
  "mcpServers": {
    "baseten-mcp": {
      "url": "https://edge.vinkius.com/vk_preview_Iok4x74580t8DfXgX0kQJnXTan6vttePrqZ5XRGU/mcp"
    }
  }
}
  • Claude
  • ChatGPT
  • Cursor
  • VS Code
  • Windsurf
  • Claude Code
  • JetBrains
  • Cline

Step-by-step instructions for each client are in the guide. How to connect

FAQ

Questions Baseten owners ask.

  • 01

    Can the AI agent run a prediction directly against my hosted model?

    Yes. By pushing a correctly formatted JSON payload to the 'predict' capability, the agent securely triggers inference on the GPU instances, returning the exact calculated response data transparently to your editor context.

  • 02

    Is my workspace and environmental secret data kept safe?

    Baseten secret fetching natively obscures variable values. When you use 'list_secrets', the agent simply evaluates the key names and identifiers existing across your environment to verify configurations without exposing plaintext passwords.

  • 03

    How do I check auto-scaling configurations for an explicitly deployed model?

    You can examine exactly how instances are managed by using 'get_deployment'. Tell the agent to target an active deployment ID and it maps the scaling limits, replica status, and container bounds out-of-the-box.