Skip to content
Vinkius

Baseten Connector for AI agents.

6 live capabilities

Manage your ML-Ops deployments and inference nodes from your workspace.

Live agent request Baseten / Connector

Waiting for input…

AI Agent

Why people use Baseten

Baseten MLOps Management for Faster Model Deployment

This Connector puts that entire lifecycle into your agent. You can check your deployment versions or run a prediction with a simple text prompt. It turns your AI into a capable ML-Ops operator that keeps your GPU lifecycle in check without the constant context switching.

  • Claude
  • ChatGPT
  • Gemini
  • Cursor
  • Visual Studio Code
  • Windsurf

What Vinkius changes

You get direct ML-Ops control inside your chat window.

Use it from Claude, ChatGPT, Cursor or another AI client you already have.

One account · 5,900+ Connectors

  1. Real-world use case 01

    Testing a new fine-tuned model

    A researcher asks their agent to run a prediction on a specific model ID using a custom JSON payload to check for accuracy.

  2. Real-world use case 02

    Debugging a failing deployment

    An SRE asks the agent to check the replica states for a specific model to see why a deployment is failing.

  3. Real-world use case 03

    Auditing environment variables

    A developer asks the agent to confirm if a specific secret is provisioned in the workspace without showing the actual key.

Complete set · 6capabilities

The complete Baseten capability set.

These are the exact actions your AI can choose when you ask it to work with Baseten.

Capability set01 / 02

01—03

3 capabilities in this set.

Part of 6 available through Baseten.

  1. 01 Capability

    List models

    See all your managed models in one place. This helps you keep track of your entire model fleet.

  2. 02 Capability

    Get model

    Pull specific details for a single model. Use this to check configurations for a specific model.

  3. 03 Capability

    Predict

    Send tensor payloads or JSON to your GPU weights for a prediction. This runs inference directly.

Capability set02 / 02

04—06

3 capabilities in this set.

Part of 6 available through Baseten.

  1. 04 Capability

    List deployments

    See active inference bounds for a specific model. It helps you see what's currently running.

  2. 05 Capability

    Get deployment

    Get the exact details of a running deployment. This is useful for auditing replica states.

  3. 06 Capability

    List secrets

    See your workspace secrets without exposing the actual values. This keeps your environment secure.

Set up in minutes

One URL. Then ask Baseten to work.

Claude and ChatGPT only need the Connector URL. Copy it once, add it in settings, and use Baseten from the conversation.

Choose your client

Live preview
Advanced clients IDE · CLI

Claude · Web + desktop

Official guide ↗

Connector URL · ready to paste

Streamable HTTP
https://edge.vinkius.com/vk_preview_Iok4x74580t8DfXgX0kQJnXTan6vttePrqZ5XRGU/mcp
  1. Step 01

    Open Connectors

    In Claude Web or Claude Desktop, open Settings and choose Connectors.

  2. Step 02

    Add the URL

    Choose Add custom connector, name it Baseten, and paste the URL above.

  3. Step 03

    Turn it on in chat

    Select +, open Connectors, and enable Baseten for the conversation.

Where the request belongs

Work Baseten can move forward.

Built around the request

This is for the ML engineer who is tired of jumping between the Baseten console and their IDE to check if a model is actually live or to run a quick test payload.

01

ML Engineer

Running test payloads against production deployments without spinning up local notebooks.

02

DevOps/SRE

Auditing running deployment resources and verifying replica states from a single command.

03

AI Researcher

Inspecting version schemas and managing inference pipeline architectures quickly.

Bring your own AI

Change the model, client or framework. Keep Baseten connected.

  • Claude
  • ChatGPT
  • Gemini
  • Cursor
  • VS Code
  • Windsurf
  • ZCode
  • Cline
  • Zed
  • Continue
  • Kiro
  • Roo Code
  • Zencoder
  • Goose
  • Void
  • Augment Code
  • Amp
  • Qodo
  • Tabnine
  • Pieces
  • Sourcegraph Cody
  • JetBrains
  • Warp
  • Amazon Q
  • Antigravity
  • BoltAI
  • Raycast
  • Jan
  • LM Studio
  • AnythingLLM
  • Open WebUI
  • Msty
  • Cherry Studio
  • LibreChat
  • TypingMind
  • Chorus
  • 5ire
  • n8n
  • LangChain
  • LlamaIndex
  • CrewAI
  • Vercel AI SDK

Before you connect

Questions about Baseten.

The practical details behind the request, access and result.

Can the Baseten MCP run my models?

Yes, it lets your agent send payloads to your GPU weights for real-time inference. You just describe the input to your agent, and it handles the prediction call.

Does it show my secret values?

No, it only lists the names of your secrets to keep your environment secure. It confirms they are mapped without exposing the actual keys.

Can I use this with Cursor or Claude?

Yes, it works with any MCP-compatible client including Claude, Cursor, Windsurf, and VS Code.

How do I run a prediction using this Connector?

Simply tell your agent what you want to predict. It will use the predict capability to send the data to your Baseten instance and show you the result.

Can it check my deployment status?

Yes, it can pull exact details on replica states and autoscaling configurations, so you can monitor your production environment easily.

Can the AI agent run a prediction directly against my hosted model?

Yes. By pushing a correctly formatted JSON payload to the 'predict' capability, the agent securely triggers inference on the GPU instances, returning the exact calculated response data transparently to your editor context.

Is my workspace and environmental secret data kept safe?

Baseten secret fetching natively obscures variable values. When you use 'list_secrets', the agent simply evaluates the key names and identifiers existing across your environment to verify configurations without exposing plaintext passwords.

How do I check auto-scaling configurations for an explicitly deployed model?

You can examine exactly how instances are managed by using 'get_deployment'. Tell the agent to target an active deployment ID and it maps the scaling limits, replica status, and container bounds out-of-the-box.

One connection away

Give your agent a direct line to Baseten.

Connect Baseten once. Keep it beside 5,900+ managed Connectors when the next task needs more.

Explore every Connector No credit card required · Free tier available