Use SambaNova with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. High-speed AI inference for Llama 3, DeepSeek, and MiniMax models via SambaNova's ultra-fast SN40L chips.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 3 capabilities
The complete SambaNova capability set.
These are the exact actions your AI can choose when you ask it to work with SambaNova.
01-03
3 capabilities in this set.
Part of 3 available through SambaNova.
- 01
Create embedding
Available on SambaStack. Create embeddings using SambaNova
- 02
Create response
Returns typed output items. Create a response using SambaNova Responses API
- 03
Create chat completion
Compatible with OpenAI Chat Completions API. Create a chat completion using SambaNova models
Observed, not estimated
866ms average. Fast in production.
SambaNova is checked daily against the live service.
- Fastest day
- 719ms
- Slowest day
- 1025ms
- 14-day trend
- Stable-4%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 3 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of SambaNova, so you can see the experience inside your AI.
It does not authenticate your account with SambaNova. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
SambaNova Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_z52tsMdkFaoCTCBeg6u7bbxNhhnjcT0r5KXU6hjj/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — SambaNova capabilities are ready to use.
{
"mcpServers": {
"sambanova-ai-inference-mcp": {
"url": "https://edge.vinkius.com/vk_preview_z52tsMdkFaoCTCBeg6u7bbxNhhnjcT0r5KXU6hjj/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions SambaNova owners ask.
- 01
Which models are available for chat completions?
You can use create_chat_completion with models like Meta-Llama-3.3-70B-Instruct, DeepSeek-V3.1, and MiniMax-M2.5 for high-speed text generation.
- 02
Can I generate embeddings for my RAG pipeline?
Yes! Use the create_embedding capability with models like E5-Mistral-7B-Instruct to create vectorized representations of your text data.
- 03
What is the difference between create_chat_completion and create_response?
create_chat_completion follows the standard OpenAI chat format, while create_response is a stateless API designed specifically for agentic workflows, returning typed output items.
Explore
More in Developer Tools
Groq AI Connector
Run large language models at unprecedented speed with custom LPU hardware that delivers real-time AI inference
ViewDeepInfra (Serverless LLM Inference) AI Connector
Run top-tier LLMs, image generation, and embeddings via DeepInfra's serverless infrastructure directly from yo
ViewGroq AI Connector
Run large language models at unprecedented speed with custom LPU hardware that delivers real-time AI inference
ViewCerebras Inference AI Connector
Access lightning-fast AI inference via Cerebras Wafer-Scale Engine — generate chat completions, manage models,
View
Suggestions
Ollama AI Connector
Run LLM models via Ollama cloud API — generate completions, chat with multimodal models, create embeddings, an
ViewAnyscale AI Connector
Orchestrate your Anyscale infrastructure — manage LLM queries, vectors, services, and cluster batch jobs direc
ViewCohere (Embed & Rerank) AI Connector
Empower RAG via Cohere — generate high-quality text embeddings, rerank documents for better accuracy, and perf
ViewHugging Face LLM AI Connector
Connect Hugging Face LLM to any AI agent via MCP.
View
