Use Cerebras Inference with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. Access lightning-fast AI inference via Cerebras Wafer-Scale Engine. generate chat completions, manage models, and run batch jobs at record speeds.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 15 capabilities
The complete Cerebras Inference capability set.
These are the exact actions your AI can choose when you ask it to work with Cerebras Inference.
01-04
4 capabilities in this set.
Part of 15 available through Cerebras Inference.
- 01
Create completion
Generate text continuations from a single prompt string
- 02
Delete file
Delete a file
- 03
Get batch
Retrieve status of a batch job
- 04
Get metrics
Retrieve Prometheus-formatted operational metrics
05-08
4 capabilities in this set.
Part of 15 available through Cerebras Inference.
- 05
Get model
Fetches details for a specific model
- 06
List batches
List all batch jobs
- 07
List files
List uploaded files
- 08
List public models
Retrieve model details without an API key
09-12
4 capabilities in this set.
Part of 15 available through Cerebras Inference.
- 09
Cancel batch
Cancel a batch job
- 10
Create batch
Create a batch job for asynchronous processing
- 11
Create chat completion
Generate conversational responses using a structured message format
- 12
Get file
Retrieve metadata for a specific file
13-15
3 capabilities in this set.
Part of 15 available through Cerebras Inference.
- 13
Get file content
Download raw content of a file
- 14
List models
Lists all currently available models
- 15
Upload file
Upload a JSONL file for Batch processing
Observed, not estimated
961ms average. Fast in production.
Cerebras Inference is checked daily against the live service.
- Fastest day
- 744ms
- Slowest day
- 1301ms
- 14-day trend
- Slowing+33%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 15 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Cerebras Inference, so you can see the experience inside your AI.
It does not authenticate your account with Cerebras Inference. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
Cerebras Inference Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_xAZFRBg8VLAlPUucEoLSiFDXRg3Jdk9xiofskaFP/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — Cerebras Inference capabilities are ready to use.
{
"mcpServers": {
"cerebras-inference-mcp": {
"url": "https://edge.vinkius.com/vk_preview_xAZFRBg8VLAlPUucEoLSiFDXRg3Jdk9xiofskaFP/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions Cerebras Inference owners ask.
- 01
How do I check which models are available for inference?
Use the list_models capability. It will return a list of all supported models, including high-performance options like Llama 3.1, which you can then use in create_chat_completion.
- 02
Can I process thousands of requests at once?
Yes. Use upload_file to provide your JSONL data and then create_batch to start an asynchronous processing job. You can monitor progress with get_batch.
- 03
Does this server support capability calling and structured outputs?
Yes. The create_chat_completion capability supports capabilities, tool_choice, and response_format parameters, allowing the model to interact with other functions or return valid JSON.
Explore
More in AI Frontier
DeepInfra (Serverless LLM Inference) AI Connector
Run top-tier LLMs, image generation, and embeddings via DeepInfra's serverless infrastructure directly from yo
ViewAnyscale AI Connector
Orchestrate your Anyscale infrastructure — manage LLM queries, vectors, services, and cluster batch jobs direc
ViewGroq AI Connector
Run large language models at unprecedented speed with custom LPU hardware that delivers real-time AI inference
ViewForefront AI Connector
Access Forefront AI models directly from your agent — generate chat completions, manage fine-tuning jobs, and
View
Suggestions
Leonardo.ai (Generative AI & Models) AI Connector
Generate high-fidelity images via Leonardo.ai — orchestrate generations, audit AI models, and manage visual as
ViewMonster API (Serverless GPU & AI Model Hosting) AI Connector
Access powerful AI models for image generation, text-to-speech, and transcription via serverless GPU infrastru
ViewMetatext AI Connector
No-code NLP and AI model management via Metatext — run inference and manage datasets.
ViewBaidu Qianfan AI Connector
Orchestrate Baidu Qianfan AI models — manage chat completions, embeddings, and prompt templates directly from
View
