Use DeepInfra with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. Run top-tier LLMs, image generation, and embeddings via DeepInfra's serverless infrastructure directly from your AI agent.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 4 capabilities
The complete DeepInfra capability set.
These are the exact actions your AI can choose when you ask it to work with DeepInfra.
01-04
4 capabilities in this set.
Part of 4 available through DeepInfra.
- 01
Run native inference
Useful for models not covered by OpenAI spec (e.g., speech-to-text, OCR, video generation, or private deployments). Run native inference for a specific model on DeepInfra
- 02
Create embedding
Create embeddings for text via DeepInfra
- 03
Generate image
Generate an image from a text prompt via DeepInfra
- 04
Create chat completion
Provide model name (e.g., deepseek-ai/DeepSeek-V3) and messages array. Create a chat completion using an LLM via DeepInfra
Observed, not estimated
843ms average. Fast in production.
DeepInfra is checked daily against the live service.
- Fastest day
- 690ms
- Slowest day
- 1045ms
- 14-day trend
- Slowing+11%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 4 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of DeepInfra, so you can see the experience inside your AI.
It does not authenticate your account with DeepInfra. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
DeepInfra Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_RMvS2UlIMhdShOjNbZfqar6XAZE9QfzTMbaejxJY/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — DeepInfra capabilities are ready to use.
{
"mcpServers": {
"deepinfra-serverless-llm-inference-mcp": {
"url": "https://edge.vinkius.com/vk_preview_RMvS2UlIMhdShOjNbZfqar6XAZE9QfzTMbaejxJY/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions DeepInfra owners ask.
- 01
Which LLM models can I use with the chat capability?
You can use any model hosted on DeepInfra, such as deepseek-ai/DeepSeek-V3 or meta-llama/Llama-3.3-70B-Instruct, by passing the model name to the create_chat_completion capability.
- 02
How do I generate images using FLUX or Stable Diffusion?
Use the generate_image capability. Simply provide the model name (e.g., black-forest-labs/FLUX-1-schnell) and your text prompt to receive the generated image URL.
- 03
What is the 'run_native_inference' capability used for?
It is used for models that don't follow the OpenAI chat/image spec, such as audio transcription (Whisper), specialized OCR models, or your own private model deployments on DeepInfra.
Explore
More in Developer Tools
Cerebras Inference AI Connector
Access lightning-fast AI inference via Cerebras Wafer-Scale Engine — generate chat completions, manage models,
ViewLeonardo.ai (Generative AI & Models) AI Connector
Generate high-fidelity images via Leonardo.ai — orchestrate generations, audit AI models, and manage visual as
ViewSambaNova (AI Inference) AI Connector
High-speed AI inference for Llama 3, DeepSeek, and MiniMax models via SambaNova's ultra-fast SN40L chips.
ViewGroq AI Connector
Run large language models at unprecedented speed with custom LPU hardware that delivers real-time AI inference
View
Suggestions
Ollama AI Connector
Run LLM models via Ollama cloud API — generate completions, chat with multimodal models, create embeddings, an
ViewHugging Face LLM AI Connector
Connect Hugging Face LLM to any AI agent via MCP.
ViewGroq AI Connector
Run large language models at unprecedented speed with custom LPU hardware that delivers real-time AI inference
ViewAnyscale AI Connector
Orchestrate your Anyscale infrastructure — manage LLM queries, vectors, services, and cluster batch jobs direc
View
