Use NVIDIA API Catalog with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. Cloud Engine proxy running native foundational completions natively utilizing active Nemotron and Llama3 architectures.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 8 capabilities
The complete NVIDIA API Catalog capability set.
These are the exact actions your AI can choose when you ask it to work with NVIDIA API Catalog.
01-04
4 capabilities in this set.
Part of 8 available through NVIDIA API Catalog.
- 01
Nvidia list foundation models
Dumps the strict array specifying explicit LLM matrix paths accessible securely natively
- 02
Nvidia list lora adapters
Evaluate explicit matrices tracking fine-tuned overrides isolating logical constraints dynamically
- 03
Nvidia vision inference
G. Llama-Vision natively). Invoke strictly multimodal abilities capturing diagnostic constraints returning inference on graphical data
- 04
Nvidia chat completion
Trigger direct NLP inference matrices directly evaluating queries over hosted LLMs
05-08
4 capabilities in this set.
Part of 8 available through NVIDIA API Catalog.
- 05
Nvidia check token quota
Poll safely dynamic credit and explicit constraint execution limits bounding inference execution
- 06
Nvidia generate embeddings
Pass parameters safely mapping explicit unstructured vectors directly using specific Embedding arrays
- 07
Nvidia get cloud status
Ping explicitly the core hosted NVIDIA matrix tracing inference endpoints evaluating latencies securely
- 08
Nvidia summarize content
Standard natively configured logical execution executing predefined abstract compression matrices smoothly
Observed, not estimated
826ms average. Fast in production.
NVIDIA API Catalog is checked daily against the live service.
- Fastest day
- 677ms
- Slowest day
- 1029ms
- 14-day trend
- Slowing+15%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 8 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of NVIDIA API Catalog, so you can see the experience inside your AI.
It does not authenticate your account with NVIDIA API Catalog. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
NVIDIA API Catalog Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_LgfbZayzkE3ZBfq9HWv3rgeK69AXP0h3anUpakUy/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — NVIDIA API Catalog capabilities are ready to use.
{
"mcpServers": {
"nvidia-api-catalog-mcp": {
"url": "https://edge.vinkius.com/vk_preview_LgfbZayzkE3ZBfq9HWv3rgeK69AXP0h3anUpakUy/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions NVIDIA API Catalog owners ask.
- 01
Can I explicitly route specific embedding vectors natively using the NVIDIA integration matrix?
Yes! Utilize generate_embeddings providing explicit logic extracting arrays natively isolating endpoints safely.
- 02
How do I explicitly explore active LLMs natively hosted inside the NVIDIA catalog bounds?
Target explicit matrices natively calling list_foundation_models returning catalog endpoints safely explicitly mapping bounds secure natively.
- 03
Does this require local Docker execution mapping explicitly NVIDIA parameters transparently?
No, this explicitly pings the hosted Cloud API. For local Docker metrics natively, switch to nvidia-nim-mcp enforcing natively local boundaries.
Explore
More in Industry Titans
Fireworks AI AI Connector
Empower LLM applications via Fireworks AI — perform ultra-fast chat completions, generate embeddings and image
ViewNVIDIA NIM AI Connector
MLOps proxy unifying explicitly local hardware limits extracting telemetry across active NVIDIA AI containers.
ViewGradient AI (LLM API & Finetuning) AI Connector
Access powerful LLMs, fine-tune models on your own data, and generate embeddings directly through your AI agen
ViewLangfuse (LLM Tracing & Evals) AI Connector
Monitor LLM apps via Langfuse — track traces, manage prompt templates, and audit evaluation scores.
View
Suggestions
Open WebUI AI Connector
Manage your Open WebUI instance — list models, handle chat completions, and manage RAG collections directly fr
ViewAI Memory Cost Analyzer AI Connector
Estimate and optimize the economic impact of AI conversation memory architectures.
ViewInworld AI AI Connector
Power your AI agents with Inworld's lifelike voices, voice cloning, and advanced character orchestration route
ViewEden AI AI Connector
Equip your AI agent to manage unified AI workflows, track providers, and monitor API usage via the Eden AI pla
View
