Helicone (LLM Observability) Connector for AI agents.
10 live capabilities
Monitor LLM costs, latency, and prompt performance in real time.
Waiting for input…
Why people use Helicone (LLM Observability)
Helicone for Real-Time LLM Cost and Latency Tracking
With this Connector, you just ask your agent 'Which feature spent the most yesterday?' It pulls the data from Helicone and gives you the answer instantly. You get to stay in your flow and make decisions based on real numbers without the manual export dance.
What Vinkius changes
You get a conversational interface for your entire LLM observability stack.
Use it from Claude, ChatGPT, Cursor or another AI client you already have.
One account · 5,900+ Connectors
- Real-world use case 01
The 'Where's the money going?' check
A product owner asks the agent to find out which feature is costing the most and gets a breakdown by tag using query_costs.
- Real-world use case 02
The 'Why is it so slow?' investigation
An engineer asks for the 10 slowest requests from the last hour to find a bottleneck in a specific provider using query_latency.
- Real-world use case 03
The 'Why did it hallucinate?' debug
A dev uses the agent to trace a specific user session to see the exact prompts that led to a bad output using query_sessions.
Complete set · 10capabilities
The complete Helicone (LLM Observability) capability set.
These are the exact actions your AI can choose when you ask it to work with Helicone (LLM Observability).
01—04
4 capabilities in this set.
Part of 10 available through Helicone (LLM Observability).
- 01 Capability
Query costs
See a breakdown of the properties driving your account spending. This helps you identify exactly which features or models are consuming your budget.
- 02 Capability
Query sessions
Enumerate the rules exporting active billing. Use this to trace multi-turn sessions and see the billing rules for those calls.
- 03 Capability
Query users
Run a validation check to route gateway history. It helps you identify your most active human clients and their interaction history.
- 04 Capability
Get prompt versions
Extract flags from prompt validations. Use this to see every version of a specific prompt and its history.
05—07
3 capabilities in this set.
Part of 10 available through Helicone (LLM Observability).
- 05 Capability
Query feedback
Inspect the internal data used to mitigate plan math issues. It lets you see the underlying logic for user feedback and critiques.
- 06 Capability
Query latency
Get a JSON payload for customer bindings on latency. It helps you pinpoint which providers are causing high Time To First Token delays.
- 07 Capability
Log feedback
Identify the active arrays for native hold parsing. Use this to pull the specific feedback logs you need to improve model grounding.
08—10
3 capabilities in this set.
Part of 10 available through Helicone (LLM Observability).
- 08 Capability
Query prompts
Retrieve cloud logging traces for vault limits. It helps you see the history and specific limits of your prompts.
- 09 Capability
List properties
Identify the active arrays spanning gateway authentication. Use this to see the properties and settings of your gateway.
- 10 Capability
Query requests
Find bounded records inside the Helicone platform. It lets you see specific request logs to debug issues.
Set up in minutes
One URL. Then ask Helicone (LLM Observability) to work.
Claude and ChatGPT only need the Connector URL. Copy it once, add it in settings, and use Helicone (LLM Observability) from the conversation.
Choose your client
Live previewAdvanced clients IDE · CLI
Claude · Web + desktop
Connector URL · ready to paste
Streamable HTTPhttps://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp - Step 01
Open Connectors
In Claude Web or Claude Desktop, open Settings and choose Connectors.
- Step 02
Add the URL
Choose Add custom connector, name it Helicone (LLM Observability), and paste the URL above.
- Step 03
Turn it on in chat
Select +, open Connectors, and enable Helicone (LLM Observability) for the conversation.
ChatGPT · Web + desktop
Connector URL · ready to paste
Streamable HTTPhttps://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp - Step 01
Open MCP settings
On desktop, open Settings and MCP servers. On web, open your workspace app or connector settings.
- Step 02
Add the URL
Choose Add server with Streamable HTTP, or create a custom MCP app, then paste the Helicone (LLM Observability) URL.
- Step 03
Save and start
Save the connection and enable Helicone (LLM Observability) in your conversation. Desktop may ask you to restart once.
Cursor · IDE configuration
Advanced setup
{
"mcpServers": {
"helicone-llm-observability": {
"url": "https://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp"
}
}
} - Step 01
Open MCP Settings
Press Cmd+Shift+P (macOS) or Ctrl+Shift+P (Windows/Linux) → search "MCP Settings"
- Step 02
Add the server config
Paste the JSON configuration above into the mcp.json file that opens
- Step 03
Save the file
Cursor will automatically detect the new Connector
- Step 04
Start using Helicone (LLM Observability)
Open Agent mode in chat and ask: "Using Helicone (LLM Observability), help me...". 10 tools available
VS Code Copilot · IDE configuration
Advanced setup
{
"mcpServers": {
"helicone-llm-observability": {
"url": "https://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp"
}
}
} - Step 01
Create MCP config
Create a .vscode/mcp.json file in your project root
- Step 02
Add the server config
Paste the JSON configuration above
- Step 03
Enable Agent mode
Open GitHub Copilot Chat and switch to Agent mode using the dropdown
- Step 04
Start using Helicone (LLM Observability)
Ask Copilot: "Using Helicone (LLM Observability), help me...". 10 tools available
Windsurf · IDE configuration
Advanced setup
{
"mcpServers": {
"helicone-llm-observability": {
"url": "https://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp"
}
}
} - Step 01
Open MCP Settings
Go to Settings → MCP Configuration or press Cmd+Shift+P and search "MCP"
- Step 02
Add the server
Paste the JSON configuration above into mcp_config.json
- Step 03
Save and reload
Windsurf will detect the new server automatically
- Step 04
Start using Helicone (LLM Observability)
Open Cascade and ask: "Using Helicone (LLM Observability), help me...". 10 tools available
Cline · IDE configuration
Advanced setup
{
"mcpServers": {
"helicone-llm-observability": {
"url": "https://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp"
}
}
} - Step 01
Open Cline MCP Settings
Click the Connectors icon in the Cline sidebar panel
- Step 02
Add remote server
Click "Add Connector" and paste the configuration above
- Step 03
Enable the server
Toggle the server switch to ON
- Step 04
Start using Helicone (LLM Observability)
Ask Cline: "Using Helicone (LLM Observability), help me...". 10 tools available
Claude Code · Terminal command
Advanced setup
claude mcp add helicone-llm-observability --transport http "https://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp" - Step 01
Install Claude Code
Run npm install -g @anthropic-ai/claude-code if not already installed
- Step 02
Add the Connector
Run the command above in your terminal
- Step 03
Verify the connection
Run claude mcp to list connected servers, or type /mcp inside a session
- Step 04
Start using Helicone (LLM Observability)
Ask Claude: "Using Helicone (LLM Observability), show me...". 10 tools are ready
Where the request belongs
Work Helicone can move forward.
This is for the engineers and product leads who are tired of hunting through logs to figure out why an AI feature is slow or expensive.
LLM Engineer
Debugs prompt performance and checks TTFT latency across different providers on a Tuesday afternoon.
Product Owner
Monitors the AI burn rate and calculates costs per feature to report to stakeholders.
Data Scientist
Analyzes user feedback and thumbs down critiques to improve model grounding and accuracy.
DevOps Engineer
Ensures the AI gateway stays up and checks the reliability of the proxy layers.
When one Connector is not enough
Carry the request into a workflow.
Combine Helicone with the systems that finish the task.
View all recipesCut AI Model Costs Without Losing Quality via MCP
Your GPT-4o bill is $4,200/month and 60% of those calls could run on Groq for $0.003 , your agent finds the waste
Monitor AI Agent Performance Using Connectors
Your agents run in production but you cannot explain why one failed at 3am , fix that
Track LLM Cost vs Quality Using Connectors
Your OpenAI bill grew from $200 to $2,400 in 2 months and you have no idea which feature caused it , because you track API spend at the account level, not at the prompt level
Build the capability set
Add more capabilities.
Each Connector adds new actions and data without changing how you work.
Browse ConnectorsPortkey
AI gateway observability: monitor logs, costs, and manage LLM configurations via agents.
Datadog AI (LLM Observability)
Monitor LLM performance via Datadog. track token usage, audit prompts, and monitor AI model metrics directly from any AI agent.
Keywords AI
Monitor and optimize your LLM API usage with a unified gateway that tracks costs, latency, and model performance across providers.
LangSmith
Observability and evaluation platform for LLM applications. monitor traces, debug agent runs, and track performance metrics across your AI stack.
LangSmith (LLM Observability & Hub)
Monitor LLM apps via LangSmith. track traces, audit prompt templates, and manage evaluation datasets.
Langfuse (LLM Tracing & Evals)
Monitor LLM apps via Langfuse. track traces, manage prompt templates, and audit evaluation scores.
Bring your own AI
Change the model, client or framework. Keep Helicone connected.
-
Claude -
ChatGPT -
Gemini -
Cursor -
VS Code -
Windsurf -
ZCode -
Cline -
Zed -
Continue -
Kiro -
Roo Code -
Zencoder -
Goose -
Void -
Augment Code -
Amp -
Qodo -
Tabnine -
Pieces -
Sourcegraph Cody -
JetBrains -
Warp -
Amazon Q -
Antigravity -
BoltAI -
Raycast -
Jan -
LM Studio -
AnythingLLM -
Open WebUI -
Msty -
Cherry Studio -
LibreChat -
TypingMind -
Chorus -
5ire -
n8n -
LangChain -
LlamaIndex -
CrewAI -
Vercel AI SDK
Before you connect
Questions about Helicone.
The practical details behind the request, access and result.
How does the Helicone MCP help with my AI budget?
It gives you a direct way to query your spending. You can ask your agent for cost breakdowns by model, user, or custom tags to see exactly where your money is going.
Can I use the Helicone MCP to find slow prompts?
Yes. You can ask your agent to find the slowest requests from a specific timeframe. It will pull the latency data and can even show you the prompt that caused the delay.
Does the Helicone MCP support multi-turn conversation tracing?
It does. You can ask your agent to pull specific sessions to see the full history of a conversation. This is great for debugging complex agentic workflows.
How do I see user feedback through the Helicone MCP?
You can ask your agent to pull thumbs up or down critiques. It retrieves the logged feedback so you can see what users liked or disliked about your AI's responses.
Can the Helicone MCP show me different prompt versions?
Yes. You can ask your agent to list all versions of a specific prompt. It will show you the history of your instructions and when each version was deployed.
Is the Helicone MCP good for monitoring my AI gateway?
It's designed for that. It connects to your Helicone account to provide a conversational way to check request logs, latency, and gateway performance.
Can I see the exact prompt that caused a specific error?
Yes. Use the query_requests capability to fetch direct prompts and outputs from the proxy logs. You can filter by status or custom tags to find the exact interaction that needs debugging.
How do I track costs for a specific customer ID?
Ask your agent to query_costs and include your customer identity in the filter. Helicone maps costs per model and user, allowing you to see exactly how much each client is burning in LLM tokens.
Can my agent log human feedback into Helicone?
Absolutely. Use the log_feedback capability to inject offline Human-in-the-Loop verdicts or text critiques directly into Helicone's database, helping you refine your model's grounding over time.
One connection away
Give your agent a direct line to Helicone.
Connect Helicone once. Keep it beside 5,900+ managed Connectors when the next task needs more.
Explore every Connector No credit card required · Free tier available