Skip to content
Vinkius

Helicone (LLM Observability) Connector for AI agents.

10 live capabilities

Monitor LLM costs, latency, and prompt performance in real time.

Live agent request Helicone (LLM Observability) / Connector

Waiting for input…

AI Agent

Why people use Helicone (LLM Observability)

Helicone for Real-Time LLM Cost and Latency Tracking

With this Connector, you just ask your agent 'Which feature spent the most yesterday?' It pulls the data from Helicone and gives you the answer instantly. You get to stay in your flow and make decisions based on real numbers without the manual export dance.

  • Claude
  • ChatGPT
  • Gemini
  • Cursor
  • Visual Studio Code
  • Windsurf

What Vinkius changes

You get a conversational interface for your entire LLM observability stack.

Use it from Claude, ChatGPT, Cursor or another AI client you already have.

One account · 5,900+ Connectors

  1. Real-world use case 01

    The 'Where's the money going?' check

    A product owner asks the agent to find out which feature is costing the most and gets a breakdown by tag using query_costs.

  2. Real-world use case 02

    The 'Why is it so slow?' investigation

    An engineer asks for the 10 slowest requests from the last hour to find a bottleneck in a specific provider using query_latency.

  3. Real-world use case 03

    The 'Why did it hallucinate?' debug

    A dev uses the agent to trace a specific user session to see the exact prompts that led to a bad output using query_sessions.

Complete set · 10capabilities

The complete Helicone (LLM Observability) capability set.

These are the exact actions your AI can choose when you ask it to work with Helicone (LLM Observability).

Capability set01 / 03

01—04

4 capabilities in this set.

Part of 10 available through Helicone (LLM Observability).

  1. 01 Capability

    Query costs

    See a breakdown of the properties driving your account spending. This helps you identify exactly which features or models are consuming your budget.

  2. 02 Capability

    Query sessions

    Enumerate the rules exporting active billing. Use this to trace multi-turn sessions and see the billing rules for those calls.

  3. 03 Capability

    Query users

    Run a validation check to route gateway history. It helps you identify your most active human clients and their interaction history.

  4. 04 Capability

    Get prompt versions

    Extract flags from prompt validations. Use this to see every version of a specific prompt and its history.

Capability set02 / 03

05—07

3 capabilities in this set.

Part of 10 available through Helicone (LLM Observability).

  1. 05 Capability

    Query feedback

    Inspect the internal data used to mitigate plan math issues. It lets you see the underlying logic for user feedback and critiques.

  2. 06 Capability

    Query latency

    Get a JSON payload for customer bindings on latency. It helps you pinpoint which providers are causing high Time To First Token delays.

  3. 07 Capability

    Log feedback

    Identify the active arrays for native hold parsing. Use this to pull the specific feedback logs you need to improve model grounding.

Capability set03 / 03

08—10

3 capabilities in this set.

Part of 10 available through Helicone (LLM Observability).

  1. 08 Capability

    Query prompts

    Retrieve cloud logging traces for vault limits. It helps you see the history and specific limits of your prompts.

  2. 09 Capability

    List properties

    Identify the active arrays spanning gateway authentication. Use this to see the properties and settings of your gateway.

  3. 10 Capability

    Query requests

    Find bounded records inside the Helicone platform. It lets you see specific request logs to debug issues.

Set up in minutes

One URL. Then ask Helicone (LLM Observability) to work.

Claude and ChatGPT only need the Connector URL. Copy it once, add it in settings, and use Helicone (LLM Observability) from the conversation.

Choose your client

Live preview
Advanced clients IDE · CLI

Claude · Web + desktop

Official guide ↗

Connector URL · ready to paste

Streamable HTTP
https://edge.vinkius.com/vk_preview_BY11i3Opi4Qq9KtIJUGQpInAaXPfnwFfbeNDvYuq/mcp
  1. Step 01

    Open Connectors

    In Claude Web or Claude Desktop, open Settings and choose Connectors.

  2. Step 02

    Add the URL

    Choose Add custom connector, name it Helicone (LLM Observability), and paste the URL above.

  3. Step 03

    Turn it on in chat

    Select +, open Connectors, and enable Helicone (LLM Observability) for the conversation.

Where the request belongs

Work Helicone can move forward.

Built around the request

This is for the engineers and product leads who are tired of hunting through logs to figure out why an AI feature is slow or expensive.

01

LLM Engineer

Debugs prompt performance and checks TTFT latency across different providers on a Tuesday afternoon.

02

Product Owner

Monitors the AI burn rate and calculates costs per feature to report to stakeholders.

03

Data Scientist

Analyzes user feedback and thumbs down critiques to improve model grounding and accuracy.

04

DevOps Engineer

Ensures the AI gateway stays up and checks the reliability of the proxy layers.

Bring your own AI

Change the model, client or framework. Keep Helicone connected.

  • Claude
  • ChatGPT
  • Gemini
  • Cursor
  • VS Code
  • Windsurf
  • ZCode
  • Cline
  • Zed
  • Continue
  • Kiro
  • Roo Code
  • Zencoder
  • Goose
  • Void
  • Augment Code
  • Amp
  • Qodo
  • Tabnine
  • Pieces
  • Sourcegraph Cody
  • JetBrains
  • Warp
  • Amazon Q
  • Antigravity
  • BoltAI
  • Raycast
  • Jan
  • LM Studio
  • AnythingLLM
  • Open WebUI
  • Msty
  • Cherry Studio
  • LibreChat
  • TypingMind
  • Chorus
  • 5ire
  • n8n
  • LangChain
  • LlamaIndex
  • CrewAI
  • Vercel AI SDK

Before you connect

Questions about Helicone.

The practical details behind the request, access and result.

How does the Helicone MCP help with my AI budget?

It gives you a direct way to query your spending. You can ask your agent for cost breakdowns by model, user, or custom tags to see exactly where your money is going.

Can I use the Helicone MCP to find slow prompts?

Yes. You can ask your agent to find the slowest requests from a specific timeframe. It will pull the latency data and can even show you the prompt that caused the delay.

Does the Helicone MCP support multi-turn conversation tracing?

It does. You can ask your agent to pull specific sessions to see the full history of a conversation. This is great for debugging complex agentic workflows.

How do I see user feedback through the Helicone MCP?

You can ask your agent to pull thumbs up or down critiques. It retrieves the logged feedback so you can see what users liked or disliked about your AI's responses.

Can the Helicone MCP show me different prompt versions?

Yes. You can ask your agent to list all versions of a specific prompt. It will show you the history of your instructions and when each version was deployed.

Is the Helicone MCP good for monitoring my AI gateway?

It's designed for that. It connects to your Helicone account to provide a conversational way to check request logs, latency, and gateway performance.

Can I see the exact prompt that caused a specific error?

Yes. Use the query_requests capability to fetch direct prompts and outputs from the proxy logs. You can filter by status or custom tags to find the exact interaction that needs debugging.

How do I track costs for a specific customer ID?

Ask your agent to query_costs and include your customer identity in the filter. Helicone maps costs per model and user, allowing you to see exactly how much each client is burning in LLM tokens.

Can my agent log human feedback into Helicone?

Absolutely. Use the log_feedback capability to inject offline Human-in-the-Loop verdicts or text critiques directly into Helicone's database, helping you refine your model's grounding over time.

One connection away

Give your agent a direct line to Helicone.

Connect Helicone once. Keep it beside 5,900+ managed Connectors when the next task needs more.

Explore every Connector No credit card required · Free tier available