# New Relic AI MCP for AI Agents AI Agent Connect

> New Relic AI (LLM Observability) lets you monitor and audit LLM telemetry directly from your AI agent. Track token costs, p95 latency, and user feedback in real-time. It connects your New Relic account to provide a clear view of how your AI infrastructure is actually performing without switching tabs. Get granular insights into model behavior and spending limits instantly.

## Overview
- **Category:** loved-by-devs
- **Price:** Free
- **Endpoint:** https://edge.vinkius.com/vk_preview_oqdaAroeFoXBv9yPI4WsHZZZZZuzqhwVoSn1YCyq/ai-agent-connect
- **Tags:** llm-monitoring, token-cost-tracking, performance-analytics, ai-observability, latency-tracking

## Description

New Relic AI connects your account to your agent so you can manage LLM observability and performance analytics through a simple chat interface. Instead of jumping between different dashboards to see how your models are behaving, you can just ask your agent for the data you need. You can pull out detailed chat completion messages to see exactly what's happening with prompts, or check your token consumption to see where your budget is actually going. It handles the heavy lifting of extracting p95 latency matrices and average response times, making sure your text generation stays fast and reliable. If you need to see how actual humans are rating your outputs, you can grab those feedback scores and chronological messages instantly to spot quality drops. You can even run custom NRQL queries to dig into multi-tenant datasets or post custom telemetry rows to keep track of internal agent states across your entire pipeline. Having this available in the Vinkius catalog means you get these enterprise-grade monitoring capabilities integrated into the tools you already use every day. It turns complex telemetry into actionable info you can grab on demand, giving you a clear picture of your AI infrastructure's health without the usual manual overhead of navigating complex monitoring suites.

## Tools

### query_llm_feedback
Retrieve Cloud logging tracing for explicit Vault limits and human ratings. This pulls in your supervisor scores.

### list_alert_policies
Inspect internal arrays that handle specific Plan Math. This helps you see what's mitigating specific logic issues.

### list_apm_apps
Run automated validation checks for explicit Gateway history. It helps you verify the routing of your history.

### custom_nrql
Execute read-only queries to extract rich Churn flags from your data. Use this for deep, custom data analysis.

### list_dashboards
Identify active arrays spanning native Gateway auth. It helps you see which dashboards are currently active.

### query_llm_errors
Identify active arrays spanning native Hold parsing for error tracking. Use this to find specific failures.

### query_llm_costs
Extract properties that drive active Account logic for cost tracking. This shows you exactly where the money goes.

### query_llm_events
Identify bounded CRM records inside the Headless New Relic Platform. This helps you see specific event records.

### query_llm_latency
Provision a JSON Payload for hard Customer bindings and latency data. Use this to see p95 response times.

### post_custom_event
Insert CustomAITelemetry rows to track internal agent states and billing rules. This helps you log custom markers.

## Prompt Examples

**Prompt:** 
```
How much did we spend on the Claude model yesterday?
```

**Response:** 
```
Checking your costs... You spent $42.50 on the Claude model yesterday. This was spread across 150,000 tokens. Would you like to see the breakdown by specific user groups?
```

**Prompt:** 
```
Are our prompts responding fast enough?
```

**Response:** 
```
I've pulled the latest latency metrics for you:

| Model | p95 Latency | Avg Response |
| :--- | :--- | :--- |
| GPT-4 | 2.4s | 1.8s |
| Claude 3 | 1.9s | 1.4s |

Most responses are under 2 seconds, but there was a spike in the 12:00 PM window.
```

**Prompt:** 
```
What is the latest human feedback for the new summary tool?
```

**Response:** 
```
Here are the last 5 feedback entries for the summary tool:

- **Rating: 5/5** - 'Great summary of the meeting.'
- **Rating: 4/5** - 'A bit too long, but accurate.'
- **Rating: 3/5** - 'Missed the last point.'
- **Rating: 5/5** - 'Perfect.'
- **Rating: 2/5** - 'Too many bullet points.'
```

## Capabilities

### Pull chat completion messages
See the literal model behavior by retrieving prompt inputs and responses.

### Track token spending
Calculate exact USD costs across your AI infrastructure by extracting cost properties.

### Check latency metrics
Get p95 latency and average response times to ensure your generation is fast.

### Grab user feedback
Retrieve human supervisor ratings and feedback messages to spot quality drops.

### Run custom NRQL queries
Execute read-only queries to pull specific insights from large AI datasets.

### Post custom telemetry
Insert generic rows to track internal agent states and behavioral markers.

### Audit app health
Enumerate active APM apps, dashboards, and alert policies to check your environment.

## Use Cases

### Checking a sudden spike in costs
An AI engineer sees a spike in costs and asks the agent to run query_llm_costs to see which model is burning the budget.

### Debugging slow response times
Users complain about slow responses, so the engineer uses query_llm_latency to find the p95 spikes in specific regions.

### Verifying human satisfaction
A team wants to see if a new prompt is working. They use query_llm_feedback to see if human supervisors are giving it high ratings.

### Identifying production errors
A production error occurs, and the DevOps lead uses query_llm_errors to identify which specific hold parsing failed.

## Benefits

- Stop manual dashboard hopping by pulling LLM telemetry directly into your chat window with query_llm_costs.
- Control your budget by seeing exact USD token consumption across your whole infrastructure using query_llm_costs.
- Fix performance bottlenecks faster by grabbing p95 latency and average response times with query_llm_latency.
- Spot quality regressions early by pulling human feedback scores and ratings using query_llm_feedback.
- Run complex data extractions without learning new UI paths by executing read-only queries with custom_nrql.
- Keep your AI environment organized by auditing APM apps and alert policies using list_apm_apps.

## How It Works

The bottom line is you get a direct line to your New Relic telemetry without leaving your chat window.

1. Subscribe to this Connector and enter your New Relic API Key and Account ID.
2. Connect your preferred AI client like Claude or Cursor.
3. Ask your agent to pull costs, check latencies, or run NRQL queries.

## Frequently Asked Questions

**Can the New Relic AI MCP help me see how much my AI agents are costing me?**
Yes, it pulls your token consumption data directly so you can see the exact USD cost across your entire infrastructure.

**How do I check if my LLM responses are getting faster with the New Relic AI MCP?**
You can ask your agent to pull p95 latency and average response times to see real-time performance trends.

**Can I use the New Relic AI MCP to see what my human supervisors think of the AI?**
It retrieves chronological feedback messages and 1-5 rating scores dumped by your human supervisors.

**Does the New Relic AI MCP let me run complex queries on my data?**
Yes, it allows you to run custom NRQL queries to extract specific insights from your multi-tenant AI datasets.

**Can I use the New Relic AI MCP to track custom internal states?**
You can post custom telemetry rows to track internal agent states and behavioral markers across your pipeline.

**Is the New Relic AI MCP safe for my data?**
It connects to your existing New Relic account and follows your established security and permissions.

**Can I check my total AI token costs through my agent?**
Yes. Use the `query_llm_costs` tool. Your agent will execute a NRQL aggregation summing the `tokenSpanCost` property from your LLM events over the last 24 hours, faceted by model, to provide a clear financial breakdown.

**How do I monitor the p95 latency of my LLM generations?**
The `query_llm_latency` tool retrieves the average duration and latency matrices for your AI providers. Your agent will report the results as a timesheet or summary, helping you identify performance bottlenecks instantly.

**Can my agent run custom NRQL queries against my telemetry data?**
Absolutely. Use the `custom_nrql` tool to provide any valid read-only NRQL string. Your agent will query New Relic's NerdGraph API and return the resulting dataset, allowing for complete flexibility in how you analyze your AI operations.