# Datadog AI (LLM Observability) MCP for AI Agents AI Agent Connect

> Datadog AI (LLM Observability) lets you monitor your AI models directly from your favorite agent. You can track token costs, find specific prompt logs, and set up alerts for when your LLM starts acting up. It brings your Datadog telemetry into a conversational interface so you don't have to keep switching tabs to check on your production models.

## Overview
- **Category:** ai-frontier
- **Price:** Free
- **Endpoint:** https://edge.vinkius.com/vk_preview_bWdhwHEHDioKsLvhV1Lvv7ec4t2gm8k5cIgCZeWx/ai-agent-connect
- **Tags:** llm-observability, token-usage, prompt-monitoring, ai-performance, telemetry, model-auditing

## Description

Imagine you're in the middle of a deployment and your agent starts giving weird responses. Usually, you'd have to jump into the Datadog dashboard, filter through mountains of logs, and try to find the exact moment things went south. This Connector changes that by bringing those metrics straight into your chat window. You can ask your agent to find specific prompt traces or check if your token spending is spiking without ever leaving your workspace. It bridges the gap between high-level monitoring and the actual work of building AI features. Since Vinkius makes it so easy to plug these connections into your existing setup, you can get these insights running in minutes. Instead of manually hunting for errors, you're just asking questions and getting answers on latency, costs, and model health.

## Tools

### list_dashboards
See all attached rules and active billing widgets. This gives you a clear view of your spending and rules.

### list_events
Find active arrays related to native Gateway authentication. It helps you track specific deployment events.

### list_incidents
Run automated checks to route Gateway history. This helps you see active outages and disruptions.

### search_llm_spans
Generate JSON payloads for customer bindings. Use this to pull specific prompt logs and traces.

### list_ai_monitors
View cloud logging for Vault limits. This lets you see your active AI monitoring status.

### query_metrics
Find CRM records inside the Datadog platform. Use this to pull token counts and latency numbers.

### submit_series
Extract properties that drive account logic. This helps you track specific property-driven logic.

### list_service_accounts
Identify active arrays for hold parsing. Use this to manage your service accounts and permissions.

### create_event
Inspect internal arrays to handle specific Plan Math scenarios. This helps you manage complex logic events.

### create_monitor
Set up validations to catch and flag high churn signals. Use this to create automated alerts for your AI.

## Prompt Examples

**Prompt:** 
```
What was our average token spend for the last 24 hours?
```

**Response:** 
```
Querying Datadog metrics for LLM observability...

**Token Usage Summary (Last 24h):**
- **Average Tokens/Request:** 450
- **Peak Usage:** 1,200 (at 03:15 AM)
- **Total Estimated Spend:** $12.40

*Note: Usage peaked during the scheduled batch processing job.*
```

**Prompt:** 
```
Find any LLM spans that mention 'permission denied' in the last hour.
```

**Response:** 
```
I searched your LLM spans for 'permission denied'. I found 3 instances:

| Time | Agent | Prompt Snippet | Status |
| :--- | :--- | :--- | :--- |
| 10:04 AM | Data-Writer | "Update user permissions..." | Failed |
| 10:12 AM | Data-Writer | "Update user permissions..." | Failed |
| 10:45 AM | Admin-Bot | "Grant access to root..." | Failed |

Would you like to see the full stack trace for any of these?
```

**Prompt:** 
```
List all my active AI monitors and tell me if any are alerting.
```

**Response:** 
```
You have 5 active AI monitors currently running:

- ✅ **LLM-Latency-High**: OK
- ⚠️ **Token-Quota-Reached**: **ALERTING** (Threshold exceeded at 02:00 PM)
- ✅ **Model-Drift-Check**: OK
- ✅ **GPU-Utilization**: OK
- ✅ **Gateway-Auth-Fail**: OK

I can help you investigate the Token Quota alert if you'd like.
```

## Capabilities

### Query real-time token counts
Get instant data on how many tokens your models are consuming.

### Search specific prompt logs
Find exact prompt logic and response traces within your telemetry.

### Create performance alerts
Set up monitors that trigger when your AI response quality drops.

### View global AI spending
See a bird's eye view of costs across different model providers.

### Track active outages
Monitor service disruptions that block your multi-agent workflows.

### View deployment marks
Identify exactly when you switched to new dynamic LLM models.

## Use Cases

### Debugging a slow prompt
An AI Engineer notices a specific request is lagging. They ask their agent to query the latency metrics for that specific model to see if it's a provider issue.

### Monthly cost auditing
A FinOps analyst wants to know the monthly spend. They ask their agent to list the dashboards for OpenAI costs to get a quick summary.

### Identifying production errors
An MLOps person sees a spike in errors. They ask the agent to search the LLM spans for 'out of bounds' to find the exact prompt that triggered it.

### Monitoring system health
An SRE wants to know if the gateway is up. They ask the agent to list active incidents to see if there are any service disruptions.

## Benefits

- Stop jumping between tabs by using query_metrics to get token usage data instantly in your chat window.
- Find the root cause of errors faster by using search_llm_spans to pull specific prompt logs and traces.
- Prevent production downtime by using create_monitor to set alerts on SLI thresholds and performance drops.
- Manage your cloud costs better by using list_dashboards to see spending across different model providers.
- Keep your team informed about outages by using list_incidents to see active disruptions in your workflow.
- Track model changes accurately by using list_events to see exactly when deployments happened in your pipeline.

## How It Works

The bottom line is you get a conversational window into your Datadog LLM telemetry without leaving your primary workspace.

1. Subscribe to this Connector on Vinkius.
2. Enter your Datadog API Key, APP Key, and Site.
3. Ask your agent to pull metrics or find logs in your workspace.

## Frequently Asked Questions

**Can the Datadog AI (LLM Observability) MCP show me how much I'm spending on OpenAI?**
Yes, it can pull your spending metrics directly. You can ask your agent to show you costs across different providers like OpenAI or Anthropic to see where your budget is going.

**How do I use Datadog AI (LLM Observability) to find specific prompt logs?**
You can simply ask your agent to search for specific keywords or errors within your prompt logs. It will pull the relevant spans and show you the exact logic and responses.

**Can I set up alerts for my LLM using Datadog AI (LLM Observability)?**
Absolutely. You can ask your agent to create new monitors that trigger alerts when your LLM latency spikes or when your token usage hits a certain threshold.

**Does the Datadog AI (LLM Observability) MCP work with Cursor?**
Yes, it works with any MCP-compatible client, including Cursor, Claude, and Windsurf. Once connected, your agent can access all your Datadog LLM telemetry.

**How does Datadog AI (LLM Observability) help with model latency?**
It allows you to query real-time latency metrics instantly. You can quickly identify which models are underperforming without having to manually filter through complex dashboards.

**Can I see my deployment history with Datadog AI (LLM Observability)?**
Yes, you can pull textual deployment marks. This helps you see exactly when you switched models or pushed new updates to your AI infrastructure.

**Can my agent check token usage for a specific LLM model?**
Yes. Use the 'query_metrics' tool with a query like 'avg:datadog.llm_observability.tokens{model:gpt-4}'. The agent will retrieve the numeric timeseries data directly from Datadog's metrics engine.

**How do I search for specific prompt text in my logs?**
Use the 'search_llm_spans' tool. Provide a search query matching your prompt identifiers. The agent will pull the explicit REST maps capturing the literal prompt logic text from your Datadog logs.

**Can I see if there are any active incidents affecting my AI services?**
Absolutely. The 'list_incidents' tool tracks outages and service disruptions in real-time. This allows your agent to identify exactly which external factors might be blocking your multi-agent orchestration pipelines.