Use Braintrust with your AI.
Connect your account once and let the AI you already use work with it, without building another integration. Automate AI evaluations with Braintrust. organize projects, test model datasets, run benchmarks, and manage prompts via any AI agent.
Developed, maintained, and hosted by Vinkius.
MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED
Waiting for input…
Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.
Complete set · 10 capabilities
The complete Braintrust capability set.
These are the exact actions your AI can choose when you ask it to work with Braintrust.
01-04
4 capabilities in this set.
Part of 10 available through Braintrust.
- 01
List experiments
Retrieve all evaluation experiments mapping model test scores and metrics
- 02
Insert dataset row
Append new test cases into a dataset matrix targeting specific evaluations
- 03
List projects
Retrieve the list of all AI evaluation projects in Braintrust
- 04
List prompts
Retrieve explicitly version-controlled system prompts isolated in Braintrust
05-07
3 capabilities in this set.
Part of 10 available through Braintrust.
- 05
Get dataset
Retrieve a specific dataset containing exact schemas bounding LLM outputs
- 06
Get prompt
Retrieve exact variable contexts and literal text templates for a prompt
- 07
List env vars
Probe the Braintrust AI Gateway configurations managing model API keys securely
08-10
3 capabilities in this set.
Part of 10 available through Braintrust.
- 08
Create experiment
Establish a new historical experiment trace to record LLM pipeline tests
- 09
Create project
Create a new project environment for tracking AI evaluations and datasets
- 10
List datasets
List isolated Ground Truth text banks used for automated evaluation scoring
Observed, not estimated
874ms average. Fast in production.
Braintrust is checked daily against the live service.
- Fastest day
- 795ms
- Slowest day
- 1072ms
- 14-day trend
- Improving-10%
Connect your client
One URL. Every client.
Activate the Connector, copy your link, and paste it into the client you already use. 10 capabilities arrive ready to run.
Preview access · not provider authentication
The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Braintrust, so you can see the experience inside your AI.
It does not authenticate your account with Braintrust. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.
Braintrust Connector
You're all set. Choose your MCP client and follow the setup instructions.
https://edge.vinkius.com/vk_preview_cr8T7CpPGy155EtBgklGtlDgFsTWLRLYfzKoHIOF/mcpClaude Desktop
Follow the steps below to connect in seconds.
- 1In Claude Desktop, open Settings → Connectors.
- 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
- 3Click Add and start a new chat — Braintrust capabilities are ready to use.
{
"mcpServers": {
"braintrust-mcp": {
"url": "https://edge.vinkius.com/vk_preview_cr8T7CpPGy155EtBgklGtlDgFsTWLRLYfzKoHIOF/mcp"
}
}
}
Claude
ChatGPT
Cursor
VS Code
Windsurf
Claude Code
JetBrains
Cline
Step-by-step instructions for each client are in the guide. How to connect
FAQ
Questions Braintrust owners ask.
- 01
Can I insert new test data dynamically tracking specific limits?
Yes. Utilizing the insert_dataset_row method, you can effortlessly inject exact JSON tracking payload mapping strings directly inside the text corpus evaluating the final results.
- 02
Does it pull out original Prompt definitions stored securely?
Certainly. The get_prompt command isolates and returns perfectly version-controlled bounding parameters slicing literal templates natively hosted under the Braintrust database.
- 03
How deeply can it inspect test regressions or scoring limits?
Using the robust list_experiments call, you can branch full arrays separating LLM version behaviors over massive iterations tracking the performance anomalies accurately.
Explore
More in Brain Trust
LangSmith (LLM Observability & Hub) AI Connector
Monitor LLM apps via LangSmith — track traces, audit prompt templates, and manage evaluation datasets.
ViewReplicate AI Connector
Equip your AI to dynamically search, run, and monitor thousands of open-source machine learning models hosted
ViewRagas AI Connector
Equip your AI with Ragas to create datasets, run RAG evaluations, and track experiment metrics directly from y
ViewDatadog AI (LLM Observability) AI Connector
Monitor LLM performance via Datadog — track token usage, audit prompts, and monitor AI model metrics directly
View
Suggestions
Humanloop (LLM Prompt Management API) AI Connector
Manage, version, and deploy LLM prompts directly from your AI agent using the Humanloop API.
ViewVectorShift (AI Workflow & RAG Automation) AI Connector
Automate AI workflows and RAG via VectorShift — manage pipelines, query knowledge bases, and deploy chatbots d
ViewRelevance AI AI Connector
Automate autonomous AI agents via Relevance AI — manage tools, trigger tasks, and monitor results directly.
ViewMetatext AI Connector
No-code NLP and AI model management via Metatext — run inference and manage datasets.
View
