ClaudeChatGPTPerplexityGeminiMicrosoft CopilotRaycastMeta AIGrokZ.aiQwenKimi
DeepSeekMistralCursorVS CodeWindsurfJetBrainsClineLovableVercel AI SDKLangChain

Use Agent Benchmark Comparison Engine with your AI.

Connect your account once and let the AI you already use work with it, without building another integration. Get a precise, weighted score for every agent.

Included with plan

Ask AI about this Connector

Developed, maintained, and hosted by Vinkius.

MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED

Waiting for input…

Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.

ChatGPTClaudeCursorPerplexityGeminiMicrosoft CopilotRaycastMeta AI

Complete set · 3 capabilities

The complete Agent Benchmark Comparison Engine capability set.

These are the exact actions your AI can choose when you ask it to work with Agent Benchmark Comparison Engine.

Capability set01 / 01

01-03

3 capabilities in this set.

Part of 3 available through Agent Benchmark Comparison Engine.

  1. 01

    Calculate agent rankings

    Performs the complete mathematical comparison and ranking of a set of agents based on provided weights

  2. 02

    Get agent performance summary

    Retrieves a high-level overview of the best-performing agents for specific use cases

  3. 03

    Validate benchmark config

    0 and metrics are within logical bounds. Ensures that a proposed set of weights and agent metrics are mathematically valid before running heavy calculations

Observed, not estimated

835ms average. Fast in production.

Agent Benchmark Comparison Engine is checked daily against the live service.

Daily averagePeak 1206ms
Aug 22Today
Fastest day
649ms
Slowest day
1206ms
14-day trend
Improving-21%

Connect your client

One URL. Every client.

Activate the Connector, copy your link, and paste it into the client you already use. 3 capabilities arrive ready to run.

Preview access · not provider authentication

The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Agent Benchmark Comparison Engine, so you can see the experience inside your AI.

It does not authenticate your account with Agent Benchmark Comparison Engine. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.

Agent Benchmark Comparison Engine Connector

You're all set. Choose your MCP client and follow the setup instructions.

Connector linkhttps://edge.vinkius.com/vk_preview_aD2zjuWHaIJRj1pXV5vT436Hv4TzDHO7e7iXDcpX/mcp

Claude Desktop

Follow the steps below to connect in seconds.

  1. 1In Claude Desktop, open Settings → Connectors.
  2. 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
  3. 3Click Add and start a new chat — Agent Benchmark Comparison Engine capabilities are ready to use.
Configuration · claude_desktop_config.jsonCopy
{
  "mcpServers": {
    "agent-benchmark-comparison-engine-mcp": {
      "url": "https://edge.vinkius.com/vk_preview_aD2zjuWHaIJRj1pXV5vT436Hv4TzDHO7e7iXDcpX/mcp"
    }
  }
}
  • Claude
  • ChatGPT
  • Cursor
  • VS Code
  • Windsurf
  • Claude Code
  • JetBrains
  • Cline

Step-by-step instructions for each client are in the guide. How to connect

Who it's for

Built for the work Agent Benchmark Comparison Engine owners hand off.

This MCP is built for technical teams that need to rigorously compare AI models. If you're constantly choosing between different LLM providers or internal agents, this capability gives you the objective data you need to make a decision.

  • 01

    ML Engineer

    Use this MCP to build automated comparison pipelines and validate benchmark configurations.

  • 02

    Data Scientist

    Run comprehensive analyses to determine which agent performs best across multiple, weighted metrics.

  • 03

    AI Product Manager

    Determine the optimal model choice for a new product feature by comparing cost, latency, and performance.

FAQ

Questions Agent Benchmark Comparison Engine owners ask.

  • 01

    Does this MCP compare models based on real-time usage?

    No. This MCP uses a deterministic mathematical framework. You must provide the performance metrics—like latency and cost—as inputs; it does not run live tests against external APIs.

  • 02

    What metrics can I use for comparison?

    The MCP is designed to handle common metrics including accuracy, latency, cost, and hallucination rates. You can assign weights to each of these factors.

  • 03

    Is the ranking customizable?

    Yes. You control the ranking by providing custom weights. You decide if latency is twice as important as cost, for example, and the MCP calculates the score accordingly.

  • 04

    What if my weights don't add up to 1.0?

    You should run the validate_benchmark_config capability first. This ensures that your proposed set of weights and metrics are mathematically sound before you run the main calculations.

  • 05

    Can I use this MCP with my existing data?

    Yes. You feed the MCP the data you've already collected. The MCP's job is to take that raw data and apply the weighted scoring formula.