ClaudeChatGPTPerplexityGeminiMicrosoft CopilotRaycastMeta AIGrokZ.aiQwenKimi
DeepSeekMistralCursorVS CodeWindsurfJetBrainsClineLovableVercel AI SDKLangChain

Use Speculative Decoding Calculator with your AI.

Connect your account once and let the AI you already use work with it, without building another integration. Optimize LLM inference speed and cost using deterministic speculative decoding metrics.

Included with plan

Ask AI about this Connector

Developed, maintained, and hosted by Vinkius.

MCP VERIFIED · PRODUCTION READY · VINKIUS GUARANTEED

Waiting for input…

Works with modern AI clients that support MCP, including ChatGPT, Claude, Cursor, and more.

ChatGPTClaudeCursorPerplexityGeminiMicrosoft CopilotRaycastMeta AI

Complete set · 3 capabilities

The complete Speculative Decoding Calculator capability set.

These are the exact actions your AI can choose when you ask it to work with Speculative Decoding Calculator.

Capability set01 / 01

01-03

3 capabilities in this set.

Part of 3 available through Speculative Decoding Calculator.

  1. 01

    Calculate operational impact

    Estimates memory and cost savings

  2. 02

    Calculate performance metrics

  3. 03

    Optimize speculation parameters

    Determines optimal draft length

Observed, not estimated

816ms average. Fast in production.

Speculative Decoding Calculator is checked daily against the live service.

Daily averagePeak 962ms
Aug 20Today
Fastest day
672ms
Slowest day
962ms
14-day trend
Slowing+17%

Connect your client

One URL. Every client.

Activate the Connector, copy your link, and paste it into the client you already use. 3 capabilities arrive ready to run.

Preview access · not provider authentication

The vk_preview_* token belongs to Vinkius preview infrastructure. It lets Claude discover and display the capabilities of Speculative Decoding Calculator, so you can see the experience inside your AI.

It does not authenticate your account with Speculative Decoding Calculator. Actions requiring credentials or live account data may not run until you activate the Connector and authorize the service.

Speculative Decoding Calculator Connector

You're all set. Choose your MCP client and follow the setup instructions.

Connector linkhttps://edge.vinkius.com/vk_preview_XYWoPCFlNTCi1nCY0hLcYF1OVsVPSrlw38p6boSh/mcp

Claude Desktop

Follow the steps below to connect in seconds.

  1. 1In Claude Desktop, open Settings → Connectors.
  2. 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
  3. 3Click Add and start a new chat — Speculative Decoding Calculator capabilities are ready to use.
Configuration · claude_desktop_config.jsonCopy
{
  "mcpServers": {
    "speculative-decoding-calculator-mcp": {
      "url": "https://edge.vinkius.com/vk_preview_XYWoPCFlNTCi1nCY0hLcYF1OVsVPSrlw38p6boSh/mcp"
    }
  }
}
  • Claude
  • ChatGPT
  • Cursor
  • VS Code
  • Windsurf
  • Claude Code
  • JetBrains
  • Cline

Step-by-step instructions for each client are in the guide. How to connect

FAQ

Questions Speculative Decoding Calculator owners ask.

  • 01

    How do I know if my speculative decoding setup is efficient?

    You can use the calculate_performance_metrics capability. It flags a configuration as inefficient if the speedup ratio is less than 1.5 or if the acceptance rate falls below 0.5.

  • 02

    Can I find the best draft length for my specific model pair?

    Yes, the optimize_speculation_parameters capability iterates through possible draft lengths to find the one that maximizes effective throughput for your given parameters.

  • 03

    How much money can I save by using this optimization?

    By using calculate_operational_impact, you can input the time saved during inference and your hardware cost per second to get an exact estimate of your cost savings.