• PASS
  • FAIL
  • REVIEW
  • HUMAN

How to set evaluation benchmarks with your AI

Connect this MCP to your AI client to pull the exact scoring thresholds needed for your quality checks.

Ask AI about this page

Short answer

How do I set evaluation benchmarks for my AI?

Your AI uses the get_dimension_benchmarks tool to pull specific scoring thresholds for different quality dimensions. This gives your agent the exact numbers it needs to decide if a score is high or low. You'll see how these thresholds map to specific quality outcomes below.

Benchmark outcomes

Where scores land.

The AI uses these thresholds to categorize your application scores.

PASS

High quality scores

The AI identifies scores that meet or exceed the required threshold. These results are flagged as strong performers.

FAIL

Low quality scores

The AI catches scores that fall below the defined minimum. These are marked as weak or insufficient.

REVIEW

Borderline scores

Scores sitting right on the edge of a threshold are flagged for your attention. This is where you decide if the AI's judgment is correct.

HUMAN

Manual overrides

When the AI flags a threshold conflict, you step in to adjust the benchmark or the final decision.

The workflow

What your AI does when the request arrives.

The AI handles the data retrieval so you can focus on the analysis.

  1. Fetch thresholds

    Your agent requests the specific scoring limits for each quality dimension.

    get_dimension_benchmarks
  2. Compare scores

    The AI compares your current application scores against the retrieved benchmarks.

    get_dimension_benchmarks
  3. Categorize results

    The agent sorts every score into pass, fail, or review buckets based on the data.

    get_dimension_benchmarks
  4. Flag outliers

    The AI highlights any scores that are statistically unusual compared to the benchmarks.

    get_dimension_benchmarks

Try it

Copy these to start.

Use these prompts to kick off the benchmarking process.

Starting points

These are opening lines. Swap out the specific dimension names to match your actual quality metrics.

Accelerator Application Quality Score Connector

You're all set. Choose your MCP client and follow the setup instructions.

Connector linkhttps://edge.vinkius.com/vk_preview_jzP4NK1jm33dd0NP6aFfcuMCufj9juqr2jZ0gJ25/mcp

Claude Desktop

Follow the steps below to connect in seconds.

  1. 1In Claude Desktop, open Settings → Connectors.
  2. 2Click “Add custom connector” and paste the connector link above as the remote MCP server URL.
  3. 3Click Add and start a new chat — Accelerator Application Quality Score capabilities are ready to use.
Configuration · claude_desktop_config.jsonCopy
{
  "mcpServers": {
    "accelerator-application-quality-score-mcp": {
      "url": "https://edge.vinkius.com/vk_preview_jzP4NK1jm33dd0NP6aFfcuMCufj9juqr2jZ0gJ25/mcp"
    }
  }
}
Copy into chat04
  • Get the benchmarks for the 'Accuracy' dimension so I can evaluate this batch.

  • Compare the scores in my 'User Experience' spreadsheet against the current thresholds.

  • What are the pass/fail limits for the 'Security' dimension right now?

  • Check the latest application scores in my GitHub repo against the quality benchmarks.

  • Claude
  • ChatGPT
  • Cursor
  • VS Code
  • Windsurf
  • Claude Code
  • JetBrains
  • Cline

Start here

Connect Accelerator Application Quality Score once, then ask.

Just click the link to connect the MCP to your client. Your login stays encrypted on our side, and you can find the full walkthrough details on the Connector page.

Connect Accelerator Application Quality Score to your AI

FAQ

How this task behaves

  • 01

    Can the AI change the benchmark thresholds?

    No. The AI can only retrieve the existing thresholds using the tool. You must update thresholds in your source system.

  • 02

    Does the AI delete any of my scoring data?

    No. The AI only reads the benchmark data to perform comparisons. It has no capability to delete or modify your records.

  • 03

    Can I use this to set new rules?

    No. The AI uses the rules that already exist. It cannot create new scoring logic or write new benchmark values.

  • 04

    What happens if a dimension doesn't have a benchmark?

    The AI will report that no threshold is defined for that specific dimension.

  • 05

    How does the AI know if a score is 'strong'?

    It compares the numerical score against the specific value returned by the get_dimension_benchmarks tool.

  • More questions about Accelerator Application Quality Score? The Connector page answers them. See everything the Accelerator Application Quality Score Connector can do

Connect Accelerator Application Quality Score to Claude, Cursor, ChatGPT & more