# Crawlbase MCP for AI Agents AI Agent Connect

> Crawlbase lets you run high-scale web scraping and headless crawling directly through your AI agent. It handles the hard stuff like JS rendering, CAPTCHA bypassing, and proxy management so you can get clean data from sites like Amazon, LinkedIn, and Google SERPs without writing custom scripts.

## Overview
- **Category:** friends-mcp
- **Price:** Free
- **Endpoint:** https://edge.vinkius.com/vk_preview_AcN0JE7lQv9RjbUnrxQnYlI9bigE4gxGmDTPSZBI/ai-agent-connect
- **Tags:** proxy, captcha-solving, html-extraction, headless-browser, data-collection, web-crawling

## Description

Instead of wrestling with headless browsers, proxy rotations, and complex selector logic, you can just tell your agent what data you need. This Connector connects your Crawlbase account to your AI client, turning web scraping into a conversation. You don't have to worry about getting blocked or dealing with messy HTML; the agent handles the heavy lifting of navigating pages and extracting structured results. It's a massive shift from manual scraping to a system where you just describe the target and the output format. When you're looking for a way to manage these complex workflows, Vinkius makes it easy to plug this into your existing setup. You'll get clean JSON, screenshots, or specific product data from nearly any corner of the web without the usual technical headaches.

## Tools

### scrape_html
Pulls HTML content from specific areas of a site using datacenter proxies. It identifies bounded routing spaces for clean extraction.

### get_screenshot_link
Creates a URL for a snapshot of a web page. Use it to get visual validation of a rendered site.

### scrape_amazon
Navigates Amazon's internal arrays to get product details. It mitigates common e-commerce constraints.

### scrape_linkedin
Matches specific structural blueprints to extract LinkedIn profiles. It verifies constraints for accurate data.

### scrape_facebook
Exports active social pages based on your defined rules. It enumerates attached rules for easy export.

### scrape_google_serp
Parses Google search results while bypassing CAPTCHAs. It identifies active arrays for search data.

### scrape_twitter
Fetches mapped limits from Twitter (X) profiles. It uses Crawlbase X extraction for social data.

### custom_scrape
Sets up custom proxies with unique headers and crawling logic. It provisions highly-available request payloads.

### scrape_js_rendered
Retrieves data from pages that require JavaScript execution. It tracks payload IDs to get dynamic content.

### scrape_json_format
Forces raw web outputs into a strictly structured JSON format. It handles the structural extraction for you.

## Prompt Examples

**Prompt:** 
```
Check the price of the Sony WH-1000XM5 on Amazon and tell me if it's under $350.
```

**Response:** 
```
I've checked the current listing for the Sony WH-1000XM5 on Amazon.

**Product Details:**
- **Title:** Sony WH-1000XM5 Wireless Noise Canceling Headphones
- **Current Price:** $348.00
- **Rating:** 4.8/5 stars

**Status:** The price is currently under your $350 limit.
```

**Prompt:** 
```
Get the top 5 organic results for 'best CRM for small business' from Google.
```

**Response:** 
```
Here are the top 5 organic results for your search:

| Rank | Title | URL |
|---|---|---|
| 1 | HubSpot CRM | hubspot.com |
| 2 | Salesforce for Small Business | salesforce.com |
| 3 | Zoho CRM | zoho.com |
| 4 | Pipedrive | pipedrive.com |
| 5 | Freshsales | freshworks.com |
```

**Prompt:** 
```
Take a screenshot of the homepage for example.com so I can see the layout.
```

**Response:** 
```
I've requested a snapshot of the example.com homepage. 

You can view the rendered image here: [Crawlbase Screenshot URL]

Let me know if you need me to check any specific elements on that page.
```

## Capabilities

### Extract HTML content
Pull raw data from specific parts of a webpage using datacenter proxies.

### Render JavaScript
Capture content from dynamic pages that require a headless browser to load.

### Convert to JSON
Turn messy web outputs into strictly structured JSON data automatically.

### Capture screenshots
Generate URLs for visual snapshots of any webpage for quick validation.

### Scrape e-commerce sites
Fetch specific product details from Amazon while avoiding common blocks.

### Parse search results
Gather Google SERP data while bypassing CAPTCHAs and limits.

### Manage custom proxies
Set up specific headers and crawling logic for high-availability requests.

## Use Cases

### Tracking competitor pricing
A growth lead wants to track Amazon prices. They ask the agent to use scrape_amazon to get the latest costs for a list of products.

### Gathering social proof
A researcher needs social proof. They use scrape_twitter to gather recent mentions of a brand and summarize the sentiment.

### Debugging web extraction
A developer is debugging a scraper. They use get_screenshot_link to see if the page is actually loading correctly or if it's stuck.

### Parsing search engine results
A data analyst needs SERP data. They use scrape_google_serp to gather organic links for a list of keywords for a SEO report.

## Benefits

- Skip the proxy headache because custom_scrape handles rotations and headers automatically for every request.
- Get clean data instantly with scrape_json_format instead of manually parsing messy HTML strings.
- Access hard-to-reach social data using scrape_linkedin and scrape_facebook without worrying about getting banned.
- See exactly what your agent sees with get_screenshot_link for visual debugging of any rendered page.
- Bypass blocks on Google results using scrape_google_serp to get clean search data without CAPTCHAs.
- Handle modern web apps easily with scrape_js_rendered to ensure dynamic content actually loads.

## How It Works

The bottom line is you get structured web data through natural language instead of manual coding.

1. Subscribe to the Crawlbase MCP on Vinkius.
2. Add your Normal and JavaScript tokens from your Crawlbase dashboard.
3. Ask your agent to scrape, crawl, or capture screenshots.

## Frequently Asked Questions

**Can Crawlbase MCP scrape LinkedIn profiles?**
Yes, it includes specific tools to retrieve LinkedIn profile data while matching structural blueprints to ensure accuracy.

**Does Crawlbase MCP handle CAPTCHAs?**
Yes, it is designed to bypass CAPTCHAs automatically, especially when parsing search engine results or high-traffic sites.

**Can I use Crawlbase MCP for Amazon products?**
Absolutely. It has a dedicated tool to navigate Amazon's internal arrays and extract product information while avoiding blocks.

**How does Crawlbase MCP handle JavaScript?**
It uses a headless engine to render JavaScript, allowing your agent to see and extract content from dynamic web applications.

**Can Crawlbase MCP export data in JSON?**
Yes, you can force the Connector to return data in a strictly structured JSON format, making it easy to use the data in other workflows.

**Does Crawlbase MCP work with Cursor and Claude?**
Yes, because it's an Connector, it connects directly to any compatible client like Claude, Cursor, Windsurf, and VS Code.

**When should I use the JavaScript (JS) Token versus the Normal Token?**
Use the Normal Token for fast, static HTML extraction. Switch to the JavaScript Token when the target site uses frameworks like React or Angular, where content is rendered dynamically in the browser. The 'scrape_js_rendered' tool requires the JS Token to function.

**Can my agent bypass CAPTCHAs while scraping Google or LinkedIn?**
Yes. Crawlbase is built to handle CAPTCHAs and blocks natively. When you use specialized tools like 'scrape_google_serp' or 'scrape_linkedin', the agent routes your requests through Crawlbase's advanced proxy infrastructure to ensure successful data extraction.

**How do I get a structured JSON response instead of raw HTML?**
Use the 'scrape_json_format' tool or the specialized scraper tools (Amazon, LinkedIn, etc.). These trigger Crawlbase's auto-extraction pipelines, which analyze the page structure and return specific data fields in a clean JSON format.