# Bright Data MCP for AI Agents AI Agent Connect

> Bright Data lets you connect your AI agent to the world's leading web data platform. It handles the heavy lifting of bypassing anti-bot protections, extracting structured search engine data, and managing your proxy infrastructure. You get clean data from sites that usually block scrapers, plus the ability to trigger and monitor large-scale data collection jobs directly through your AI chat interface.

## Overview
- **Category:** developer-tools
- **Price:** Free
- **Endpoint:** https://edge.vinkius.com/vk_preview_Z06xKXBLbhmRgabDfSd6BBkkFtoyph9gsiWpikY1/ai-agent-connect
- **Tags:** proxy-management, serp-api, web-unlocker, data-extraction, browser-automation

## Description

Bright Data lets you connect your AI agent to a massive web data platform. You can get past anti-bot systems and pull structured data from search engines without the usual headache of managing proxies. It handles the heavy lifting of web data engineering, so you can focus on the results. Because Vinkius hosts this Connector, you can plug it into your existing workflow and start getting clean data from protected sources. It turns your agent into a data engineer that can navigate the messy reality of the modern web. You can manage your entire proxy infrastructure, from checking zone statuses to retrieving credentials for external scripts. It also lets you browse a marketplace of pre-collected datasets for industries like retail or social media. You can scrape a specific URL or trigger a massive collection job for thousands of pages. This tool makes the process feel like a simple conversation. You get to skip the manual setup of headless browsers and the constant battle against IP blocks. You can monitor active browser sessions to see how they are performing or check your account status to ensure your keys are still valid. It's about getting the data you need, in the format you want, without the technical friction of traditional scraping. You can pull everything from real-time search results to structured data from complex platforms like LinkedIn and Amazon.

## Tools

### get_dataset_snapshot
This tool lets you download structured JSON records from a completed collection. It is the best way to grab your final data once a job finishes.

### list_datasets
This tool lets you browse all available datasets in the Bright Data marketplace. Use it to find the correct dataset ID for your specific project.

### send_request
This tool scrapes a specific URL or extracts structured search results from a zone. It is your primary way to bypass anti-bot protections and get data.

### get_all_zones
This tool lists every proxy zone currently configured in your account. It helps you see what is available before you try to send a request.

### get_dataset_progress
This tool shows the current status, record count, and error count for a running job. You can poll it to see when your data is ready for download.

### get_account_status
This tool checks your API key validity and account standing. Use it first if your other requests are returning authentication errors.

### get_zone_info
This tool provides the configuration and status for a specific proxy zone. It is a great way to verify a zone is active before you try to use it.

### get_zone_passwords
This tool retrieves proxy credentials for use in external tools like Selenium or Playwright. It is useful for when you need a direct connection outside of this Connector.

### list_browser_sessions
This tool lists your active Scraping Browser sessions. You can use it to check bandwidth usage or see how long a session has been running.

### trigger_dataset
Use this to start an async scraping job for over 100 sources like LinkedIn, Amazon, or Instagram. It handles the heavy lifting of initiating the data pull.

## Prompt Examples

**Prompt:** 
```
Can you find the top 5 results for 'best laptops 2024' on Google and put them in a table?
```

**Response:** 
```
I've pulled the top 5 search results from Google for you. Here is the structured data:

| Rank | Title | URL |
|---|---|---|
| 1 | Best Laptops 2024 - Tech Review | https://example.com/tech1 |
| 2 | Top 10 Laptops of the Year | https://example.com/tech2 |
| 3 | Laptop Buying Guide 2024 | https://example.com/tech3 |
| 4 | Best Work Laptops | https://example.com/tech4 |
| 5 | Student Laptop Deals | https://example.com/tech5 |
```

**Prompt:** 
```
I need to start a scrape of LinkedIn posts from this URL: https://linkedin.com/pulse/example-post. Let me know when it's finished.
```

**Response:** 
```
I have triggered the LinkedIn Posts dataset for that URL. I will monitor the progress for you and send a message as soon as the snapshot is ready for download.
```

**Prompt:** 
```
What proxy zones do I have active right now?
```

**Response:** 
```
You have the following proxy zones currently configured:

* **web_unlocker_1**: Web Unlocker zone
* **serp_api_zone**: SERP API zone
* **retail_proxy_main**: Dedicated retail scraper zone
```

## Capabilities

### Bypass anti-bot protections
Get around sophisticated blocks on high-traffic websites to access content that is usually hidden.

### Extract search results
Pull structured data from Google, Bing, and other search engines into a clean format.

### Trigger large datasets
Start automated scraping jobs for over 100 different sources like Amazon or Instagram.

### Monitor scraping progress
Check the status of active data collection jobs in real-time to see when they are finished.

### Manage proxy zones
Configure and verify your proxy infrastructure through natural language instead of a dashboard.

### Track browser sessions
See what's happening inside your active Scraping Browser sessions to monitor bandwidth and duration.

## Use Cases

### Tracking competitor rankings
A researcher needs to track a competitor's ranking on Google. They ask the agent to use send_request on a search URL to get the top 10 results.

### Bulk scraping social media
A dev needs to scrape 1,000 LinkedIn posts. They use trigger_dataset and then poll get_dataset_progress until the agent confirms completion.

### Verifying proxy health
An analyst wants to see if their proxy keys are still valid. They ask the agent to check the account status to ensure they are ready for use.

### Finding retail datasets
A data scientist needs to find a specific dataset for Amazon products. They use list_datasets to browse the options and find the right ID.

## Benefits

- Skip the captcha wall by accessing protected content through specialized proxy zones.
- Automate social media monitoring by triggering datasets for platforms like LinkedIn.
- Get clean search data from Google or Bing using the SERP API.
- Keep your proxy infrastructure organized by checking your zone status through chat.
- Save time on data engineering by pulling finished records directly into your session.
- Debug your scraping workflows in real-time by viewing active browser sessions.

## How It Works

The bottom line is you get high-quality web data without the headache of managing the infrastructure.

1. Connect your Bright Data API key to the Connector.
2. Tell your agent which website or dataset you need to target.
3. Receive the structured data or status updates directly in your chat.

## Frequently Asked Questions

**Does the Bright Data MCP help with Google search results?**
Yes, it allows your agent to pull structured data from search engines like Google and Bing using the SERP API. This means you get clean results instead of raw HTML.

**How does Bright Data help bypass bot detection?**
It uses specialized proxy zones and web unlockers to get past captchas and IP blocks on protected sites. This allows your agent to access data that would normally be blocked.

**Can I use this to scrape LinkedIn or Instagram?**
Yes, you can trigger specific datasets for those platforms to collect posts or profile data automatically. It handles the complexity of the platform's anti-scraping measures.

**Is it easy to manage proxies with this Connector?**
You can list your zones and check their status using natural language instead of clicking through a dashboard. It makes managing your proxy infrastructure much faster.

**How do I get the data after a scrape is done?**
Your agent can pull the final JSON records for you once the collection status is marked as ready. You don't have to manually log in to find your files.

**Can I use this for my own custom scraping scripts?**
You can use the tool to fetch proxy credentials for use in your own Selenium or Playwright setups. It's a great way to get secure, high-quality proxies for your custom code.

**What is the workflow for scraping a LinkedIn post?**
Use `trigger_dataset` with the LinkedIn Posts dataset ID (`gd_lyy3tktm25m4avu764`) and a direct post URL (e.g., `https://www.linkedin.com/feed/update/urn:li:activity:...`). This returns a `snapshot_id`. Then poll `get_dataset_progress` every 15–30 seconds until the status is `ready` — LinkedIn scraping typically takes 60–120 seconds. Finally, call `get_dataset_snapshot` to retrieve the structured data including post content, author details, reactions, and comment count.

**Why does send_request fail with a zone error?**
The `send_request` tool requires an active proxy zone (Web Unlocker or SERP API). If your account has no zones configured, the request will fail. Run `get_all_zones` first to check your available zones. If the list is empty, you need to create a zone in the [Bright Data dashboard](https://brightdata.com/cp/zones) before using `send_request`.

**Which URL formats does the LinkedIn Posts dataset accept?**
The LinkedIn Posts dataset (`gd_lyy3tktm25m4avu764`) only accepts direct post URLs matching the pattern `linkedin.com/(pulse|posts|feed/update)`. Valid examples: `https://www.linkedin.com/feed/update/urn:li:activity:1234567890` or `https://www.linkedin.com/posts/username_title-activity-1234567890`. Search result pages, profile pages, and hashtag pages are not supported.

**How do I find the right dataset ID for my use case?**
Use the `list_datasets` tool to browse all 100+ available datasets. Common IDs: LinkedIn Posts (`gd_lyy3tktm25m4avu764`), LinkedIn People (`gd_l1viktl72bvl7bjuj0`), LinkedIn Companies (`gd_l1vikfnt1wgvvqz95w`), Amazon Products (`gd_l7q7dkf244hwjntr0`), Instagram Profiles (`gd_l1vikfch901nx3by4`), Google Maps (`gd_m8ebnr0q2qlklc02fz`).