# Internet Archive Wayback MCP for AI Agents AI Agent Connect

> Internet Archive Wayback MCP lets your AI agent access over 800 billion archived web pages from the last 25 years. It allows you to verify if a URL was preserved, find the most recent snapshot, or pull detailed capture histories including status codes and MIME types. Use it to track how websites evolved, recover deleted content, or perform digital forensics on historical web data.

## Overview
- **Category:** knowledge-management
- **Price:** Free
- **Endpoint:** https://edge.vinkius.com/vk_preview_hnm6MKLSX94NqzQjsRaRbJo66cdvIxnhQfEOtu5s/ai-agent-connect
- **Tags:** web-archiving, url-history, snapshot-analysis, digital-preservation, internet-history, cdx-api

## Description

The Internet Archive Wayback MCP connects your AI agent to the Wayback Machine, giving it access to over 800 billion archived pages. You can ask your agent to pull a specific snapshot, check if a URL was ever live, or grab a list of every time a page was saved. It's a massive win for anyone who needs to verify past claims, track how a brand's design has changed over decades, or hunt for subdomain footprints. Because Vinkius hosts this Connector, it plugs right into your existing setup without needing any API keys or complex configuration. You get straight answers about a site's history, from its first appearance to its most recent save, making it a powerhouse for deep research and digital forensics.

## Tools

### get_captures_by_mime_type
The get_captures_by_mime_type tool finds specific resource types like PDFs or images in the archive. This helps you narrow down results when you only need a certain file format.

### get_captures_by_status
The get_captures_by_status tool analyzes site availability patterns by filtering for specific HTTP status codes. Use this to see how often a page returned errors over time.

### get_captures_by_year
The get_captures_by_year tool isolates snapshots from a specific year for temporal analysis. It lets you see what a site looked like during a specific calendar year.

### get_cdx_captures
The get_cdx_captures tool pulls detailed capture history including timestamps and file sizes. Use this for a deep dive into every time a URL was saved.

### check_availability
The check_availability tool verifies if a URL was archived and finds its most recent snapshot. This is the quickest way to see if a page exists in the archive.

### get_captures_collapsed
The get_captures_collapsed tool provides a clean list of unique page captures without redundant entries. This is useful for getting a clear overview of a domain's history.

### get_capture_count
The get_capture_count tool shows the total number of times a specific URL was saved. This helps you measure the frequency of a site's preservation.

### get_first_capture
The get_first_capture tool finds the very first time a page was preserved in the archive. Use this to identify when a site first appeared online.

### get_latest_capture
The get_latest_capture tool retrieves the most recent archived version of a specific URL. This gives you the newest available snapshot of a page.

### get_subdomain_captures
The get_subdomain_captures tool discovers all archived subdomains for a given domain. Use this to map out the full footprint of a site's historical presence.

## Prompt Examples

**Prompt:** 
```
Did this site exist in 2014?
```

**Response:** 
```
✅ **Snapshot Found!**

The site was archived in 2014. Here are the details:

*   **Closest Snapshot:** 2014-05-12
*   **Status:** 200 OK
*   **View Snapshot:** [https://web.archive.org/web/20140512...](https://web.archive.org/web/20140512...)

I can also pull the full capture history for this year if you'd like.
```

**Prompt:** 
```
Show me all the PDFs archived for example.com.
```

**Response:** 
```
I found the following PDF resources in the archive for **example.com**:

| Date | URL | Status |
| :--- | :--- | :--- |
| 2023-01-10 | /docs/manual.pdf | 200 |
| 2022-06-15 | /reports/annual.pdf | 200 |
| 2021-11-20 | /whitepaper.pdf | 301 |

Would you like me to find the most recent version of any of these?
```

**Prompt:** 
```
Find the first time archive.org was saved.
```

**Response:** 
```
The earliest recorded capture for **archive.org** in the Wayback Machine is:

*   **First Capture Date:** 1998-01-01
*   **Original URL:** https://archive.org
*   **Status Code:** 200

This shows the site has been preserved for over 25 years.
```

## Capabilities

### Verify URL snapshots
Check if a specific page exists in the archive and find its most recent version.

### Pull full CDX histories
Get a complete log of timestamps, status codes, and file sizes for every save.

### Isolate snapshots by year
Filter the archive to see what a site looked like during a specific calendar year.

### Filter captures by status
Analyze site availability patterns by looking for specific HTTP status codes.

### Filter captures by MIME type
Find specific resource types like PDFs, images, or stylesheets in the history.

### Find earliest captures
Identify the very first time a page was preserved in the archive.

### Find most recent captures
Retrieve the newest available snapshot of a specific URL.

### Count total captures
Measure how frequently a specific page has been saved over time.

### Deduplicate page results
Get a clean list of unique page captures without redundant entries.

### Discover archived subdomains
Map out the full footprint of a domain by finding all archived subdomains.

## Use Cases

### Verifying a deleted claim
A journalist asks the agent to find the 2015 version of a news site to confirm a quote that has since been removed.

### Tracking a brand's rebranding
A designer uses the Connector to see how a company's homepage changed from 2018 to 2024 to study design trends.

### Digital forensics on a phishing site
A security pro finds the first capture of a malicious domain to see the original setup and infrastructure.

### SEO audit of a legacy site
A developer uses get_captures_by_status to see how often a site returned 404s over the last year.

## Benefits

- You can verify historical claims by using check_availability to see exactly what a page looked like on a specific date.
- You can track site evolution over time by using get_captures_by_year to see how a brand changed its messaging.
- You can identify deleted content for legal purposes using get_first_capture to find the original state of a site.
- You can perform deep web forensics using get_cdx_captures to pull every status code and MIME type ever recorded.
- You can map out an entire domain's footprint using get_subdomain_captures to find every archived subdomain.
- You can clean up messy data sets with get_captures_collapsed to see unique pages instead of redundant entries.

## How It Works

The bottom line is you get instant access to decades of web history without manual searching.

1. Subscribe to the Internet Archive Wayback MCP in the Vinkius marketplace.
2. Connect your AI client to the new MCP.
3. Ask your agent to find specific web snapshots or history logs.

## Frequently Asked Questions

**Can I find deleted pages with the Internet Archive Wayback MCP?**
Yes, this Connector allows you to find snapshots of pages that have been removed from the live web, provided they were captured by the Wayback Machine in the past.

**Does the Internet Archive Wayback MCP work for live sites?**
No, this Connector is specifically for accessing archived data. It does not browse the live internet or interact with current website content.

**How do I find old images with the Internet Archive Wayback MCP?**
You can use the MIME type filter to specifically look for image files like JPEGs or PNGs that were saved during previous captures.

**Can I see a company's old logos with the Internet Archive Wayback MCP?**
Yes, by requesting snapshots from

**How far back does the Wayback Machine go?**
The Wayback Machine has archived web pages since 1996. However, coverage varies significantly — major websites have captures going back 20+ years, while smaller or newer sites may have fewer or no captures. Use get_first_capture to find the earliest archived version of any URL.

**Can I find captures that returned 404 errors?**
Yes! Use get_captures_by_status with status_code="404". This returns all archived versions where the page returned a Not Found error. This is useful for tracking when pages were removed or URLs changed structure.

**Can I discover all subdomains of a website that have been archived?**
Yes! Use get_subdomain_captures with the base domain (e.g., "example.com"). This returns captures for all subdomains like www.example.com, blog.example.com, api.example.com, etc. It's useful for mapping the full archival footprint of an organization's web presence.