---
name: firecrawl
description: Use when extracting web data, scraping pages, crawling websites, searching the web, or interacting with dynamic content. Agents should reach for Firecrawl when they need to gather structured data from URLs, handle JavaScript-rendered content, extract specific information from web pages, or monitor websites for changes.
metadata:
    mintlify-proj: firecrawl
    version: "1.0"
---

# Firecrawl Skill

## Product Summary

Firecrawl is a web data API for AI agents that converts any URL into clean markdown, structured JSON, or other formats. It handles proxies, anti-bot detection, JavaScript rendering, and dynamic content automatically. Use it to scrape individual pages, crawl entire websites, search the web, interact with dynamic content, or monitor pages for changes. The primary docs are at https://docs.firecrawl.dev.

**Key files and commands:**
- API base URL: `https://api.firecrawl.dev/v2`
- Authentication: `Authorization: Bearer fc-YOUR-API-KEY` header
- SDKs: Python (`pip install firecrawl`), Node.js (`npm install firecrawl`), CLI (`npx firecrawl-cli@latest init`)
- Core endpoints: `/scrape`, `/crawl`, `/search`, `/interact`, `/parse`, `/map`, `/monitor`, `/agent`
- MCP server (keyless): `https://mcp.firecrawl.dev/v2/mcp`

## When to Use

Reach for Firecrawl when:
- **Scraping a single URL**: Extract markdown, HTML, JSON, screenshots, or structured data from any webpage
- **Crawling a website**: Recursively gather content from all pages on a domain with filtering and depth control
- **Searching the web**: Find pages and get full content from results in one call
- **Extracting structured data**: Use JSON schemas or the `product` format to pull specific fields from pages
- **Handling dynamic content**: Pages require JavaScript rendering, form filling, or browser interaction
- **Monitoring pages**: Track when content changes and get notified via webhooks
- **Processing documents**: Parse PDFs, DOCX, XLSX, and HTML files from URLs or local disk
- **Batch operations**: Scrape multiple URLs in parallel with a single job

Do not use Firecrawl for: authentication setup, account management, or dashboard-only operations.

## Quick Reference

### Essential Endpoints

| Endpoint | Purpose | Use When |
|----------|---------|----------|
| `POST /v2/scrape` | Extract content from a single URL | You need data from one page |
| `POST /v2/crawl` | Recursively crawl a website | You need data from all pages on a domain |
| `POST /v2/search` | Search the web and get full page content | You need to find and extract from search results |
| `POST /v2/batch/scrape` | Scrape multiple URLs in parallel | You have a list of URLs to process |
| `POST /v2/interact/browser` | Create a browser session for interaction | You need to click, fill forms, or navigate |
| `POST /v2/parse` | Parse local files (PDF, DOCX, etc.) | You have files on disk to convert |
| `POST /v2/map` | Discover all URLs on a website | You need a complete URL inventory |
| `POST /v2/monitor` | Schedule recurring checks | You need to track content changes |

### Scrape Formats

Request multiple formats in a single call. Common formats:
- `markdown` — clean markdown output
- `html` — cleaned HTML
- `rawHtml` — original unmodified HTML
- `json` — structured data (requires schema or prompt)
- `screenshot` — base64-encoded image
- `links` — all outbound links
- `product` — deterministic product extraction (title, price, variants)
- `branding` — brand colors, fonts, typography
- `question` — answer a natural-language question about the page
- `highlights` — extract relevant source text

### SDK Installation

```bash
# Python
pip install firecrawl

# Node.js
npm install firecrawl

# CLI
npx firecrawl-cli@latest init --all --browser
```

### Basic Usage Pattern

```python
from firecrawl import Firecrawl

firecrawl = Firecrawl(api_key='fc-YOUR_API_KEY')

# Scrape a single URL
result = firecrawl.scrape('https://example.com', formats=['markdown', 'json'])

# Crawl a website
crawl_result = firecrawl.crawl('https://example.com', limit=100)

# Search the web
search_result = firecrawl.search('AI agents', limit=5)
```

### Crawl Configuration Parameters

| Parameter | Type | Default | Purpose |
|-----------|------|---------|---------|
| `url` | string | required | Starting URL to crawl |
| `limit` | integer | 10000 | Max pages to crawl |
| `maxDiscoveryDepth` | integer | none | Max link-discovery hops from root |
| `includePaths` | string[] | none | Regex patterns for paths to include |
| `excludePaths` | string[] | none | Regex patterns for paths to exclude |
| `crawlEntireDomain` | boolean | false | Follow sibling/parent URLs, not just children |
| `allowSubdomains` | boolean | false | Follow links to subdomains |
| `allowExternalLinks` | boolean | false | Follow links to external sites (one hop) |
| `sitemap` | string | "include" | Sitemap handling: "include", "skip", or "only" |
| `maxConcurrency` | integer | team limit | Concurrent scrapes at once |
| `delay` | number | none | Delay in seconds between scrapes |
| `scrapeOptions` | object | none | Options for each page (formats, proxy, etc.) |

## Decision Guidance

### When to Use X vs Y

| Scenario | Use | Why |
|----------|-----|-----|
| Single URL extraction | `/scrape` | Faster, simpler, returns immediately |
| Multiple URLs (list) | `/batch/scrape` | Parallel processing, single job |
| Entire website | `/crawl` | Discovers links automatically, handles sitemaps |
| Find pages + extract | `/search` | One call gets results and full content |
| Local file conversion | `/parse` | Works with files on disk, not URLs |
| Dynamic content | `/interact` + `/scrape` | Scrape first, then interact with result |
| Structured extraction | `json` format + schema | Deterministic; use `product` format for products |
| Change detection | `/monitor` | Recurring checks with webhook notifications |
| Polling vs webhook | Webhook | Async, real-time, no polling overhead |

### Scrape vs Parse

- **Scrape**: URL-based, handles JavaScript, proxies, anti-bot. Use for web pages.
- **Parse**: File-based (local disk), no JavaScript. Use for PDFs, DOCX, XLSX, HTML files.

### Crawl Scope Control

- **Children only (default)**: `website.com/blogs/` crawls `website.com/blogs/post-1` but not `website.com/other/`
- **Entire domain**: Set `crawlEntireDomain: true` to include siblings and parents
- **Subdomains**: Set `allowSubdomains: true` to follow `blog.website.com` from `website.com`
- **External links**: Set `allowExternalLinks: true` to follow one hop off-domain (homepage links skipped)

## Workflow

### Typical Scrape Task

1. **Identify the URL and data needed**: Determine what page to scrape and what format you need (markdown, JSON, screenshot, etc.).
2. **Choose formats**: Request multiple formats in one call if needed (e.g., `markdown` + `json` with schema).
3. **Set options**: Configure caching (`maxAge`), location, proxy, PII redaction, or other parameters.
4. **Make the request**: Call `/scrape` with the URL and options.
5. **Check response status**: Verify `success: true` and check `data.metadata.statusCode` for the page's HTTP status.
6. **Extract data**: Access `data.markdown`, `data.json`, `data.screenshot`, etc. based on requested formats.
7. **Handle errors**: If `success: false`, check the error message and retry with appropriate backoff.

### Typical Crawl Task

1. **Define scope**: Set `limit`, `includePaths`, `excludePaths`, and depth/domain parameters.
2. **Configure scrape options**: Specify formats, proxy, caching, and extraction schemas for all pages.
3. **Submit crawl**: Call `/crawl` with the starting URL and configuration.
4. **Poll or wait**: Use SDK's `crawl()` method (waits) or `startCrawl()` + polling, or subscribe to webhooks.
5. **Handle pagination**: Results are capped at 10MB; follow `next` URL if present and status is not `completed`.
6. **Process results**: Iterate over `data` array; use `GET /crawl/{id}/errors` for failed pages.
7. **Check expiration**: Results expire after 24 hours; download or process immediately.

### Batch Scrape Task

1. **Prepare URL list**: Collect all URLs to scrape.
2. **Set common options**: Define formats, proxy, caching, and extraction schemas.
3. **Submit batch**: Call `/batch/scrape` with URL list and options.
4. **Poll or wait**: Use SDK's `batch_scrape()` (waits) or `start_batch_scrape()` + polling.
5. **Retrieve results**: Access `data` array with results for each URL.
6. **Check errors**: Use `GET /batch/scrape/{id}/errors` for failed URLs.

## Common Gotchas

- **Crawl scope defaults to children only**: By default, `website.com/blogs/` does not crawl `website.com/other/`. Set `crawlEntireDomain: true` to include siblings and parents.
- **Regex patterns match pathname, not full URL**: `includePaths` and `excludePaths` match only the path part of the URL by default. Set `regexOnFullURL: true` to match the full URL including query parameters.
- **Sitemap is included by default**: The crawler uses the sitemap to discover URLs. Set `sitemap: "skip"` to use HTML links only, or `sitemap: "only"` to crawl only the sitemap.
- **Results expire after 24 hours**: Crawl and batch scrape results are only available via API for 24 hours. Download or process immediately.
- **Credit costs vary by format**: Base scrape costs 1 credit. JSON mode adds 4 credits, question/highlights add 4 credits each, PII redaction adds 1 credit, PDF parsing adds 1 credit per page, audio/video extraction adds 4 credits.
- **Batch scrape and crawl are async**: Both return a job ID. Use polling or webhooks to get results; do not expect immediate responses.
- **Caching is on by default**: `maxAge` defaults to 2 days. Set `maxAge: 0` to force fresh content, but this bypasses cache and takes longer.
- **robots.txt is respected**: The crawler respects robots.txt by default. Set `ignoreRobotsTxt: true` (Enterprise only) to bypass.
- **External link homepage skipping**: With `allowExternalLinks: true`, links to external site homepages are intentionally skipped to avoid crawling unrelated sites.
- **Non-deterministic crawl results**: Crawl results may vary between runs because pages are scraped concurrently. Set `maxConcurrency: 1` or `sitemap: "only"` for reproducibility.
- **Webhook signature verification required**: Always verify the `X-Firecrawl-Signature` header using HMAC-SHA256 before processing webhooks.
- **API key in headers, not query params**: Always use the `Authorization: Bearer` header; never pass the API key in the URL.
- **429 rate limit responses**: When you hit rate limits, the API returns 429. Implement exponential backoff; SDKs do this automatically.

## Verification Checklist

Before submitting work with Firecrawl:

- [ ] **API key is set**: Verify the API key is in the `Authorization` header or environment variable.
- [ ] **URL is valid**: Confirm the URL is accessible and not behind authentication (unless using `/interact`).
- [ ] **Formats are correct**: Ensure requested formats are valid (e.g., `markdown`, `json`, `screenshot`).
- [ ] **JSON schema is valid**: If using `json` format, validate the schema is proper JSON Schema.
- [ ] **Crawl scope is intentional**: Confirm `limit`, `includePaths`, `excludePaths`, and domain parameters match your intent.
- [ ] **Response status is checked**: Verify `success: true` and check `data.metadata.statusCode` for the page's HTTP status.
- [ ] **Errors are handled**: Implement retry logic for transient failures (408, 409, 5xx); do not retry 401 or 429 without backoff.
- [ ] **Results are downloaded**: For crawl/batch scrape, download results within 24 hours before they expire.
- [ ] **Webhook signature is verified**: If using webhooks, verify the `X-Firecrawl-Signature` header.
- [ ] **Credit usage is monitored**: Check `data.creditsCost` in responses to track spending.
- [ ] **Pagination is handled**: For crawl results, follow `next` URL if present until `status` is `completed` and `next` is absent.

## Resources

- **Full documentation index**: https://docs.firecrawl.dev/llms.txt (comprehensive page-by-page navigation for agents)
- **API Reference**: https://docs.firecrawl.dev/api-reference/v2-introduction
- **Scrape Feature Docs**: https://docs.firecrawl.dev/features/scrape
- **Crawl Feature Docs**: https://docs.firecrawl.dev/features/crawl

---

> For additional documentation and navigation, see: https://docs.firecrawl.dev/llms.txt