webclaw

The web scraper API your AI agent deserves.

webclaw extracts public web pages into Markdown and structured JSON for your agents. Reduce page clutter, retrieve source content, and request the fields your application needs.

  • Cancel anytime
  • Self-host forever
  • Open source core

pages extracted

bot walls bypassed

websites scraped

,

github stars

Try it

Paste a URL. See what your model would read.

Sign in and it runs for real — 3 free runs a day, no card.Plans from $19/mo

also on an account: extract · research · batch

Global reach

Read public pages around the world

Retrieve public pages from sites around the world. Content and access can vary by region, language and website. Check the returned content against your application’s needs.

One API, every job

Why webclaw

Everything an agent needs to read the web

Formats for your workflow

Markdown, JSON, plain text and raw HTML, plus structured extraction and an LLM mode that removes page clutter.

Handles protected pages

The API attempts protected-page handling and JavaScript rendering when needed. Check returned content and handle upstream errors in your application.

No browser to babysit

Send a URL and receive readable content. JavaScript rendering and page loading are handled by the API, so you can focus on your application.

12 MCP tools, one install

Scrape, crawl, search, extract, research and more, over stdio — inside Claude, Cursor, Windsurf, Codex or any other MCP client.

Open source, self-hostable

An AGPL-3.0 Rust core you can run yourself. CLI, REST server and MCP server, on your own hardware without hosted API credits.

Firecrawl-compatible

Compatible /v2/scrape, /v2/crawl, and /v2/search routes can reduce migration work. Check supported options before switching.

Before vs after webclaw

Stop building your own scraper. Call one web scraper API instead.

Without webclaw

You wire up parsers, headless browsers and proxy pools, then babysit the stack every time a site ships a redesign and your selectors snap.

Your requests hit a bot wall, come back as an empty shell of HTML, and your agent quietly reasons over nothing.

You dump raw page markup into the model and watch nav bars, cookie banners and script tags eat the context window.

Without retrieved sources, a model may rely on outdated training data when answering questions about prices, documentation or news.

One page is JavaScript, the next is a PDF, a third is a DOCX — so you bolt on a different library for every format and glue them by hand.

With webclaw

Request Markdown, structured JSON, or LLM-oriented text through one API, then check the result against your application’s content requirements.

webclaw attempts protected-page handling and JavaScript rendering when needed. Upstream restrictions can still prevent extraction.

Removing page markup and boilerplate can reduce the text you send to your model. The reduction depends on the page and output format.

Give agents retrieved page content and source URLs. Set cache options to suit your freshness needs and verify the citations they produce.

Extract content from HTML, PDFs, and DOCX through the API. The open-source MCP server offers 12 tools built on the same core extraction library.

webclaw offers compatible /v2/scrape, /v2/crawl, and /v2/search routes. Check supported request options and response fields in the migration guide before switching.

See the full comparison

Your sources

Start with the public sites your workflow needs. Test your target pages and verify the returned content.

NikeAirbnbEtsyShopifyIKEATargetTripAdvisorZillowIMDbRotten TomatoesYelpGoodreadsWikipediaGitHubProduct HuntY CombinatorStripeSpotifyNikeAirbnbEtsyShopifyIKEATargetTripAdvisorZillowIMDbRotten TomatoesYelpGoodreadsWikipediaGitHubProduct HuntY CombinatorStripeSpotify

Use cases

Who is pointing webclaw at the web

For AI engineers

Give your agent live web data instead of stale training data.

For growth and sales

Track competitors' pricing, changelogs and docs as they change.

For data teams

Keep a dataset fresh with scheduled crawls and change detection.

Pricing

Pay for pages, not seats

Standard page operations use credits per processed page. Research has a separate plan allowance. See endpoint costs in the docs.

Starter

$19/mo

Card required · cancel anytime

  • Credits 5,000/mo
  • Research 3 runs/mo
  • Concurrency 5
  • Support Email

Growth · popular

$49/mo

  • Credits 25,000/mo
  • Research 10 runs/mo
  • Concurrency 20
  • Support Priority

Pro

$99/mo

  • Credits 100,000/mo
  • Research 20 runs/mo
  • Concurrency 50
  • Support Priority

Scale

$399/mo

  • Credits 500,000/mo
  • Research 60 runs/mo
  • Concurrency 100
  • Support Priority + Slack

Cancel anytime · credit packs from $20 · self-host the open-source core for free

See how credits work

FAQ

Your questions, answered

For anything else, we are one message away in Discord.

Book a call

Webclaw extracts content from public web pages for applications and AI workflows. Request Markdown, text, structured data, or an LLM format that removes page markup and boilerplate.

Partners

Backing open web extraction

SerpApiColdProxyNodeMavenMangoProxySerpApiColdProxyNodeMavenMangoProxySerpApiColdProxyNodeMavenMangoProxySerpApiColdProxyNodeMavenMangoProxy

Get started

Ship your agent today. Build with web data.

Cancel anytime. Use the migration guide to connect through the supported Firecrawl-compatible routes.

Star on GitHub

AGPL-3.0Built in Rust2,343 stars on GitHub