The web scraper API your AI agent deserves.
webclaw extracts public web pages into Markdown and structured JSON for your agents. Reduce page clutter, retrieve source content, and request the fields your application needs.
- Cancel anytime
- Self-host forever
- Open source core
pages extracted
bot walls bypassed
websites scraped
github stars
Try it
Paste a URL. See what your model would read.
Sign in and it runs for real — 3 free runs a day, no card.Plans from $19/mo
also on an account: extract · research · batch
Global reach
Read public pages around the world
Retrieve public pages from sites around the world. Content and access can vary by region, language and website. Check the returned content against your application’s needs.
One API, every job
Why webclaw
Everything an agent needs to read the web
Formats for your workflow
Markdown, JSON, plain text and raw HTML, plus structured extraction and an LLM mode that removes page clutter.
Handles protected pages
The API attempts protected-page handling and JavaScript rendering when needed. Check returned content and handle upstream errors in your application.
No browser to babysit
Send a URL and receive readable content. JavaScript rendering and page loading are handled by the API, so you can focus on your application.
12 MCP tools, one install
Scrape, crawl, search, extract, research and more, over stdio — inside Claude, Cursor, Windsurf, Codex or any other MCP client.
Open source, self-hostable
An AGPL-3.0 Rust core you can run yourself. CLI, REST server and MCP server, on your own hardware without hosted API credits.
Firecrawl-compatible
Compatible /v2/scrape, /v2/crawl, and /v2/search routes can reduce migration work. Check supported options before switching.
Before vs after webclaw
Stop building your own scraper. Call one web scraper API instead.
Without webclaw
You wire up parsers, headless browsers and proxy pools, then babysit the stack every time a site ships a redesign and your selectors snap.
Your requests hit a bot wall, come back as an empty shell of HTML, and your agent quietly reasons over nothing.
You dump raw page markup into the model and watch nav bars, cookie banners and script tags eat the context window.
Without retrieved sources, a model may rely on outdated training data when answering questions about prices, documentation or news.
One page is JavaScript, the next is a PDF, a third is a DOCX — so you bolt on a different library for every format and glue them by hand.
With webclaw
Request Markdown, structured JSON, or LLM-oriented text through one API, then check the result against your application’s content requirements.
webclaw attempts protected-page handling and JavaScript rendering when needed. Upstream restrictions can still prevent extraction.
Removing page markup and boilerplate can reduce the text you send to your model. The reduction depends on the page and output format.
Give agents retrieved page content and source URLs. Set cache options to suit your freshness needs and verify the citations they produce.
Extract content from HTML, PDFs, and DOCX through the API. The open-source MCP server offers 12 tools built on the same core extraction library.
webclaw offers compatible /v2/scrape, /v2/crawl, and /v2/search routes. Check supported request options and response fields in the migration guide before switching.
See the full comparisonYour sources
Start with the public sites your workflow needs. Test your target pages and verify the returned content.
Use cases
Who is pointing webclaw at the web
For AI engineers
Give your agent live web data instead of stale training data.
For growth and sales
Track competitors' pricing, changelogs and docs as they change.
For data teams
Keep a dataset fresh with scheduled crawls and change detection.
Pricing
Pay for pages, not seats
Standard page operations use credits per processed page. Research has a separate plan allowance. See endpoint costs in the docs.
Starter
$19/mo
Card required · cancel anytime
- Credits 5,000/mo
- Research 3 runs/mo
- Concurrency 5
- Support Email
Growth · popular
$49/mo
- Credits 25,000/mo
- Research 10 runs/mo
- Concurrency 20
- Support Priority
Pro
$99/mo
- Credits 100,000/mo
- Research 20 runs/mo
- Concurrency 50
- Support Priority
Scale
$399/mo
- Credits 500,000/mo
- Research 60 runs/mo
- Concurrency 100
- Support Priority + Slack
Cancel anytime · credit packs from $20 · self-host the open-source core for free
See how credits workWebclaw extracts content from public web pages for applications and AI workflows. Request Markdown, text, structured data, or an LLM format that removes page markup and boilerplate.
Get started
Ship your agent today. Build with web data.
Cancel anytime. Use the migration guide to connect through the supported Firecrawl-compatible routes.
AGPL-3.0Built in Rust2,343 stars on GitHub














