BackCrawl4AI
🐙

Crawl4AI

Open Source

Freeweb scrapingragllmdata extractionopen source

Crawl4AI is an open-source web crawling and scraping library purpose-built for feeding AI pipelines — it converts pages directly into clean, LLM-ready markdown and structured JSON rather than raw HTML, with built-in support for JavaScript-heavy sites, chunking strategies, and extraction schemas. It's designed as a fast, self-hostable alternative to paid scraping APIs for teams building RAG systems, AI agents, or LLM training/fine-tuning datasets that need reliable, cleaned web content at scale. With nearly 75,000 GitHub stars, it's a popular default choice in the open-source RAG tooling ecosystem.