MarketplaceWeb CrawlingBrowsertrix Crawler
Web CrawlingOpen Source Developer Tools

Browsertrix Crawler

Run a high-fidelity browser-based web archiving crawler in a single Docker container

Open Source

About

Browsertrix Crawler is a standalone browser-based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. Browsertrix Crawler uses Puppeteer to control one or more Brave Browser browser windows in parallel. Data is captured through the Chrome Devtools Protocol (CDP) in the browser.

Open Source Health

Not enough history
Stars
1,134
Forks
151
License
AGPL-3.0
Last commit
29 days ago
TypeScript

Related Categories