Browsertrix Crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
About
Browsertrix Crawler is a standalone browser-based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. Browsertrix Crawler uses Puppeteer to control one or more Brave Browser browser windows in parallel. Data is captured through the Chrome Devtools Protocol (CDP) in the browser.
Open Source Health
- Stars
- 1,134
- Forks
- 151
- License
- AGPL-3.0
- Last commit
- 29 days ago
Resources & Links
Related Categories
Vendor
Webrecorder
Publisher of ArchiveWeb.page and Browsertrix Crawler
Quick Links
Open Source
Stars
1.1K
Forks
151
Contributors
0
Open Issues
147
More by Webrecorder
Related Products
Deep Research
An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine…
Shared categories
select.rs
A Rust library to extract useful data from HTML documents, suitable for web scraping.
Shared categories
Scrapely
A pure-python HTML screen-scraping library
Shared categories
Faster Than Requests
Faster requests on Python 3
Shared categories
MechanicalSoup
A Python library for automating interaction with websites.
Shared categories