Linkcrawler
Cross-platform persistent and distributed web crawler :link:
About
linkcrawler is persistent because the queue is stored in a remote database that is automatically re-initialized if interrupted. linkcrawler is distributed because multiple instances of linkcrawler will work on the remotely stored queue, so you can start as many crawlers as you want on separate machines to speed along the process. linkcrawler is also fast because it is threaded and uses connection pools.
Open Source Health
- Stars
- 113
- Forks
- 9
- License
- MIT
- Last commit
- 9 years ago
Related Categories
Vendor
Schollz
Publisher of Hostyoself, e2ecp and PAKE
Quick Links
Open Source
More by Schollz
Related Products
Deep Research
An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine…
Shared categories
select.rs
A Rust library to extract useful data from HTML documents, suitable for web scraping.
Shared categories
Scrapely
A pure-python HTML screen-scraping library
Shared categories
Faster Than Requests
Faster requests on Python 3
Shared categories
MechanicalSoup
A Python library for automating interaction with websites.
Shared categories