Web CrawlingOpen Source Developer Tools

Linkcrawler

Cross-platform persistent and distributed web crawler :link:

Open Source

About

linkcrawler is persistent because the queue is stored in a remote database that is automatically re-initialized if interrupted. linkcrawler is distributed because multiple instances of linkcrawler will work on the remotely stored queue, so you can start as many crawlers as you want on separate machines to speed along the process. linkcrawler is also fast because it is threaded and uses connection pools.

Open Source Health

Not enough history
Stars
113
Forks
9
License
MIT
Last commit
9 years ago
Go

Related Categories