Web CrawlingOpen Source Developer Tools

Corpuscrawler

Crawler for linguistic corpora

Open Source

About

Corpus Crawler is designed for linguistic researchers and developers working on language-processing software. It crawls publicly accessible web pages in specific languages, removes HTML markup, and outputs the text into plaintext files. The tool adheres to the Robots Exclusion Standard and is optimized to minimize server load, making it suitable for gathering data for minority languages.

Open Source Health

Not enough history
Stars
218
Forks
53
License
Not stated
Last commit
1 years ago
Python

Related Categories