About
Corpus Crawler is designed for linguistic researchers and developers working on language-processing software. It crawls publicly accessible web pages in specific languages, removes HTML markup, and outputs the text into plaintext files. The tool adheres to the Robots Exclusion Standard and is optimized to minimize server load, making it suitable for gathering data for minority languages.
Open Source Health
- Stars
- 218
- Forks
- 53
- License
- Not stated
- Last commit
- 1 years ago
Related Categories
Vendor
Organize the world's information and make it universally accessible.
Quick Links
Open Source
More by Google
Related Products
Deep Research
An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine…
Shared categories
select.rs
A Rust library to extract useful data from HTML documents, suitable for web scraping.
Shared categories
Scrapely
A pure-python HTML screen-scraping library
Shared categories
Faster Than Requests
Faster requests on Python 3
Shared categories
MechanicalSoup
A Python library for automating interaction with websites.
Shared categories