Olm Datasets
Pipeline for pulling and processing online language model pretraining data from the web
About
This repo enables you to pull a large and up-to-date text corpus from the web. It uses state-of-the-art processing methods to produce a clean text dataset that you can immediately use to pretrain a large language model, like BERT, GPT, or BLOOM. The main use-case for this repo is the Online Language Modelling Project, where we want to keep a language model up-to-date by pretraining it on the latest Common Crawl and Wikipedia dumps every month or so. You can see the models for the OLM project here: https://huggingface.co/olm. They actually get better performance than their original static counterparts.
Open Source Health
- Stars
- 179
- Forks
- 21
- License
- Apache-2.0
- Last commit
- 3 years ago
Resources & Links
Related Categories
Vendor
Hugging Face
Publisher of Speech To Speech, Diffusers and Optimum
Quick Links
Open Source
More by Hugging Face
Related Products
PyTables
A Python package to manage extremely large amounts of data
Top category match
Wiktextract
Wiktionary dump file parser and multilingual data extractor
Top category match
Yahoo Finance
Python module to get stock data from Yahoo! Finance
Top category match
Avsc
Avro for JavaScript :zap:
Top category match
Picard
A set of command line tools (in Java) for manipulating high-throughput sequencing (HTS) data and formats such as SAM/BAM/CRAM and VCF.
Top category match