SciBERT
A BERT model for scientific text.
About
Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-quality, large-scale labeled scientific data. SciBERT leverages unsupervised pretraining on a large multi-domain corpus of scientific publications to improve performance on downstream scientific NLP tasks. We evaluate on a suite of tasks including sequence tagging, sentence classification and dependency parsing, with datasets from a variety of scientific domains. We demonstrate statistically significant improvements over BERT and achieve new state-of-the-art results on several of these tasks. The code and pretrained models are available at https://github.com/allenai/scibert/.
Open Source Health
- Stars
- 1,714
- Forks
- 232
- License
- Apache-2.0
- Last commit
- 5 years ago
Related Categories
Vendor
Allenai
Publisher of Dont Stop Pretraining, Objaverse Xl and Procthor
Quick Links
Open Source
More by Allenai
Related Products
Dont Stop Pretraining
Code associated with the Don't Stop Pretraining ACL 2020 paper
More from this vendor
Textrank
TextRank implementation for Python 3.
Top category match
KoBERT
Korean BERT pre-trained cased (KoBERT)
Top category match
TAADpapers
Must-read Papers on Textual Adversarial Attack and Defense
Top category match
AutoPhrase
AutoPhrase: Automated Phrase Mining from Massive Text Corpora
Top category match