BPEmb
Pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE)
About
BPEmb is a collection of pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE) and trained on Wikipedia. Its intended use is as input for neural models in natural language processing.
Open Source Health
- Stars
- 1,224
- Forks
- 100
- License
- MIT
- Last commit
- 2 years ago
Related Categories
Vendor
Benjamin Heinzerling
Open-source project on GitHub: BPEmb
Quick Links
Open Source
Related Products
Turbovec
A vector index built on TurboQuant, written in Rust with Python bindings
Top category match
NGT
Nearest Neighbor Search with Neighborhood Graph and Tree for High-dimensional Data
Top category match
Chromem Go
Embeddable vector database for Go with Chroma-like interface and zero third-party dependencies. In-memory with optional persistence.
Top category match
Vecs
Postgres/pgvector Python Client
Top category match
Keras-TextClassification
Macropodus擅长自然语言处理,深度学习与tensorflow,LLM,等方面的知识,Macropodus关注机器翻译,神经网络,知识图谱,语音识别,排序算法,推荐算法,tensorflow,语言模型,nlp,机器学习,人工智能,pytorch,自然语言处理,数据挖掘,分类领域.
Top category match