shibing624/text2vec

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/shibing624/text2vec)

shibing624 / text2vec

text2vec, text to vector. 文本向量表征工具，把文本转化为向量矩阵，实现了Word2Vec、RankBM25、Sentence-BERT、CoSENT等文本表征、文本相似度计算模型，开箱即用。

☆4,974

Alternatives and similar repositories for text2vec

Users that are interested in text2vec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

shibing624 / similarities
View on GitHub
Similarities: a toolkit for similarity calculation and semantic search. 相似度计算、匹配搜索工具包，支持亿级数据文搜文、文搜图、图搜图，python3开发，开箱即用。
☆903Mar 5, 2026Updated 4 months ago
GanymedeNil / document.ai
View on GitHub
基于向量数据库与GPT3.5的通用本地知识库方案(A universal local knowledge base solution based on vector database and GPT3.5)
☆3,671May 12, 2023Updated 3 years ago
chatchat-space / Langchain-Chatchat
View on GitHub
Langchain-Chatchat（原Langchain-ChatGLM）基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain…
☆38,457Nov 10, 2025Updated 8 months ago
FlagOpen / FlagEmbedding
View on GitHub
Retrieval and Retrieval-augmented LLMs
☆11,968Apr 22, 2026Updated 3 months ago
ymcui / Chinese-LLaMA-Alpaca
View on GitHub
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
☆18,944Apr 19, 2026Updated 3 months ago
Proton VPN Special Offer - Get 70% off • Ad
Special partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
LianjiaTech / BELLE
View on GitHub
BELLE: Be Everyone's Large Language model Engine（开源中文对话大模型）
☆8,274Oct 16, 2024Updated last year
wangyuxinwhy / uniem
View on GitHub
unified embedding model
☆876Sep 1, 2023Updated 2 years ago
zai-org / ChatGLM-6B
View on GitHub
ChatGLM-6B: An Open Bilingual Dialogue Language Model | 开源双语对话语言模型
☆41,022Jun 27, 2024Updated 2 years ago
zai-org / ChatGLM2-6B
View on GitHub
ChatGLM2-6B: An Open Bilingual Chat LLM | 开源双语对话语言模型
☆15,542Jun 27, 2024Updated 2 years ago
shibing624 / textgen
View on GitHub
TextGen: Implementation of Text Generation models, include LLaMA, BLOOM, GPT2, BART, T5, SongNet and so on. 文本生成模型，实现了包括LLaMA，ChatGLM，BLO…
☆981Sep 14, 2024Updated last year
shibing624 / pycorrector
View on GitHub
pycorrector is a toolkit for text error correction. 文本纠错，实现了Kenlm，T5，MacBERT，ChatGLM3，Qwen2.5等模型应用在纠错场景，开箱即用。
☆6,494Jun 4, 2026Updated last month
hiyouga / ChatGLM-Efficient-Tuning
View on GitHub
Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调
☆3,719Oct 12, 2023Updated 2 years ago
mymusise / ChatGLM-Tuning
View on GitHub
基于ChatGLM-6B + LoRA的Fintune方案
☆3,745Nov 25, 2023Updated 2 years ago
IDEA-CCNL / Fengshenbang-LM
View on GitHub
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系，成为中文AIGC和认知智能的基础设施。
☆4,125Jun 8, 2026Updated last month
AI Agents on DigitalOcean Gradient AI Platform • Ad
Build production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
yangjianxin1 / Firefly
View on GitHub
Firefly: 大模型训练工具，支持训练Qwen2.5、Qwen2、Yi1.5、Phi-3、Llama3、Gemma、MiniCPM、Yi、Deepseek、Orion、Xverse、Mixtral-8x7B、Zephyr、Mistral、Baichuan2、Llma2、…
☆6,646Oct 24, 2024Updated last year
yanqiangmiffy / Chinese-LangChain
View on GitHub
中文langchain项目|小必应，Q.Talk，强聊，QiangTalk
☆2,829Jun 20, 2023Updated 3 years ago
ymcui / Chinese-BERT-wwm
View on GitHub
Pre-Training with Whole Word Masking for Chinese BERT（中文BERT-wwm系列模型）
☆10,223Apr 19, 2026Updated 3 months ago
baichuan-inc / Baichuan-7B
View on GitHub
A large-scale 7B pretraining language model developed by BaiChuan-Inc.
☆5,652Jul 18, 2024Updated 2 years ago
dongrixinyu / JioNLP
View on GitHub
中文 NLP 预处理、解析工具包，准确、高效、易用 A Chinese NLP Preprocessing & Parsing Package www.jionlp.com
☆3,855Jun 5, 2026Updated last month
Embedding / Chinese-Word-Vectors
View on GitHub
100+ Chinese Word Vectors 上百种预训练中文词向量
☆12,230Oct 30, 2023Updated 2 years ago
PaddlePaddle / PaddleNLP
View on GitHub
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
☆12,958May 23, 2026Updated 2 months ago
HarderThenHarder / transformers_tasks
View on GitHub
⭐️ NLP Algorithms with transformers lib. Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SF…
☆2,420Sep 29, 2023Updated 2 years ago
Facico / Chinese-Vicuna
View on GitHub
Chinese-Vicuna: A Chinese Instruction-following LLaMA-based Model —— 一个中文低资源的llama+lora方案，结构参考alpaca
☆4,119Apr 18, 2025Updated last year
Deploy to Railway using AI coding agents - Free Credits Offer • Ad
Use Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
bojone / CoSENT
View on GitHub
比Sentence-BERT更有效的句向量方案
☆373Nov 9, 2022Updated 3 years ago
lm-sys / FastChat
View on GitHub
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
☆39,497May 1, 2026Updated 2 months ago
yuanzhoulvpi2017 / zero_nlp
View on GitHub
中文nlp解决方案(大模型、数据、模型、训练、推理)
☆3,831Aug 5, 2025Updated 11 months ago
ymcui / Chinese-LLaMA-Alpaca-2
View on GitHub
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
☆7,132Apr 19, 2026Updated 3 months ago
QwenLM / Qwen
View on GitHub
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
☆21,467Mar 5, 2026Updated 4 months ago
CVI-SZU / Linly
View on GitHub
Chinese-LLaMA 1&2、Chinese-Falcon 基础模型；ChatFlow中文对话模型；中文OpenLLaMA模型；NLP预训练/指令微调数据集
☆3,046Apr 14, 2024Updated 2 years ago
liucongg / ChatGLM-Finetuning
View on GitHub
基于ChatGLM-6B、ChatGLM2-6B、ChatGLM3-6B模型，进行下游具体任务微调，涉及Freeze、Lora、P-tuning、全参微调等
☆2,774Dec 12, 2023Updated 2 years ago
hiyouga / LlamaFactory
View on GitHub
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
☆73,452Updated this week
brightmart / nlp_chinese_corpus
View on GitHub
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
☆9,906Feb 6, 2026Updated 5 months ago
Wordpress hosting with auto-scaling - Free Trial Offer • Ad
Fully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
fighting41love / funNLP
View on GitHub
中英文敏感词、语言检测、中外手机/电话归属地/运营商查询、名字推断性别、手机号抽取、身份证抽取、邮箱抽取、中日文人名库、中文缩写库、拆字词典、词汇情感值、停用词、反动词表、暴恐词表、繁简体转换、英文模拟中文发音、汪峰歌词生成器、职业名称词库、同义词库、反义词库、否定词库、汽…
☆81,963May 10, 2024Updated 2 years ago
LlamaChinese / Llama-Chinese
View on GitHub
Llama中文社区，实时汇总最新Llama学习资料，构建最好的中文Llama大模型开源生态，完全开源可商用
☆14,745Apr 6, 2025Updated last year
zai-org / ChatGLM3
View on GitHub
ChatGLM3 series: Open Bilingual Chat LLMs | 开源双语对话语言模型
☆13,669Jan 13, 2025Updated last year
lonePatient / awesome-pretrained-chinese-nlp-models
View on GitHub
Awesome Pretrained Chinese NLP Models，高质量中文预训练模型&大模型&多模态模型&大语言模型集合
☆5,571Jun 19, 2026Updated last month
wenda-LLM / wenda
View on GitHub
闻达：一个LLM调用平台。目标为针对特定环境的高效内容生成，同时考虑个人和中小企业的计算资源局限性，以及知识安全和私密性问题
☆6,166Jan 23, 2025Updated last year
eosphoros-ai / DB-GPT
View on GitHub
open-source agentic AI data assistant for the next generation of AI + Data products.
☆19,539Updated this week
AiHubCN / Awesome-Chinese-LLM
View on GitHub
整理开源的中文大语言模型，以规模较小、可私有化部署、训练成本较低的模型为主，包括底座模型，垂直领域微调及应用，数据集与教程等。
☆22,692May 10, 2026Updated 2 months ago