shibing624/similarities

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/shibing624/similarities)

shibing624 / similarities

Similarities: a toolkit for similarity calculation and semantic search. 相似度计算、匹配搜索工具包，支持亿级数据文搜文、文搜图、图搜图，python3开发，开箱即用。

☆903

Alternatives and similar repositories for similarities

Users that are interested in similarities are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

shibing624 / text2vec
View on GitHub
text2vec, text to vector. 文本向量表征工具，把文本转化为向量矩阵，实现了Word2Vec、RankBM25、Sentence-BERT、CoSENT等文本表征、文本相似度计算模型，开箱即用。
☆4,974Feb 14, 2026Updated 5 months ago
shibing624 / nerpy
View on GitHub
🌈 NERpy: Implementation of Named Entity Recognition using Python. 命名实体识别工具，支持BertSoftmax、BertSpan等模型，开箱即用。
☆118Feb 19, 2024Updated 2 years ago
shibing624 / pytextclassifier
View on GitHub
pytextclassifier is a toolkit for text classification. 文本分类，LR，Xgboost，TextCNN，FastText，TextRNN，BERT等分类模型实现，开箱即用。
☆524Sep 25, 2024Updated last year
shibing624 / textgen
View on GitHub
TextGen: Implementation of Text Generation models, include LLaMA, BLOOM, GPT2, BART, T5, SongNet and so on. 文本生成模型，实现了包括LLaMA，ChatGLM，BLO…
☆982Sep 14, 2024Updated last year
shibing624 / pycorrector
View on GitHub
pycorrector is a toolkit for text error correction. 文本纠错，实现了Kenlm，T5，MacBERT，ChatGLM3，Qwen2.5等模型应用在纠错场景，开箱即用。
☆6,495Updated this week
Wordpress hosting with auto-scaling - Free Trial Offer • Ad
Fully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
shibing624 / pke_zh
View on GitHub
pke_zh, python keyphrase extraction for chinese(zh). 中文关键词或关键句提取工具，实现了KeyBert、PositionRank、TopicRank、TextRank等算法，开箱即用。
☆216Mar 27, 2024Updated 2 years ago
OFA-Sys / Chinese-CLIP
View on GitHub
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
☆5,979Mar 31, 2026Updated 3 months ago
shibing624 / ChatPDF
View on GitHub
RAG for Local LLM, chat with PDF/doc/txt files, ChatPDF. 纯原生实现RAG功能，基于本地LLM、embedding模型、reranker模型实现，支持GraphRAG，无须安装任何第三方agent库。
☆854Apr 2, 2025Updated last year
shibing624 / SearchGPT
View on GitHub
SearchGPT: Building a quick conversation-based search engine with LLMs.
☆45Jan 5, 2025Updated last year
FlagOpen / FlagEmbedding
View on GitHub
Retrieval and Retrieval-augmented LLMs
☆11,990Apr 22, 2026Updated 3 months ago
shibing624 / ChatPilot
View on GitHub
ChatPilot: Chat Agent Web UI，实现Chat对话前端，支持Google搜索、文件网址对话（RAG）、代码解释器功能，复现了Kimi Chat(文件，拖进来；网址，发出来)。
☆600Jan 27, 2026Updated 6 months ago
dongrixinyu / JioNLP
View on GitHub
中文 NLP 预处理、解析工具包，准确、高效、易用 A Chinese NLP Preprocessing & Parsing Package www.jionlp.com
☆3,855Jun 5, 2026Updated last month
shibing624 / MedicalGPT
View on GitHub
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型，实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。
☆5,667Jun 3, 2026Updated last month
hiyouga / ChatGLM-Efficient-Tuning
View on GitHub
Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调
☆3,720Oct 12, 2023Updated 2 years ago
Serverless GPU API endpoints on Runpod - Get Bonus Credits • Ad
Skip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
HarderThenHarder / transformers_tasks
View on GitHub
⭐️ NLP Algorithms with transformers lib. Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SF…
☆2,421Sep 29, 2023Updated 2 years ago
mymusise / ChatGLM-Tuning
View on GitHub
基于ChatGLM-6B + LoRA的Fintune方案
☆3,746Nov 25, 2023Updated 2 years ago
bojone / CoSENT
View on GitHub
比Sentence-BERT更有效的句向量方案
☆373Nov 9, 2022Updated 3 years ago
yangjianxin1 / Firefly
View on GitHub
Firefly: 大模型训练工具，支持训练Qwen2.5、Qwen2、Yi1.5、Phi-3、Llama3、Gemma、MiniCPM、Yi、Deepseek、Orion、Xverse、Mixtral-8x7B、Zephyr、Mistral、Baichuan2、Llma2、…
☆6,649Oct 24, 2024Updated last year
yuanzhoulvpi2017 / quick_sentence_transformers
View on GitHub
sentence-transformers to onnx 让sbert模型推理效率更快
☆166Mar 11, 2022Updated 4 years ago
shibing624 / github-hot
View on GitHub
Tracking the hot Github repos and update daily 每天自动追踪Github热门项目
☆52Updated this week
chatchat-space / Langchain-Chatchat
View on GitHub
Langchain-Chatchat（原Langchain-ChatGLM）基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain…
☆38,482Nov 10, 2025Updated 8 months ago
ymcui / Chinese-LLaMA-Alpaca
View on GitHub
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
☆18,945Apr 19, 2026Updated 3 months ago
wangyuxinwhy / uniem
View on GitHub
unified embedding model
☆876Sep 1, 2023Updated 2 years ago
GPU virtual machines on DigitalOcean Gradient AI • Ad
Get to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
PaddlePaddle / PaddleNLP
View on GitHub
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
☆12,963May 23, 2026Updated 2 months ago
ymcui / Chinese-BERT-wwm
View on GitHub
Pre-Training with Whole Word Masking for Chinese BERT（中文BERT-wwm系列模型）
☆10,224Apr 19, 2026Updated 3 months ago
yanqiangmiffy / InstructGLM
View on GitHub
ChatGLM-6B 指令学习|指令数据|Instruct
☆651Apr 10, 2023Updated 3 years ago
baichuan-inc / Baichuan-7B
View on GitHub
A large-scale 7B pretraining language model developed by BaiChuan-Inc.
☆5,651Jul 18, 2024Updated 2 years ago
zai-org / VisualGLM-6B
View on GitHub
Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型
☆4,156Aug 23, 2024Updated last year
TigerResearch / TigerBot
View on GitHub
TigerBot: A multi-language multi-task LLM
☆2,259Dec 28, 2024Updated last year
zai-org / ChatGLM2-6B
View on GitHub
ChatGLM2-6B: An Open Bilingual Chat LLM | 开源双语对话语言模型
☆15,539Jun 27, 2024Updated 2 years ago
LianjiaTech / BELLE
View on GitHub
BELLE: Be Everyone's Large Language model Engine（开源中文对话大模型）
☆8,279Oct 16, 2024Updated last year
IDEA-CCNL / Fengshenbang-LM
View on GitHub
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系，成为中文AIGC和认知智能的基础设施。
☆4,125Jun 8, 2026Updated last month
Wordpress hosting with auto-scaling - Free Trial Offer • Ad
Fully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
shibing624 / relext
View on GitHub
RelExt: A Tool for Relation Extraction from Text. 文本实体关系抽取工具。
☆50Jun 9, 2022Updated 4 years ago
zai-org / ChatGLM-6B
View on GitHub
ChatGLM-6B: An Open Bilingual Dialogue Language Model | 开源双语对话语言模型
☆41,014Jun 27, 2024Updated 2 years ago
CLUEbenchmark / CLUEDatasetSearch
View on GitHub
搜索所有中文NLP数据集，附常用英文NLP数据集
☆4,459Nov 21, 2022Updated 3 years ago
liucongg / ChatGLM-Finetuning
View on GitHub
基于ChatGLM-6B、ChatGLM2-6B、ChatGLM3-6B模型，进行下游具体任务微调，涉及Freeze、Lora、P-tuning、全参微调等
☆2,773Dec 12, 2023Updated 2 years ago
lonePatient / awesome-pretrained-chinese-nlp-models
View on GitHub
Awesome Pretrained Chinese NLP Models，高质量中文预训练模型&大模型&多模态模型&大语言模型集合
☆5,575Jun 19, 2026Updated last month
yuanzhoulvpi2017 / zero_nlp
View on GitHub
中文nlp解决方案(大模型、数据、模型、训练、推理)
☆3,833Aug 5, 2025Updated 11 months ago
esbatmop / MNBVC
View on GitHub
MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化，也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志…
☆4,244Jul 13, 2026Updated 2 weeks ago