A scalable data preprocessing framework built on PySpark for LLM training
☆26Dec 9, 2025Updated 8 months ago
Alternatives and similar repositories for Kaiyuan-Spark
Users that are interested in Kaiyuan-Spark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- llms related stuff , including code, docs☆13Feb 25, 2025Updated last year
- Using Feature Decomposition method to accelerate GNN inference☆13Sep 27, 2021Updated 4 years ago
- LLM training in simple, raw C/CUDA☆15Dec 5, 2024Updated last year
- This is a project created and completed by team BOOM(Beihang OO masters).This is a superscalar processor with a 13-stage out-of-order dua…☆18Sep 29, 2024Updated last year
- Modified Score-Entropy-Discrete-Diffusion to do a character level ml model and integrate with Oxen☆22Apr 26, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆27Jan 8, 2024Updated 2 years ago
- ☆14May 12, 2025Updated last year
- Awesome-RL-Reasoning☆17Aug 26, 2026Updated last week
- 这是一个从零学习CUDA课程☆13Nov 3, 2024Updated last year
- Survey on Knowledge Graph☆15Dec 5, 2018Updated 7 years ago
- Official code implementation for the ACL 2025 paper: 'CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis'☆32May 19, 2025Updated last year
- Summary of system papers/frameworks/codes/tools on training or serving large model☆57Dec 17, 2023Updated 2 years ago
- 保存有关DDPM直播的资料☆20Apr 7, 2024Updated 2 years ago
- run claude code in podman containers☆23Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- L1 Data, L1 Instruction and L2 Unified Cache Design☆16Aug 21, 2026Updated 2 weeks ago
- Verilog Code and Logisim simulation of a Weighted Round Robit Arbiter circuit using digital components☆20Mar 10, 2018Updated 8 years ago
- Improved AppImage of opencode which is able to work on any linux system [Maintainer=@Samueru-sama]☆16Updated this week
- ☆19May 11, 2024Updated 2 years ago
- ☆14Feb 18, 2024Updated 2 years ago
- ☆19Jan 31, 2025Updated last year
- [ICCV2025] Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning☆25Nov 13, 2025Updated 9 months ago
- A small repository demonstrating the use of Webdataset and Imagenet☆17Dec 19, 2023Updated 2 years ago
- ☆17Jul 18, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆51Apr 11, 2025Updated last year
- ☆61Dec 6, 2024Updated last year
- Awesome Chinese Corpus Datasets and Models.☆21Oct 28, 2019Updated 6 years ago
- 乱序双发处理器,在2024年计算机系统能力大赛CPU赛道(龙芯杯)获二等奖,全国第四☆24Aug 20, 2024Updated 2 years ago
- Various test models in WNNX format. It can view with `pip install wnetron && wnetron`☆12Jun 22, 2022Updated 4 years ago
- Official implementation of 'A Large-Scale Exploration of mu-Transfer' (CoRR 2024)☆31Jun 5, 2025Updated last year
- Notes and slides for Stanford CS231n 2021 & 2022 in English. I merged the contents together to get a better version. Assignments are not …☆27Aug 29, 2026Updated last week
- ☆19Nov 11, 2025Updated 9 months ago
- DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery☆23Sep 24, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆19May 9, 2025Updated last year
- A SysY Compiler written by Java for the Compiler Technology Course in BUAA☆20Sep 18, 2023Updated 2 years ago
- ☆41May 26, 2026Updated 3 months ago
- Official PyTorch implementation of QwT—“Quantization without Tears” (CVPR 2025): fast, accurate, and hassle-free post-training network qu…☆32Sep 30, 2025Updated 11 months ago
- kernelbench.com — GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.☆75Updated this week
- ☆26Apr 6, 2026Updated 5 months ago
- [ACL'25] UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench☆36Aug 12, 2025Updated last year