Curate better data for LLMs
☆1,069Mar 19, 2024Updated 2 years ago
Alternatives and similar repositories for lilac
Users that are interested in lilac are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.☆3,375Updated this week
- Go ahead and axolotl questions☆12,537Updated this week
- Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verifi…☆3,409Updated this week
- Easily use and train state of the art late-interaction retrieval methods (ColBERT) in any RAG pipeline. Designed for modularity and ease-…☆3,964May 17, 2025Updated last year
- Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets☆5,141Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Generate textbook-quality synthetic LLM pretraining data☆507Oct 19, 2023Updated 2 years ago
- Tools for merging pretrained large language models.☆7,390Sep 12, 2026Updated 3 weeks ago
- Structured Outputs☆15,908Sep 21, 2026Updated 2 weeks ago
- structured outputs for llms☆13,987Updated this week
- data cleaning and curation for unstructured text☆327Aug 6, 2024Updated 2 years ago
- A lightweight, low-dependency, unified API to use all common reranking and cross-encoder models.☆1,639Dec 20, 2025Updated 9 months ago
- Generate Synthetic Data Using OpenAI, MistralAI or AnthropicAI☆221Apr 29, 2024Updated 2 years ago
- Robust recipes to align language models with human and AI preferences☆5,691Sep 23, 2026Updated 2 weeks ago
- DSPy: The framework for programming—not prompting—language models☆38,557Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Customizable implementation of the self-instruct paper.☆1,052Mar 7, 2024Updated 2 years ago
- an ambient intelligence library☆6,202Sep 23, 2026Updated 2 weeks ago
- Adding guardrails to large language models.☆7,497Updated this week
- Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs☆3,836May 28, 2026Updated 4 months ago
- Automatically evaluate your LLMs in Google Colab☆698May 7, 2024Updated 2 years ago
- A guidance language for controlling large language models.☆21,787May 21, 2026Updated 4 months ago
- A lightweight library for generating synthetic instruction tuning datasets for your data without GPT.☆828Jul 15, 2025Updated last year
- Chat language model that can use tools and interpret the results☆1,595Jun 30, 2026Updated 3 months ago
- Large Language Model Text Generation Inference☆10,880Mar 21, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build cu…☆10,685Updated this week
- Create Custom LLMs☆1,873Jun 27, 2026Updated 3 months ago
- Easily embed, cluster and semantically label text datasets☆609Mar 28, 2024Updated 2 years ago
- Synthetic data curation for post-training and structured data extraction☆1,742Sep 29, 2026Updated last week
- Simple speculative decoding technique, integrated in vLLM and transformers☆616Aug 23, 2024Updated 2 years ago
- Training LLMs with QLoRA + FSDP☆1,553Nov 9, 2024Updated last year
- DataDreamer: Prompt. Generate Synthetic Data. Train & Align Models. 🤖💤☆1,122Feb 2, 2025Updated last year
- 20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.☆13,695Updated this week
- Data and tools for generating and inspecting OLMo pre-training data.☆1,553Aug 24, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A Bulletproof Way to Generate Structured JSON from Language Models☆4,937Feb 24, 2024Updated 2 years ago
- clean up your LLM datasets☆113May 30, 2023Updated 3 years ago
- Use late-interaction multi-modal models such as ColPali in just a few lines of code.☆851Jan 28, 2025Updated last year
- LLMs build upon Evol Insturct: WizardLM, WizardCoder, WizardMath☆9,479Jun 7, 2025Updated last year
- Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean…☆15,549Updated this week
- ColBERT: state-of-the-art neural search (SIGIR'20, TACL'21, NeurIPS'21, NAACL'22, CIKM'22, ACL'23, EMNLP'23)☆3,945Oct 14, 2025Updated 11 months ago
- Stanford NLP Python library for Representation Finetuning (ReFT)☆1,588Mar 5, 2026Updated 7 months ago