SVGBench: A challenging LLM benchmark that tests knowledge, coding, physical reasoning capabilities of LLMs.
☆73Feb 12, 2026Updated 5 months ago
Alternatives and similar repositories for SVGBench
Users that are interested in SVGBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FamilyBench evaluation tool for testing the relational reasoning capabilities of Large Language Models (LLMs).☆47May 4, 2026Updated 2 months ago
- ☆19Jan 10, 2026Updated 6 months ago
- Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a sm…☆72Apr 16, 2026Updated 3 months ago
- This is a training method to produce a split brain model☆14Mar 7, 2025Updated last year
- A Field-Theoretic Approach to Unbounded Memory in Large Language Models☆20Apr 15, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Public Goods Game (PGG) Benchmark: Contribute & Punish is a multi-agent benchmark that tests cooperative and self-interested strategies a…☆41Apr 10, 2025Updated last year
- Benchmarking vision language vision on face tasks☆16Mar 30, 2025Updated last year
- Try out HallOumi, a state-of-the-art claim verification model in a simple UI!☆41Apr 2, 2025Updated last year
- ☆16Feb 1, 2025Updated last year
- A benchmark for conversational bargaining by language models. In each 20‑round match one LLM plays buyer, one plays seller, and both hold…☆44Jun 23, 2026Updated last month
- Benchmark that evaluates LLMs using 759 NYT Connections puzzles extended with extra trick words☆230Updated this week
- Attend - to what matters.☆17Feb 22, 2025Updated last year
- ☆15May 17, 2023Updated 3 years ago
- See how code changes affect your UI across viewports, themes, and languages☆15Jul 17, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Learning problem-solving, logic/set, math, physics, economics through functional programming using Haskell☆19Oct 16, 2015Updated 10 years ago
- ☆136May 2, 2025Updated last year
- Line-numbered patch format. Non-sequential, llm and stream-friendly☆15Nov 7, 2024Updated last year
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆15Jun 28, 2025Updated last year
- ☆15May 13, 2025Updated last year
- Official repository for "NoLiMa: Long-Context Evaluation Beyond Literal Matching"☆202Jul 17, 2025Updated last year
- function for agents in OpenWebUI☆17Jun 15, 2025Updated last year
- Analyze Reddit posts☆32Jun 5, 2026Updated last month
- The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation☆41May 4, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLM…☆89Dec 9, 2025Updated 7 months ago
- “A locally hosted, memory-aware AI microservice—designed for cultural continuity, decentralized intelligence, and ethical autonomy.”☆27May 1, 2025Updated last year
- Validated, private, shareable knowledge-graph memory for AI — per-tenant, write-gated, PostgreSQL-authoritative, served over MCP.☆18Jul 19, 2026Updated last week
- Recursive self-correcting intelligence framework☆17Nov 13, 2025Updated 8 months ago
- ☆16Oct 28, 2025Updated 9 months ago
- "Efficient Neural Theorem Proving via Fine-grained Proof Structure Analysis" (ICML 2025) official implementation.☆16Jun 8, 2025Updated last year
- ☆212Dec 20, 2024Updated last year
- [SIGGRAGH'25] Official repository of VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control☆34Dec 5, 2025Updated 7 months ago
- Benchmark evaluating LLMs on their ability to create and resist disinformation. Includes comprehensive testing across major models (Claud…☆33Mar 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆25Jan 28, 2026Updated 6 months ago
- ☆17Oct 27, 2024Updated last year
- ☆12Jun 1, 2026Updated last month
- Remote controlling your keyboard since 2016.☆10Jun 16, 2016Updated 10 years ago
- An Automatic Theorem Prover for Hilbert System, generating nearly-minimal proofs.☆14Jan 21, 2025Updated last year
- Developed a high-performance triangle-to-quad conversion operator, formulated as a maximum-weight matching problem on the triangle adjace…☆17May 28, 2026Updated 2 months ago
- Zero-config local LLM optimization for Ollama, LM Studio, and Apple Silicon MLX. Reduces TTFT by 40%, wall time for local agents by 46%, …☆32Jun 30, 2026Updated 3 weeks ago