SVGBench: A challenging LLM benchmark that tests knowledge, coding, physical reasoning capabilities of LLMs.
☆75Feb 12, 2026Updated 6 months ago
Alternatives and similar repositories for SVGBench
Users that are interested in SVGBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FamilyBench evaluation tool for testing the relational reasoning capabilities of Large Language Models (LLMs).☆46May 4, 2026Updated 3 months ago
- ☆19Jan 10, 2026Updated 7 months ago
- Academic papers and works related to SWE-bench and SWE-agents☆15Dec 8, 2025Updated 8 months ago
- Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a sm…☆73Apr 16, 2026Updated 4 months ago
- This is a training method to produce a split brain model☆14Mar 7, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A Field-Theoretic Approach to Unbounded Memory in Large Language Models☆20Apr 15, 2025Updated last year
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated 11 months ago
- Public Goods Game (PGG) Benchmark: Contribute & Punish is a multi-agent benchmark that tests cooperative and self-interested strategies a…☆41Apr 10, 2025Updated last year
- Benchmarking vision language vision on face tasks☆16Mar 30, 2025Updated last year
- Various LLM Benchmarks☆26Feb 20, 2026Updated 5 months ago
- Try out HallOumi, a state-of-the-art claim verification model in a simple UI!☆41Apr 2, 2025Updated last year
- Hallucinations (Confabulations) Document-Based Benchmark for RAG. Includes human-verified questions and answers.☆249Aug 7, 2025Updated last year
- A benchmark for conversational bargaining by language models. In each 20‑round match one LLM plays buyer, one plays seller, and both hold…☆46Jun 23, 2026Updated last month
- Benchmark that evaluates LLMs using 759 NYT Connections puzzles extended with extra trick words☆238Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- wavedrom to verilog converter☆17Sep 14, 2021Updated 4 years ago
- Attend - to what matters.☆17Feb 22, 2025Updated last year
- [ICLR 2024]: Is Self-Repair a Silver Bullet for Code Generation?☆15May 2, 2024Updated 2 years ago
- See how code changes affect your UI across viewports, themes, and languages☆15Jul 17, 2025Updated last year
- deep hermes, but decides how to respond based on its OWN decision, no need for system prompts.☆44Apr 1, 2025Updated last year
- ☆136May 2, 2025Updated last year
- Coview: The Right way to browse the web. Sends dom diffs for collaborative browsing. Blazing fast, at a fraction of bandwidth.☆19May 17, 2026Updated 3 months ago
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".☆25Jul 17, 2026Updated last month
- ☆15May 13, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official repository for MiniAppBench. Contains the complete pipeline and codebase for LLM-powered interactive HTML generation and agentic…☆24Mar 9, 2026Updated 5 months ago
- Official repository for "NoLiMa: Long-Context Evaluation Beyond Literal Matching"☆201Jul 17, 2025Updated last year
- [AAAI 2025] The official code of the paper "InverseCoder: Unleashing the Power of Instruction-Tuned Code LLMs with Inverse-Instruct"(http…☆16Jul 10, 2024Updated 2 years ago
- ☆19Dec 23, 2025Updated 7 months ago
- function for agents in OpenWebUI☆17Jun 15, 2025Updated last year
- SketchINR: A First Look into Sketches as Implicit Neural Representations [CVPR 2024]☆12Aug 19, 2024Updated last year
- Analyze Reddit posts☆32Jun 5, 2026Updated 2 months ago
- Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLM…☆90Dec 9, 2025Updated 8 months ago
- “A locally hosted, memory-aware AI microservice—designed for cultural continuity, decentralized intelligence, and ethical autonomy.”☆27May 1, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Validated, private, shareable knowledge-graph memory for AI — per-tenant, write-gated, PostgreSQL-authoritative, served over MCP.☆19Updated this week
- ☆17Jul 12, 2025Updated last year
- Expense tracking web app☆10May 6, 2026Updated 3 months ago
- rknn rust ffi binding☆19Feb 26, 2026Updated 5 months ago
- Recursive self-correcting intelligence framework☆17Nov 13, 2025Updated 9 months ago
- ☆16Oct 28, 2025Updated 9 months ago
- A series-symbol (S2) dual-modality data generation mechanism, enabling the unrestricted creation of high-quality time series data paired …☆84Updated this week