tinybirdco / llm-benchmarkView on GitHub
We assessed the ability of popular LLMs to generate accurate and efficient SQL from natural language prompts. Using a 200 million record dataset from the GH Archive uploaded to Tinybird, we asked the LLMs to generate SQL based on 50 prompts.
☆87Oct 7, 2026Updated this week

Alternatives and similar repositories for llm-benchmark

Users that are interested in llm-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?