Local and hosted model benchmark
Test your models against each other.
Run the same work on every model. Compare empirical scores, keep human preference separate, and inspect the exact setup behind each result.
Saved local run
One suite. Six models. Separate evidence.
qwen2.5-coder-1.5b-instruct
qwen3.5-2b-mlx
These values come from a saved local run on July 20, 2026. The Bench Score uses 43 checkable tasks per model. Community Score is blank because this run has no blind votes. Speed is shown as context and does not affect rank.
How the result is built
A score with enough detail to rerun it.
Choose the models
Connect Ollama, LM Studio, or an OpenAI compatible endpoint. Different providers can run in the same benchmark.
Run frozen tasks
Each run records the task definitions, prompts, settings, seeds, attempts, diagnostics, and raw outputs.
Compare the evidence
Bench Score contains empirical results. Blind votes build a separate Community Score when subjective judgment is needed.
Built as a real desktop app
Rust engine. Tauri shell. Local first.
Runs work without an account. MongoDB adds shared observations, blind comparisons, model identity, and community leaderboards.