A. Smyntyna / Code
ENRU
The Benchmarks

Local and hosted model benchmark

Test your models against each other.

Run the same work on every model. Compare empirical scores, keep human preference separate, and inspect the exact setup behind each result.

EmpiricalKnown answers, hidden tests, and working artifacts
TraceableModel, artifact, runtime, settings, seed, and hardware
ComparableThe same frozen task definitions for every model

Saved local run

One suite. Six models. Separate evidence.

RUN 20260720-181238SUITE V1.0LM STUDIO
01

ternary-bonsai-8b-mlx

LM StudioLocal run43 tasks
Bench Score86Empirical tasks only
Community ScoreNot ratedNo blind votes yet
Reasoning90%
Code77%
Communication86%
Writing checks100%
79.5 tok/s context100% task coverage
02

qwen2.5-coder-1.5b-instruct

LM StudioLocal run43 tasks
Bench Score67Empirical tasks only
Community ScoreNot ratedNo blind votes yet
Reasoning89%
Code80%
Communication36%
Writing checks50%
149.9 tok/s context100% task coverage
03

qwen3.5-2b-mlx

LM StudioLocal run43 tasks
Bench Score64Empirical tasks only
Community ScoreNot ratedNo blind votes yet
Reasoning88%
Code70%
Communication38%
Writing checks0%
136.6 tok/s context100% task coverage

These values come from a saved local run on July 20, 2026. The Bench Score uses 43 checkable tasks per model. Community Score is blank because this run has no blind votes. Speed is shown as context and does not affect rank.

How the result is built

A score with enough detail to rerun it.

01

Choose the models

Connect Ollama, LM Studio, or an OpenAI compatible endpoint. Different providers can run in the same benchmark.

02

Run frozen tasks

Each run records the task definitions, prompts, settings, seeds, attempts, diagnostics, and raw outputs.

03

Compare the evidence

Bench Score contains empirical results. Blind votes build a separate Community Score when subjective judgment is needed.

Built as a real desktop app

Rust engine. Tauri shell. Local first.

Runs work without an account. MongoDB adds shared observations, blind comparisons, model identity, and community leaderboards.