Benchmarks¶
Performance of Shifty’s validate pipeline (inference + validation) across
real building models, tracked over each release.
Wall-clock validate time is dominated by a fixed setup cost — preparing
the shapes graph — that is paid no matter how small the data graph is. A
16-triple Brick model still takes ~3.7 s. So the raw number mostly reflects
model size, and a single average hides which part of the engine a release
improved.
Each bar below is the measured time to validate one model, averaged over the corpus, split into the three things that time is spent on.
Show the numbers
Per-model results¶
These tables show exact validate times for the two most recent releases.
The % column flags regressions or improvements per model, and the
geomean row summarises the overall change between the two versions.
Regenerating benchmark data¶
./benchmark/run_history.sh
uv run benchmark/process_results.py
cd docs && make html
run_history.sh benchmarks every release tag, then the current checkout as a
final HEAD entry. Tagged results are reused when they already exist, so a
repeat run only re-measures HEAD — but the first run measures every tag and
takes hours. To refresh just the HEAD entry after a code change:
BENCH_ONLY_HEAD=1 ./benchmark/run_history.sh
cd docs && make html
Because HEAD is built from the working tree, uncommitted changes are
included; the run logs the commit it started from and whether the tree was
dirty. BENCH_HEAD=0 restricts a run to tags only, and BENCH_ITERS
controls how many samples each measurement takes (default 3, median reported).
Shapes and models always come from the current checkout, so changing a fixture
invalidates every previously recorded result — delete
benchmark/results/v*/ and re-run the full history when that happens.