BIM-Guard Evaluation

BIM-Guard Evaluation

Evaluation harnesses and empirical validation for BIM-Guard, an automated code-compliance platform for OpenBIM (IFC, BCF, IDS). This site presents what the evaluation repository can and cannot back up.

Read this first. Several numbers that appeared in earlier drafts (inter-annotator κ, the LLM-judge threshold sweep, cross-jurisdiction F1 = 100%) come from scripts that simulate their inputs. They are not evidence about BIM-Guard, annotators or the judge, and this site labels them accordingly. Only the results marked reproduced should be cited.

What is currently supported

22 / 22
Architecture engine benchmark cases correct (13 TP, 9 TN, 0 FP, 0 FN)
60 / 60
NLP annotation test cases passing, fully reproducible
Wilson 95%
Confidence intervals on every real proportion; accuracy CI is [85.1%, 100%]

The 22-case benchmark is small and procedurally generated, so the interval matters more than the point estimate: it shows what 22 cases can and cannot establish. See Results.

How this site is organised

Verification status legend

Status Meaning
reproduced Re-run from source inputs on the current codebase; matches the claim.
archived-artifact Original output is committed and hash-pinned, but not re-executed.
single-run Real result from one execution, no repeats, no variance measured.
not-verified Asserted somewhere with no corresponding artifact.
simulated Inputs generated from gold data or hard-coded values; not evidence.
retired-domain Real at the time, but the measured capability has since been removed from the product.