BIM-Guard Evaluation

Rendered from research/CLAIMS.md.

Claims-to-Evidence Ledger

This is a claim-by-claim accounting of every headline number this repository (or the thesis it backs) asserts, what evidence backs it, where that evidence lives, and how confidently it can be verified today. It exists because a self-audit (research/BIMGUARD AI — Dual Repository Pre-Submission Audit.md) found several of these numbers unverifiable from the repository as it stood on 2026-09-21–24, and the correct response to that finding is to make every claim traceable, not to quietly drop the ones that were briefly hard to find.

Verification status legend:

Last updated: 2026-10-03 (§5, §6, §8 downgraded to simulated; the real extraction evidence is the e2e run in eval/results/e2e/).

Active claims (what the thesis's present-tense validation narrative should rest on)

bim-guard permanently retired its Piping/Corrosion and Seismic domains on 2026-09-21 (see the retired-domain note above). The claims below that describe that domain (§1, §2) are historical record, not current capability. The claims currently verifiable against the live product are §4 (NLP annotation, 60/60, fully reproducible), §7 (Architectural compliance engines benchmark, 22/22, fully reproducible with Wilson score 95% CIs), and, with the caveats stated in §3 and §5, the rule-extraction and LLM-judge harnesses. Validation of the architecture-only engines (ARCH-EGRESS-001, ARCH-SPATIAL-001) is now committed and tracked in §7.


1. 38-model IFC validation sweep — 223,516 clashes

Claim Automated validation sweep processed 38 real-world IFC models, generated 49,736 halo volumes, and found 223,516 clashes (211,581 minor / 6,699 major / 5,236 critical), across bim-guard's since-retired GC-001/CC-001/MC-001/MM-001/XM-001 corrosion engines.
Producing script eval/test_all_38_models.py — removed from this repository on 2026-09-25 (called bim-guard modules deleted in the domain retirement; see below).
Run date 2026-09-18
Artifact research/archive/retired_corrosion_piping_seismic_domain/appendix_b/run_20260918/validation_sweep_summary.json
SHA-256 6303c1cfa6a08aef6a83d92d367bde5bc47a5c9d137f9db80ece356273deec0b
Status retired-domain (was archived-artifact 2026-09-24 — see history below)

History: these artifacts (and the 7 derived tables + 4 figures) were committed after the 2026-09-18 run, then deleted from the working tree by commit 750ae73 ("Remove obsolete research tables and test results", 2026-09-21) as part of an unrelated cleanup pass. They were never removed from git history. The Dual-Repository Pre-Submission Audit was written against the working tree at that point and correctly reported the claim as unverifiable from the repo contents. That finding was superseded on 2026-09-24 by restoring the artifacts from 750ae73^ — see the restoration commit and .../appendix_b/run_20260918/PROVENANCE.md.

Superseded again, 2026-09-25: while extending that restoration into the repository rebuild, testing revealed that bim-guard permanently removed the entire Piping/Corrosion and Seismic domain this sweep measured, on 2026-09-21 — the same week as both deletions above (bim-guard commit 3157c1b, "Remove piping and seismic analysis domains; app is architecture-only"). This is not a bug and not something to repair: it is a deliberate, documented product pivot. The restored artifacts and the eval script that produced them have accordingly been moved to research/archive/retired_corrosion_piping_seismic_domain/, and this claim's status changes from archived-artifact (implying pending re-verification) to retired-domain (there is nothing left to re-verify against — the capability is gone by design). The number itself is unchanged and still real: 223,516 clashes were genuinely found, on 2026-09-18, by code that existed at that time.

Caveats visible in this same data, disclosed here rather than left for someone else to find:


2. 25-element synthetic validation dataset

Claim Corrosion-engine validation also ran against a 25-element synthetic dataset (pipe segments, fittings, and fasteners with corrosion-relevant materials — e.g. SS_316_passive, Copper, Galvanized_steel).
Producing code generate_synthetic_elements(n=25), app/modules/ifc_reader/ifc_parser.py:335 (bim-guard)
Status retired-domain. Re-checked 2026-09-25: the function still exists and is still called (app/modules/orchestrator.py:472, as a no-IFC-uploaded demo-data fallback), but no corrosion engine remains in bim-guard to validate this data against — the claim as originally stated ("corrosion-engine validation ran against...") describes a validation activity that can no longer happen. The function's continued existence for an unrelated generic-demo purpose does not rescue the original claim.

3. Gold-standard rule set — 29 rules, Code Part 9.8 (stairs)

Claim 29 hand-annotated ground-truth rules (plus 5 explicitly excluded clauses) for OBC Part 9, §9.8.2–9.8.4.7 (stairs), used to score rule-extraction recall.
Artifact eval/eval_gold_code_9_8_stairs.py — GOLD_RULES (29 entries), EXCLUDED_CLAUSES (5 entries)
SHA-256 f1801caf5371d8faa15ca7e923736a361a03bfab84bae930d6d848fa20596b06
Status single-run annotation — n=1 annotator, no adjudication, no inter-annotator agreement computed. See LIMITATIONS.md.

Resolved 2026-09-24: the module's docstring previously cited its source as data/uploads/..._pdf_stairs_mock.pdf — a file that does not exist anywhere in this repository. The actual source has been verified directly by text extraction: sources/OBC_2023.Volume_1_P_9.pdf, page 32 of 302 onward, whose text opens with "9.8.2. Stair Dimensions / 9.8.2.1. Stair Width" and matches SOURCE_TEXT verbatim. The module's docstring now cites this file with its SHA-256 and page range, and states the annotator (single annotator, no adjudication, no IAA — see LIMITATIONS.md).


3a. Rule extraction vs. human gold set — OBC 2023 §9.8 (end to end)

Claim Through the live BIM-Guard UI, extraction on the human-annotated §9.8 clauses, scored against the Label Studio gold set: see the four runs below. Runs 1–3 are scored on gold project1_human_2026-10-02b.json (89 rules, 117 clauses), run 4 on project1_human_2026-10-03.json (116 rules, 129 clauses). The two gold versions are not comparable; compare runs only within a version.
Producing script eval/e2e/extraction_confusion_e2e.py → eval/score_extraction_vs_human.py
Artifacts eval/results/e2e/run1_browser/, run2_playwright/, run3_variance_a/, run4_corrected_gold/ (confusion.json/.md, drafts.json); figure docs/publication/figures/fig_extraction_confusion_run1.png (eval/plot_extraction_confusion.py)
Status single-run, high variance: run 1 is an outlier (see below). Do not quote run 1 alone.
Run Gold Drafts Clause TP/FP/FN/TN Lenient P / R / F1 Normalized F1 Strict F1
run 1 (run1_browser, UI by hand; openai/gpt-5.6-luna-pro, observed in the request/backend log, not recorded in the output) 02b 104 37 / 0 / 13 / 67 98.3 / 64.0 / 77.6 57.1 23.1
run 2 (run2_playwright; openai/gpt-5.6-luna-pro, per run2_playwright/run.log) 02b 48 10 / 1 / 40 / 66 74.1 / 22.5 / 34.5 20.3 11.9
run 3 (run3_variance_a, openai/gpt-5.6-luna-pro, 2026-10-02 22:01 UTC, after an app rebuild at 21:50 UTC) 02b 53 13 / 2 / 37 / 65 71.4 / 16.9 / 27.3 8.9 8.9
run 4 (run4_corrected_gold, openai/gpt-5.6-luna-pro, 2026-10-03, corrected clause text as a new document) 10-03 42 11 / 1 / 48 / 69 84.6 / 19.0 / 31.0 13.9 12.5

All rows were re-scored on 2026-10-03 with the current scorer, which (a) maps table-derived drafts to their table and spelled-out counts to their sentence, and (b) reports rules that restate an already-matched human rule (one rule per element for "stairs and ramps") in a separate Redundant column instead of as false positives. This changed runs 2 and 3 slightly (previously lenient F1 32.2% and 26.8%; clause FP 3 and 4); run 1 is unchanged.

Runs 1, 2 and 4 used the same model (openai/gpt-5.6-luna-pro; an earlier version of this section said GPT-6.1 Sol Pro, which the run evidence contradicts), yet gave 104 vs 48 vs 42 drafts. Runs 2–4 each produced drafts for only part of §9.8 (e.g. run 4 has none for 9.8.5, 9.8.6 or 9.8.10): with a pasted clause file BIM-Guard's Smart TOC finds a single section, so the whole text reaches the model as one chunk and the output is cut short. Recall is therefore dominated by run-to-run coverage, not per-clause accuracy; precision stays high (74–98% lenient) across runs. The extraction call accepts no seed or temperature. Per-run metadata: eval/results/e2e/README.md.

Gold-set caveats (disclosed, not hidden):


4. NLP annotation test suite — 60/60

Claim 60-point automated test suite across 6 linguistic/DocLang annotation capabilities, currently 60/60 passing.
Producing script eval/score_nlp_annotation.py
SHA-256 see eval/score_nlp_annotation.py in the running tree at time of check
Status reproduced — re-run on 2026-09-24, confirmed 60/60.

Note: this is a deterministic, hermetic, hand-written assertion suite over nlp_annotation/ with no LLM calls and no external dependencies — it is the one measurement in this repository that is trivially and fully reproducible. It measures rule-based pattern coverage, not real-world extraction accuracy; see LIMITATIONS.md for what it does not tell you.


5. LLM-as-judge rule-generation quality & calibration

Claim (withdrawn) Threshold sweep τ ∈ {2,3,4,5}, N=5 repeated draws, Pearson r = 0.9929 / Spearman ρ = 0.9702 against human experts; "τ = 4 is optimal".
Producing script eval/score_judge_sensitivity.py
Status simulated (downgraded 2026-10-03). The 20 calibration cases and their "judge_draws" are hard-coded in the script (its own comment: "simulated / recorded"). No live judge call produced them and the "human" scores are not traceable to a named rater. The correlation therefore measures the script's constants, and τ = 4 is not shown to be optimal.
To restore Run eval/eval_harness.py with a real judge on the gold cases, record the raw draws and an independent human rating per case, then recompute.

6. Inter-annotator agreement (IAA)

Claim (withdrawn) 30-task study, Architect vs. Computational BIM Specialist with Senior Adjudication; Cohen's κ = 0.957 (entity), 1.000 (deontic), Fleiss' κ = 0.971.
Producing script eval/score_iaa.py (scoring code is sound)
Corpus research/annotations/dual_annotator_corpus.json, written by research/annotations/generate_corpus.py, which hard-codes both annotators' labels and an "ann2_alt" variation per clause.
Status simulated (downgraded 2026-10-03). The corpus was generated by code ("realistic annotator agreement variations"), not collected from two people, so the κ values describe the generator. The scorer itself can be reused on real data.
Real data today The only human annotation is the single-annotator Label Studio set (research/label_studio/), no adjudication, no IAA.
To restore Have a second person annotate a subset in Label Studio independently, then run score_iaa.py on the two exports.

7. Architectural compliance engines benchmark (ARCH-EGRESS-001, ARCH-SPATIAL-001)

Claim 22-case grounded benchmark evaluating the active architecture compliance engines across egress travel distances, storey exit counts, emergency escape windows, daylight glazing ratios, and fire separation ratings, reporting 100% accuracy (13 TP, 9 TN, 0 FP, 0 FN) with Wilson score 95% confidence intervals.
Producing script eval/score_arch_engines.py (supported by eval/generate_arch_test_models.py and eval/stats_util.py)
Baseline eval/baselines/score_arch_engines.baseline.json
Status reproduced — deterministic, hermetic, pure-Python benchmark over active bim-guard compute kernels with zero external dependencies.

8. Cross-jurisdiction generalization benchmark (OBC 2024 vs. SBC-201-2007)

Claim (withdrawn) F1 = 100% on OBC (29 rules) and SBC (28 rules); generalization gap ΔF1 = 0.
Producing script eval/score_cross_code.py
Status simulated (downgraded 2026-10-03). The script runs no extractor. _simulate_rule_extraction_benchmark copies each gold rule into a "candidate", corrupts a random 3.5%, and scores the result against the same gold rules; TN is hard-coded to 5 and the "PASS" and "generalizes without overfitting" text is printed unconditionally. It says nothing about BIM-Guard or about transfer between codes. The SBC gold rule set (eval_gold_sbc_chapter10.py) is real annotation work and can be used for a real run.
To restore Extract the SBC clauses with the real pipeline and score with score_rule_extraction.py-style matching.

9. Procedural whole-building topological IFC4 modeling

Claim Fully schema-valid procedural IFC4 building model generator with spatial containment hierarchy and explicit IfcRelSpaceBoundary topological relationships connecting rooms to walls, doors, windows, and stairs for network egress path calculations.
Producing script eval/generate_arch_test_models.py
Artifact eval/fixtures/procedural_benchmark_building.ifc
Status reproduced — verified by direct ifcopenshell.open() parsing (4 connected spaces, 10 IfcRelSpaceBoundary records).

Open items tracked, not yet resolved