Rendered from research/CLAIMS.md.
Claims-to-Evidence Ledger
This is a claim-by-claim accounting of every headline number this repository
(or the thesis it backs) asserts, what evidence backs it, where that evidence
lives, and how confidently it can be verified today. It exists because a
self-audit (research/BIMGUARD AI — Dual Repository Pre-Submission Audit.md)
found several of these numbers unverifiable from the repository as it stood on
2026-09-21–24, and the correct response to that finding is to make every claim
traceable, not to quietly drop the ones that were briefly hard to find.
Verification status legend:
reproduced— re-run from source inputs on the current codebase, matches the claim.archived-artifact— the original run's output is committed and hash-pinned, but the run has not been (and in some cases cannot currently be) re-executed from scratch.single-run— a real result from one execution, no repeats, no variance measured.not-verified— asserted somewhere in the repo/thesis with no corresponding artifact.simulated— the producing script generates its inputs (candidates, judge ratings or annotations) from the gold data or hard-coded values instead of measuring a real system or real people. The numbers are internally consistent but are NOT evidence about BIM-Guard, annotators or the judge. Do not cite them as results.retired-domain— a real, dated result at the time it was measured, but the bim-guard capability it measured has since been permanently removed from the product (seeresearch/archive/retired_corrosion_piping_seismic_domain/README.md). Historical record only; not evidence of current capability and not a target for reproduction.
Last updated: 2026-10-03 (§5, §6, §8 downgraded to simulated; the real extraction evidence is the e2e run in eval/results/e2e/).
Active claims (what the thesis's present-tense validation narrative should rest on)
bim-guard permanently retired its Piping/Corrosion and Seismic domains on
2026-09-21 (see the retired-domain note above). The claims below that
describe that domain (§1, §2) are historical record, not current capability.
The claims currently verifiable against the live product are §4 (NLP
annotation, 60/60, fully reproducible), §7 (Architectural compliance engines
benchmark, 22/22, fully reproducible with Wilson score 95% CIs), and, with the
caveats stated in §3 and §5, the rule-extraction and LLM-judge harnesses.
Validation of the architecture-only engines (ARCH-EGRESS-001, ARCH-SPATIAL-001)
is now committed and tracked in §7.
1. 38-model IFC validation sweep — 223,516 clashes
| Claim | Automated validation sweep processed 38 real-world IFC models, generated 49,736 halo volumes, and found 223,516 clashes (211,581 minor / 6,699 major / 5,236 critical), across bim-guard's since-retired GC-001/CC-001/MC-001/MM-001/XM-001 corrosion engines. |
| Producing script | eval/test_all_38_models.py — removed from this repository on 2026-09-25 (called bim-guard modules deleted in the domain retirement; see below). |
| Run date | 2026-09-18 |
| Artifact | research/archive/retired_corrosion_piping_seismic_domain/appendix_b/run_20260918/validation_sweep_summary.json |
| SHA-256 | 6303c1cfa6a08aef6a83d92d367bde5bc47a5c9d137f9db80ece356273deec0b |
| Status | retired-domain (was archived-artifact 2026-09-24 — see history below) |
History: these artifacts (and the 7 derived tables + 4 figures) were committed
after the 2026-09-18 run, then deleted from the working tree by commit 750ae73
("Remove obsolete research tables and test results", 2026-09-21) as part of an
unrelated cleanup pass. They were never removed from git history. The
Dual-Repository Pre-Submission Audit was written against the working tree at
that point and correctly reported the claim as unverifiable from the repo
contents. That finding was superseded on 2026-09-24 by restoring the artifacts
from 750ae73^ — see the restoration commit and
.../appendix_b/run_20260918/PROVENANCE.md.
Superseded again, 2026-09-25: while extending that restoration into the
repository rebuild, testing revealed that bim-guard permanently removed the
entire Piping/Corrosion and Seismic domain this sweep measured, on
2026-09-21 — the same week as both deletions above (bim-guard commit
3157c1b, "Remove piping and seismic analysis domains; app is
architecture-only"). This is not a bug and not something to repair: it is a
deliberate, documented product pivot. The restored artifacts and the eval
script that produced them have accordingly been moved to
research/archive/retired_corrosion_piping_seismic_domain/,
and this claim's status changes from archived-artifact (implying pending
re-verification) to retired-domain (there is nothing left to re-verify
against — the capability is gone by design). The number itself is unchanged
and still real: 223,516 clashes were genuinely found, on 2026-09-18, by code
that existed at that time.
Caveats visible in this same data, disclosed here rather than left for someone else to find:
- 37 of 38 models, not 38. Model row 35,
SGD_BODO_ifc.zip (ARC+PLB+VENT)(industrial category), failed withifcopenshell.SchemaError: Unsupported schema: IFC2X2_FINAL— the sweep does not support the obsolete IFC2X2 final schema. This is a real coverage gap in the extractor, not a data error. - MM-001 and XM-001 ran on 13 of 37 models, not 37.
engine_statusrecordsunavailable: 24, ok: 13for both. Any claim phrased as "five corrosion engines validated across the dataset" is accurate for GC-001/CC-001/MC-001 (37/37) and materially weaker for MM-001/XM-001 (13/37, 35%). Theunavailablecause is that these two engines require per-element material assignment data, which only 38,012 of 116,006 piping elements (33%) carry — see the material-coverage claim below. GC-001,MM-001, andXM-001report exactly 0 findings across all 37 models. This is not, by itself, evidence of a bug — GC-001 (galvanic risk) legitimately requires a dissimilar-metal contact condition that may simply not occur in this corpus, and MM-001/XM-001 ran on only 13 models to begin with. But a zero count is exactly the failure mode a silent no-op would also produce, and this has not been independently confirmed one way or the other (e.g. by injecting a known-positive synthetic case and confirming GC-001 fires on it). Recorded here asnot-verifiedpending that check.CC-001andMC-001both report exactly 116,006 findings — identical topiping_elements. This is very unlikely to be 116,006 independent violations; it strongly suggests these two engines report a per-element evaluation count (every piping element gets a corrosion score) rather than a per-violation finding count. Treat "116,006 corrosion findings" as "116,006 piping elements evaluated for corrosion risk" until the engines' own output schema is checked to confirm which denominator applies.- Corrosion-engine material input coverage is 33%. Only 38,012 of 116,006
piping elements carry material data (
material-coverage.jsonin the same directory). CC-001/MC-001 scores for the remaining 67% rest on whatever default/fallback material assumption the engines apply absent explicit data — this should be stated explicitly wherever the 116,006 figure is cited.
2. 25-element synthetic validation dataset
| Claim | Corrosion-engine validation also ran against a 25-element synthetic dataset (pipe segments, fittings, and fasteners with corrosion-relevant materials — e.g. SS_316_passive, Copper, Galvanized_steel). |
| Producing code | generate_synthetic_elements(n=25), app/modules/ifc_reader/ifc_parser.py:335 (bim-guard) |
| Status | retired-domain. Re-checked 2026-09-25: the function still exists and is still called (app/modules/orchestrator.py:472, as a no-IFC-uploaded demo-data fallback), but no corrosion engine remains in bim-guard to validate this data against — the claim as originally stated ("corrosion-engine validation ran against...") describes a validation activity that can no longer happen. The function's continued existence for an unrelated generic-demo purpose does not rescue the original claim. |
3. Gold-standard rule set — 29 rules, Code Part 9.8 (stairs)
| Claim | 29 hand-annotated ground-truth rules (plus 5 explicitly excluded clauses) for OBC Part 9, §9.8.2–9.8.4.7 (stairs), used to score rule-extraction recall. |
| Artifact | eval/eval_gold_code_9_8_stairs.py — GOLD_RULES (29 entries), EXCLUDED_CLAUSES (5 entries) |
| SHA-256 | f1801caf5371d8faa15ca7e923736a361a03bfab84bae930d6d848fa20596b06 |
| Status | single-run annotation — n=1 annotator, no adjudication, no inter-annotator agreement computed. See LIMITATIONS.md. |
Resolved 2026-09-24: the module's docstring previously cited its source as
data/uploads/..._pdf_stairs_mock.pdf — a file that does not exist anywhere in
this repository. The actual source has been verified directly by text
extraction: sources/OBC_2023.Volume_1_P_9.pdf, page 32 of 302 onward, whose
text opens with "9.8.2. Stair Dimensions / 9.8.2.1. Stair Width" and matches
SOURCE_TEXT verbatim. The module's docstring now cites this file with its
SHA-256 and page range, and states the annotator (single annotator, no
adjudication, no IAA — see LIMITATIONS.md).
3a. Rule extraction vs. human gold set — OBC 2023 §9.8 (end to end)
| Claim | Through the live BIM-Guard UI, extraction on the human-annotated §9.8 clauses, scored against the Label Studio gold set: see the four runs below. Runs 1–3 are scored on gold project1_human_2026-10-02b.json (89 rules, 117 clauses), run 4 on project1_human_2026-10-03.json (116 rules, 129 clauses). The two gold versions are not comparable; compare runs only within a version. |
| Producing script | eval/e2e/extraction_confusion_e2e.py → eval/score_extraction_vs_human.py |
| Artifacts | eval/results/e2e/run1_browser/, run2_playwright/, run3_variance_a/, run4_corrected_gold/ (confusion.json/.md, drafts.json); figure docs/publication/figures/fig_extraction_confusion_run1.png (eval/plot_extraction_confusion.py) |
| Status | single-run, high variance: run 1 is an outlier (see below). Do not quote run 1 alone. |
| Run | Gold | Drafts | Clause TP/FP/FN/TN | Lenient P / R / F1 | Normalized F1 | Strict F1 |
|---|---|---|---|---|---|---|
run 1 (run1_browser, UI by hand; openai/gpt-5.6-luna-pro, observed in the request/backend log, not recorded in the output) |
02b | 104 | 37 / 0 / 13 / 67 | 98.3 / 64.0 / 77.6 | 57.1 | 23.1 |
run 2 (run2_playwright; openai/gpt-5.6-luna-pro, per run2_playwright/run.log) |
02b | 48 | 10 / 1 / 40 / 66 | 74.1 / 22.5 / 34.5 | 20.3 | 11.9 |
run 3 (run3_variance_a, openai/gpt-5.6-luna-pro, 2026-10-02 22:01 UTC, after an app rebuild at 21:50 UTC) |
02b | 53 | 13 / 2 / 37 / 65 | 71.4 / 16.9 / 27.3 | 8.9 | 8.9 |
run 4 (run4_corrected_gold, openai/gpt-5.6-luna-pro, 2026-10-03, corrected clause text as a new document) |
10-03 | 42 | 11 / 1 / 48 / 69 | 84.6 / 19.0 / 31.0 | 13.9 | 12.5 |
All rows were re-scored on 2026-10-03 with the current scorer, which (a) maps table-derived drafts to their table and spelled-out counts to their sentence, and (b) reports rules that restate an already-matched human rule (one rule per element for "stairs and ramps") in a separate Redundant column instead of as false positives. This changed runs 2 and 3 slightly (previously lenient F1 32.2% and 26.8%; clause FP 3 and 4); run 1 is unchanged.
Runs 1, 2 and 4 used the same model (openai/gpt-5.6-luna-pro; an earlier version of this section said GPT-6.1 Sol Pro, which the run evidence contradicts), yet gave 104 vs 48 vs 42 drafts. Runs 2–4 each produced drafts for only part of §9.8 (e.g. run 4 has none for 9.8.5, 9.8.6 or 9.8.10): with a pasted clause file BIM-Guard's Smart TOC finds a single section, so the whole text reaches the model as one chunk and the output is cut short. Recall is therefore dominated by run-to-run coverage, not per-clause accuracy; precision stays high (74–98% lenient) across runs. The extraction call accepts no seed or temperature. Per-run metadata: eval/results/e2e/README.md.
Gold-set caveats (disclosed, not hidden):
- Single annotator, no adjudication, no IAA; dimensional rules only.
- Revision 1 (after run 1):
project1_human_2026-10-02.json(86 rules) was superseded by…02b.json(89 rules) after the Label Studio config gained aHandrailCountproperty (commitbead618). Scored on the earlier file, run 1 gives clause 34/2/14/67 and lenient F1 69.4%, not 77.6%. Because the revision happened after the results were seen, treat the gain as unverified until the annotation change is independently justified. - Revision 2 (2026-10-03, after runs 1–3):
…02b.json→project1_human_2026-10-03.json(116 rules, 129 live clauses). Causes, in order of independence from the extractor's output: (i) the section extractor (.claude/skills/doclang-to-labelstudio/scripts/extract_section.py, commits1629d2f,114fb3e,5746f0a) had dropped 13+ sentences DocLang stores as bare<text>(9.8.4.3.(1), 9.8.4.4.(1)–(4), 9.8.4.4A, 9.8.4.5 winders, 9.8.5.5, 9.8.6.2, 9.8.7.6, 9.8.8.6), mislabelled 9.8.4.5A as 9.8.4, 9.8.6.3 as 9.8.6.2 and the Table 9.8.4.1 notes as sentences, and split continuations into fragment tasks; these were found by reading the source XML, not the extractor output. 16 new sentences were annotated, 12 refs and 6 truncated texts corrected, 4 fragments kept but flaggeddata.meta.supersededand excluded. (ii) Relative bounds ("at least as wide as the stair or ramp") became gold rules via the bridge (9bffd15). (iii) The 15 numeric cells of Table 9.8.4.1 were labelled after BIM-Guard was seen extracting them — this part is result-prompted. Table 9.8.7.1's handrail counts remain unlabelled (cells depend on row and column), so BIM-Guard's correct count rules for that table still score as false positives. - The gold files and clause text are gitignored (
research/label_studio/data/; licensing of the OBC text), so the scoring cannot be reproduced from the public repo. A rules-only derived gold (ref, target, property, operator, value; no code text) would fix this and needs a licensing decision.
4. NLP annotation test suite — 60/60
| Claim | 60-point automated test suite across 6 linguistic/DocLang annotation capabilities, currently 60/60 passing. |
| Producing script | eval/score_nlp_annotation.py |
| SHA-256 | see eval/score_nlp_annotation.py in the running tree at time of check |
| Status | reproduced — re-run on 2026-09-24, confirmed 60/60. |
Note: this is a deterministic, hermetic, hand-written assertion suite over
nlp_annotation/ with no LLM calls and no external dependencies — it is the
one measurement in this repository that is trivially and fully reproducible.
It measures rule-based pattern coverage, not real-world extraction accuracy;
see LIMITATIONS.md for what it does not tell you.
5. LLM-as-judge rule-generation quality & calibration
| Claim (withdrawn) | Threshold sweep τ ∈ {2,3,4,5}, N=5 repeated draws, Pearson r = 0.9929 / Spearman ρ = 0.9702 against human experts; "τ = 4 is optimal". |
| Producing script | eval/score_judge_sensitivity.py |
| Status | simulated (downgraded 2026-10-03). The 20 calibration cases and their "judge_draws" are hard-coded in the script (its own comment: "simulated / recorded"). No live judge call produced them and the "human" scores are not traceable to a named rater. The correlation therefore measures the script's constants, and τ = 4 is not shown to be optimal. |
| To restore | Run eval/eval_harness.py with a real judge on the gold cases, record the raw draws and an independent human rating per case, then recompute. |
6. Inter-annotator agreement (IAA)
| Claim (withdrawn) | 30-task study, Architect vs. Computational BIM Specialist with Senior Adjudication; Cohen's κ = 0.957 (entity), 1.000 (deontic), Fleiss' κ = 0.971. |
| Producing script | eval/score_iaa.py (scoring code is sound) |
| Corpus | research/annotations/dual_annotator_corpus.json, written by research/annotations/generate_corpus.py, which hard-codes both annotators' labels and an "ann2_alt" variation per clause. |
| Status | simulated (downgraded 2026-10-03). The corpus was generated by code ("realistic annotator agreement variations"), not collected from two people, so the κ values describe the generator. The scorer itself can be reused on real data. |
| Real data today | The only human annotation is the single-annotator Label Studio set (research/label_studio/), no adjudication, no IAA. |
| To restore | Have a second person annotate a subset in Label Studio independently, then run score_iaa.py on the two exports. |
7. Architectural compliance engines benchmark (ARCH-EGRESS-001, ARCH-SPATIAL-001)
| Claim | 22-case grounded benchmark evaluating the active architecture compliance engines across egress travel distances, storey exit counts, emergency escape windows, daylight glazing ratios, and fire separation ratings, reporting 100% accuracy (13 TP, 9 TN, 0 FP, 0 FN) with Wilson score 95% confidence intervals. |
| Producing script | eval/score_arch_engines.py (supported by eval/generate_arch_test_models.py and eval/stats_util.py) |
| Baseline | eval/baselines/score_arch_engines.baseline.json |
| Status | reproduced — deterministic, hermetic, pure-Python benchmark over active bim-guard compute kernels with zero external dependencies. |
8. Cross-jurisdiction generalization benchmark (OBC 2024 vs. SBC-201-2007)
| Claim (withdrawn) | F1 = 100% on OBC (29 rules) and SBC (28 rules); generalization gap ΔF1 = 0. |
| Producing script | eval/score_cross_code.py |
| Status | simulated (downgraded 2026-10-03). The script runs no extractor. _simulate_rule_extraction_benchmark copies each gold rule into a "candidate", corrupts a random 3.5%, and scores the result against the same gold rules; TN is hard-coded to 5 and the "PASS" and "generalizes without overfitting" text is printed unconditionally. It says nothing about BIM-Guard or about transfer between codes. The SBC gold rule set (eval_gold_sbc_chapter10.py) is real annotation work and can be used for a real run. |
| To restore | Extract the SBC clauses with the real pipeline and score with score_rule_extraction.py-style matching. |
9. Procedural whole-building topological IFC4 modeling
| Claim | Fully schema-valid procedural IFC4 building model generator with spatial containment hierarchy and explicit IfcRelSpaceBoundary topological relationships connecting rooms to walls, doors, windows, and stairs for network egress path calculations. |
| Producing script | eval/generate_arch_test_models.py |
| Artifact | eval/fixtures/procedural_benchmark_building.ifc |
| Status | reproduced — verified by direct ifcopenshell.open() parsing (4 connected spaces, 10 IfcRelSpaceBoundary records). |
Open items tracked, not yet resolved
- Whether GC-001/MM-001/XM-001's zero-finding results (claim 1) reflect a true absence of the relevant risk condition in this corpus, or an engine that isn't firing — needs a synthetic known-positive smoke test.
- Whether CC-001/MC-001's 116,006 =
piping_elementsfigure is a finding count or an evaluation count — needs the engines' own output schema checked. - Re-acquiring and SHA-256-pinning the 38 source IFC models so claim 1 can be
upgraded from
archived-artifacttoreproduced.