BIM-Guard Evaluation

Results

Architectural compliance engines reproduced

ARCH-EGRESS-001 and ARCH-SPATIAL-001 evaluated on 22 procedurally generated IFC scenarios covering egress travel distance, storey exit count, emergency escape windows, daylight glazing ratio and fire-separation rating. The benchmark is deterministic and hermetic; ground truth comes from the generator (eval/generate_arch_test_models.py), not from human labelling.

Architectural Compliance Engines Validation Performance (ARCH-EGRESS-001 & ARCH-SPATIAL-001)
MetricPoint Estimate95% Wilson CITP / FPFN / TNTotal Cases
Accuracy100.0%[85.1%, 100.0%]13 / 00 / 922
Sensitivity / Recall100.0%[77.2%, 100.0%]——13 Positives
Specificity100.0%[70.1%, 100.0%]——9 Negatives
Precision100.0%[77.2%, 100.0%]——13 Flagged
F1 Score100.0%[77.2%, 100.0%]———
Balanced Accuracy100.0%[73.6%, 100.0%]———

Source: table_2_arch_confusion_matrix.tex (generated by eval/generate_publication_artifacts.py)

Confusion matrix for the architecture engines (13 TP, 9 TN, 0 FP, 0 FN).
Confusion matrix for the architecture engines (13 TP, 9 TN, 0 FP, 0 FN).

How to read this. 100% on 22 cases is a statement about these cases. The Wilson interval, with a lower bound of 85.1% for accuracy, is the honest summary of what the sample supports. Because the ground truth is synthetic, it shows the engines implement the rules as specified, not that the rules match real-world inspector judgement. Replacing it with human labels is the next step (see Limitations).

Claim and artifact: §7 of the claims ledger.

Extraction evidence

The real rule-extraction evidence is the end-to-end run against the human-labelled OBC 9.8 gold set, committed in eval/results/e2e/. It has a single annotator and no adjudication; see Limitations.

Simulated harnesses simulated

Not evidence. The tables below are kept so the repository's history is transparent, and because the harnesses are useful scaffolding for real data. Their inputs are generated by the scripts themselves, so the numbers are internally consistent but say nothing about BIM-Guard, human annotators or an LLM judge. Do not cite them as results. Details: claims ledger §5, §6 and §8.

Inter-annotator agreement (simulated annotators)

The 30-task corpus is generated by research/annotations/generate_corpus.py with hard-coded labels for both "annotators".

Inter-Annotator Agreement (IAA) on Dual Building Code Regulatory Corpus (N=30)
Annotation TargetCohen's κ (1 vs 2)95% Bootstrap CIFleiss' κ (Multi-Rater)95% Bootstrap CI
Target IFC Entity0.957[0.864, 1.000]0.971[0.909, 1.000]
Property Name0.683[0.506, 0.820]0.786[0.657, 0.879]
Deontic Modality1.000[1.000, 1.000]1.000[1.000, 1.000]
Measurement Unit1.000[1.000, 1.000]1.000[1.000, 1.000]
Span Extraction MetricPrecisionRecallF1 Score95% Bootstrap CI
Span IoU ≥ 0.500.7780.7780.778[0.691, 0.824]
Span IoU ≥ 0.750.6670.6670.667[0.574, 0.722]

Source: table_1_iaa_metrics.tex (generated by eval/generate_publication_artifacts.py)

Simulated IAA profile. Not a measurement of human agreement.
Simulated IAA profile. Not a measurement of human agreement.
Cross-jurisdiction generalization (simulated extraction)

eval/score_cross_code.py perturbs gold rules instead of running an extractor, so F1 = 100% is guaranteed by construction.

Cross-Jurisdictional Generalization: Ontario Building Code (OBC 2024) vs. Saudi Building Code (SBC-201-2007)
Benchmark MetricOBC 2024 (Part 9)SBC-201-2007 (Chapter 8)Generalization Gap (Δ)
Gold Standard Rules2928Balanced Sample
Target IFC Entity Accuracy100.0%100.0%Δ = 0.0%
Operator Extraction Accuracy100.0%100.0%Δ = 0.0%
Precision [95% CI]100.0%100.0%Δ = 0.0%
Recall / Sensitivity [95% CI]100.0%100.0%Δ = 0.0%
F1 Score [95% CI]100.0%100.0%\mathbf{Δ F_1 = 0.0000}

Source: table_3_cross_code_generalization.tex (generated by eval/generate_publication_artifacts.py)

LLM-judge threshold sweep (hard-coded ratings)

eval/score_judge_sensitivity.py scores hard-coded judge ratings. It does not show that τ = 4 is optimal.

LLM-as-Judge Threshold Sensitivity Sweep (τ ∈ {2, 3, 4, 5}) on Regulatory Verification
CutoffTPFPFNTNPrecisionRecallSpecificityF1 Score
τ = 21340376.5%100.0%42.9%86.7%
τ = 31310692.9%100.0%85.7%96.3%
τ = 410037100.0%76.9%100.0%87.0%
τ = 530107100.0%23.1%100.0%37.5%

Source: table_4_judge_sensitivity.tex (generated by eval/generate_publication_artifacts.py)

Judge threshold sweep over simulated ratings. Illustrative only.
Judge threshold sweep over simulated ratings. Illustrative only.