Measured, not promised
Every figure below is copied from the project’s registry, with the split it was measured on. Nothing here is a benchmark against another product, and nothing was re-measured for this page.
| Metric | Value | Split | Source |
|---|---|---|---|
| IoU | 0.637 | held-out test, threshold 0.85 | registry.py:737 |
| Precision | 0.811 | held-out test, threshold 0.85 | registry.py:737 |
| Recall | 0.748 | held-out test, threshold 0.85 | registry.py:738 |
| IoU at the previous threshold 0.25 | 0.607 | held-out test | registry.py:738 |
| Precision at the previous threshold 0.25 | 0.664 | held-out test | registry.py:738 |
| Best validation IoU during training | 0.6063 | validation | model_registry.json:123 |
| Validation IoU at threshold 0.85 | 0.5962 | validation | model_registry.json:123 |
The headline figure was last changed in the repository on 2026-08-15.
Known weaknesses
Model and threshold
| Registry key | crack_segmentation |
|---|---|
| Architecture | SegFormer-B5 (run crack_segformer_b5, 40 epochs) |
| Input size | 1024 px |
| Decision threshold | 0.85 |
| Labels | crack |
| Status | installed (registry status; the weights file is not in the Git repository) |
History
- Replaced SegFormer-B2, which reached test-split IoU 0.515.
- The classical heuristic baseline reaches IoU 0.045 on the same data.
- Raising the threshold from 0.25 to 0.85 gained +0.030 IoU and +0.147 precision with no retraining.
What was it trained on?
crack_segmentation_kaggle: 11,298 crack image/mask pairs aggregating six public crack corpora (CFD, Crack500, GAPs384, DeepCrack, Rissbilder, Volker) (11,298 images). The 0.85 threshold was chosen by sweep on the validation split. IoU, precision and recall below are on the held-out test split.
| Source | www.kaggle.com/datasets/lakshaymiddha/crack-segmentation-dataset |
|---|---|
| Licence | Catalogue says: "See Kaggle dataset page; aggregates CFD, Crack500, GAPs384, DeepCrack, Rissbilder, Volker" (no single licence recorded) registry.py:94 |
Weights trained on a dataset carry that dataset’s terms, which can be stricter than the code licence. The weights are not in the Git repository.
What happens when it cannot answer?
Without usable weights the crack path does not report AI. README on main: 'Anything without usable weights still reports `heuristic`, never as AI.' core/detection.py reports a registered key whose weights are not installed.
Sources: README.md:108 · detection.py:521
Provenance
The registry records the SHA-256 of the weights file whose metrics were measured. At load, the installed file is hashed and compared; a different file is reported as a mismatch, so these figures are never attached to weights they do not describe.
sha256 3391eca15d43bcc2810c3fcfe7e9db9f99dddb5de4633d5963c705c8e1cc0c23Tests named for this capability in the registry: tests/test_honesty.py::TestDetectionReportsWhatItActuallyUsed.

Other model cards: Corrosion severity · Solar cell defects (electroluminescence) · Solar thermal anomalies · AI defect detection