Measured, not promised
Every figure below is copied from the project’s registry, with the split it was measured on. Nothing here is a benchmark against another product, and nothing was re-measured for this page.
| Metric | Value | Split | Source |
|---|---|---|---|
| Severe-pixel recall | 0.788 | validation | registry.py:764 |
| Mean IoU (four ordinal grades) | 0.5769 | validation | model_registry.json:223 |
| Pixel accuracy | 0.8497 | validation | model_registry.json:223 |
| IoU, good | 0.871 | validation | model_registry.json:223 |
| IoU, fair | 0.498 | validation | model_registry.json:223 |
| IoU, poor | 0.45 | validation | model_registry.json:223 |
| IoU, severe | 0.489 | validation | model_registry.json:223 |
The headline figure was last changed in the repository on 2026-08-25.
Known weaknesses
Model and threshold
| Registry key | corrosion_severity_segmentation |
|---|---|
| Architecture | SegFormer-B2 (run corrosion_segformer_b2_cs, 80 epochs, log-inverse class weighting) |
| Input size | 512 px |
| Labels | good, fair, poor, severe |
| Status | installed (registry status; the weights file is not in the Git repository) |
History
- A first attempt, a single-class YOLO11l corrosion detector on 498 images, reached mAP50 0.254 with recall 0.257 (three corrosion sites in four unfound). It was rejected rather than registered.
What was it trained on?
Virginia Tech Corrosion Condition State: 440 annotated 512x512 frames of steel bridge elements, photographed with a drone and a DSLR and rated against AASHTO/BIRM condition states (440 images). 330 train / 66 validation / 44 test. The archive ships train and test only, so validation is carved out of train by salted hash.
| Source | Not recorded in the catalogue |
|---|---|
| Licence | CC0 (stated in training/configs/corrosion_segformer_b2_cs.yaml line 1 and the training/datasets/prepare.py docstring at line 1095; this dataset has no entry in training/datasets/registry.py on main) corrosion_segformer_b2_cs.yaml:1 |
Weights trained on a dataset carry that dataset’s terms, which can be stricter than the code licence. The weights are not in the Git repository.
What happens when it cannot answer?
No heuristic fallback. With the model absent it raises CorrosionGradingRefused: '<key> is registered but its weights are not installed at <path>. There is no heuristic severity grade to fall back to, so nothing is returned.'
Sources: detection.py:519
Provenance
The registry records the SHA-256 of the weights file whose metrics were measured. At load, the installed file is hashed and compared; a different file is reported as a mismatch, so these figures are never attached to weights they do not describe.
sha256 edcecd84cc6839d76b7db3e5a6475ec694636fc6cffe545ca4718c7147b5002fTests named for this capability in the registry: tests/test_corrosion_severity.py::TestItActuallyPredictsTheWorstGrade, tests/test_corrosion_severity.py::TestRefusalRatherThanAGuessedGrade, tests/test_corrosion_severity.py::TestTheGradeIsAnArgmaxNotAThreshold, tests/test_corrosion_severity.py::TestTheScaleIsOrdinalAndSaysSo.
Other model cards: Crack segmentation · Solar cell defects (electroluminescence) · Solar thermal anomalies · AI defect detection