Beyond Clean Accuracy: Corruption Robustness and Failure Modes of CNN and Transformer Models for Magnetic Tile Defect Segmentation
Authors
| Issue | Vol. 1 No. 1 (2025) |
| Published | 27 August 2026 |
| Section | Articles |
Abstract
Reliable magnetic-tile defect segmentation requires models that remain stable when image quality degrades, yet clean benchmark performance alone may not reveal such behavior. This study evaluates the corruption robustness of U-Net + ResNet18, DeepLabV3+ + MobileNetV2, and SegFormer-B0 using 199 independent magnetic-tile scene groups under group-aware five-fold cross-validation. Four test-only corruptions, Gaussian Noise, Gaussian Blur, Brightness Increase, and Contrast Reduction, were applied at five severity levels. Clean Macro Dice reached 0.5965, 0.5710, and 0.6247 for U-Net, DeepLabV3+, and SegFormer-B0, respectively, with no significant difference among models (Friedman p = 0.1693). Under corruption, mean Dice decreased to 0.1912, 0.2026, and 0.2935, corresponding to mean relative performance degradations of 67.95%, 64.52%, and 53.02%. Mean corrupted Dice differed significantly among models (Friedman p = 0.001118), with SegFormer-B0 achieving significantly higher paired mean corrupted Dice than both CNN baselines after Holm correction. Gaussian Noise further revealed distinct false-alarm behavior on defect-free scenes. These findings show that corruption-aware evaluation reveals reliability differences and failure modes that were not apparent from clean segmentation assessment alone.
