Skip to content

Computer vision · Structural health monitoring

Multi-defect concrete inspection with deep learning

I benchmarked YOLO-seg, U-Net and FPN models for crack and multi-defect segmentation, then tested them on real inspection photos from an industry partner. The field test showed where benchmark accuracy fails on site.

Context
MSc dissertation, Ulster University · industry partner Amphora Consulting (Belfast)
Period
2025
My role
Sole researcher: data preparation, model training, field validation, practitioner survey
Five inspection photographs with U-Net crack masks drawn in red along the cracks
U-Net (ResNet-50) crack segmentation on inspection images (dissertation Fig. 18).
0.807
Box mAP50
YOLOv8x-seg · Crack-Seg test
0.630
mIoU (F1 0.775)
U-Net ResNet-50 · validation split
0.933
Field recall
YOLO11x-seg · 14 of 15 cracked images
27 min
689 images, full pipeline + report
≈2.38 s per image

Chapter 01

The engineering challenge

Visual inspection of reinforced concrete is slow, subjective and often done at height. In a survey I ran, the 21 professionals who responded said a medium-sized building typically takes 4–7 days to inspect.

Published crack detectors report high benchmark scores. Site photos, however, are full of pipes, cables, coatings and shadows that look like cracks, so benchmark accuracy says little about field reliability.

Chapter 02

Input data

Crack-Seg benchmark

3,717 train · 112 val · 200 test images, 416 × 416

DACL10k

9,920 bridge-inspection images, 18 defect classes, polygon masks

Industry site photos

689 images from live projects in Belfast: pipes, coatings, shadows

Practitioner survey

21 professionals across the UK, KSA, UAE and Egypt

Three site photos with crack masks from three YOLO models side by side
External validation: the same site images through YOLO11x-seg, YOLO11n-seg and YOLOv8x-seg.

Chapter 03

The AI & engineering workflow

Stage 1 of 5

Data

Crack-Seg · DACL10k (9,920 images, 18 classes) · 689 industry site photos

What I did

  • Trained and compared three YOLO-seg variants (YOLOv8x, YOLO11x, YOLO11n) on the Crack-Seg dataset (3,717 / 112 / 200 images).
  • Trained a U-Net with a ResNet-50 encoder for pixel-accurate crack masks suitable for width and area measurement.
  • Extended to 18-class multi-defect segmentation on DACL10k with an FPN–EfficientNet-B4, using a Dice + weighted BCE loss, warm-up and cosine schedules, and mixed precision.
  • Ran an external validation on real site images from Amphora and classified every prediction by hand as TP/FP/TN/FN.
  • Designed and analysed a practitioner survey (UK, KSA, UAE, Egypt) to anchor the work in inspection practice.
PythonPyTorch 2.5Ultralytics YOLOsegmentation-models-pytorchAlbumentationsOpenCVRTX 4060 (8 GB)

Input → output

The field test, one photo at a time.

U-Net · ResNet-50 output on the same photo
Site photo: Beam soffit by a window
InputU-Net · ResNet-50
Site photo
Model output

What to look for

Beam soffit by a window

YOLO11x-seg marks the window mullions; U-Net keeps to the crack on the beam.

Real outputs from the dissertation's external validation on industry site photographs (Belfast). Drag the divider or use the slider to compare with the untouched input.

Chapter 04

Technical result

  • U-Net gave the cleanest crack boundaries (mIoU 0.630, precision 0.762, recall 0.787). It is the right choice when crack width and area must be measured.
  • On site imagery YOLO11x-seg was the most robust detector: it caught 14 of 15 cracked images and correctly cleared 13 crack-free images, where YOLOv8x-seg cleared only 1.
  • The 18-class model extended coverage to rust, spalling, drainage and equipment, which reduces false crack alarms. Thin classes remained weak, and I documented the reasons.
  • The integrated pipeline processed 689 industry images in 27 min 22 s. The same survey put a medium-sized building's manual inspection at 4–7 days.

YOLO-seg on the Crack-Seg test split

Box vs mask mAP50. Box localisation is strong; pixel masks are harder.

0.000.250.500.751.000.807YOLOv8x-seg · Box mAP50: 0.8070.674YOLOv8x-seg · Mask mAP50: 0.674YOLOv8x-seg0.804YOLO11x-seg · Box mAP50: 0.8040.639YOLO11x-seg · Mask mAP50: 0.639YOLO11x-seg0.792YOLO11n-seg · Box mAP50: 0.7920.658YOLO11n-seg · Mask mAP50: 0.658YOLO11n-seg
  • Box mAP50
  • Mask mAP50
View data table
CategoryBox mAP50Mask mAP50
YOLOv8x-seg0.8070.674
YOLO11x-seg0.8040.639
YOLO11n-seg0.7920.658

Source: MSc dissertation (Ulster University, 2025), Table 6

Field validation on 44 real site photos

Image-level results after manual verification (15 cracked, 29 crack-free).

0.000.250.500.751.000.93YOLO11x-seg · Recall: 0.930.47YOLO11x-seg · Precision: 0.470.61YOLO11x-seg · Accuracy: 0.61YOLO11x-seg0.67YOLO11n-seg · Recall: 0.670.33YOLO11n-seg · Precision: 0.330.43YOLO11n-seg · Accuracy: 0.43YOLO11n-seg0.93YOLOv8x-seg · Recall: 0.930.33YOLOv8x-seg · Precision: 0.330.34YOLOv8x-seg · Accuracy: 0.34YOLOv8x-seg
  • Recall
  • Precision
  • Accuracy

YOLOv8x-seg matched YOLO11x-seg on recall but flagged 28 of 29 crack-free images, mostly pipes, shadows and coatings.

View data table
CategoryRecallPrecisionAccuracy
YOLO11x-seg0.930.470.61
YOLO11n-seg0.670.330.43
YOLOv8x-seg0.930.330.34

Source: MSc dissertation (Ulster University, 2025), Table 7

DACL10k per-class IoU: this study vs published benchmark

FPN–EfficientNet-B4 at 384 px, no auxiliary head (mIoU 0.317 vs 0.414 benchmark best).

0.000.200.400.600.80PEquipment0.59PEquipment · This study: 0.59PEquipment · Flotzinger et al. (2023): 0.68Graffiti0.57Graffiti · This study: 0.57Graffiti · Flotzinger et al. (2023): 0.59Hollowareas0.55Hollowareas · This study: 0.55Hollowareas · Flotzinger et al. (2023): 0.54Bearing0.47Bearing · This study: 0.47Bearing · Flotzinger et al. (2023): 0.68Drainage0.47Drainage · This study: 0.47Drainage · Flotzinger et al. (2023): 0.52EJoint0.46EJoint · This study: 0.46EJoint · Flotzinger et al. (2023): 0.47Rust0.36Rust · This study: 0.36Rust · Flotzinger et al. (2023): 0.41Weathering0.36Weathering · This study: 0.36Weathering · Flotzinger et al. (2023): 0.42ACrack0.35ACrack · This study: 0.35ACrack · Flotzinger et al. (2023): 0.47Spalling0.29Spalling · This study: 0.29Spalling · Flotzinger et al. (2023): 0.37Efflorescence0.25Efflorescence · This study: 0.25Efflorescence · Flotzinger et al. (2023): 0.34JTape0.24JTape · This study: 0.24JTape · Flotzinger et al. (2023): 0.36Wetspot0.21Wetspot · This study: 0.21Wetspot · Flotzinger et al. (2023): 0.23Restformwork0.20Restformwork · This study: 0.20Restformwork · Flotzinger et al. (2023): 0.34ExposedRebars0.17ExposedRebars · This study: 0.17ExposedRebars · Flotzinger et al. (2023): 0.39Rockpocket0.10Rockpocket · This study: 0.10Rockpocket · Flotzinger et al. (2023): 0.27Crack0.07Crack · This study: 0.07Crack · Flotzinger et al. (2023): 0.29
  • This study
  • Flotzinger et al. (2023)

Thin, low-pixel classes (crack is ~0.3% of pixels) suffered most at reduced resolution. The dissertation reports this gap openly.

View data table
CategoryThis studyFlotzinger et al. (2023)
PEquipment0.590.68
Graffiti0.570.59
Hollowareas0.550.54
Bearing0.470.68
Drainage0.470.52
EJoint0.460.47
Rust0.360.41
Weathering0.360.42
ACrack0.350.47
Spalling0.290.37
Efflorescence0.250.34
JTape0.240.36
Wetspot0.210.23
Restformwork0.200.34
ExposedRebars0.170.39
Rockpocket0.100.27
Crack0.070.29

Source: MSc dissertation evaluation outputs (per-class metrics)

Throughput per configuration

Seconds per image on an RTX 4060 (lower is faster).

0.000.651.301.952.60Crack U-Net0.53Crack U-Net · s / image: 0.53Crack-only YOLO0.77Crack-only YOLO · s / image: 0.77Multi-class segmentation2.04Multi-class segmentation · s / image: 2.04Full classification + report2.38Full classification + report · s / image: 2.38
View data table
Categorys / image
Crack U-Net0.53
Crack-only YOLO0.77
Multi-class segmentation2.04
Full classification + report2.38

Source: MSc dissertation (Ulster University, 2025), Table 8

Limitations, stated plainly

  • The multi-class model underperformed the published benchmark (0.317 vs 0.414 mIoU). Likely causes: 384 px inputs on an 8 GB GPU, no auxiliary classification head, and limited tuning.
  • The field set is small (44 images). It is a robustness check, not a statistically powered benchmark.
Fig. 01Failure analysis: shadows from steel members and hanging cables flagged as cracks.
Fig. 02Multi-defect predictions on industry inspection images (Fig. 19).
Fig. 03FPN–EfficientNet-B4 multi-defect predictions on industry imagery.
Fig. 04YOLO11x-seg validation batch on Crack-Seg.
Fig. 05U-Net training and validation loss: fast convergence with a small generalisation gap.

Chapter 05

Practical impact & relevance

  • Shows how to choose a model by where it will be deployed. Benchmark mAP alone would have picked the wrong detector for site use.
  • Pairs a fast detector with a precise segmenter and a broad multi-class model. This is the architecture later taken forward in the AECAI product.
EvidenceWhere every figure on this page comes from

Each metric is taken from a primary record, not from a CV. Hover a metric to see its source. The underlying files are available for review at interview.

  • MSc dissertation (Ulster University, 2025)
  • MSc dissertation evaluation outputs (per-class metrics)
  • MSc dissertation figures
  • Ulster University MSc transcript (BLD811: 74)

Site photographs were supplied by Amphora Consulting for this research and appear in the submitted dissertation.

Next case study

AECAI: from inspection models to a working product