This study analyzes the performance of two semantic segmentation architectures, DeepLabV3+ (CNN-based) and SegFormer (Transformer-based), applied to precision viticulture, with the aim of understanding their strengths and limitations in real-world scenarios. The investigation focused on the influence of some key parameters (Atrous Rate, segmentation channels, batch size) on two datasets: San Donaci, large but low resolution, and Zeuli, smaller but high definition. The training results show how the size and resolution of the dataset significantly affect performance. SegFormer demonstrated superior efficiency with the MiT-B2 encoder, while DeepLabV3+ produced the best mIoU results with EfficientNet-B4. In testing, all models proved to be effective for the leaves and background classes, but encountered more difficulties with grapes bunches and trunks, often fragmented or poorly visible due to chromatic similarities and lighting conditions. In general, SegFormer with MiT b2 proved to be an interesting choice due to its precision in segmenting edges. In future, combining the two datasets and exploring hybrid CNN-Transformer architectures that exploit the strengths of each could improve the robustness of the segmentation.

A study on the influence of training parameters in CNN and Transformer architectures for semantic segmentation in vineyards / Bono, A., Guaragnella, C., D'Orazio, T.. - (2025), pp. 484-489. (2025 IEEE International Workshop on Metrology for Agriculture and Forestry, MetroAgriFor 2025 ita 2025) [10.1109/metroagrifor66923.2025.11512443].

A study on the influence of training parameters in CNN and Transformer architectures for semantic segmentation in vineyards

Bono, Annaclaudia;Guaragnella, Cataldo;
2025

Abstract

This study analyzes the performance of two semantic segmentation architectures, DeepLabV3+ (CNN-based) and SegFormer (Transformer-based), applied to precision viticulture, with the aim of understanding their strengths and limitations in real-world scenarios. The investigation focused on the influence of some key parameters (Atrous Rate, segmentation channels, batch size) on two datasets: San Donaci, large but low resolution, and Zeuli, smaller but high definition. The training results show how the size and resolution of the dataset significantly affect performance. SegFormer demonstrated superior efficiency with the MiT-B2 encoder, while DeepLabV3+ produced the best mIoU results with EfficientNet-B4. In testing, all models proved to be effective for the leaves and background classes, but encountered more difficulties with grapes bunches and trunks, often fragmented or poorly visible due to chromatic similarities and lighting conditions. In general, SegFormer with MiT b2 proved to be an interesting choice due to its precision in segmenting edges. In future, combining the two datasets and exploring hybrid CNN-Transformer architectures that exploit the strengths of each could improve the robustness of the segmentation.
2025
2025 IEEE International Workshop on Metrology for Agriculture and Forestry, MetroAgriFor 2025
A study on the influence of training parameters in CNN and Transformer architectures for semantic segmentation in vineyards / Bono, A., Guaragnella, C., D'Orazio, T.. - (2025), pp. 484-489. (2025 IEEE International Workshop on Metrology for Agriculture and Forestry, MetroAgriFor 2025 ita 2025) [10.1109/metroagrifor66923.2025.11512443].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11589/306543
Citazioni
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact