The volume of visual data generated by pervasive cameras and sensing systems has surpassed the zettabyte scale, yet deploying deep neural network object detectors remains constrained by the need for large amounts of manually annotated bounding boxes. In domain-specific scenarios such as agricultural monitoring, where targets are small and environments visually complex, expert annotation becomes a critical scalability bottleneck. We propose a classifier-guided self-training framework that models the box-level accept/reject decision of a human annotator. Starting from a small subset of images annotated by a human, i.e., the labeled seed, an object detector generates candidate bounding boxes on unlabeled images, which are validated by a lightweight crop-level classifier trained to emulate human decisions. Only validated detections are retained as pseudo-labels, explicitly controlling pseudo-label quality and stabilizing iterative retraining. We evaluate the approach on a real-world insect detection dataset collected from stationary orchard cameras. The learned box-level classifier achieves up to 99.4% accuracy and 98.8% F1-score in distinguishing valid detections from false positives. Integrated into self-training, classifier guidance keeps the pseudo-label precision drift close to 1 across iterations. With only 5% labeled data, the framework reaches precision and recall up to 94% and 87% (F1 up to 91%), achieving performance comparable (in terms of F1-score) to a fully supervised baseline trained with 10× more labeled data. These results demonstrate that learned box-level validation can reliably approximate human annotation behavior and significantly improve the stability of self-training in challenging domain-specific settings.

Can a Classifier Replace the Human? Box-Level Surrogates for Low-Annotation Object Detection

Betti Sorbelli, Francesco
;
Das, Papiya;Palazzetti, Lorenzo
;
Pinotti, Cristina M.
2026

Abstract

The volume of visual data generated by pervasive cameras and sensing systems has surpassed the zettabyte scale, yet deploying deep neural network object detectors remains constrained by the need for large amounts of manually annotated bounding boxes. In domain-specific scenarios such as agricultural monitoring, where targets are small and environments visually complex, expert annotation becomes a critical scalability bottleneck. We propose a classifier-guided self-training framework that models the box-level accept/reject decision of a human annotator. Starting from a small subset of images annotated by a human, i.e., the labeled seed, an object detector generates candidate bounding boxes on unlabeled images, which are validated by a lightweight crop-level classifier trained to emulate human decisions. Only validated detections are retained as pseudo-labels, explicitly controlling pseudo-label quality and stabilizing iterative retraining. We evaluate the approach on a real-world insect detection dataset collected from stationary orchard cameras. The learned box-level classifier achieves up to 99.4% accuracy and 98.8% F1-score in distinguishing valid detections from false positives. Integrated into self-training, classifier guidance keeps the pseudo-label precision drift close to 1 across iterations. With only 5% labeled data, the framework reaches precision and recall up to 94% and 87% (F1 up to 91%), achieving performance comparable (in terms of F1-score) to a fully supervised baseline trained with 10× more labeled data. These results demonstrate that learned box-level validation can reliably approximate human annotation behavior and significantly improve the stability of self-training in challenging domain-specific settings.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11391/1629114
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact