Publications by Guido Borghi

Explore our research publications: papers, articles, and conference proceedings from AImageLab.

Tip: type @ to pick an author and # to pick a keyword.

Active filters (Clear): Author: Guido Borghi

Adaptive-LwF: continual training of morphing attack detector without forgetting

Authors: Pellegrini, Lorenzo; Borghi, Guido; Franco, Annalisa; Maltoni, Davide

Published in: Frontiers in Imaging

2026 Articolo su rivista

An Investigation on Incremental Learning from Unbalanced Streamed Data

Authors: Borghi, Guido; Graffieti, Gabriele; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

2026 Relazione in Atti di Convegno

BoltNet: An Ultra-Lightweight Convolutional Network for On-Device Plant Species Identification

Authors: Rossi, Daniel; Borghi, Guido; Vezzani, Roberto

Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar … (Read full abstract)

Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar species, and long-tailed observations require real model capacity, while field use constrains memory, latency, and power. Model size is only part of the deployment cost: intermediate activations held in memory during inference and platformdependent execution behavior matter too, so compact recognition must be assessed on target hardware rather than through complexity metrics alone. We present BoltNet, an ultra-lightweight fully convolutional architecture combining a Spatial Redistribution Bottleneck and Logit PreSampling to improve the tradeoff between predictive performance and model size in high-cardinality classification, and report the AccuracyCompression Tradeoff as a complementary diagnostic. On Pl@ntNet300K, BoltNet reaches 0.682 F1-score with 341K parameters (1.37 MB), the highest F1-score among evaluated models below 2 MB and close to substantially larger convolutional backbones. Model-only measurements on a Raspberry Pi 5, Jetson Orin Nano, and Hailo-8 characterize execution across CPU, GPU, and NPU platforms, where BoltNet is the most consistently efficient model, with the best FPS/W on the GPU and NPU and second-best on the CPU. Results on AIDERv2 and CLRS provide secondary evidence of transfer across environmental image-classification tasks. Code available at: https://codeberg.org/danielrossi/BoltNet.

2026 Relazione in Atti di Convegno

Evaluating Age Estimation Robustness Under Realistic Facial Occlusions

Authors: Tanveer, Waqar; Franco, Annalisa; Borghi, Guido; Fernández-Robles, Laura; Fidalgo, Eduardo

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Facial age estimation has shown notable progress under controlled conditions. However, in unconstrained real-world environments, accurate age estimation remains challenging. … (Read full abstract)

Facial age estimation has shown notable progress under controlled conditions. However, in unconstrained real-world environments, accurate age estimation remains challenging. This difficulty becomes more severe when facial images contain partial occlusions, as these obstructions hide important age-related information. Moreover, there is no publicly available occluded age estimation dataset to improve performance in real-world scenarios. To overcome this issue, we propose and publicly release three new datasets, FG-NET-O8, APPA-REAL-O8, and MORPH-O8, derived from existing benchmarks. These datasets contain eight types of realistic occlusions, providing a comprehensive testbed for age estimation under occlusions. These occlusions are generated using multiple diffusion-based methods, including Stable Diffusion Realistic Vision, Blended Latent Diffusion, and Fooocus, while preserving the facial identity of each subject. We also design and conduct a human survey to evaluate the quality of the generated occlusions. Furthermore, we test five state-of-the-art age estimation approaches to analyze the impact of real-world occlusions on age estimation performance. Experimental results demonstrate that all approaches exhibit severe performance degradation for nearly all occlusion types across all three datasets.

2026 Relazione in Atti di Convegno

Fake3DGS: A Benchmark for 3D Manipulation Detection in Neural Rendering

Authors: Nucci, Davide Di; Catalini, Riccardo; Borghi, Guido; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Recent advances in 3D reconstruction and neural rendering, particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D … (Read full abstract)

Recent advances in 3D reconstruction and neural rendering, particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D scenes and re-render them as highly realistic images. Therefore, security concerns arise regarding the authenticity of 3D content. Despite this threat, 3D fake detection remains largely unexplored in the literature, and most existing work is limited to 2D space. Therefore, in this paper, we formalize the concept of 3D fake detection and introduce Fake3DGS, a dataset of 3D Gaussian splatting scenes and corresponding rendered views, where fake images are produced by controlled manipulations of geometry, appearance, and spatial layout, while preserving high visual realism. Using this benchmark, we demonstrate that current state-of-the-art 2D detectors struggle to distinguish between original and 3D manipulated images. To bridge this gap, we introduce a 3D-aware detection method that leverages multi-view coherence and features derived from the Gaussian splatting representation. Experimental results demonstrate a substantial improvement in recognizing modified 3D content, underscoring the validity of the new dataset and the necessity for authenticity assessment techniques that extend beyond 2D evidence. Code and data are publicly released (https://github.com/iot-unimore/Fake3DGS) for future investigations.

2026 Relazione in Atti di Convegno

GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation

Authors: Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Giampiero; Kawana, Yuki; Vezzani, Roberto

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single … (Read full abstract)

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed

2026 Relazione in Atti di Convegno

PopEYE - Infrared Ocular Image Dataset for Eye State and Gaze-Direction Classification

Authors: Gibertoni, Giovanni; Borghi, Guido; Rovati, Luigi

The PopEYE dataset is a specialized collection of 14,976 near-infrared (NIR) images of the human eye region, specifically designed to … (Read full abstract)

The PopEYE dataset is a specialized collection of 14,976 near-infrared (NIR) images of the human eye region, specifically designed to support the development and benchmarking of computer vision algorithms for eye-state detection and coarse gaze-direction classification. Each image is provided in a fixed resolution of 772 × 520 pixels in 8-bit grayscale PNG format. The acquisition was performed frontally using a custom-developed Maxwellian-view optical configuration, consisting of a board-level CMOS camera and a specialized lens system where the subject's eye is precisely positioned at the focal point. This setup ensures a high-contrast representation of the anterior segment, making the pupil, iris, limbus, and portions of the sclera and eyelids clearly distinguishable under stable 850 nm infrared illumination. The dataset is categorized into six mutually exclusive classes identified through manual annotation supported by fixed visual aids and an expert system algorithm. The classification includes a correct positioning class for eyes open and properly aligned for clinical measurements (8,160 images), a closed class representing full eye closures such as blinks or sustained lid closure (1,790 images), and four directional classes representing gaze shifts relative to the central optical axis, specifically up (1,379 images), down (1,015 images), left (1,296 images), and right (1,336 images). The data captures the natural anatomical variability of 22 subjects and incorporates common real-world artifacts such as specular reflections from NIR sources and partial pupil occlusions by eyelashes or eyelids. By providing standardized labels and high-resolution NIR imagery, PopEYE serves as a robust resource for training machine learning models intended for real-time patient monitoring during ophthalmic examinations.

2026 Banca dati

Quality-driven Adaptive Morphing Attack Detection in Operational Scenarios via Online Learning

Authors: Domenico, Nicolò Di; Franco, Annalisa; Borghi, Guido; Maltoni, Davide

Morphing Attack Detection (MAD) systems often suffer from performance degradation when deployed in operational environments, such as airports, that differ … (Read full abstract)

Morphing Attack Detection (MAD) systems often suffer from performance degradation when deployed in operational environments, such as airports, that differ from the training domain. We propose an adaptive differential MAD framework that continuously refines a pre-trained detector using live bona fide samples acquired at the gate. The system is memoryless, so no samples are stored in memory to mitigate privacy concerns about the collection of personal data. To prevent the loss of discriminative power caused by bona fide-only adaptation, the method generates synthetic morph samples on-the-fly by combining the current operational subject with identities from an external public or synthetic face dataset. The adaptation process further relies on a quality-aware bona fide selection strategy and a controlled balancing mechanism for synthetic morph generation. Experimental results show that the proposed method improves target-domain specialization while maintaining robustness against morph attacks.

2026 Relazione in Atti di Convegno

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Authors: Simoni, Alessandro; Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Gianpiero; Kawana, Yuki; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods … (Read full abstract)

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models’ generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.

2026 Relazione in Atti di Convegno

TakuNet: Energy-Efficient Models for Real-Time Aerial Disaster Response and Monitoring on Edge Devices

Authors: Rossi, Daniel; Filippini, Gianluca; Torlai, Andrea; Borghi, Guido; Vezzani, Roberto

Published in: IMAGE AND VISION COMPUTING

In this work, we present TakuNet, a family of ultra-lightweight convolutional neural networks designed for realtime aerial image classification on … (Read full abstract)

In this work, we present TakuNet, a family of ultra-lightweight convolutional neural networks designed for realtime aerial image classification on resource-constrained embedded devices. The proposed TakuNetV2 architecture enhances feature extraction and generalization capabilities through the incorporation of a denser stem coupled with hybrid feature extractor blocks, wherein diverse convolutional operations are synergistically combined to yield richer spatial representations without compromising latency or parameter efficiency. We extensively evaluate the TakuNet family on three public aerial image classification datasets against well-known light-weight and ultra-lightweight architectures, and measured relative performance on five heterogeneous embedded platforms, spanning from CPUs to GPUs, and the Hailo-8 NPU. TakuNet achieves state-of-the-art accuracy and energy efficiency, outperforming competing models in frames per second per additional watt consumed, confirming its suitability for battery-powered edge devices. Additionally, this paper introduces the Astrial platform, highlighting its role in enabling efficient deep learning inference on industrial-grade edge applications. Although current NPU hardware and compiler limitations pose challenges, TakuNet sets a new benchmark for efficient, high-performance embedded artificial intelligence in aerial surveillance and emergency response. Code, models, and weights are publicly available https://github.com/DanielRossi1/TakuNetV2.

2026 Articolo su rivista
2 3 »

Page 1 of 10 • Total publications: 92