Publications by Roberto Vezzani

Explore our research publications: papers, articles, and conference proceedings from AImageLab.

Tip: type @ to pick an author and # to pick a keyword.

Active filters (Clear): Author: Roberto Vezzani

A Workflow for Cost- and Time-Aware Refueling Itinerary Optimization

Authors: Savarese, Marco; Zaccagnino, Carmine; De Blasi, Antonio; Salici, Giacomo; Cascianelli, Silvia; Vezzani, Roberto; Grazia, Carlo Augusto

The complete workflow of the RI-PIENO framework is presented, a system for refueling itinerary optimization that extends the original PIENO … (Read full abstract)

The complete workflow of the RI-PIENO framework is presented, a system for refueling itinerary optimization that extends the original PIENO design. While prior work introduced the conceptual modules of RI-PIENO, their operational pipeline was not described in detail. This study makes the workflow explicit, covering the end-to-end process from CAN Bus data acquisition and stop detection to the construction of daily trip graphs, refueling optimization, and mileage prediction. By clarifying the sequence of operations, the contribution provides a reproducible and extensible foundation for future research and development.

2026 Relazione in Atti di Convegno

An Investigation on Incremental Learning from Unbalanced Streamed Data

Authors: Borghi, Guido; Graffieti, Gabriele; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

2026 Relazione in Atti di Convegno

BoltNet: An Ultra-Lightweight Convolutional Network for On-Device Plant Species Identification

Authors: Rossi, Daniel; Borghi, Guido; Vezzani, Roberto

Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar … (Read full abstract)

Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar species, and long-tailed observations require real model capacity, while field use constrains memory, latency, and power. Model size is only part of the deployment cost: intermediate activations held in memory during inference and platformdependent execution behavior matter too, so compact recognition must be assessed on target hardware rather than through complexity metrics alone. We present BoltNet, an ultra-lightweight fully convolutional architecture combining a Spatial Redistribution Bottleneck and Logit PreSampling to improve the tradeoff between predictive performance and model size in high-cardinality classification, and report the AccuracyCompression Tradeoff as a complementary diagnostic. On Pl@ntNet300K, BoltNet reaches 0.682 F1-score with 341K parameters (1.37 MB), the highest F1-score among evaluated models below 2 MB and close to substantially larger convolutional backbones. Model-only measurements on a Raspberry Pi 5, Jetson Orin Nano, and Hailo-8 characterize execution across CPU, GPU, and NPU platforms, where BoltNet is the most consistently efficient model, with the best FPS/W on the GPU and NPU and second-best on the CPU. Results on AIDERv2 and CLRS provide secondary evidence of transfer across environmental image-classification tasks. Code available at: https://codeberg.org/danielrossi/BoltNet.

2026 Relazione in Atti di Convegno

Fake3DGS: A Benchmark for 3D Manipulation Detection in Neural Rendering

Authors: Nucci, Davide Di; Catalini, Riccardo; Borghi, Guido; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Recent advances in 3D reconstruction and neural rendering, particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D … (Read full abstract)

Recent advances in 3D reconstruction and neural rendering, particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D scenes and re-render them as highly realistic images. Therefore, security concerns arise regarding the authenticity of 3D content. Despite this threat, 3D fake detection remains largely unexplored in the literature, and most existing work is limited to 2D space. Therefore, in this paper, we formalize the concept of 3D fake detection and introduce Fake3DGS, a dataset of 3D Gaussian splatting scenes and corresponding rendered views, where fake images are produced by controlled manipulations of geometry, appearance, and spatial layout, while preserving high visual realism. Using this benchmark, we demonstrate that current state-of-the-art 2D detectors struggle to distinguish between original and 3D manipulated images. To bridge this gap, we introduce a 3D-aware detection method that leverages multi-view coherence and features derived from the Gaussian splatting representation. Experimental results demonstrate a substantial improvement in recognizing modified 3D content, underscoring the validity of the new dataset and the necessity for authenticity assessment techniques that extend beyond 2D evidence. Code and data are publicly released (https://github.com/iot-unimore/Fake3DGS) for future investigations.

2026 Relazione in Atti di Convegno

GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation

Authors: Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Giampiero; Kawana, Yuki; Vezzani, Roberto

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single … (Read full abstract)

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed

2026 Relazione in Atti di Convegno

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Authors: Simoni, Alessandro; Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Gianpiero; Kawana, Yuki; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods … (Read full abstract)

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models’ generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.

2026 Relazione in Atti di Convegno

TakuNet: Energy-Efficient Models for Real-Time Aerial Disaster Response and Monitoring on Edge Devices

Authors: Rossi, Daniel; Filippini, Gianluca; Torlai, Andrea; Borghi, Guido; Vezzani, Roberto

Published in: IMAGE AND VISION COMPUTING

In this work, we present TakuNet, a family of ultra-lightweight convolutional neural networks designed for realtime aerial image classification on … (Read full abstract)

In this work, we present TakuNet, a family of ultra-lightweight convolutional neural networks designed for realtime aerial image classification on resource-constrained embedded devices. The proposed TakuNetV2 architecture enhances feature extraction and generalization capabilities through the incorporation of a denser stem coupled with hybrid feature extractor blocks, wherein diverse convolutional operations are synergistically combined to yield richer spatial representations without compromising latency or parameter efficiency. We extensively evaluate the TakuNet family on three public aerial image classification datasets against well-known light-weight and ultra-lightweight architectures, and measured relative performance on five heterogeneous embedded platforms, spanning from CPUs to GPUs, and the Hailo-8 NPU. TakuNet achieves state-of-the-art accuracy and energy efficiency, outperforming competing models in frames per second per additional watt consumed, confirming its suitability for battery-powered edge devices. Additionally, this paper introduces the Astrial platform, highlighting its role in enabling efficient deep learning inference on industrial-grade edge applications. Although current NPU hardware and compiler limitations pose challenges, TakuNet sets a new benchmark for efficient, high-performance embedded artificial intelligence in aerial surveillance and emergency response. Code, models, and weights are publicly available https://github.com/DanielRossi1/TakuNetV2.

2026 Articolo su rivista

3D Pose Nowcasting: Forecast the future to improve the present

Authors: Simoni, A.; Marchetti, F.; Borghi, G.; Becattini, F.; Seidenari, L.; Vezzani, R.; Del Bimbo, A.

Published in: COMPUTER VISION AND IMAGE UNDERSTANDING

Technologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last … (Read full abstract)

Technologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last few years. A critical component useful for realizing this collaborative paradigm is the understanding of human and robot 3D poses using non-invasive systems. Therefore, in this paper, we propose a novel vision-based system leveraging depth data to accurately establish the 3D locations of skeleton joints. Specifically, we introduce the concept of Pose Nowcasting, denoting the capability of the proposed system to enhance its current pose estimation accuracy by jointly learning to forecast future poses. The experimental evaluation is conducted on two different datasets, providing accurate and real-time performance and confirming the validity of the proposed method on both the robotic and human scenarios.

2025 Articolo su rivista

AURALYS: smart glasses to improve audio selection and perception in educational and working contexts

Authors: Filippini, Gianluca; Borghi, Guido; Giliberti, Enrico; Damiani, Paola; Vezzani, Roberto

2025 Relazione in Atti di Convegno

BRUM: Robust 3D Vehicle Reconstruction from 360° Sparse Images

Authors: Di Nucci, Davide; Tomei, Matteo; Borghi, Guido; Ciuffreda, Luca; Vezzani, Roberto; Cucchiara, Rita

2025 Relazione in Atti di Convegno
2 3 »

Page 1 of 13 • Total publications: 129