Publications

Explore our research publications: papers, articles, and conference proceedings from AImageLab.

Tip: type @ to pick an author and # to pick a keyword.

Segment-wise Anomaly Detection via Compression Tokens in Industrial Production Lines

Authors: Salici, Giacomo; Köhler, Stefan; Fiorina, Andrea; Zannella, Franco; Porrello, Angelo; Calderara, Simone

We present a predictive maintenance approach for industrial production lines based on multivariate segment-wise time-series analysis. To address the high … (Read full abstract)

We present a predictive maintenance approach for industrial production lines based on multivariate segment-wise time-series analysis. To address the high cost of collecting anomalous samples, we propose a novelty detection framework in which a transformer autoencoder is trained in a semi-supervised fashion exclusively on nominal sequences, and anomaly scores are derived from reconstruction error at test time. We introduce a set of learnable “compression tokens” into the transformer encoder; these tokens serve as the bottleneck from which the decoder reconstructs the input. We compare this model against an MLP-based autoencoder baseline; the results show that the novelty-detection model remains strong, with near-perfect performance under time-aware and device-aware validation, which are the conditions that most faithfully simulate deployment.

2026 Relazione in Atti di Convegno

Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing

Authors: Barsellotti, Luca; Sundermeyer, Martin; Segu, Mattia; Araslanov, Nikita; Ferjad Naeem, Muhammad; Cornia, Marcella; Xian, Yongqin; Berman, Maxim

Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have … (Read full abstract)

Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computational cost of pixel decoding, textual modality fusion, and object decoding to make these architectures more suitable for mobile devices, real-time on-device inference at high frame rates remains an open challenge. In this paper, we introduce SegFS, a dual-stream fast-slow framework that significantly improves efficiency without sacrificing accuracy. On sparse keyframes, an open-vocabulary object-based model predicts instance-level representations. These representations are then projected back into the backbone feature space to condition a lightweight fast network, which efficiently relocalizes and segments the instances in subsequent frames. By shifting instance propagation from object decoding to feature-space conditioning, our approach decouples multimodal semantic understanding from dense mask prediction and enables efficient temporal propagation. The proposed fast branch achieves up to 14x lower latency than the mobile-oriented MOBIUS model, while maintaining competitive segmentation performance on standard OV-VIS benchmarks.

2026 Relazione in Atti di Convegno

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Authors: Zaccagnino, Carmine; Quattrini, Fabio; Simsar, Enis; Tintoré Gazulla, Marta; Cucchiara, Rita; Tonioni, Alessio; Cascianelli, Silvia

Published in: PROCEEDINGS OF MACHINE LEARNING RESEARCH

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering … (Read full abstract)

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors predominantly support global or single-instruction edits and struggle with multi-instance scenarios, where multiple parts of a reference input must be edited independently without semantic interference. We identify this limitation as a consequence of globally conditioned velocity fields and joint attention mechanisms, which entangle concurrent edits. To address this issue, we introduce Instance-Disentangled Attention, a mechanism that partitions joint attention operations, enforcing binding between instance-specific textual instructions and spatial regions during velocity field estimation. We evaluate our approach on both natural image editing and a newly introduced benchmark of text-dense infographics with region-level editing instructions. Experimental results demonstrate that our approach promotes edit disentanglement and locality while preserving global output coherence, enabling single-pass, instance-level editing.

2026 Relazione in Atti di Convegno

Sketch2Stitch: GANs for Abstract Sketch-Based Dress Synthesis

Authors: Farooq Khan, Faizan; Mohamed Bakr, Eslam; Morelli, Davide; Cornia, Marcella; Cucchiara, Rita; Elhoseiny, Mohamed

In the realm of creative expression, not everyone possesses the gift of effortlessly translating their imaginative visions into flawless sketches. … (Read full abstract)

In the realm of creative expression, not everyone possesses the gift of effortlessly translating their imaginative visions into flawless sketches. More often than not, the outcome resembles an abstract, perhaps even slightly distorted representation. The art of producing impeccable sketches is not only challenging but also a time-consuming process. Our work is the first of this kind in transforming abstract, sometimes deformed garment sketches into photorealistic catalog images, to empower the everyday individual to become their own fashion designer. We create Sketch2Stitch, a dataset featuring over 65,000 abstract sketch images generated from garments of DressCode and VITONHD, two benchmark datasets in the virtual try-on task. Sketch2Stitch is the first dataset in the literature to provide abstract sketches in the fashion domain. We propose a StyleGAN-based generative framework that bridges freehand sketching with photorealistic garment synthesis. We demonstrate that our framework allows users to sketch rough outlines and optionally provide color hints, producing realistic designs in seconds. Experimental results demonstrate, both quantitatively and qualitatively, that the proposed framework achieves superior performance against various baselines and existing methods on both subsets of our dataset. Our work highlights a pathway toward AI-assisted fashion design tools, democratizing garment ideation for students, independent designers, and casual creators.

2026 Relazione in Atti di Convegno

SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation

Authors: Testa, Zeno; Baraldi, Lorenzo; Díaz-Rodríguez, Natalia; Furnari, Antonino

2026 Relazione in Atti di Convegno

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Authors: Simoni, Alessandro; Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Gianpiero; Kawana, Yuki; Vezzani, Roberto

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods … (Read full abstract)

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models’ generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.

2026 Relazione in Atti di Convegno

SVGauge: Towards Human-Aligned Evaluation for SVG Generation

Authors: Zini, Leonardo; Frigieri, Elia; Aloscari, Sebastiano; Generali, Marcello; Dodi, Lorenzo; Dosen, Robert; Baraldi, Lorenzo

Published in: LECTURE NOTES IN COMPUTER SCIENCE

Generated Scalable Vector Graphics (SVG) images demand evaluation criteria tuned to their symbolic and vectorial nature – criteria that existing … (Read full abstract)

Generated Scalable Vector Graphics (SVG) images demand evaluation criteria tuned to their symbolic and vectorial nature – criteria that existing metrics such as FID, LPIPS, or CLIPScore fail to satisfy. In this paper, we introduce SVGauge, the first human-aligned, reference-based metric for text-to-SVG generation. SVGauge jointly measures (i) visual fidelity, obtained by extracting SigLIP image embeddings and refining them with PCA and whitening for domain alignment, and (ii) semantic consistency, captured by comparing BLIP-2-generated captions of the SVGs against the original prompts in the combined space of SBERT and TF-IDF. Evaluation on the proposed SHE benchmark shows that SVGauge attains the highest correlation with human judgments and reproduces system-level rankings of eight zero-shot LLM-based generators more faithfully than existing metrics. Our results highlight the necessity of vector-specific evaluation and provide a practical tool for benchmarking future text-to-SVG generation models.

2026 Relazione in Atti di Convegno

TakuNet: Energy-Efficient Models for Real-Time Aerial Disaster Response and Monitoring on Edge Devices

Authors: Rossi, Daniel; Filippini, Gianluca; Torlai, Andrea; Borghi, Guido; Vezzani, Roberto

Published in: IMAGE AND VISION COMPUTING

In this work, we present TakuNet, a family of ultra-lightweight convolutional neural networks designed for realtime aerial image classification on … (Read full abstract)

In this work, we present TakuNet, a family of ultra-lightweight convolutional neural networks designed for realtime aerial image classification on resource-constrained embedded devices. The proposed TakuNetV2 architecture enhances feature extraction and generalization capabilities through the incorporation of a denser stem coupled with hybrid feature extractor blocks, wherein diverse convolutional operations are synergistically combined to yield richer spatial representations without compromising latency or parameter efficiency. We extensively evaluate the TakuNet family on three public aerial image classification datasets against well-known light-weight and ultra-lightweight architectures, and measured relative performance on five heterogeneous embedded platforms, spanning from CPUs to GPUs, and the Hailo-8 NPU. TakuNet achieves state-of-the-art accuracy and energy efficiency, outperforming competing models in frames per second per additional watt consumed, confirming its suitability for battery-powered edge devices. Additionally, this paper introduces the Astrial platform, highlighting its role in enabling efficient deep learning inference on industrial-grade edge applications. Although current NPU hardware and compiler limitations pose challenges, TakuNet sets a new benchmark for efficient, high-performance embedded artificial intelligence in aerial surveillance and emergency response. Code, models, and weights are publicly available https://github.com/DanielRossi1/TakuNetV2.

2026 Articolo su rivista

Tecniche avanzate di Intelligenza Artificiale per l’apprendimento continuo e robusto su dati strutturati

Authors: Menabue, Martin

I metodi di Intelligenza Artificiale hanno raggiunto risultati notevoli in diversi ambiti, ma la loro applicazione efficace a dati dinamici … (Read full abstract)

I metodi di Intelligenza Artificiale hanno raggiunto risultati notevoli in diversi ambiti, ma la loro applicazione efficace a dati dinamici e strutturati rimane una sfida significativa. Questa tesi indaga tecniche avanzate di IA per l’apprendimento continuo e robusto in scenari in cui i dati evolvono nel tempo e presentano complesse dipendenze. La ricerca esplora diverse direzioni complementari per affrontare le limitazioni dei modelli attuali in termini di adattabilità e resilienza. In primo luogo, vengono studiati metodi di apprendimento continuo per consentire alle reti neurali di apprendere da flussi sequenziali di dati senza dimenticare le conoscenze acquisite in precedenza. Viene proposto un approccio basato sulla distillazione che sfrutta i Vision Transformer, in cui le rappresentazioni di attenzione vengono trasferite tra modelli teacher e student, migliorando la stabilità. Inoltre, viene sviluppata una strategia di prompt learning basata sugli embedding del modello CLIP, che seleziona dinamicamente prompt specifici per ciascun task, migliorando le prestazioni. La seconda linea di ricerca della tesi riguarda il federated learning, un contesto distribuito in cui le informazioni strutturate emergono naturalmente dalla collaborazione tra i client. Viene introdotto un nuovo meccanismo di difesa contro gli attacchi backdoor, che sfrutta le proprietà spettrali delle rappresentazioni locali dei dati per identificare e mitigare i partecipanti malevoli attraverso tecniche di sintesi e allineamento dei dati. Infine, la tesi analizza attacchi backdoor adattivi e le relative difese, sottolineando come tali vulnerabilità rappresentino una minaccia critica per i processi e le infrastrutture industriali. Nel complesso, il lavoro contribuisce alla progettazione di modelli di IA capaci di adattamento continuo, collaborazione sicura e sfruttamento efficace delle informazioni strutturali per applicazioni reali e industriali.

2026 Tesi di dottorato

The aporetic dialogs of Modena on gender differences: Is it all about testosterone? Episode III: Mathematics

Authors: Brigante, G.; Costantino, F.; Bellelli, A.; Boni, S.; Furini, C.; Cucchiara, R.; Simoni, M.

Published in: ANDROLOGY

This report is the transcript of what was discussed in a convention at the Endocrinology Unit in Modena, Italy, in … (Read full abstract)

This report is the transcript of what was discussed in a convention at the Endocrinology Unit in Modena, Italy, in the form of the aporetic dialogs of ancient Greece. It is the third episode of a series of four discussions on the differences between males and females, with a multidisciplinary approach. In this work, the role of testosterone in gender differences in the aptitude for mathematics is explored. First, the definitions of mathematical abilities were provided together with any gender difference in the distribution of females and males in science, technology, engineering, and mathematics subjects. A clear predominance of males is evident at most science, technology, engineering, and mathematics education levels, especially in advanced academic careers. Then, the discussants were divided into two groups: group 1, which illustrated the thesis that testosterone promotes the development of logical‒mathematical skills, and group 2, which, in contrast, asserted the inconsistency of a direct role of testosterone in improving cognitive abilities and that socio-cultural factors should be considered on the basis of this gender gap. In the end, an expert referee (a female engineer) tried to resolve the aporia: are the two theories equivalent or is one superior?.

2026 Articolo su rivista

Page 7 of 114 • Total publications: 1132