AI • Vision • Research

Signals from the machine perception frontier.

Selected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.

Curated from primary sources · refreshed hourly

Latest briefings

Follow the feed ↗
arXiv Computer Vision

From Visual Cues to Spoken Narration: Rethinking Audio Description

arXiv:2609.01725v1 Announce Type: new Abstract: Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visual event) and when (position for inserting the AD) to narrate, to achiev…

Read at source
arXiv Computer Vision

UAV Thermal Imagery for Inert Ordnance Screening: Multi Campaign Dataset Development,Object Detection, and Practical Recommendations

arXiv:2609.01738v1 Announce Type: new Abstract: Unexploded ordnance (UXO) continues to restrict civilian access, agricultural activity, infrastructure recovery, and environmental remediation in contaminated areas around the world. This study created a multi campaign UAV thermal image data set of inert ordnance, develo…

Read at source
arXiv Computer Vision

ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

arXiv:2609.01740v1 Announce Type: new Abstract: Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compresse…

Read at source
arXiv Computer Vision

Evidential Deep Learning for Multi-Modal Anti-UAV Detection

arXiv:2609.01742v1 Announce Type: new Abstract: Anti-UAV systems increasingly fuse multiple sensors, yet their detection heads provide no per-modality reliability signal. This study evaluates whether evidential deep learning (EDL) heads, Dempster-Shafer (DS) evidence fusion, and uncertainty-driven temporal sensor gati…

Read at source
arXiv Computer Vision

AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $\alpha$-Corrected Binary Cross Entropy and Factorized Latent Supervision

arXiv:2609.01757v1 Announce Type: new Abstract: Vision-Language Pretrained Models (VLPMs) offer a scalable path to open-vocabulary chest radiology understanding, yet two aspects remain underexplored: how structured clinical semantics extracted from medical reports can reduce in-batch noise during contrastive learning…

Read at source
arXiv Computer Vision

DESA-TTA: Dynamic EMA and Source Anchoring for Test-Time Adaptation

arXiv:2609.01795v1 Announce Type: new Abstract: Vision-language object detectors (VLODs) achieve strong zero-shot performance but remain vulnerable to distribution shifts during deployment. Mean-teacher methods for test-time adaptation (TTA) can improve robustness by updating a student model using teacher-generated ps…

Read at source
arXiv Computer Vision

Consistency as Regularization for Unsupervised Shadow Removal

arXiv:2609.01806v1 Announce Type: new Abstract: Shadow removal is an important preprocessing step for many vision tasks, yet existing supervised methods require paired shadow and shadow-free images, while unsupervised approaches often still rely on shadow masks or shadow-free references. We propose ShadowCLR, an unsup…

Read at source
arXiv Computer Vision

Integrated Laser Scanning and Image-Based Topology Optimization Techniques for Detection and Quantification of Visible and Subsurface Structural Defects

arXiv:2609.01808v1 Announce Type: new Abstract: Reliable characterization of structural defects requires methods capable of resolving both directly observable surface damage and damage that is not visible from the inspected surface. This study presents two complementary non-contact, vision-based approaches for the det…

Read at source