AI • Vision • Research

Signals from the machine perception frontier.

Selected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.

Curated from primary sources · refreshed hourly

Latest briefings

Follow the feed ↗
arXiv Computer Vision

M2LG-DG: A Multi-modal Local-Global Domain Generalization Framework for Cross-site Major Depressive Disorder Classification

arXiv:2609.09186v1 Announce Type: new Abstract: Classification models based on resting-state functional magnetic resonance imaging (rs-fMRI) often show lower performance at imaging sites not included during model development, which can limit their use in clinical settings. Domain generalization (DG) addresses this iss…

Read at source
arXiv Computer Vision

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

arXiv:2609.09187v1 Announce Type: new Abstract: Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not…

Read at source
arXiv Computer Vision

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

arXiv:2609.09206v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately reflect the ac…

Read at source
arXiv Computer Vision

Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

arXiv:2609.09300v1 Announce Type: new Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize within a single model. We introduce Video-MOPD-8B, an open-weight model dedicated to video un…

Read at source
arXiv Computer Vision

The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library

arXiv:2609.09368v1 Announce Type: new Abstract: We present the Living Library, an end-to-end framework for transforming fragmented digital archives into governed, conversational, in-person exhibit experiences. Developed and deployed at the Theodore Roosevelt Presidential Library, the framework comprises four layers: d…

Read at source
arXiv Computer Vision

OmniPoint: Universal Monocular Metric Pointcloud from Any Camera

arXiv:2609.09394v1 Announce Type: new Abstract: Recovering metric 3D geometry from monocular images is a fundamental computer vision task, yet current methods remain heavily fragmented by fixed camera model assumptions and inflexible input schemes. We present OmniPoint, a unified framework designed to generalize metri…

Read at source
arXiv Computer Vision

Infra-Bench CLS: A Global, Open-Source Benchmark for Critical Infrastructure Classification with Earth Observation Foundation Models

arXiv:2609.09482v1 Announce Type: new Abstract: Critical infrastructure location data is often incomplete and unevenly distributed globally, especially in developing regions. Earth observation foundation models are proposed as a new step in enabling us to more efficiently understand the natural and built environment…

Read at source
arXiv Computer Vision

RoMa-$\Omega$: What Feed-Forward 3D Models Know About Image Matching

arXiv:2609.09507v1 Announce Type: new Abstract: Learned image matching has experienced significant progress in recent years, culminating in robust and accurate matchers such as RoMa, whose robustness is often attributed to its use of frozen DINO features. In a parallel development, feed-forward reconstruction models…

Read at source