AI • Vision • Research

Signals from the machine perception frontier.

Selected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.

Curated from primary sources · refreshed hourly

Latest briefings

Follow the feed ↗
arXiv Computer Vision

ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

arXiv:2609.16284v1 Announce Type: new Abstract: Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an expo…

Read at source
arXiv Computer Vision

Sequence Recognition in Bharatnatyam dance

arXiv:2609.16306v1 Announce Type: new Abstract: Bharatanatyam is the oldest Indian Classical Dance (ICD) which is learned and practiced across India and the world. Adavu is the core of this dance form. There exist 15 Adavus and 58 variations. Each Adavu variation comprises a well-defined set of motions and postures (c…

Read at source
arXiv Computer Vision

Racing in Volume with Flow Ensembles

arXiv:2609.16310v1 Announce Type: new Abstract: Streaming 4D reconstruction has been demonstrated only indoors, on dense camera rigs surrounding subjects that move at human pace. Outdoor 4D reconstruction exists but relies either on cameras mounted on the moving vehicle itself, or on limited-coverage arrays observing…

Read at source
arXiv Computer Vision

Reasoning with Image Generation

arXiv:2609.16409v1 Announce Type: new Abstract: Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direc…

Read at source
arXiv Computer Vision

Counterfactual Reasoning for Robust Visual Question Answering

arXiv:2609.16567v1 Announce Type: new Abstract: Modern Visual Question Answering (VQA) models often exploit spurious correlations in training data, leading to poor out-of-distribution (OOD) generalization due to language bias. Although counterfactual learning has shown promise, existing methods can be improved to bett…

Read at source