AI • Vision • Research

Signals from the machine perception frontier.

Selected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.

Curated from primary sources · refreshed hourly

Latest briefings

Follow the feed ↗
arXiv Computer Vision

Synthetic Leprosy Image Generation Using Mask-Conditioned Latent Diffusion and Transfer Learning from Large Chronic Wound Datasets

arXiv:2609.13226v1 Announce Type: new Abstract: Machine learning for neglected tropical diseases is limited by data, not algorithms: public annotated image sets for leprosy (Hansen's disease) number in the hundreds, orders of magnitude below what generative models require. We ask whether a model trained on abundant ch…

Read at source
arXiv Computer Vision

What Does the Encoder Actually Decide? A Controlled Comparison of Vision Backbones on Joint Tree Segmentation and Stereo Depth

arXiv:2609.13232v1 Announce Type: new Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a vision backbone chosen by reputation rather than measurement. Holding dataset, decoders, losses, schedule, and evaluati…

Read at source
arXiv Computer Vision

Occlusal Geometry in Closed Form for Orthodontic Report Generation

arXiv:2609.13237v1 Announce Type: new Abstract: Orthodontic report generation from intraoral data is normally cast as multimodal captioning, yet the released Bite2Text scan pairs are supplied already registered in occlusion, which makes several core occlusal quantities directly measurable rather than inferable. The sy…

Read at source
arXiv Computer Vision

Pixel-wise Planarity for High-Precision Monocular Plane Segmentation

arXiv:2609.13246v1 Announce Type: new Abstract: Plane segmentation from a single RGB image remains challenging due to imprecise region grouping and geometrically inconsistent supervision, often leading to over-segmentation and false planar detections. We propose instead a pixel-wise planarity prediction framework for…

Read at source
arXiv Computer Vision

(How) Do MLLMs Report Bistable Images Like Humans?

arXiv:2609.13254v1 Announce Type: new Abstract: Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and wh…

Read at source
arXiv Computer Vision

Interpretable Temporal Video Reasoning with EventGraph and EventField

arXiv:2609.13258v1 Announce Type: new Abstract: We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyph view. On a calibrated EPIC-KITCHENS subset of 10 videos and 50 temporal reasoning questions, EventField+Glyph achieves 0…

Read at source