AI • Vision • Research

Signals from the machine perception frontier.

Selected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.

Curated from primary sources · refreshed hourly

Latest briefings

Follow the feed ↗
arXiv Computer Vision

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

arXiv:2609.19236v1 Announce Type: new Abstract: Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no objective metrics to quantify scope-navigation performance to evaluate whe…

Read at source
arXiv Computer Vision

Open-vocabulary 3D object detection with promptable segmentation

arXiv:2609.19358v1 Announce Type: new Abstract: Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved…

Read at source
arXiv Computer Vision

Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models

arXiv:2609.19384v1 Announce Type: new Abstract: Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independe…

Read at source
arXiv Computer Vision

ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

arXiv:2609.19463v1 Announce Type: new Abstract: We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which rep…

Read at source
arXiv Computer Vision

AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend

arXiv:2609.19518v1 Announce Type: new Abstract: We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that prog…

Read at source
arXiv Computer Vision

A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

arXiv:2609.19555v1 Announce Type: new Abstract: Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existi…

Read at source
arXiv Computer Vision

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

arXiv:2609.19592v1 Announce Type: new Abstract: This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection mo…

Read at source