<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>Vision Ledger</title><link>https://weblog.computervisionstack.com/</link><description>Selected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.</description><language>en</language><lastBuildDate>Wed, 02 Sep 2026 12:57:03 +0000</lastBuildDate><item><title>A Cone-Constrained Bilinear Decomposition for Total Scaled-Gradient Variation Models</title><link>https://arxiv.org/abs/2609.00036</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00036</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00036v1 Announce Type: new Abstract: The total scaled-gradient variation (TSGV) regularizer, derived from sparse modeling of piecewise-linear structures, has been shown to preserve edges and corners in image restoration. However, its highly nonconvex and nonlinear nature poses severe computational challenge…</description></item><item><title>Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving</title><link>https://arxiv.org/abs/2609.00111</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00111</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00111v1 Announce Type: new Abstract: We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion plannin…</description></item><item><title>ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training</title><link>https://arxiv.org/abs/2609.00188</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00188</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00188v1 Announce Type: new Abstract: Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of…</description></item><item><title>A Lagrangian View of Flow Matching</title><link>https://arxiv.org/abs/2609.00198</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00198</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00198v1 Announce Type: new Abstract: Modern explicit-time generative models, such as Flow Matching [Lipman et al., 2023] and Rectified Flow [Liu et al., 2023], are typically derived top-down via Optimal Transport and the continuity equation. This standard Eulerian approach focuses on the macroscopic transpo…</description></item><item><title>Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation</title><link>https://arxiv.org/abs/2609.00206</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00206</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00206v1 Announce Type: new Abstract: Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a compositional safety blind spot: videos composed of seemingly benign components can convey harmful meaning when interpreted as a whole. We refer to this phenomenon as Distri…</description></item><item><title>Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM</title><link>https://arxiv.org/abs/2609.00231</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00231</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00231v1 Announce Type: new Abstract: Existing research on object hallucination in multimodal large language models (MLLMs) predominantly attributes the problem to language priors such as over-reliance on textual co-occurrence statistics. We challenge this view by presenting quantitative evidence for a compl…</description></item><item><title>Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning</title><link>https://arxiv.org/abs/2609.00232</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00232</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00232v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmarks. Yet existing evaluations largely assume that every task is valid and answerable. In real-world OCR scenarios, this a…</description></item><item><title>CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction</title><link>https://arxiv.org/abs/2609.00242</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00242</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00242v1 Announce Type: new Abstract: Long-tail autonomous driving failures are often framed as rare-object recognition errors. We argue that this view is incomplete: the decision-critical question is not only whether a model recognizes an unusual object, but whether it infers how that object changes the ego…</description></item><item><title>CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space</title><link>https://arxiv.org/abs/2609.00272</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00272</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00272v1 Announce Type: new Abstract: Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Multimodal scenarios involve images produced by fundamentally different sensing processes, such as multispectral imaging, RG…</description></item><item><title>StreamScout: Learning When to Look Deeper for Streaming Video Understanding</title><link>https://arxiv.org/abs/2609.00291</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00291</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00291v1 Announce Type: new Abstract: Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a bounded memory, yet access that memory using the same fixed-cost procedure for every query…</description></item><item><title>TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening</title><link>https://arxiv.org/abs/2609.00300</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00300</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00300v1 Announce Type: new Abstract: Reducing the review of clearly cancer-negative screening mammograms could lower radiologist workload without compromising cancer detection. We propose a closed-loop threshold-aware training strategy in which the dismissal threshold is recalculated during training and use…</description></item><item><title>Where Should Experience Live? Hierarchical Hebbian Memory for Continual Vision Transformers</title><link>https://arxiv.org/abs/2609.00358</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00358</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00358v1 Announce Type: new Abstract: Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, limiting their ability to organize newly acquired information across different memory timescales. This work proposes \textit{Hierarchical Hebbian Memory}, a three-l…</description></item><item><title>Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation</title><link>https://arxiv.org/abs/2609.00369</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00369</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00369v1 Announce Type: new Abstract: Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constra…</description></item><item><title>FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos</title><link>https://arxiv.org/abs/2609.00377</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00377</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00377v1 Announce Type: new Abstract: We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tool…</description></item><item><title>SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling</title><link>https://arxiv.org/abs/2609.00396</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00396</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00396v1 Announce Type: new Abstract: Histopathological whole slide images (WSIs) are central to cancer diagnosis, but their gigapixel scale, tissue heterogeneity, weak slide-level supervision, sparse diagnostic regions, and multi-scale evidence make robust automated analysis challenging. Multiple instance l…</description></item><item><title>Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models</title><link>https://arxiv.org/abs/2609.00411</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00411</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00411v1 Announce Type: new Abstract: Modern face recognition (FR) owes much of its success to deep neural networks that learn to extract compact identity embeddings from face images. These models are typically trained for identity discrimination, producing embeddings that are highly effective for biometric…</description></item><item><title>Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT</title><link>https://arxiv.org/abs/2609.00447</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00447</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00447v1 Announce Type: new Abstract: Accurate 3D abnormality segmentation in chest CT requires dense spatial supervision, but obtaining expert voxel-level labels is costly. Radiology reports, however, are routinely generated during clinical interpretation and contain instance-specific descriptions that can…</description></item><item><title>SAM3-LoRA: Parameter-Efficient Adaptation of a Concept-Promptable Foundation Model for Multi-Class Structural Defect Segmentation</title><link>https://arxiv.org/abs/2609.00469</link><guid isPermaLink="true">https://arxiv.org/abs/2609.00469</guid><pubDate>Wed, 02 Sep 2026 04:00:00 +0000</pubDate><category>arXiv Computer Vision</category><description>arXiv:2609.00469v1 Announce Type: new Abstract: Promptable segmentation foundation models such as SAM3 accept an open-vocabulary text concept and return every instance matching it, but adapting them to a specialized domain by full fine-tuning is computationally prohibitive for the organizations that would benefit most…</description></item></channel></rss>
