A connectomics milestone: Mapping the complete male fruit fly brain
General Science
Read at sourceSelected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.
General Science
Read at sourcearXiv:2609.01659v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an answer. In autonomous driving, the answer is a continuous action. Thus its reasoning must share the same spatiotemporal structure as the physical world. This sur…
Read at sourcearXiv:2609.01683v1 Announce Type: new Abstract: Vision models deployed on microcontrollers (MCUs) are quantized to integer-only arithmetic and run in inference-only runtimes that do not carry the machinery backpropagation needs: the standard tool for adapting a model to the distribution shift (sensor noise, blur, ligh…
Read at sourcearXiv:2609.01691v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FAIRLENS, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three high-stakes domains: hiring, legal, and health…
Read at sourcearXiv:2609.01725v1 Announce Type: new Abstract: Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visual event) and when (position for inserting the AD) to narrate, to achiev…
Read at sourcearXiv:2609.01738v1 Announce Type: new Abstract: Unexploded ordnance (UXO) continues to restrict civilian access, agricultural activity, infrastructure recovery, and environmental remediation in contaminated areas around the world. This study created a multi campaign UAV thermal image data set of inert ordnance, develo…
Read at sourcearXiv:2609.01740v1 Announce Type: new Abstract: Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compresse…
Read at sourcearXiv:2609.01742v1 Announce Type: new Abstract: Anti-UAV systems increasingly fuse multiple sensors, yet their detection heads provide no per-modality reliability signal. This study evaluates whether evidential deep learning (EDL) heads, Dempster-Shafer (DS) evidence fusion, and uncertainty-driven temporal sensor gati…
Read at sourcearXiv:2609.01743v1 Announce Type: new Abstract: Edge vision models are difficult to deploy on resource-constrained hardware, making low-bit post-training quantization (PTQ) attractive. In practice, standard FP32 training often produces heavy-tailed activation distributions whose outliers destabilize activation quantiz…
Read at sourcearXiv:2609.01749v1 Announce Type: new Abstract: Modern generative models, such as GANs, diffusion architectures, and autoregressive systems, now produce facial images that are nearly indistinguishable from authentic photographs. This capability makes detecting forged images increasingly difficult, raising serious conc…
Read at sourcearXiv:2609.01757v1 Announce Type: new Abstract: Vision-Language Pretrained Models (VLPMs) offer a scalable path to open-vocabulary chest radiology understanding, yet two aspects remain underexplored: how structured clinical semantics extracted from medical reports can reduce in-batch noise during contrastive learning…
Read at sourcearXiv:2609.01778v1 Announce Type: new Abstract: Large-scale video retrieval requires embedding models to encode long and diverse videos under tight visual-input and inference budgets. Existing methods typically sample a small, fixed set of frames at their original resolution, limiting temporal coverage and ignoring fr…
Read at sourcearXiv:2609.01786v1 Announce Type: new Abstract: Hyperspectral image classification still relies heavily on random pixel splits within a single scene. The Salinas dataset, randomly split, is among the most widely used datasets for comparing different architectures. However, under a random split method, a large fraction…
Read at sourcearXiv:2609.01787v1 Announce Type: new Abstract: Vision Transformers (ViT) excel in semantic understanding but fail to discriminate between object instances (e.g., identical embeddings for two dogs), limiting their use in instance-level tasks such as object detection and instance segmentation. We propose Contrastive Vi…
Read at sourcearXiv:2609.01795v1 Announce Type: new Abstract: Vision-language object detectors (VLODs) achieve strong zero-shot performance but remain vulnerable to distribution shifts during deployment. Mean-teacher methods for test-time adaptation (TTA) can improve robustness by updating a student model using teacher-generated ps…
Read at sourcearXiv:2609.01800v1 Announce Type: new Abstract: Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) is a task largely dominated by deep neural networks (DNNs). Most SAS-ATR models use convolutional neural network (CNN) architectures whereas transformer-based architectures have had much less representa…
Read at sourcearXiv:2609.01806v1 Announce Type: new Abstract: Shadow removal is an important preprocessing step for many vision tasks, yet existing supervised methods require paired shadow and shadow-free images, while unsupervised approaches often still rely on shadow masks or shadow-free references. We propose ShadowCLR, an unsup…
Read at sourcearXiv:2609.01808v1 Announce Type: new Abstract: Reliable characterization of structural defects requires methods capable of resolving both directly observable surface damage and damage that is not visible from the inspected surface. This study presents two complementary non-contact, vision-based approaches for the det…
Read at source