arXiv:2609.13190v1 Announce Type: new Abstract: Continuous cuffless blood pressure (BP) monitoring using photoplethysmography (PPG) offers a promising solution for personalized healthcare. However, existing methods have two major limitations. Handcrafted feature-based approaches rely on precise fiducial point detectio…
arXiv:2609.13225v1 Announce Type: new Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing wh…
arXiv:2609.13226v1 Announce Type: new Abstract: Machine learning for neglected tropical diseases is limited by data, not algorithms: public annotated image sets for leprosy (Hansen's disease) number in the hundreds, orders of magnitude below what generative models require. We ask whether a model trained on abundant ch…
arXiv:2609.13228v1 Announce Type: new Abstract: Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-ca…
arXiv:2609.13232v1 Announce Type: new Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a vision backbone chosen by reputation rather than measurement. Holding dataset, decoders, losses, schedule, and evaluati…
arXiv:2609.13233v1 Announce Type: new Abstract: Thin structures such as tree branches are among the hardest cases for stereo matching: a branch is only a few pixels wide, the background is cluttered, and dense ground truth for real branches is nearly impossible to label by hand. We make three contributions. First, EMC…
arXiv:2609.13237v1 Announce Type: new Abstract: Orthodontic report generation from intraoral data is normally cast as multimodal captioning, yet the released Bite2Text scan pairs are supplied already registered in occlusion, which makes several core occlusal quantities directly measurable rather than inferable. The sy…
arXiv:2609.13239v1 Announce Type: new Abstract: Diffusion models represent one of the most advanced paradigms in generative modeling. Leveraging their development, a growing number of style transfer methods based on diffusion models have been proposed. However, among these methods, multi-image style transfer approache…
arXiv:2609.13240v1 Announce Type: new Abstract: The AffectiveArt Multidimensional Art Emotion Understanding task asks to jointly predict an artwork's fine-grained emotion (12 classes, 1549:1 head-to-tail ratio), binary valence/arousal, and five attribute-grounded descriptions -- sub-tasks that exhibit strong empirical…
arXiv:2609.13245v1 Announce Type: new Abstract: Speculative Jacobi Decoding (SJD) is an important approach for accelerating autoregressive image generation. Although SJD has shown superior performance, recent studies point out that it usually suffers from a token ambiguity issue during token verification but its reaso…
arXiv:2609.13246v1 Announce Type: new Abstract: Plane segmentation from a single RGB image remains challenging due to imprecise region grouping and geometrically inconsistent supervision, often leading to over-segmentation and false planar detections. We propose instead a pixel-wise planarity prediction framework for…
arXiv:2609.13250v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) cannot process every frame of a long video because of limitations in visual-token and computational budgets. Three primary approaches have been proposed to enhance their long-video understanding capabilities: (i) Retraining an MLL…
arXiv:2609.13251v1 Announce Type: new Abstract: Commercial and advertising images are frequently affected by poor framing, partially cropped subjects, truncated text or logos, and insufficient context, all of which can reduce subject clarity, i.e., the ability of an image to clearly communicate its primary subject. Im…
arXiv:2609.13254v1 Announce Type: new Abstract: Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and wh…
arXiv:2609.13255v1 Announce Type: new Abstract: Facial state analysis plays a crucial role in understanding human expressions, psychological modeling, and human computer interaction. Traditional unimodal vision-based methods are often limited by environmental sensitivity and weak interpretability. Multimodal facial st…
arXiv:2609.13257v1 Announce Type: new Abstract: Test-time scaling (TTS) can improve generation only when additional compute produces better candidates and the system can reliably identify them. This distinction is especially important for video world models, where a wider sample pool may contain stronger rollouts with…
arXiv:2609.13258v1 Announce Type: new Abstract: We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyph view. On a calibrated EPIC-KITCHENS subset of 10 videos and 50 temporal reasoning questions, EventField+Glyph achieves 0…
arXiv:2609.13259v1 Announce Type: new Abstract: Virtual Try-On (VTON) aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turning this preference-oriented goal into an actionable objective relies on a scoring function aligned with human taste. Howeve…