MilleMiglia: A realistic instance generator for middle-mile logistics
Algorithms & Theory
Read at sourceSelected computer vision, artificial intelligence, and applied research briefings with direct links to the original work.
Algorithms & Theory
Read at sourceA new focus on fiction and memoir aims to help the MIT community celebrate the power of storytelling and strengthen social connection.
Read at sourcearXiv:2609.19230v1 Announce Type: new Abstract: Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we present SonoCorpus, an open resource unifying 456,963 images and 1,626,085 ex…
Read at sourcearXiv:2609.19236v1 Announce Type: new Abstract: Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no objective metrics to quantify scope-navigation performance to evaluate whe…
Read at sourcearXiv:2609.19354v1 Announce Type: new Abstract: Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-s…
Read at sourcearXiv:2609.19358v1 Announce Type: new Abstract: Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved…
Read at sourcearXiv:2609.19377v1 Announce Type: new Abstract: Recovering numerical series from line plots requires accurate axis calibration and reliable curve extraction. We present LinePilot Digitizer (LinePilot), which combines continuous color-based curve recovery with three calibration modes: LinePilot (standard), LinePilot (e…
Read at sourcearXiv:2609.19384v1 Announce Type: new Abstract: Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independe…
Read at sourcearXiv:2609.19393v1 Announce Type: new Abstract: Work zones alter lane geometry through temporary traffic controls and closures that may be absent from on-board maps, challenging autonomous vehicle (AV) perception and planning. Generalization is also limited by scarce public datasets with structured geometric supervisi…
Read at sourcearXiv:2609.19421v1 Announce Type: new Abstract: Gaussian Splatting has significantly improved the quality of novel view synthesis with explicit Gaussian representation. However, we observed that existing 3D Gaussian Splatting methods (3DGS) often suffer from surface collapse issues on reflective regions, and thus prod…
Read at sourcearXiv:2609.19444v1 Announce Type: new Abstract: Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, and atubular glomeruli. Supervised classification requires labeled examp…
Read at sourcearXiv:2609.19451v1 Announce Type: new Abstract: The Mobile Unified Multimodal Understanding (MUMU) Challenge requires a single efficient model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. We present Efficient Unified Multimodal Understanding (EUMU), the winnin…
Read at sourcearXiv:2609.19463v1 Announce Type: new Abstract: We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which rep…
Read at sourcearXiv:2609.19483v1 Announce Type: new Abstract: Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as predic…
Read at sourcearXiv:2609.19518v1 Announce Type: new Abstract: We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that prog…
Read at sourcearXiv:2609.19542v1 Announce Type: new Abstract: Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present PerSeM, a training-free persistent semantic memory framework for long-ho…
Read at sourcearXiv:2609.19555v1 Announce Type: new Abstract: Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existi…
Read at sourcearXiv:2609.19592v1 Announce Type: new Abstract: This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection mo…
Read at source