本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。
统计
今日共更新1056篇论文,其中:
- 自然语言处理144篇
- 信息检索19篇
- 计算机视觉210篇
自然语言处理
1. 【2610.12427】FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
链接:https://arxiv.org/abs/2610.12427
作者:Yuxuan Hu,Weikang Shi,Yang Bo,Xudong Lu,Xintong Guo,Shuhan Li,Yuyang He,Huankang Guan,Peiwen Sun,Yunqiao Yang,Wenbo Li,Rui Liu,Hongsheng Li
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Large Language Models, Video Large Language, Large Language, existing benchmarks focus, Streaming Video Large
备注:
点击查看摘要
Abstract:Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios. Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events. We introduce FastBench to evaluate high-dynamic perception in real-world video streams. Its trajectory-grounded pipeline combines QA generation from high-FPS clips, filtering of questions answerable at 2 FPS, answer verification using SAM3 and CoTracker3 trajectories, and three rounds of human inspection. FastBench contains 306 QA pairs across eight domains, six capabilities, and forward, instant, and backward temporal scopes, with human-annotated evidence intervals. We also present ProactiveFrame, a training-free baseline that adjusts incoming frame rates through text tokens. A dual-tier sliding window retains recent high-FPS observations while downsampling older ones into sparse history. Experiments reveal substantial limitations: the strongest model, Gemini-3.5-Flash, scores only 50.7%. Denser sampling improves Qwen3-VL-8B from 32.9% at 2 FPS to 44.6% at 24 FPS, but gains saturate as history is compressed. ProactiveFrame outperforms sparse uniform sampling by 5.4 and 1.5 percentage points, yet remains well below oracle-guided focusing, showing that current VLMs struggle to determine from the stream alone when finer temporal perception is needed. FastBench provides a testbed for high-dynamic streaming video understanding. Code and data: this https URL.
2. 【2610.12417】WOVEN: Weaving Visual World Modeling into Multimodal LLMs
链接:https://arxiv.org/abs/2610.12417
作者:Zheyu Fan,Yue Zhang,Mingkai Deng,Kangrui Wang,Qineng Wang,Canyu Chen,Jie Hao,Xing Fan,Chenlei Guo,Eric P. Xing,Mohit Bansal,Manling Li
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Multimodal large language, Multimodal large, large language models, visual transition reasoning, struggle with spatial
备注:
点击查看摘要
Abstract:Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
3. 【2610.12410】Predicting Alignment Generalization with Value Representations
链接:https://arxiv.org/abs/2610.12410
作者:Andy Liu,Mehar Bhatia,Karolina Stanczak,Mona Diab,Vered Shwartz,Daniel Fried
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:LLM developers post-train, behavioral traits, alignment, developers post-train, exhibit prosocial
备注:
点击查看摘要
Abstract:LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.
4. 【2610.12403】ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
链接:https://arxiv.org/abs/2610.12403
作者:Hongxing Li,Dingming Li,Yixin Li,Yong Du,Wenqi Zhang,Weiming Lu,Jun Xiao,Yueting Zhuang,Yongliang Shen
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Skill-augmented agents improve, improve sample efficiency, agents improve sample, Skill-augmented agents, reusable strategies
备注: Code: [this https URL](https://github.com/ZJU-REAL/ViSkill)
点击查看摘要
Abstract:Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at this https URL.
5. 【2610.12402】SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
链接:https://arxiv.org/abs/2610.12402
作者:Hongxing Li,Jinyue Su,Dingming Li,Wenqi Zhang,Weiming Lu,Jun Xiao,Yueting Zhuang,Yongliang Shen
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Existing spatial reasoning, Existing spatial, reading off relations, relations already visible, test spatial perception
备注: Code: [this https URL](https://github.com/ZJU-REAL/SpaceCast-Bench) Dataset: [this https URL](https://huggingface.co/datasets/hongxingli/SpaceCast-Bench)
点击查看摘要
Abstract:Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
6. 【2610.12390】Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search
链接:https://arxiv.org/abs/2610.12390
作者:Ziming Dai,Dabiao Ma,Ziheng Guo,Jack Dong,Zimu Zhou
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Industrial risk-control systems, systems typically rely, substantial valuable information, valuable information remains, information remains embedded
备注:
点击查看摘要
Abstract:Industrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manual feature engineering is labor-intensive, while requiring a large language model (LLM) to process every real-time input may not meet practical deployment requirements. To address this challenge, we propose LLM-BlockFE, an LLM-guided offline feature construction framework that converts long text into executable feature programs, thereby avoiding LLM calls during online inference. LLM-BlockFE constructs feature programs by incrementally appending immutable code blocks and evaluates candidate features using a downstream model. To address the tendency of conventional greedy search to become trapped in suboptimal solutions, our method introduces a block-level rollback mechanism based on depth-calibrated credit allocation and advances multiple independent search trajectories in an interleaved manner, reducing redundant exploration by sharing fixed descriptions of each trajectory's exploration direction. After the search, the resulting programs are frozen and deployed to extract structured features for downstream prediction models. Across two public and two private datasets, LLM-BlockFE achieves absolute AUC improvements of 0.0069 to 0.0358 over the strongest baseline on each dataset in the full-dataset comparison. Post-launch monitoring across five deployed financial risk-control applications shows absolute KS improvements of 0.02 to 1.56 percentage points over the existing manually designed strategy.
7. 【2610.12376】Latent Core Tokenizer: Compress, but Meaningfully
链接:https://arxiv.org/abs/2610.12376
作者:Felermino D. M. A. Ali,Millicent Ochieng,Ogbemi Ekwejunor-Etchie,Ade Famoti,Jacki O'Neill,Debjit Paul
类目:Computation and Language (cs.CL)
关键词:Latent Core Tokenizer, Minimum Description Length, Core Tokenizer, commonly optimized, necessarily distribute
备注: Under review
点击查看摘要
Abstract:Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.
8. 【2610.12375】OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
链接:https://arxiv.org/abs/2610.12375
作者:Babak Barazandeh,Connor Swanson,Chinmay Kulkarni,Nikhil Mungel
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
关键词:incident triage, LLM agents work, deployed in applications, applications from trip, trip planners
备注:
点击查看摘要
Abstract:Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior or evaluating logs post-hoc. The first adds cost and latency to every step; the second delivers its verdict after the run, when tokens are burned and damage is done. To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step. We study this problem in three regimes of decreasing access: full reference access (historical runs and tool schemas), intermediate access (only tool schemas), and no prior knowledge (only step logs as generated). Expectation of OnTrack's monitoring capabilities reduces as data access drops, ranging from plan violation detection to identifying loops, stalls, and repeated tool calls. Finally, we evaluate OnTrack using SWE-bench trajectories. Based on the first 8 steps, our method ranks failing trajectories below succeeding ones better than content similarity approaches (+0.057 AUROC). With an abort policy, we save about 18% of compute that would be burned on failing runs, where 83% of interrupted runs were actually heading to failure (5 out of 6 aborts were correct).
9. 【2610.12367】Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
链接:https://arxiv.org/abs/2610.12367
作者:Yuhan Liu,Xiyao Ma,Zhongkai Sun,Xu Han,Chengyuan Ma,Benjamin Z. Yao,Chenlei Guo
类目:Computation and Language (cs.CL)
关键词:LLM downstream performance, improve LLM downstream, substantially improve LLM, reusable procedural guidance, procedural guidance added
备注:
点击查看摘要
Abstract:Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026). Prior work retrieves skills from a bank by semantic relevance, then uses them as inference-time patches or for model distillation. The individual utility of each skill, however, is largely neglected. We first show that, in on-policy distillation where skill-conditioned policies serve as teachers, fewer than 25% of retrieved skills provide useful distillation signals. We then propose SGUID, a method for selecting a compact subset of skills for distillation. SGUID retains a skill only if it consistently yields effective learning signals during training. The selected skills are then distilled to produce a better model. Our results show that not all skills are worth distilling. Across four models from the Olmo and Qwen families, distilling 6 selected skills matches or exceeds full-bank distillation in mean avg@12 on three of the four models, and on all four after a second round that distills 3 newly selected skills, while the full banks are up to 11x larger. Importantly, SGUID supports stable model-skill co-evolution: after a distillation round, a new candidate bank is curated from the updated model's rollouts, and SGUID selects which skills to internalize next. In the second round, this loop selects 3 new skills and improves Qwen3-8B from 64.3% to 66.3%. The selection step is essential for stability: on Qwen3-4B, naively updating the model with unfiltered skills degrades performance, including a 0.3 percentage point drop on HMMT25, whereas SGUID improves HMMT25 by 0.5 points after the first round and 1.1 points after the second. These results identify skill selection as the key mechanism for stable model-skill co-evolution.
10. 【2610.12361】Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
链接:https://arxiv.org/abs/2610.12361
作者:Saisab Sadhu,Shreeyans Arora,Pratinav Seth
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
关键词:Large language models, Large language, language models increasingly, justify legal decisions, treated as evidence
备注:
点击查看摘要
Abstract:Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.
11. 【2610.12360】Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
链接:https://arxiv.org/abs/2610.12360
作者:Kaiser Sun,Bernal Jimenez Gutierrez,Hongjun Liu,Jingyu Zhang,Jie Gao,Mark Dredze,Daniel Khashabi
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:agent prior beliefs, prior beliefs, retrieved evidence contradicts, acknowledge uncertainty, incorrect conclusion
备注: EMNLP 2026 Camera Ready
点击查看摘要
Abstract:When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.
12. 【2610.12345】Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution
链接:https://arxiv.org/abs/2610.12345
作者:Haohui Wang,Jiahao Xu,Wangzhi Zhan,Tong Zeng,Dongqi Fu,Hong Li,Swastik Roy,Naren Ramakrishnan,Chris North,Jian Kang,Yujun Yan,Dawei Zhou
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Supervised fine-tuning, adapts pretrained large, pretrained large language, large language models, downstream tasks
备注:
点击查看摘要
Abstract:Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel notion named prior barrier to quantify how strongly the pretrained model supports competing concepts over the target concept. We observe that prior barriers follow a long-tail distribution, placing head and tail concepts at different starting points for SFT: head concepts face lower prior barriers, whereas tail concepts require additional instructions to overcome their higher prior barriers. Our theoretical analysis further derives a predictive risk bound for SFT under long-tail prior barriers, explicitly characterizing how the prior barrier and accumulated SFT evidence jointly determine predictive performance. Motivated by this prior barrier-dependent demand, we propose PASS, an adaptive SFT instruction selection method that constructs reference-derived concepts and estimates the distinguishing evidence provided by each instruction, and adaptively allocates the selection budget toward concepts that remain insufficiently covered under the current selection. In this way, PASS jointly considers which instructions can provide useful evidence and where additional supervision is needed under a limited budget. Experiments show that our method consistently outperforms seven state-of-the-art instruction selection methods on four backbone-budget settings. An ablation study further shows that PASS's adaptive allocation consistently improves over uniform allocation.
13. 【2610.12341】Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
链接:https://arxiv.org/abs/2610.12341
作者:Kaisen Yang,Qingle Liu,Kejin Wang,Yicheng Zhao,Jieming Li,Shenghan Zheng,Ruize Yang,Bojun Yang,Heng Gong,Xiang Gao,Lanyue Zhang,Kaiyu Zhong,Zhuo Liu,Shaoxuan Li,Chengxi Li,Yong Yan,Weixuan Zhang,Tianwei Luo,Situ Wang,Youjie Zheng,Sihan Zhao,Shengyuan Wang,Huan-ang Gao,Jiazheng Xu,Xiaohui Xie,Wentao Han,Hongning Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:samples remain challenging, limited samples remain, Adversarial Heuristic Learning, heuristic learning, remain challenging
备注:
点击查看摘要
Abstract:Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
14. 【2610.12338】VFold: Symmetry-Aware Cross-Layer Value Cache Compression
链接:https://arxiv.org/abs/2610.12338
作者:Neha Verma,Sungwon Kim,Kenton Murray,Kevin Duh
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large Language Model, accelerates Large Language, states accelerates Large, Language Model, Large Language
备注:
点击查看摘要
Abstract:While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.
15. 【2610.12327】SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
链接:https://arxiv.org/abs/2610.12327
作者:Qitong Wang,Xinwei Niu,Mingluo Su,Shanwei Zhao,Shiai Zhu,Huan Wang
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:inference incurs significant, incurs significant latency, large language model, inference incurs, significant latency
备注:
点击查看摘要
Abstract:The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding. Nevertheless, typical methods in this line compute the Hessian using pre-collected natural sequences, whereas the model is fed self-generated tokens during decoding, creating a distribution shift between the two sequences. The Hessian calculated on the natural sequence is different from that calculated on the generated sequence. We observe that this discrepancy causes the activation distribution during generation to deviate from that used for pruning, further hurting the pruned model performance. Moreover, most existing LLM pruning methods that bring actual speedup primarily target the sparse matrix-matrix (SpMM) multiplication, providing limited support for the sparse matrix-vector (SpMV) operations, which dominate decoding. To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding. Specifically, at the algorithmic axis, SparseDecoding constructs calibration matrices from layer-wise activations collected during the dense-model autoregressive generation, excluding prefill, thereby aligning the pruning objective with the decoding activations. At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal. Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.
16. 【2610.12313】Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
链接:https://arxiv.org/abs/2610.12313
作者:Saisab Sadhu,Aadit Sengupta,Vinay kumar Sankarapu,Pratinav Seth
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
关键词:Large language model, Large language, language model compliance, model compliance systems, systems are deployed
备注:
点击查看摘要
Abstract:Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.
17. 【2610.12274】HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
链接:https://arxiv.org/abs/2610.12274
作者:Haolin Yang,Jipeng Zhang,Jian Xie,Shuaishuai Gong,Sirui Han,Yike Guo
类目:Computation and Language (cs.CL)
关键词:executing probe queries, static queries, probe queries, map questions directly, inspecting schemas
备注:
点击查看摘要
Abstract:Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.
18. 【2610.12248】EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
链接:https://arxiv.org/abs/2610.12248
作者:Heeseung Kim
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
关键词:Wearable augmented reality, continuous real-world interaction, Wearable augmented, timely spoken guidance, provide timely spoken
备注: Accepted to EMNLP 2026 (Main Conference). 25 pages, 12 figures, 11 tables. Project page: [this https URL](https://egocentricvoice.github.io/)
点击查看摘要
Abstract:Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. While proactive video assistants, spoken dialog systems, and egocentric task understanding have each advanced rapidly, existing systems do not address the joint problem of deciding when to speak and what to say from continuous first-person streams. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation and speech resynthesis, and convert each video session into a format where the model must decide at each moment whether to remain silent or provide spoken guidance. We fine-tune an omni-modal LLM with our data, and further improve its proactive intervention behavior with direct preference optimization. Experiments across closed and open-source models show that existing systems rarely produce well-timed, meaningful proactive interventions, while EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.
19. 【2610.12243】NativeScope: Relation-Localized Retrieval over Native Topology with a Correct Anchor
链接:https://arxiv.org/abs/2610.12243
作者:Long Wang
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:semantic similarity, ranks text chunks, Dense retrieval, percent, text chunks
备注: 14 pages, 4 figures, and 6 tables. Includes an appendix with reproduction information and an evidence inventory
点击查看摘要
Abstract:Dense retrieval usually ranks text chunks by their semantic similarity to a question. This ignores structure that many data systems already store, including section membership, session boundaries, and native order. We propose NativeScope, a scope-then-rank method for queries with a known anchor and relation. It represents a query as q - (A, r, B). The anchor A and relation r select native units through belonging, before, or after operators, and the target term B ranks only chunks that overlap the selected scope. An internal variant, NS-FullQ, ranks the same candidates with the full question. We evaluate both methods on 200 controlled document and memory records derived from QASPER and LongMemEval under a 1,024-token budget. NativeScope attains native-unit recall of 89.28 percent for documents and 72.50 percent for memories, improving over instance-wide Dense RAG by 42.75 and 22.00 percentage points. NS-FullQ reaches 87.78 percent and 68.50 percent; its differences from NativeScope are inconclusive, locating the primary gain in relational scoping rather than the shorter ranking query. With automatic Top-1 anchors, memory recall falls to 35.50 percent. NativeScope is therefore effective when anchor coordinates and native relations are reliable, but hard scoping inherits errors from the localization interface.
20. 【2610.12242】okenRouter: Efficient Serving System for Token-Level LLM Routing
链接:https://arxiv.org/abs/2610.12242
作者:Tianyu Fu,Tengxuan Liu,Ruoxi Wang,Yixin Dong,Yi Ge,Yichen You,Yu Wang
类目:Computation and Language (cs.CL)
关键词:cost-quality Pareto frontier, Large language model, Large language, cost-quality Pareto, Pareto frontier
备注: Accepted by NeurIPS 2026
点击查看摘要
Abstract:Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at this https URL.
21. 【2610.12235】Language Models as AI Research World Models
链接:https://arxiv.org/abs/2610.12235
作者:Zijun Wang,Zewen Liu,Minhua Lin,Zhaotian Weng,Zhan Shi,Bing He,Yisi Sang,Dakuo Wang,Benoit Dumoulin,Wei Jin,Yuyin Zhou,Cihang Xie,Hanqing Lu
类目:Computation and Language (cs.CL)
关键词:research agents automate, cycle of proposing, opening a path, agents automate, automate the cycle
备注:
点击查看摘要
Abstract:AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement. Yet their ability to propose experiments outpaces their capacity to execute them in real environments, making outcome prediction a key capability for sustained self-improvement under limited experimental budgets. We investigate language models as Research World Models (RWMs), which predict the outcomes of candidate interventions across research environments. Our evaluation draws on over 2,600 experimental records from nine research environments spanning pretraining, post-training, and inference, representing more than 171,000 H100 GPU-hours of experimentation. Research knowledge acquired from real experimental experience improves RWM predictions of unseen interventions within the same environment (Spearman +0.27), and can be reused across environments. For example, using only pretraining experience from OLMo3, Marin, and Nanochat, an RWM reduces selection regret in the Qwen3 environment by 78% compared with zero-experience setting. These benefits extend to multi-round Autoresearch under a fixed selection budget: RWMs with in-env and cross-env research knowledge increase the best gain achieved by 15.8% and 11.6%, respectively. Ablations across 13 language models used as RWMs show that adding research knowledge can improve intervention ranking more than changing models or increasing reasoning effort alone. These findings support language models as RWMs and motivate accumulating experimental data for future RWM training.
22. 【2610.12214】DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion
链接:https://arxiv.org/abs/2610.12214
作者:Heeseung Kim
类目:ound (cs.SD); Computation and Language (cs.CL)
关键词:spoken dialog models, dialog models enable, models enable simultaneous, Recent full-duplex spoken, enable simultaneous listening
备注: 41 pages, 11 figures, 18 tables. Preprint, under review. Project page: [this https URL](https://diffuplex.github.io/)
点击查看摘要
Abstract:Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-LISTEN consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-SPEAK can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-LISTEN and DiffuPlex-SPEAK achieve $1.46\times$ and $1.59\times$ deployment-path wall-clock speedups and $1.61\times$ and $1.80\times$ Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that LISTEN preserves speech naturalness and conversational quality, while SPEAK retains conversational quality with some degradation in speech naturalness.
23. 【2610.12207】SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
链接:https://arxiv.org/abs/2610.12207
作者:Thomas Gebhart,Russell J. Funk
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Pre-trained transformer models, technological progress, models, Pre-trained transformer, text
备注:
点击查看摘要
Abstract:Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.
24. 【2610.12161】When KL Regularization Misfires in Group Policy Optimization
链接:https://arxiv.org/abs/2610.12161
作者:Fei Ding
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Calibrated Policy Optimization, improve group policy, group policy optimization, policy optimization, regularization sometimes improve
备注: 27 pages
点击查看摘要
Abstract:Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
25. 【2610.12144】Language-Specific Effects of Tokenizer Choice in Multilingual Language Models
链接:https://arxiv.org/abs/2610.12144
作者:Clara Meister,Gül Sena Altıntaş,Antoine Bosselut
类目:Computation and Language (cs.CL)
关键词:choice affects multilingual, multilingual language modeling, affects multilingual language, Tokenizer choice affects, tokenizer choice matters
备注:
点击查看摘要
Abstract:Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.
26. 【2610.12133】Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read
链接:https://arxiv.org/abs/2610.12133
作者:Zhiyun Shi
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:prompt prefix shared, cached for retrieval, shared across requests, long conversation, document cached
备注: 14 pages, 5 figures
点击查看摘要
Abstract:Many key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation. The prevailing approach scores KV entries by rehearsal: the model rereads the context and keeps the entries it attends to, assuming that the more completely a cache rehearses its context, the better it remembers it. We show that under tight budgets this assumption backfires: rehearse everything, remember nothing. At a 3% keep ratio, rereading the whole context keeps 31.5 of 96.5 points on RULER, and on LongBench's natural-text tasks it falls below methods that rehearse nothing at all. The cause is that a cache keeps what it rehearses: rereading spreads the budget across the whole context, so the answer's own entries survive at little more than chance. Like a student before an exam, a cache remembers more by testing itself than by rereading. Two principles follow: rehearse what will be read, and rehearse as much as there is. We instantiate them as Attic-KV (Attic for short), a training-free rehearsal in which the model quizzes itself with question-answer pairs that quote the context, alongside anchor tokens in a content-adaptive amount. Changing only the rehearsal lifts three hosts that score it in three different ways: Attic alone is the best training-free method in all eight settings we test on RULER and LongBench's natural-text tasks, and plugged into the gradient-based KVgrad and the trained RestoreKV+, it raises them by up to 17.1 and 28.1 points. Its advantage grows as the budget shrinks, reaching 41.9 points over full rereading at a 3% keep ratio, and it compresses faster than rereading the whole context.
27. 【2610.12118】A persistent accuracy ceiling in automated verbal deception detection
链接:https://arxiv.org/abs/2610.12118
作者:Riccardo Loconte,Jonas Festor,Zane Fatjanova,Mariam Bolkvadze,Bennett Kleinberg
类目:Computation and Language (cs.CL)
关键词:verbal deception detection, human verbal deception, evidence remains fragmented, Automated methods, deception detection
备注:
点击查看摘要
Abstract:Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines. We systematically reviewed 25 years of research (289 reports, 6,136 classification models) and meta-analyzed 3,653 models nested within 97 datasets. Pooled accuracy was 74.4% (95% CI: 71.2%-77.4%) with substantial heterogeneity. Accuracy was driven by methodological quality (ground truth, data source, class balance, evaluation procedure) more than by model complexity: the adoption of embeddings and large language models has not translated into improved predictive performance. Only 12.46% of reports used data with verifiable ground-truth, and only 23.96% of models were evaluated on independent data. The pooled accuracy aligns with meta-analyses of manual approaches, suggesting a ceiling of 70-75%, unlikely to be lifted by current research conventions.
28. 【2610.12083】All Verdicts are Not Equal: Rethinking LLM Judge Reliability
链接:https://arxiv.org/abs/2610.12083
作者:Vineet Kumar,Darshita Rathore,Anindya Moitra
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:remains poorly understood, systemic reliability remains, reliability remains poorly, standard paradigm, remains poorly
备注: Accepted at AACL IJCNLP (Main) 2026
点击查看摘要
Abstract:LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formalize these multi-faceted failure modes, we introduce the trustworthy verdict rate (T ), a unified metric capturing the joint probability that an evaluation is reproducible, order-invariant, and accurate. UsingT , we derive a theoretical upper bound on accuracy imposed by position bias and show that reliability is item-specific rather than modellevel. Finally, we demonstrate that shifting from pairwise win-rate to holistic rubric scoring improves trustworthiness more than any single-format prompting intervention, offering an actionable framework for robust NLP evaluation.
29. 【2610.12064】ILM: An AI-Powered Storytelling Educational Tool
链接:https://arxiv.org/abs/2610.12064
作者:Suhaila Mohammed,Abdelaziz Serour,Allison Lahnala
类目:Computation and Language (cs.CL)
关键词:Digital technologies, existing platforms provide, platforms provide limited, structured knowledge, provide limited support
备注:
点击查看摘要
Abstract:Digital technologies have made Islamic narratives more accessible, but existing platforms provide limited support for structured learning and comprehension of these stories, particularly in Arabic and multilingual settings. We present ILM, an interactive educational platform for Stories of the Prophets that combines Arabic natural language processing, structured knowledge representation, and retrieval-based question generation. Admin-approved Arabic narratives are processed by a Knowledge Graph (KG) Constructor Engine that identifies entities and narrative relationships and stores them as structured knowledge, enabling learners to explore stories through a visual story map and answer entity- and relation-based questions generated from the KG. Separately, a multilingual retrieval pipeline retrieves relevant passages from the original narratives to generate multiple-choice and open-ended comprehension questions. For open-ended questions, an LLM-as-a-Judge evaluates learners' answers against the retrieved passages and reference answers to determine correctness. The platform also incorporates Quranic content as a separate enrichment layer, allowing selected narratives to be supplemented with source-supported information. By combining structured knowledge with passage-based retrieval, ILM supports narrative exploration, comprehension, and assessment across Arabic and multilingual content. The system demonstrates the feasibility of combining structured knowledge representation and retrieval-based generation to support interactive learning of Islamic narratives. A demo is available at this http URL.
30. 【2610.12061】When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation
链接:https://arxiv.org/abs/2610.12061
作者:Yiruo Cheng,Shen Huang,Xiaoshuai Song,Jiejun Tan,Guanting Dong,Pengjun Xie,Ji-Rong Wen,Zhicheng Dou
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large language model, demonstrated strong capabilities, Large language, reasoning, language model
备注:
点击查看摘要
Abstract:Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key challenge is therefore to determine when existing reasoning remains sufficient and when a new reasoning step is needed, without relying on costly generation-based verification. We find that decreases in the likelihood of subsequent reference actions after removing additional reasoning closely track whether those actions remain recoverable given earlier reasoning, providing an effective and lightweight signal for estimating cross-turn action support. Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning. RACE introduces a Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC) procedure that progressively identifies reasoning turns whose removal has limited impact on the current and subsequent reference actions. The resulting removal signals are incorporated into both supervised fine-tuning and agentic reinforcement learning, enabling the policy to learn when to reason and when to act directly. Extensive experiments on four representative agent benchmarks show that RACE substantially reduces reasoning cost while maintaining or improving task performance.
31. 【2610.12030】Natural Language to First-Order Logic LLM-based Autoformalization
链接:https://arxiv.org/abs/2610.12030
作者:Andrea Brunello,Cristian Curaba,Luca Geatti,Michele Mignani,Angelo Montanari,Nicola Saccomanno
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Large Language, Language Models, interest in autoformalization, renewed interest
备注: Accepted to EMNLP-2026
点击查看摘要
Abstract:Large Language Models (LLMs) have renewed interest in autoformalization. Yet, when First-Order Logic (FOL) is considered as the target formalism, the field still lacks a unified task formulation and a systematic survey. This paper addresses this gap: we first provide a principled definition for the FOL-autoformalization task by distinguishing Ontology Extraction from Logical Translation, showing how their conflation obscures (cross-study) evaluation; we review existing datasets, evaluation metrics, and LLM-based methods, including fine-tuning, prompting, and verification-based refinement; we identify open challenges in benchmarking, semantic evaluation, ontology-aware methods, and end-to-end applications.
32. 【2610.12023】InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers
链接:https://arxiv.org/abs/2610.12023
作者:Jonathan Ivey,Aimee Liang,Arthur Y.S. Wang,Madeline Mandell,Ziang Xiao,Anjalie Field
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:elicit open-ended responses, social science research, market research, science research, public polling
备注: Preprint. 23 pages
点击查看摘要
Abstract:Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
33. 【2610.12022】Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology
链接:https://arxiv.org/abs/2610.12022
作者:Zhaoxin Yu,Qingchao Kong,Dajun Zeng,Wenji Mao
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Large language models, attributing external events, Large language, agents' social behaviors, attribution
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly deployed in sociotechnical systems where social attribution, the reasoning process attributing external events to the causes and reasons of agents' social behaviors, plays a critical role. These processes involve judgments of social cause, responsibility, and blame/credit to agents. Although attributional models are well-studied in social psychology and cognition through Attribution Theory, social attribution remains underexplored in AI, particularly LLM social reasoning. This paper provides the first systematic exploration of LLM social attribution. Our work focuses on responsibility and blame attributions, examining current LLMs' judgments and their underlying internal mechanisms. Guided by attribution theory, we construct a social attribution benchmark consisting of a Vignette subset based on classic scenarios from attribution theory research and a Reality subset based on real-world social narratives, yielding 7,639 responsibility/blame judgment questions. On this basis, we evaluate 32 representative LLMs and 5 basic non-LLM baselines. To further explore the internal mechanisms underlying the LLM judgment process, we develop a probing-based methodology to investigate the latent-space representations of 5 key attribution dimensions and the consistency of their influences on LLM judgments compared to those in human social attribution. Our research findings reveal that current LLMs exhibit measurable but incomplete agreement with human responsibility and blame judgments, and meanwhile, this agreement is positively correlated with model size. Some attribution dimensions are systematically decodable from specific positions in LLM hidden states, and their influences on the final judgment are consistent with those indicated by human Attribution Theory. The dataset and associated code are available at this https URL.
34. 【2610.12002】Agentic-TTT: Training test-time policy for test-time training
链接:https://arxiv.org/abs/2610.12002
作者:Jiahao Lu,Mohan Kankanhalli
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:designated open problems, make striking improvements, adapts an LLM, IMO competitions, LLM parameters
备注:
点击查看摘要
Abstract:Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions. Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions. On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training. Together, these results point toward autonomous self-improvement: models that can decide how to learn from their own deployment experience.
35. 【2610.11993】DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
链接:https://arxiv.org/abs/2610.11993
作者:Yupeng Xie,Zhenyang Wang,Jiayi Zhu,Yinghao Tang,Zhouan Shen,Yiyu Chen,Yuyu Luo
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:integrates data visualization, data video understanding, Data video, video understanding, Data
备注: 46 pages, 22 figures, 14 tables
点击查看摘要
Abstract:Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at this https URL.
36. 【2610.11978】Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks
链接:https://arxiv.org/abs/2610.11978
作者:Xing Li,Qingcheng Chang,Jinzhong Ning,Changfeng Xu,Shenlong Zhang,Yijia Zhang,Ling Luo,Hongfei Lin
类目:Computation and Language (cs.CL)
关键词:generating text, returns a choice, specialized decision model, System, LLMs
备注: 10 pages, 2 figures, 7 tables
点击查看摘要
Abstract:Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.
37. 【2610.11966】MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation
链接:https://arxiv.org/abs/2610.11966
作者:Mengdi Liu,Wenjue Chen,Wenyue Chen,Cheng Yang,Fanqi Kong,Zhangyang Gao,Xiaoxue Cheng,Yiheng Li,Yujian Yuan,Keliang Li,Hong Chang,Shiguang Shan,Chenglin Wu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)
关键词:fundamental engine, engine of scientific, remains difficult, Research idea innovation, scientific progress
备注:
点击查看摘要
Abstract:Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility. While recent LLM-based approaches have made progress through carefully designed prompts or agent pipelines, they are constrained by predefined, static ideation workflows. To address this limitation, we propose MindFlow, a framework that explicitly formulates ideation as a graph-structured Flow in Mind, which is composed of modular thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples thinking flows to generate candidate ideas. This open-ended problem is optimized using a tournament-based relative ranking, enabling the controller to progressively favor higher-quality thinking flows. We further introduce an evaluation protocol that jointly assesses problem finding and problem solving, going beyond title- or abstractonly judgments. Across diverse topics, MindFlow shows its superiority as an explicit, controllable and optimizable research idea innovator.
38. 【2610.11959】MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
链接:https://arxiv.org/abs/2610.11959
作者:Xiaomi LLM-Core Team:Zongming Qiao,Ziyue Hua,Zirui Ou,Zihao Yue,Zihan Jiang,Zhuo Huang,Zhiyang Chen,Zhixian Zheng,Zhipeng Xu,Zhengrui Ma,Yuyang Hu,Yuhang Dong,Yuechen Zhang,Yudong Wang,Yuanxin Liu,Yixin Yang,Yishuo Cai,Yikai Zhao,Yihan Yan,Yifan Zhang,Yifan Song,Xiyu Wei,Xing Zhang,Xin Zhang,Xiaoqian Liu,Xiaodong Ji,Xiangwei Deng,Xueyu Guo,Wenhan Ma,Weimin Xiong,Weikun Wang,Weiji Zhuang,Shuo Liu,Shuhuai Ren,Shuhao Gu,Shimao Chen,Shijie Cao,Shihua Yu,Shicheng Li,Shengjie Zhou,Shaolei Zhang,Rang Li,Qiying Wang,Qingkai Fang,Qianli Chen,Minzheng Wang,Liwen Wang,Linli Yao,Linghao Zhang,Liangyu Cheng,Liang Zhao,Lei Li,Jinhao Dong,Jinyu Xiang,Jianyu Wei,Jiangshan Duo,Huaqiu Liu,Huanjie Fan,Hongyi Guan,Hongshen Xu,Hao Tian,Hanyu Li,Hailin Zhang,Gang Wang,Fuli Luo,Feng Wei,Dong Zhang,Dawei Zhu,Chiheng Lou,Chenhong He,Chenhao He,Chenghua Liu,Bowen Ye,Bowen Shen,Boshen Xu,Bo Yang,Bingquan Xia,Bangjun Xiao,Baixuan Xu,Zhouxiang Mao,Zhiyang Zhang,Zhixiang Xu,Zhenru Lin,Zhengju Tang,Zhaojun Huang,Yuzhe Weng,Yuxing Xiang,Yuxiao Li,Yuheng Yang,Yuhang Wang,Yuchen Liu,Yuanyuan Tian,Yuanliang Dong,Yu Cheng,Yongzhe He,Yongshun Liang,Yong Wang,Yiyan Wang,Yitian Gong
类目:Computation and Language (cs.CL)
关键词:advancing large foundation, large foundation models, Reinforcement learning, central training paradigm, paradigm for advancing
备注:
点击查看摘要
Abstract:Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
39. 【2610.11948】When History Helps and Hurts: Selective History Use across Multimodal Turns
链接:https://arxiv.org/abs/2610.11948
作者:Shuoyang Sun,Kerui Gu,Hao Fang,Shaoli Huang,Bin Chen
类目:Computation and Language (cs.CL)
关键词:Reliable multimodal interaction, multimodal interaction depends, interaction depends, conflicting new observations, Reliable multimodal
备注:
点击查看摘要
Abstract:Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.
40. 【2610.11922】Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
链接:https://arxiv.org/abs/2610.11922
作者:Jimmy Lin,Sahel Sharifymoghaddam,Lingwei Gu,Nour Jedidi
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:Project Greenhouse represents, modest computational resources, Project Greenhouse, Greenhouse represents, build fully open
备注:
点击查看摘要
Abstract:Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.
41. 【2610.11920】Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents
链接:https://arxiv.org/abs/2610.11920
作者:Yichen Liu,Chunfeng Yuan,Haowei Liu,Wenjuan Li,Zefeng Lin,Bing Li,Xu Chen,Weiming Hu
类目:Computation and Language (cs.CL)
关键词:personalized conversational agents, storing past interactions, retrieving relevant information, memory, working memory
备注: 17 pages, 7 figures. Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
点击查看摘要
Abstract:For persistent and personalized conversational agents, memory systems can enable them to remember, update, and reason over long histories by storing past interactions and retrieving relevant information. Existing memory systems typically follow two paradigms: flat-structured memory and graph-based memory. The former is lightweight but leaves event relations and state updates implicit, while the latter explicitly models memory structure but incurs additional construction cost and introduces irrelevant relations over long histories. To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory. QGMem converts long dialogue histories into event-indexed atomic memory units that preserve individual experiences and consolidates related units into dynamic memory traces that retain state trajectories and current states. When a query arrives, hybrid memory retrieval gathers complementary candidate memories, and query-aware reranking activates the most relevant units as a compact working memory. To expose relational dependencies in the working memory and support conflict-aware reasoning, QGMem organizes the working memory as a local graph, which is then encoded as a graph token and provided to the LLM together with the textual working memory to improve evidence utilization during answer generation. Experiments across six benchmarks validate the framework and show consistent gains in retrieval, multi-hop evidence composition, conflict resolution, and ultra-long dialogue reasoning with compact contexts and moderate inference cost.
42. 【2610.11915】Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning
链接:https://arxiv.org/abs/2610.11915
作者:Xunlei Chen,Qinghui Gong,Jingkun Xue,Qihe Liu,Shijie Zhou,Fei Ye
类目:Computation and Language (cs.CL)
关键词:Machine unlearning, remove unwanted knowledge, large language models, language models aims, large language
备注:
点击查看摘要
Abstract:Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior. Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance. In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities. PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints. Both stages serve a common goal: satisfying the forgetting requirements while preserving fluent generation and performance on non-target tasks. Across three benchmarks, PTP-U achieves the strongest forgetting-retention trade-off among evaluated methods, reaching 81.22%-91.03% forgetting while preserving 94.20% non-target utility on average. At matched forgetting, PTP-U consistently retains higher non-target utility.
43. 【2610.11901】Can Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMs
链接:https://arxiv.org/abs/2610.11901
作者:Xing Li,Jinzhong Ning,Yijia Zhang,Liang Yang,Hongfei Lin
类目:Computation and Language (cs.CL)
关键词:detection requires identifying, conversational context, requires identifying, identifying an author, author attitude
备注: 8 pages, 1 figure, 5 tables
点击查看摘要
Abstract:Stance detection requires identifying an author's attitude toward a given target, sometimes based on conversational context. Jev, a specialized decision model designed for structured decision-making, offers an alternative to general-purpose large language models (LLMs). In this work, we evaluate Jev on two stance detection datasets, VAST (English texts) and ZS-CSD (Chinese conversations), comparing it with four general-purpose LLMs and two fine-tuned models. Results show that Jev achieves competitive performance on VAST, matching GPT-5.6 and outperforming the other general-purpose LLMs. However, it falls behind stronger LLMs on ZS-CSD, particularly in distinguishing favor from against. Further analysis suggests that this limitation may be related to understanding reply relationships and stance direction rather than conversation length alone. These findings highlight both the potential and limitations of Jev for stance detection.
44. 【2610.11899】Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System
链接:https://arxiv.org/abs/2610.11899
作者:Irene Weber(University of Applied Sciences Kempten, Germany)
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
关键词:Large language models, Large language, increasingly embedded, embedded as components, retrieval-augmented generation
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent. Whether these labels denote genuine architectural forms or serve as branding has not been assessed systematically. In the sources surveyed, labels do carry architectural content, most clearly in vendor usage: copilot denotes a router-worker architecture operating a host application under step-by-step user confirmation, while the more recent shift to the label agent coincides with AI-planned multi-step execution of which the user sees only the outcome. The coding agents of four major providers share one architecture, a reason-and-act loop delegating to subagents. This survey describes seven recurring forms---LLM chats, custom agents, retrieval-augmented generation (RAG), AI-enhanced workflows, copilots, coding agents, and, in part, agentic RAG---in a common vocabulary of agents and tools. Each is characterized along four structural dimensions (agentic RAG only partially): the architectural pattern, the control of execution and the point of user intervention, the number of agent calls per task, and tool use. An illustrative corpus of 22 systems from research publications and vendor documentation grounds the descriptions and shows where they reach their limit.
Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
ACMclasses:
I.2.11; D.2.11; I.2.7
Cite as:
arXiv:2610.11899 [cs.CL]
(or
arXiv:2610.11899v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2610.11899
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
45. 【2610.11854】GRPODropout: Less is More for Online Reinforcement Learning Rollouts
链接:https://arxiv.org/abs/2610.11854
作者:Hexuan Deng,Zihao Yan,Xuebo Liu,Shuo Nie,Yue Wang,Chen Wang,Zhaohua Zhang,Tianwen Jiang,Qiuyong Xiao,Jihong Zhang,Min Zhang
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:large language model, language model reasoning, diversity weakens exploration, Reinforcement learning, sampling diversity weakens
备注:
点击查看摘要
Abstract:Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at this https URL.
46. 【2610.11845】Detecting Spin in Clinical Trials with Large Language Models
链接:https://arxiv.org/abs/2610.11845
作者:Tjaš Ajdovec,Marko Robnik-Šikonja,Simon Šuster
类目:Computation and Language (cs.CL)
关键词:includes reporting practices, clinical trials includes, trials includes reporting, includes reporting, reporting practices
备注: 5 pages, 1 figure, 2 tables. Accepted at the 29th International Multiconference Information Society (IS 2026), AI in Healthcare track, Ljubljana, Slovenia. Code: [this https URL](https://github.com/ta5946/Spin)
点击查看摘要
Abstract:Spin in clinical trials includes reporting practices that distort the presentation of results. This is particularly critical in medicine, where spin is present in more than 50% of randomized controlled trials that fail to reach statistical significance. The comparison of primary and reported outcomes is crucial for detecting several types of spin, including outcome switching. We used 300 pairs of outcomes labeled with semantic similarity to develop a system for automatic detection of outcome switching. We evaluated baseline text similarity models and open-source LLMs using generated similarity scores and the Youden index to determine the classification threshold. The proposed approach involves prompt engineering, classification based on token probabilities, and majority voting for the final decision. The results on the test set of 2,496 examples with an F1 score of 0.78 and an accuracy of 0.90 outperform baseline text similarity models but trail behind fine-tuned versions of BERT. We used LLMs to generate natural language explanations for the classified instances and manually assessed their quality.
47. 【2610.11794】Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
链接:https://arxiv.org/abs/2610.11794
作者:Haoyu Zhao,Zhengxu Yu,Zhiyuan He,Meng Fang,Rasul Tutunov,Haitham Bou-Ammar,Weilin Luo,Jun Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:unfamiliar environments requires, act in unfamiliar, multiple world models, world models, environments requires agents
备注:
点击查看摘要
Abstract:Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
48. 【2610.11790】Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
链接:https://arxiv.org/abs/2610.11790
作者:Nicolás Vera Zúñiga
类目:Computation and Language (cs.CL)
关键词:Byte Latent Transformer, Latent Transformer, Byte-level language models, Byte-level language, Byte Latent
备注: 13 pages, 4 figures, 6 tables. Code, logs and results: [this https URL](https://github.com/nicoveraz/segresearch) (archived: doi: [https://doi.org/10.5281/zenodo.23238215](https://doi.org/10.5281/zenodo.23238215) )
点击查看摘要
Abstract:Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch. BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict. We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution. Under tight patch budgets, entropy-triggered layouts skip these positions, and accuracy on them collapses. In Meta's BLT-1B with patch starts on 10% of bytes, the entropy rule puts a patch start at 16% of the computed results in GSM8K solutions and gets 19.0% of them exactly right; a boundary after each "=" at the same patch count gets 51.8%, and entropy combined with a label-free boundary-dependence signal gets 67.1% (default layout at 26% of bytes: 76.8%). The gap survives adapting BLT-1B to the budget with low-rank fine-tuning (32.9% vs 72.7%, three runs per rule, paired p 1e-200) and grows with model size in byte models trained from scratch at a 10% budget: at 1M, 12M and 50M parameters, boundary dependence beats entropy on final answers by -1.6, +10.1 and +19.8 points, and at 50M it gets 35.9% of computed results against 13.9% (3 seeds each). BLT's entropy-jump rule helps neither target at 50M. The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%. The effect is specific to computed values: copies and lookups gain little, and values the model cannot compute gain nothing. Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels: combined with entropy it beats the hand-written rule on computed results.
49. 【2610.11787】From Sparse Representations to Behavioral Insights for Multimodal Depression Assessment
链接:https://arxiv.org/abs/2610.11787
作者:Guimin Hu,Zihao Song,Jiachen Luo,Jiayuan Xie,Ruichu Cai
类目:Computation and Language (cs.CL)
关键词:behavioral patterns, offers a promising, multimodal behavioral representations, behavioral, multimodal behavioral
备注:
点击查看摘要
Abstract:Multimodal depression assessment offers a promising approach to analyzing behavioral patterns associated with depression. However, existing methods often rely on dense and opaque multimodal representations, making it difficult to interpret the behavioral patterns underlying their predictions. In this work, we introduce BehavDep, a sparse factor-based framework that decomposes multimodal behavioral representations into sparse latent factors and associates them with behaviorally meaningful concepts through a semantic bridge. To address the mismatch between user-level annotations and heterogeneous video-level behaviors, BehavDep further learns video-level depression tendency scores under weak supervision and aggregates information across multiple observations for user-level assessment. Extensive experiments demonstrate that BehavDep achieves the best overall assessment performance while revealing complementary modality contributions, heterogeneous behavioral patterns across observations, and prediction responses to concept-level editing. These results show that BehavDep provides a structured and interpretable approach to analyzing multimodal behavioral representations for depression assessment.
50. 【2610.11776】DPPM: Dual-Path Parametric Memory for Personalized Language Models
链接:https://arxiv.org/abs/2610.11776
作者:Yuhao Chen,Shuochen Liu,Jiayao Shi,Jian Hong,Chen Cheng,Xinyun Ding,Tao Wang,Ya Li,Quan Liu,Tong Xu
类目:Computation and Language (cs.CL)
关键词:Long-term personalization requires, personalization requires language, track users' preferences, Long-term personalization, requires language models
备注: 12 pages, 5 figures
点击查看摘要
Abstract:Long-term personalization requires language models to use interaction history to track users' preferences across sessions. Parametric memory encodes this interaction history into model parameters or adapters, reducing the need to include it in the inference context. However, independent context compilation leaves cross-session integration unspecified, while recurrent updates can attenuate earlier evidence. To address these challenges, we propose Dual-Path Parametric Memory (DPPM). Its Evidence path directly pools representations of the interaction history to preserve earlier evidence, while its Delta path sequentially updates an associative state to capture changes. Fusing both outputs produces history-conditioned LoRA adapters that combine evidence accumulation with ordered revision. Across multiple backbones, DPPM outperforms the evaluated baselines, achieving 54.22% on PersonaMem-v2 and 86.79% on PrefEval. These results suggest that DPPM provides a simple and effective design choice for cross-session personalized parametric memory.
51. 【2610.11775】RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
链接:https://arxiv.org/abs/2610.11775
作者:Ilya Lasy,Nora Yinuo Cai,Kola Ayonrinde
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:modular expert networks, Sparse Mixture, scale more efficiently, efficiently than dense, active for processing
备注: 33 pages (12 non-appendix pages), 7 figures, published as a conference paper at ICML 2026
点击查看摘要
Abstract:Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with ${\sim}65\%$ higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.
52. 【2610.11766】Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
链接:https://arxiv.org/abs/2610.11766
作者:Haitong Jiang,Chunlin Liu,Sihan Tang,Chan Wu,Xiaoqing Su,Yuhong Feng
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:commonly summarize harmful-output, summarize harmful-output behavior, attack success rate, language models commonly, models commonly summarize
备注: 12 pages, 2 figures. Code and experiment inputs: [this https URL](https://github.com/kevinjiang0121-cyber/IRIS)
点击查看摘要
Abstract:Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates. Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern. A complementary contrast comes from FormalLogic, where high recovery can still coincide with frequent harmful assistance. Together, these results show that non-harmful outcomes are not equally informative about model safety, motivating the joint reporting of intent recovery and ASR in LLM safety evaluation. Code and experiment inputs are available at this https URL.
53. 【2610.11765】hinking Inertia: LLMs Keep Thinking When Told Not To
链接:https://arxiv.org/abs/2610.11765
作者:Dianqiao Lei,Kevin Qinghong Lin,Pan Lu,Philip Torr,James Zou
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Large Language, Language Models, increasingly ship, thinking modes
备注: Accepted by NeurIPS 2026. Website: [this https URL](https://thinking-inertia.github.io) GitHub: [this https URL](https://github.com/thinking-inertia/code)
点击查看摘要
Abstract:Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.
54. 【2610.11716】4-Tensor Attention Model for Semantic Physical Reality
链接:https://arxiv.org/abs/2610.11716
作者:Jongwook Kim,Sangheon Yun
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:robot planning, video generation, generation and robot, normalizes attention jointly, softmax normalizes attention
备注: 36 pages, 4 figures
点击查看摘要
Abstract:We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning. A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window. Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update. To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer. On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3. At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M. On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
55. 【2610.11715】Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
链接:https://arxiv.org/abs/2610.11715
作者:Peter Devine,Nick Ryan,Benjamin Sirb,Alex Chiocchi
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:language model carry, document-specific LoRA adapters, Mapping hypernetwork, prior work, billion parameters
备注: 14 pages, 3 figures, 1 table
点击查看摘要
Abstract:Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one. On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8% top-5 teacher-forced accuracy against 63.4% and 83.5% for the base model, with nothing in the context window but a three-word instruction. Once the hypernetwork is trained, a single forward pass turns any document into an adapter for such a model, which could be served alone for speed or alongside the document in the window to raise accuracy further.
Comments:
14 pages, 3 figures, 1 table
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:
arXiv:2610.11715 [cs.AI]
(or
arXiv:2610.11715v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2610.11715
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Peter Devine [view email] [v1]
Thu, 8 Oct 2026 11:19:13 UTC (111 KB)
56. 【2610.11695】Structured Sentiment Analysis Using Sequence Labeling as Dependency Graph Parsing
链接:https://arxiv.org/abs/2610.11695
作者:Muhammad Imran,Ana Ezquerro,Carlos Gómez-Rodríguez,Anders Søgaard,David Vilares
类目:Computation and Language (cs.CL)
关键词:structured sentiment analysis, nodes represent spans, fine-grained sentiment graph, sentiment analysis, sentiment holders
备注:
点击查看摘要
Abstract:This study addresses the problem of structured sentiment analysis, whose goal is to obtain a fine-grained sentiment graph where the nodes represent spans of sentiment holders, targets, and expressions, while the arcs define the relationships among them. Our proposed approach casts the task as dependency graph parsing, but departs from traditional parsing methods by solving it through sequence labeling. To do so, we leverage recent advances in linearized graph encodings that allow each word in the input to be assigned a label, effectively capturing the structure of the dependency graph. We conducted experiments on seven datasets spanning five languages (English, Spanish, Norwegian, Basque, and Catalan), showing performance competitive with leading, more complex single-model approaches.
57. 【2610.11678】RACE: Diagnosing Verifier Brittleness in Agentic Evaluation
链接:https://arxiv.org/abs/2610.11678
作者:Radhika Gaonkar
类目:Computation and Language (cs.CL)
关键词:large language model, Verifier scores, language model, benchmark metrics, metrics and training
备注:
点击查看摘要
Abstract:Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability. It may instead reflect a change in the evaluation. We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible. In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent. On public $\tau^2$-bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate. In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within $\pm$0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation. Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome. TRACE thus separates what a score change says about the agent from what it says about the measurement.
58. 【2610.11659】DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
链接:https://arxiv.org/abs/2610.11659
作者:Anhao Zhao,Haoran Xin,Junlong Tong,Yingqi Fan,Xuan Lu,Ping Nie,Wenjie Li,Xiaoyu Shen
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:supervises student-generated trajectories, On-policy distillation, supervises student-generated, Vanilla OPD, student-generated trajectories
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. A parameter beta controls this weighting, and the highest-scoring tokens are retained. Across 4 teacher-student pairs and 7 mathematical reasoning benchmarks, we compare DIAL-OPD with 9 baselines. Retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants, with mean accuracy gains reaching 5.25 percentage points over Vanilla OPD, and doubles AIME25 Pass@16 from 13.33% to 26.67%. It also achieves up to an 18% relative improvement in mean accuracy over the strongest token-selection baseline at matched retention ratios. With a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, showing that effective supervision allocation can outweigh teacher scaling. Further analysis shows that moderate beta balances suppressing low-low tokens against preserving useful disagreements. Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.
59. 【2610.11655】Harness Evolution Hits a Ceiling: When Weight Training Should Begin
链接:https://arxiv.org/abs/2610.11655
作者:Yuan Tian,Bing Hu,Hao Wang,Binghang Lu,Fang Wu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:long-horizon LLM agent, Improving a long-horizon, long-horizon LLM, LLM agent, content failures
备注: 21 pages, 7 figures, 15 tables
点击查看摘要
Abstract:Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.
60. 【2610.11646】Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
链接:https://arxiv.org/abs/2610.11646
作者:Christopher Witzl,Tobias Bocklet,Korbinian Riedhammer
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Liang hyphenation algorithm, morphologically rich language, Liang hyphenation, hyphenation algorithm, morphologically rich
备注:
点击查看摘要
Abstract:German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.
61. 【2610.11638】UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue Interactions
链接:https://arxiv.org/abs/2610.11638
作者:Mengze Hong,Zeyang Lei,Wenbo Shang,Xia Zeng,Xiying Zhao,Qi Zhu,Chen Jason Zhang,Di Jiang,Taiming Fu,Qiongyi Zhou,Qinghe Chang,Fubao Zhang,Chenxuan Ma,Minlong Peng,Jinfeng Huang,Zineng Zhou,Jindou Wu,Muge Qi,Sijun He,Xin Cui,Di Liang,Yuan Hua,Davey Chen
类目:Computation and Language (cs.CL)
关键词:Evaluating user experience, gained increasing attention, automated computational methods, Evaluating user, increasing attention
备注:
点击查看摘要
Abstract:Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.
62. 【2610.11599】Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing
链接:https://arxiv.org/abs/2610.11599
作者:Kazuki Nakajima,Takayuki Mizuno
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Digital Libraries (cs.DL); Social and Information Networks (cs.SI)
关键词:Journals and conferences, large language models, conferences have begun, text written, written using large
备注:
点击查看摘要
Abstract:Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects the screening of scientific manuscripts. We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026. We then trained detectors under maintenance scenarios ranging from a detector retrained on every new version to one trained once and never updated. Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it. Detectors trained on later versions can also miss rewrites of earlier ones. Vocabulary differences between versions largely track where detection transfers and where it fails. In the two screening scenarios we simulated, screens covering all 23 versions either flagged one in eight human-written abstracts or missed one in three rewrites of the newest version. Indeed, a commercial detector missed most rewrites of the version just after the sharpest boundary while flagging almost no human-written abstracts. Research-integrity policy should therefore treat the benchmark accuracy of a detector as provisional, to be re-verified with every LLM release, including earlier versions.
63. 【2610.11592】Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog
链接:https://arxiv.org/abs/2610.11592
作者:Hadeel Al-Negheimish,Jasna Ilieva,Yoon Kim
类目:Computation and Language (cs.CL)
关键词:theoretically process long, Current frontier LLMs, process long contexts, Current frontier, theoretically process
备注: Findings of EMNLP 2026
点击查看摘要
Abstract:Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.
64. 【2610.11586】Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts
链接:https://arxiv.org/abs/2610.11586
作者:Umaira Izhar,Gunjan Arora,Pushpendra Singh
类目:Computation and Language (cs.CL)
关键词:Existing evaluation methods, assess factual correctness,safety, primarily assess factual, providing limited insight, situated healthcare reasoning
备注:
点击查看摘要
Abstract:Existing evaluation methods for healthcare LLMs primarily assess factual correctness,safety, and fluency, while providing limited insight into whether generated interactions reflect culturally situated healthcare reasoning. This limitation is particularly important in maternal health, where care decisions are shaped by social and relational norms. We introduce MH-INDIC, a culturally grounded evaluation framework for maternal-health interactions in urban and semi-urban North Indian contexts that operationalises cultural behaviour through ten dimensions of maternal-health reasoning. Using a 26-item survey administered to 102 pregnant and postpartum women from urban and semi-urban North India, we evaluate ten LLMs. We distinguish population level cultural alignment from profile-level behavioural variation. Although several models approximate the human population-level distribution, all evaluated systems exhibit substantially lower variation across demographic and household profiles than the human cohort, revealing a gap between aggregate alignment and profile-conditioned sensitivity. As a downstream application of MH-INDIC, we use the strongest-aligned proprietary and open-source models to generate culturally conditioned maternal-health dialogues under zero-shot, self-conditioned, and human-grounded prompting. Human-grounded conditioning produces stronger profile alignment and dialogue quality ratings, suggesting that measured cultural profiles can improve the cultural grounding of generated interactions
65. 【2610.11585】Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection
链接:https://arxiv.org/abs/2610.11585
作者:Vinko Sabolčec,Bettina Messmer,Yassine Turki,Martin Jaggi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Recent advances, pretraining highlight, advances in large, highlight the role, large language model
备注:
点击查看摘要
Abstract:Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data. This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant. Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels. Our 1B, 3B and 8B scale experiments show that our approach maintains the downstream LLM benchmark performance of existing multilingual model-based filtering baselines, without harming regional and cultural knowledge benchmarks. To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.
66. 【2610.11578】Chronos Enables Code Agents to Reason over Software Evolution
链接:https://arxiv.org/abs/2610.11578
作者:Xin Yin,Yiang Zhang,Zhiyuan Peng,Chao Ni,Zhe Cui,Xiaohua Xin
类目:oftware Engineering (cs.SE); Computation and Language (cs.CL)
关键词:codebase current state, Historical pull requests, Historical pull, compatibility constraints, pull requests record
备注: 22 pages, 3 figures
点击查看摘要
Abstract:Historical pull requests record the design decisions, compatibility constraints, and implementation patterns behind a codebase's current state. Experience relevant to a new task can span related changes whose descriptions emphasize different concerns. We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents. Chronos distills merged pull requests into structured experience cards and connects them through a typed graph of code-level, developer-intent, and organizational relations. Semantic search identifies entry cards, and weighted multi-hop expansion retrieves connected changes for selective reading. The same memory guides candidate generation and patch selection: a patch-focused change agent and a validation-strategy agent each develop a patch, and an evolution steward consults history to select between them. On SWE-Bench Verified, the full workflow improves SWE-Agent across all six evaluated LLM backbones, raising the mean resolution rate from 69.2% to 72.9% and reaching 79.8% with MiniMax M2.5. With the same backbone, it raises resolution rates from 48.3% to 51.7% on SWE-Bench Pro and from 41.0% to 43.5% on FEA-Bench Lite. Both experience-guided single-agent variants also outperform the base agent. In a human evaluation on 100 tasks with ten cards retrieved per task, graph-grounded retrieval increases the mean number of useful cards from 1.24 to 2.87 over flat semantic retrieval. These results demonstrate the value of PR relations for retrieving useful repository experience and of the evaluated workflows for applying that experience during patch generation and selection.
67. 【2610.11575】Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
链接:https://arxiv.org/abs/2610.11575
作者:Yunkai Chai,Tong Zhu,Xiaoye Qu,Xuyang Hu,Guanjie Chen,Qipeng Guo,Yu Cheng
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:models scale parameter, scale parameter capacity, parameter capacity efficiently, fixed compute budget, models scale
备注: 13 pages, 4 figures
点击查看摘要
Abstract:Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-$k$ baseline by $1.6$ points on average across downstream tasks.
68. 【2610.11566】Incremental Open-Ended Deep Research with Structured Harness
链接:https://arxiv.org/abs/2610.11566
作者:Meilin Chen,Hongyuan Bao
类目:Computation and Language (cs.CL)
关键词:Existing Open-Ended Deep, Open-Ended Deep Research, systems primarily generate, Incremental Open-Ended Deep, primarily generate reports
备注:
点击查看摘要
Abstract:Existing Open-Ended Deep Research (OEDR) systems primarily generate reports from scratch, making them inefficient for scenarios where research reports need to be continuously maintained as new information emerges. We introduce \textbf{Incremental Open-Ended Deep Research (Incremental-OEDR)}, a research setting that treats a report as an evolving research state and incrementally updates it by preserving valid knowledge, revising outdated or incomplete content, and incorporating newly available information. To support this setting, we propose \textbf{Structured Harness}, which represents reports as structured collections of outlines, sections, and supporting evidence, and provides structured retrieval, a persistent structured evidence pool, and structured generation for selective report updating and evidence reuse. We further establish a temporal evaluation framework spanning ten years, with \emph{Single-Step Task} and \emph{Long-Chain Task} to evaluate incremental updates over both individual transitions and long-term update chains. Extensive Experiments on DeepResearch Bench and DeepConsult under both the Open-source Configuration (OC) and Proprietary Configuration (PC) show that Incremental-OEDR maintains competitive report quality while substantially improving report continuity and reducing research costs. As shown in Figure~\ref{fig:profile}, it achieves up to 0.51 higher content-level ROUGE-L F1, 0.63 higher outline-level EM F1, 33\% lower token consumption, and 61\% fewer search calls than OEDR on DeepResearch Bench. For more details, please refer to our project page: this https URL.
69. 【2610.11559】SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
链接:https://arxiv.org/abs/2610.11559
作者:Hexuan Deng,Yue Wang,Wenyu Jiang,Cheng Yang,Haolin Yang,Zhaohua Zhang,Chenchen Zhao,Beiduo Chen,Muxi Chen,Sa Zhu,Geyuan Zhu,Jianhuan Zhuo,Qiuyong Xiao,Tianwen Jiang,Jihong Zhang,Xuebo Liu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
关键词:application of LLM, Claude Code, existing benchmarks remain, Code and Codex, LLM agents
备注:
点击查看摘要
Abstract:Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.
70. 【2610.11544】Prosody-to-Text: Predicting text from low-pass filtered speech
链接:https://arxiv.org/abs/2610.11544
作者:David Porteš,Aleš Horák
类目:Computation and Language (cs.CL)
关键词:remains largely overlooked, remains largely, largely overlooked, established task, opposite direction
备注: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
点击查看摘要
Abstract:While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked. We find this unfortunate, because this opposite direction could lead to some very interesting use cases. Therefore, in this paper, we make the first steps in the prosody-to-text direction by inves- tigating how much of the original sentence can be recovered from its prosodic pattern. To this end, we fine-tune the Whis- per model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances be- ing recovered perfectly, and 40% of utterances having Word Error Rate at or below 25%. We also find that, given the correct prefix, the next token was predicted correctly in 79% of cases. Our results suggest that the relationship between low-frequency speech features and lexical content is much stronger than previously thought, and we believe that direct- ing more attention to this topic might open the door to new applications, such as using prosody to guide text generation of modern LLMs
71. 【2610.11543】Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation
链接:https://arxiv.org/abs/2610.11543
作者:Yuxin Meng,Ruixu Zhang,Junjie Wang,Yuhan Suo,Yuhan Sun,Ruining Hu,Yiyao Yu,Yubin Wang,Shouwei Ruan,Bin Wang,Yue Liao,Yuxiang Zhang,Yujiu Yang
类目:Computation and Language (cs.CL)
关键词:Functional Web generation, imperfect implementations underexplored, existing methods largely, methods largely focus, repairing imperfect implementations
备注:
点击查看摘要
Abstract:Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.
72. 【2610.11542】Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper
链接:https://arxiv.org/abs/2610.11542
作者:Masaaki Nakatsu(AO, Inc. / OrbLabs AG),Reno Wang(AO, Inc.)
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Multi-agent LLM systems, stateful counterpart waste, counterpart hidden acceptance, LLM systems negotiating, hidden acceptance condition
备注: 29 pages, 3 figures. The Gatekeeper, agents, constitution, lexicon, 30 run logs and analysis scripts are released (see Appendix F). Companion paper: [arXiv:2610.09772](https://arxiv.org/abs/2610.09772)
点击查看摘要
Abstract:Multi-agent LLM systems negotiating with a stateful counterpart waste model calls in three ways: polite loops that never meet the counterpart's hidden acceptance condition, malformed outputs that trigger retries, and compliance deadlocks in which the counterpart demands something the agent must refuse. We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text. The testbed has a known solution: it measures whether the stack executes a constitution-aligned strategy against swarm drift and recovers from deadlock, not whether it discovers anything. In five runs per configuration (30 runs; Gemini 2.5 Pro agents, Claude Haiku 4.5 Gatekeeper) we find: (i) the constitution and Director make an acceptable framing possible but not reliable - 0/5 baseline unlocks versus 1/5 and 2/5 with the constitution; when the swarm unlocks it does so in one turn with 7-8 calls and about 15k tokens (67-73% below baseline); when it does not, it costs 17-38% more; (ii) the Monitor and hard gate do not reduce unlocks and leave an audit trail; (iii) under a honeytrap-to-compliance deadlock, LLM-only steering escapes 0 of 5 times while atomic purge plus a canonical strike escapes 5 of 5 (Fisher $p = 0.008$) at the same call budget, with zero calls for the strike. LLM-written strikes failed the deterministic pre-flight 5 of 5 times although an LLM Monitor had approved four. Pre-registered hypotheses on average call and token reduction were not supported. Cost is bounded in every arm by deterministic stop rules; the stack adds recovery at no extra model cost.
73. 【2610.11520】When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing
链接:https://arxiv.org/abs/2610.11520
作者:Minnie Kabra,Benjamin Lecouteux,Maximin Coavoux
类目:Computation and Language (cs.CL)
关键词:task recently proposed, intermediate neural networks, speech parsing, recently proposed, consists in predicting
备注: to appear in Findings of EMNLP 2026
点击查看摘要
Abstract:End-to-end speech parsing, a task recently proposed, consists in predicting both the transcription and the syntactic tree for a spoken utterance. Existing architectures for speech parsing often utilise intermediate neural networks. In this work, we examine the effectiveness of intermediate neural networks (NN) for parsing, and, specifically, what role do they play. We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing. We demonstrate that intermediate NN units help reduce the representational gap when the pre-trained encoder is frozen. We do a comprehensive evaluation of speech parsing on French, and medium-low resource languages Slovenian and Naija. We further investigate the impact of the training data size and intermediate layers of the pretrained speech encoder on speech parsing.
74. 【2610.11519】Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
链接:https://arxiv.org/abs/2610.11519
作者:Xiaobing Chen,Zhiqi Pang
类目:Computation and Language (cs.CL)
关键词:post-training reasoning models, Reinforcement learning, on-policy distillation, reasoning models, learning with verifiable
备注:
点击查看摘要
Abstract:Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.
75. 【2610.11510】Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs? A Representation-Level Analysis
链接:https://arxiv.org/abs/2610.11510
作者:Abdu Sallouh,Nicholas Popovič,Michael Färber
类目:Computation and Language (cs.CL)
关键词:Modern Standard Arabic, Modern Standard, default to Modern, Standard Arabic, Large language models
备注: EMNLP 2026
点击查看摘要
Abstract:Large language models (LLMs) often default to Modern Standard Arabic (MSA) when generating Arabic, even when prompted with dialectal Arabic. A natural explanation is that their internal representations are dominated by MSA. We test this hypothesis by adapting the language-dominance framework of Shani and Basirat (2025) (this https URL) to 26 Arabic varieties. Across layers and model families, we find no evidence that MSA acts as a dominant internal representation for Arabic dialects. Instead, dialect representations form a dense and highly overlapping space: normalized mutual information drops sharply compared to patterns reported for more distinct languages. Moreover, the strongest separability effects are not confined to intermediate layers, but can shift toward later layers depending on the architecture. These findings challenge a common interpretation of MSA-biased generation: output preference does not necessarily reveal internal representational dominance. Analyses of multilingual and dialectal LLMs should therefore distinguish generation bias from the geometry of internal representations.
76. 【2610.11501】Beyond Sequences: Distilling Structured Decision Memory for LLM Recommendation
链接:https://arxiv.org/abs/2610.11501
作者:Leikun Liang,Guoshuai Wang,Xingsheng He,Yushan Han,Yunyi Xuan,Xiaoxiao Xu,Lin Qu
类目:Computation and Language (cs.CL)
关键词:large language models, model single-type behaviors, views or purchases, prevailing approaches, language models
备注:
点击查看摘要
Abstract:Despite the adoption of large language models (LLMs) in recommendation systems, prevailing approaches mostly model single-type behaviors (e.g., views or purchases). Even when incorporating multiple behaviors, existing methods flatten heterogeneous actions into homogeneous token sequences, ignoring their distinct decision-making roles. This flattening fails to capture semantic hierarchies and contextual nuances in complex decision-making, such as trade-offs between price and quality. Consequently, performance degrades in critical ``difficult-choice'' scenarios involving highly similar items. To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence. MARI maintains a Decision Memory Bank (DMB) that archives users' past rationales as Structured Decision Memories (SDMs): concise records of goals, constraints, and trade-offs. These SDMs are generated offline via Post-Hoc Decision Distillation from heterogeneous behaviors and user-generated content. By retrieving relevant SDMs to augment LLM reasoning, MARI achieves interpretability and scalability without the prohibitive cost of processing long raw sequences. Extensive experiments show MARI significantly outperforms state-of-the-art baselines on standard next-item prediction and a newly introduced Difficult Choice Prediction task, incurring low latency overhead by decoupling memory construction from online inference. Qualitative analyses reveal actionable, human-readable insights into user decision-making, marking a concrete step toward reasoning-aware recommendation systems.
77. 【2610.11469】SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
链接:https://arxiv.org/abs/2610.11469
作者:Jeonghyo Song,YoungJoon Yoo
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:large language model, Recent large vision-language, large vision-language models, Recent large, diverse image-text tasks
备注: Accepted to EMNLP 2026 Findings
点击查看摘要
Abstract:Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.
78. 【2610.11464】Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
链接:https://arxiv.org/abs/2610.11464
作者:Xing Zhang,Guanghui Wang,Yanwei Cui,Ziyuan Li,Wei Qiu,Bing Zhu,Peiyang He
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:bare LLM judge, LLM judge, self-improving agent loop, bare LLM, verifier
备注: Accepted at the NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development
点击查看摘要
Abstract:We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.
79. 【2610.11461】Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
链接:https://arxiv.org/abs/2610.11461
作者:Shiao Zhu,Lianbo Liu,Sizhen Lyu,Yuzhe Wang,Sheng Li,Takahiro Shinozaki
类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:Natural-language style descriptions, large language models, style descriptions provide, Natural-language style, provide an interpretable
备注:
点击查看摘要
Abstract:Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
80. 【2610.11451】SAIL: Scientific Agentic Intelligence via a Science-Aware Loop
链接:https://arxiv.org/abs/2610.11451
作者:SAIL Model Team:Boyuan Sun,Bryan Dai,Che Liu,Chi Liu,Derek Li,Hongming Piao,Mengzhuo Chen,Xidong Wang,Yan Shu,Yinda Chen,Ziyang Zeng
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
关键词:multi-step research workflows, introduce SAIL, open model, SAIL, research workflows
备注: 16 pages, technical report
点击查看摘要
Abstract:We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows. SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze its task failures and construct training tasks that address the underlying capability gaps. The diagnosis examines search and evidence selection in literature tasks, scientific assumptions and reasoning in coding, and planning and revision in longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools. We repeat this loop over multiple development cycles and train SAIL through supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning. SAIL achieves competitive performance across scientific research tasks with substantially fewer parameters than leading open-weight models.
81. 【2610.11436】Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation
链接:https://arxiv.org/abs/2610.11436
作者:Hongliang Liu
类目:Computation and Language (cs.CL); Cryptography and Security (cs.CR)
关键词:explicitly wrong final, answer judge instructed, final commitment, wrong final, answer judge
备注:
点击查看摘要
Abstract:An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference. We show that adding one colon to a candidate can violate this requirement depending on the presentation of the structured judging request. Numeric references certify the error, and paired interventions distinguish the candidate edit from the integration's presentation choices. On 200 previously unused DROP and GSM8K source clusters, the edit increased Jev's false acceptance from 1.0% to 26.0% with three output labels and from 3.0% to 26.5% with the published four-label grading instruction under sorted JSON keys. Both candidate variants were rejected under insertion presentation. These interactions passed the prespecified statistical correction even though Jev met the control thresholds under both grading configurations and presentations. Most excess acceptances occurred among candidates assigned larger numerical errors. GPT-6 Sol produced no observed cue-condition false acceptances, with missing responses unresolved. The result shows that basic judging competence can coexist with sharply different vulnerability to a fixed candidate edit across logically equivalent request presentations. The reference-aware grammar and compound ordering change limit the finding's operational scope and leave its internal cause unmeasured.
82. 【2610.11430】BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text
链接:https://arxiv.org/abs/2610.11430
作者:Roshan Balaji,Pavan Kumar S,Vasudev Gupta,Sreejith N,Keerthana Sridhar,Nirav Bhatt
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:domain-specific Large Language, limited context windows, Large Language Models, vast biomedical knowledge, encoded vast biomedical
备注:
点击查看摘要
Abstract:While domain-specific Large Language Models (LLMs) have encoded vast biomedical knowledge, their limited context windows often hinder a deep understanding of nuanced relationships within and across texts. To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies. BioBigBird leverages a sparse attention mechanism to process sequences up to 4096 tokens, and its training incorporates a multi-stage process to mitigate noise from the large-scale pre-training corpus. We further enhance its performance by employing a multi-task learning (MTL) framework that jointly optimizes for Named Entity Recognition and Relation Extraction. Comprehensive evaluations on the BLURB benchmark reveal that our MTL-enhanced BioBigBird achieves highly competitive results against state-of-the-art models. Our work contributes an effective methodology for developing powerful, long-context language models for specialized domains, demonstrating the value of extended sequence processing for complex text analysis. Our models are publicly available at this https URL.
83. 【2610.11389】Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation
链接:https://arxiv.org/abs/2610.11389
作者:Navam Obeysekara,Nevidu Jayatilleke
类目:Computation and Language (cs.CL)
关键词:Neural Machine Translation, highly fluent outputs, producing highly fluent, Machine Translation, Neural Machine
备注: 11 pages, 1 figure, 7 tables, Accepted paper at the 13th Conference on Computational Linguistics and Speech Processing (ROCLING) 2026
点击查看摘要
Abstract:Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source. This vulnerability is acute in low-resource settings like Sinhala-to-English, where weak cross-lingual alignment leads to hallucinations. This paper introduces a framework for reference-free hallucination detection in this language pair. We present a 45,000-sample synthetic dataset generated through a probabilistic chain of five linguistically motivated corruption strategies, with a semantic rescue mechanism that uses character-level similarity to distinguish hallucinations from morphological variants. We fine-tune mDeBERTa-v3 for token-level sequence labelling, reaching a token-level F1 of 0.841 +/- 0.001 over three seeds on a source-disjoint test set, and study a three-signal ensemble integrating neural risk scores, sequence log-probabilities, and cross-lingual semantic embeddings (LaBSE). A source-ablation control shows that the detector relies on the Sinhala source rather than on surface artefacts of the corruption process: shuffling or removing the source reduces sentence-level AUROC from 0.970 to chance. We benchmark eight NMT systems spanning five model families and find that detector firings vary by an order of magnitude across architectures.
84. 【2610.11373】From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning
链接:https://arxiv.org/abs/2610.11373
作者:Fengyuan Liu,Yue Wang,Hangxi Guo,Fengyuan Liu,Chenxu Wu,Yanguang Liu,Mengnan Du
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:degrading existing capabilities, learning remains challenging, remains challenging, acquire new skills, skills and knowledge
备注:
点击查看摘要
Abstract:Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids costly parameter updates while achieving competitive or even superior performance to reinforcement learning methods such as GRPO on individual knowledge-intensive and reasoning tasks. This raises a natural question: \textit{Can prompt optimization, as an efficient adaptation approach, be directly applied to continual learning?} Our analysis shows that, under sequential task adaptation, it suffers from catastrophic forgetting, while optimized prompts accumulate rules that overfit to local task distributions. To address these limitations, we propose \emph{Evolving Functional REpertoires} (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones. On a three-task continual-learning stream, EFRE achieves a final average performance 7.50 percentage points higher than GRPO. Moreover, after adaptation to the Bio task, its performance on FinQA decreases by only 1.56 percentage points, compared with 25.10 percentage points for the base prompt optimization method. We further instantiate EFRE in a minimal agent system and observe consistent improvements across different backbone models. Overall, these results demonstrate EFRE's strong performance in continual learning for large language models and highlight its substantial potential for continual learning in advanced agent systems.
85. 【2610.11371】SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
链接:https://arxiv.org/abs/2610.11371
作者:Zhi Rao,Yucheng Zhou,Qianran Sun,Yiqing Huang,Longcan Yuan,Jiayi Hou,Chengwen Yao,Lin Cheng,Donghui Sun,Xiaoxin Chen,Jun Wan
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:Contemporary decoder-only large, demonstrated strong capabilities, decoder-only large language, Contemporary decoder-only, large language models
备注:
点击查看摘要
Abstract:Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{this https URL}{GitHub}, together with models of different sizes to support future academic research.
86. 【2610.11363】UniData: Universal Multimodal Instruction Generation Pipeline
链接:https://arxiv.org/abs/2610.11363
作者:Jiaqi Tang,Yi-Feng Wu,Yuting Zhang,Hao Lu,Bowen Fu,Qing-Guo Chen,Xiaogang Xu,Yuwei Hu,Shiyin Lu,Wei Wei,Lei Zhang,Zhao Xu,Weihua Luo,Qifeng Chen,Ying-Cong Chen
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Multimodal Large Language, Large Language Models, Large Language, Language Models, real-world scenarios
备注: Accepted by EMNLP 2026 Findings
点击查看摘要
Abstract:Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they often face limitations in modality support and struggle with generating multi-round instructions. To address these problems, we introduce UniData, a universal instruction generation pipeline, to transform simple user requirements into multi-round, multimodal instructions. Specifically, UniData first expands user requirements into multiple diverse events. Using these events, UniData then integrates an any-to-any large model for multimodal instruction generation. Finally, UniData enhances data quality by correcting irrelevant and redundant inference flow, leveraging correlations between instruction rounds. To train this pipeline, we also build UniDataset, a dataset comprising 20,000 entries across nine modalities for improved multimodal generation. Our experiments demonstrate that UniData achieves SOTA performance in data quality and can also enhance the understanding and generation capabilities of other multimodal models.
87. 【2610.11354】AdaptEvo: Adaptive Agent Learning with Evolving Supervision
链接:https://arxiv.org/abs/2610.11354
作者:Shijun Wan,Jiancong Xie,Hang Xu,Jin Duan,Qixiong Wang,Xi Xiang,Maofei Que,Yahui Liu,Zhongyu Wei,Mu Chuan
类目:Computation and Language (cs.CL)
关键词:Rule-governed contextual decision, Rule-governed contextual, contextual decision tasks, decision tasks require, tasks require models
备注: 21 pages, 4 figures
点击查看摘要
Abstract:Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence. Written rules can leave gaps in decision guidance and process evaluation, while reference judgments vary in their support from the rules and evidence. To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics. Its Training module uses Confidence-Adaptive GRPO (CA-GRPO) to balance outcome and process rewards according to reference confidence. Its Evolution module synthesizes reusable decision knowledge from recurring failures across training cases and refines process rubrics to detect overlooked errors. To support empirical evaluation, we construct an industrial multimodal content moderation dataset comprising a training set and In-Period and Out-of-Period test sets, with the latter collected under changed rules. Using Qwen3.6-35B-A3B, AdaptEvo achieves 61.9% exact-label accuracy and 72.2% binary decision accuracy on In-Period, exceeding GRPO by 7.5 and 3.7 percentage points, respectively. On Out-of-Period, the policy trained with CA-GRPO retains exact-label accuracy gains over the base model across evaluated checkpoints without injected decision knowledge, while GRPO declines with continued training. CA-GRPO also outperforms the tested fixed reward mixtures on both Out-of-Period metrics.
88. 【2610.11352】RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
链接:https://arxiv.org/abs/2610.11352
作者:Gukhyeon Lee,SangKeun Lee
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Verifiable Rewards, Reinforcement Learning, Learning with Verifiable, trained with Reinforcement, Language models
备注: AACL-IJCNLP 2026
点击查看摘要
Abstract:Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-aware training methods for LMs, which incorporate objectives for uncertainty estimation into training, improve calibration but still exhibit overconfidence under distribution shift, while sacrificing reasoning performance. To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence. Specifically, RL-ARC leverages reasoning confidence as an auxiliary signal for calibrating answer confidence, applying it as reasoning-guided regularization for correct cases and as an overconfidence penalty for incorrect cases. Comprehensive results across ID and OOD settings show that, beyond improving calibration, RL-ARC enables reasoning models to adaptively estimate confidence based on the given question without substantially sacrificing reasoning performance, thereby highlighting the importance of reasoning confidence for training reliable reasoning models.
89. 【2610.11351】Deception by Omission: Language Models Knowingly Hide Their Mistakes
链接:https://arxiv.org/abs/2610.11351
作者:Lucas Florin,Amelie Knecht,Ulysse Schaller,Thilo Hagendorff
类目:Computation and Language (cs.CL)
关键词:Large language models, Large language, increasingly act, Large, agentic rollouts
备注:
点击查看摘要
Abstract:Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed. Users then depend on the model to report what went wrong. An honest model discloses its mistakes, while a deceptive one conceals them. However, it is unclear how current LLMs behave in such situations. In this study, we prefill LLM trajectories with synthetic mistakes. The trajectories resemble real deployments in chat and agentic settings. Models fail to disclose their mistake in 36.4% of chat and 67.1% of agentic rollouts. In 2.4% and 5.3% of rollouts, respectively, they are aware of the mistake in their chain of thought but still deceptively conceal it. Rates vary by model: for instance, Gemini 3.5 Flash knowingly conceals mistakes in up to 19.9% of agentic rollouts. In 11.9% of chat and 51.8% of agentic rollouts, models show no awareness of mistakes, even though they reliably spot them when reviewing the same transcript as an outside observer. Our results show that, as agents take on more tasks with less oversight, users cannot rely on them to self-report possible mistakes. Developers should instead use independent monitors that review agent trajectories, or specifically train models to check their past actions and disclose what they find.
90. 【2610.11337】ype-Checking for Pattern-Based Tree Transformations
链接:https://arxiv.org/abs/2610.11337
作者:C. Aiswarya,Sahil Mhaskar,M. Praveen
类目:Formal Languages and Automata Theory (cs.FL); Computation and Language (cs.CL)
关键词:source pattern, cdot, introduce and study, pattern, target pattern
备注: 26 pages, 9 figures, full version of a preprint accepted at FSTTCS 2026
点击查看摘要
Abstract:We introduce and study pattern-based tree transformations. As an illustrating example, consider a source pattern $(x \cdot y) + (x \cdot z)$ and a target pattern $x \cdot (y + z)$ as a pair. This source pattern matches any expression $e$ of the form $(e_1 \cdot e_2) + (e_1 \cdot e_3)$ (by substituting $x$ with $e_1$, $y$ with $e_2$, and $z$ with $e_3$) and the pair transforms it into the expression $e_1 \cdot (e_2 + e_3)$ as dictated by the target pattern. Note that in this example, the set of expressions that match the source pattern is not a regular tree language. We propose a model of tree transformations given by a finite representation of a (possibly infinite) set of such (source pattern, target pattern) pairs. The expressive power of this model comes at the cost of undecidability of checking equivalence. Nevertheless, we show that the type-checking problem is decidable for our model of pattern-based tree transformations. The type-checking problem asks whether applying a given transformation to trees having a given regular property (type) preserves the property. Our decision procedure is by a reduction to the emptiness problem of alternating tree automata.
Comments:
26 pages, 9 figures, full version of a preprint accepted at FSTTCS 2026
Subjects:
Formal Languages and Automata Theory (cs.FL); Computation and Language (cs.CL)
ACMclasses:
F.1.1; F.4.1; F.4.2; D.3.1; F.4.3
Cite as:
arXiv:2610.11337 [cs.FL]
(or
arXiv:2610.11337v1 [cs.FL] for this version)
https://doi.org/10.48550/arXiv.2610.11337
Focus to learn more
arXiv-issued DOI via DataCite</p>
91. 【2610.11332】ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
链接:https://arxiv.org/abs/2610.11332
作者:Houcheng Jiang,Mao Zheng,Mingyang Song,Qiyong Zhong,Jie Sun,Tianyu Zhang,Junfeng Fang
类目:Computation and Language (cs.CL)
关键词:resulting capability degradation, Structured pruning reduces, hinder subsequent on-policy, Structured pruning, deployment cost
备注:
点击查看摘要
Abstract:Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning. RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions. Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons. Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery. These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.
92. 【2610.11316】MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface
链接:https://arxiv.org/abs/2610.11316
作者:Jianpeng Cheng,Guangyu Sun,Aashu Singh,Benyu Zhang,Haixing Dai,Hossein Mansour,Jiangfan Zhang,Shlok Kumar Mishra,Wei Sun,Xuanming Cui,Yanli Liu,Qi Guo,Max Xiangjun Fan,Jun Xiao
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:free-form text generation, models output constrained, output constrained decisions, System One models, text generation
备注:
点击查看摘要
Abstract:System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder. To scale effectively across both small closed-set ( 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment. We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.
93. 【2610.11314】From Retrieval to Reconstruction: Constructing Evolvable Cognitive Memory for Long-Term Dialogue
链接:https://arxiv.org/abs/2610.11314
作者:Zirui Liao,Zhengxian Wu,Zhuohong Chen,Yunyao Yu,Xiaoyu Liu,Yifan Xu,Haoqian Wang
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Large Language, Language Models, agents require memory, require memory systems
备注: Accepted to EMNLP 2026 (main conference). 21 pages
点击查看摘要
Abstract:Large Language Models (LLMs) serving as long-term dialogue agents require memory systems that support reliable reasoning over extended interactions. However, existing Retrieval-Augmented Generation (RAG) frameworks typically treat memory as passive storage, making it difficult to distinguish source-attributed beliefs from unattributed event/fact records and to connect evidence dispersed across sessions. We introduce CogMem, a cognitive memory architecture based on the PEC$^2$F (Person-Event-Concept-Claim-Fact) graph schema. Dedicated Claim nodes preserve the source and target of subjective statements, while Fact and Event nodes represent semantic and episodic knowledge. Dialogue turns are incrementally converted into provenance-aware graph records, consolidated into higher-level facts, and reconciled into temporally scoped Claim views when the same source provides conflicting updates. For retrieval, a rule-based controller driven by LLM intent parsing composes four deterministic graph operators---anchoring, traversal, intersection, and evidence grounding---to reconstruct query-relevant context. Experiments on LoCoMo and LongMemEval show strong performance, especially on multi-hop, temporal, and knowledge-update tasks. Ablations and a semantic-collapse probe support complementary contributions from epistemic separation, consolidation, and agentic retrieval. Code: this https URL.
94. 【2610.11305】BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models
链接:https://arxiv.org/abs/2610.11305
作者:Shuai Guo,Yidong Cui
类目:Computation and Language (cs.CL)
关键词:receiving genuinely relevant, receiving directional user, genuinely relevant evidence, directional user pressure, relevant fact
备注: 18 pages, 13 figures, 18 tables. Includes appendices
点击查看摘要
Abstract:A language model may revise the same proposition after receiving genuinely relevant evidence or after receiving directional user pressure that adds no relevant fact. The observable response shift alone therefore does not reveal which source drove the change. We introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition. BeliefScope crosses Evidence and Pressure with factor-specific local controls and measures response changes through probability reports, categorical judgments, and action recommendations on channel-appropriate scales. To determine when these observable contrasts support reliable attribution, we evaluate the observation design under controlled synthetic conditions. Known-truth recovery and targeted ablations establish where Evidence- and Pressure-related effects can be separated, while semi-synthetic stress tests map how that recoverability changes as the observation process becomes noisier and more heterogeneous. Across a 36-family Qwen/Llama study, with targeted 12-family checks that also include Gemma3-12B, the resulting profiles show substantial evaluation-context dependence: broad model-level differences can change under matched controls, decoding, or response interfaces, while some narrower within-model patterns remain stable. Instruction interventions further show that reduced target-aligned Pressure following can reflect either stable resistance or movement in the opposite direction. BeliefScope summarizes these measurements as a conditional belief-response profile that keeps diagnostic effects tied to the evaluation conditions under which they are observed, together with explicit validity boundaries for that diagnosis.
95. 【2610.11291】When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
链接:https://arxiv.org/abs/2610.11291
作者:Siyan Zhao,Yonggan Fu,Jindong Jiang,Shih-Yang Liu,Song Bian,Byung-Kwan Lee,Sharath Turuvekere Sreenivas,Wenliang Dai,Hanrong Ye,Aditya Grover,Pavlo Molchanov
类目:Computation and Language (cs.CL)
关键词:transferring teacher capabilities, OPD, popular for transferring, Semi-OPD outperforms OPD, teacher-student pairs
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.
96. 【2610.11287】REMORY: Learning Residual Memory for Context Compaction
链接:https://arxiv.org/abs/2610.11287
作者:Hanchen Xia,Baoyou Chen,Yutang Ge,Naihao Deng,Senqiao Yang,Zilong Dong,Weihao Yuan,Siyu Zhu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:finite context window, Long-horizon agents compact, context window, subsequent decision, finite context
备注:
点击查看摘要
Abstract:Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.
97. 【2610.11275】Phonological Interference in Multilingual Speech Models
链接:https://arxiv.org/abs/2610.11275
作者:Moran Yanuka,Raja Giryes,Moris Alper
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:smallest sound units, Phoneme-level models transcribe, language, distinguish words, phonemes
备注: Preprint
点击查看摘要
Abstract:Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.
98. 【2610.11270】Gated Memory: Admission-Controlled Memory Formation for Conversational AI
链接:https://arxiv.org/abs/2610.11270
作者:Preeti Saraswat,Divya Neelagiri,Ajay Manoj
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:persistent vector stores, Personalized conversational, long-term memory systems, vector stores, conversational AI relies
备注:
点击查看摘要
Abstract:Personalized conversational AI relies on long-term memory systems that extract facts from user utterances and store them in persistent vector stores. Despite progress in retrieval, deduplication, and lifecycle management, the formation stage, the moment a fact is first written to storage has received almost no principled attention. We identify this as the binding constraint on memory quality in production systems. Critical contextual signals, such as the distinction between a permanent user attribute and a transient situation, exist only in the original utterance and are irreversibly lost the moment extraction produces a subject-relation-object triple. No downstream process can recover them. We propose Gated Memory, a lightweight, modular formation framework that interposes two decision checkpoints between conversation and storage: an admission gate that evaluates every candidate fact against the full utterance context before extraction runs, and a conditional enrichment stage that grounds admitted facts through an entity scope taxonomy with privacy constraints. The gate evaluates only the current exchange while using prior turns as read-only reference context, and produces a structured formation record. Admitted content is decomposed into atomic facts, each categorized, tagged with provenance (directly stated versus inferred), scoped to its condition of applicability, and grounded in resolved time and place, subject to a constraint that no entity absent from the context may be asserted. On the LoCoMo-10 benchmark with atypical emotional density in utterance data, Gated Memory achieves an overall +2.6% relative improvement in LLM-judge accuracy over a strong baseline with identical retrieval and generation, establishing formation quality as a measurable constraint on memory performance.
99. 【2610.11247】Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
链接:https://arxiv.org/abs/2610.11247
作者:Lei Zhao,Qichao Zhao,Bowen Zuo,Qishi Zhan
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:enables effective capability, effective capability transfer, On-policy distillation, enables effective, language models
备注: 50 pages. Code: [this https URL](https://github.com/leizhao7/opd-learning-signals)
点击查看摘要
Abstract:On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA $0.98$ across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested. Code is available at this https URL.
100. 【2610.11245】Read What Matters: Query-Adaptive Quantization for KV Caches
链接:https://arxiv.org/abs/2610.11245
作者:Siddharth Bhandari,Lucas Gretta,Krishna Balasubramanian,Shiva Kasiviswanathan
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Performance (cs.PF)
关键词:future queries, query, KV-cache entries, bits, budget
备注: 50 pages, 3 figures, 12 tables
点击查看摘要
Abstract:KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places. We study this mismatch using separate budgets for retained bits and bits fetched per query. ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions. For each query, it allocates key-channel prefixes using the query, computes attention from the reconstructed keys, and then allocates value-token prefixes using that attention. Stored entries remain unchanged. Each stage optimizes a calibrated distortion objective under a fixed budget; we prove exact allocation under diminishing refinement gains and relate these objectives to attention-output error. We also exhibit a finite-dimensional attention family where query-dependent access strictly outperforms every query-independent reader at the same read budget, even with unrestricted competing encoders and decoders. Across six base models, reading four bits on average from an eight-bit cache increases C4 perplexity by at most 0.66%, using about one quarter of the logical reads and half the retained capacity of a 16-bit cache. It is consistently more accurate than storing and fully reading four bits at the same payload-read budget. Retaining more bits than each query fetches is aimed at long-context decoding, where the cache bytes moved per step, rather than the weights, dominate cost. Long-context question answering and retrieval on two instruction-tuned models provide additional quality evidence. On the tested 8K-token, batch-one, single-layer workload on an NVIDIA A10G, a restricted eight-bit ReadKV reader with a two-bit mean payload-read budget has 39% lower latency than the tested TurboQuant codec.
Comments:
50 pages, 3 figures, 12 tables
Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL); Performance (cs.PF)
Cite as:
arXiv:2610.11245 [cs.LG]
(or
arXiv:2610.11245v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2610.11245
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
101. 【2610.11233】MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
链接:https://arxiv.org/abs/2610.11233
作者:Leran Chen,Lingnan Kong,Zile Cai
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Multimedia (cs.MM)
关键词:evidence, core challenge, challenge in short-video, short-video fact-checking, fact-checking is identifying
备注: 33 pages, 2 figures, 20 tables
点击查看摘要
Abstract:A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.
102. 【2610.11223】SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning
链接:https://arxiv.org/abs/2610.11223
作者:Weizhe Xu,Jialiang Fan,Mengyu Liu,Fanxin Kong
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Logic in Computer Science (cs.LO)
关键词:Large Reasoning Language, Reasoning Language Models, enable multi-step reasoning, Language Models, constraint violations unresolved
备注: Video: [this https URL](https://youtu.be/dbU7WskCNgY)
点击查看摘要
Abstract:Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.
103. 【2610.11216】he Lattice of Transition Laws
链接:https://arxiv.org/abs/2610.11216
作者:T. Y. Tsui,Jiatao Gu,Lingjie Liu
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:specialising in discrete, categories of generative, diffusion specialising, specialising, models
备注:
点击查看摘要
Abstract:Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens. Recent work seeks to combine the advantages of the two models, and each hybrid fixes its decoding schedule by design. In this paper, we ask whether the performance of decoding schedules of one model can be predicted before decoding at a fixed number of steps. We describe diffusion, AR, and models in between as paths on one corruption lattice, and define the cost of a schedule as the dependence its parallel steps discard. The cost shows that the fewest steps of a zero-cost schedule are set by the geometry of the data, in the same way for tokens and for continuous fields. In particular, for data that are Markov on a graph and dependent along its paths, the fewest steps equal the graph's treedepth, which is logarithmic in the length of a sequence and linear in the side length of a grid. With fewer steps than the treedepth, every schedule pays a positive cost, whose ranking we predict before decoding with a kernel of pairwise dependence estimated from pretrained weights. Across text generation, image generation, and video generation, we verify most of the predictions about the rankings of different schedules under different metrics and benchmarks. This work therefore provides a design principle for decoding for future AR models, diffusion models, and anything in between. Our code is available at this https URL.
104. 【2610.11214】Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
链接:https://arxiv.org/abs/2610.11214
作者:Kaicheng Xiao,Liran Dong,Haotian Li,Guoliang Xing
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:KV-cache quantization compresses, KV-cache quantization, representative approaches, approaches to tackling, tackling the storage
备注:
点击查看摘要
Abstract:KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.
105. 【2610.11196】Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
链接:https://arxiv.org/abs/2610.11196
作者:Yulin Sun,Kele Xu,Yong Dou
类目:ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large audio-language models, alter text-reasoning decisions, Large audio-language, exploit multimodal evidence, listening is unnecessary
备注: 24 pages, 5 figures, 20 tables
点击查看摘要
Abstract:Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
106. 【2610.11183】RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
链接:https://arxiv.org/abs/2610.11183
作者:Shunyuan Zhou,Hao Chen,Tianyu Wang,Goose Lin,Zaiyuan Wang,Haiying Zhao
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:previously gave correctly, guarantee factual correctness, previously gave, evidence, answer
备注:
点击查看摘要
Abstract:Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly. Standard accuracy measures obscure this behavior by combining answer replacement with preexisting errors. We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation. The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text. We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set. We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations. Instructions that prioritize documents consistently produce higher MR than those permitting reliance on prior knowledge. Averaged over models and positions, the gap ranges from 10.9 to 13.5 percentage points across the three QA datasets. Mean MR follows End $$ Beginning $$ Middle under both policies, although individual models do not uniformly follow this ordering. A separate paired audit of 500 questions and two checkpoints supports increased harmful override without establishing a corresponding improvement in beneficial correction. These findings distinguish evidence adherence from factual reliability and motivate evaluating whether retrieved evidence preserves, replaces, or corrects a model's answers.
107. 【2610.11160】LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs
链接:https://arxiv.org/abs/2610.11160
作者:Xiaobing Yu,Peijie Qiu,Jin Yang,Xuanzhao Dong,Weiwei Ma,Zhaoqi An,Xiaoqi Zhao,Xiaofeng Liu
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM)
关键词:LLMs requires storing, requires storing thousands, Lifelong editing, editing of LLMs, LLMs requires
备注: EMNLP 2026 Main Conference Long Paper
点击查看摘要
Abstract:Lifelong editing of LLMs requires storing thousands of edits after acquisition. A widely used family of approaches attaches one LoRA adapter per edit, which preserves behavior but grows linearly in storage. To address this challenge, we propose LadderEdit, a method that compresses each LoRA adapter after it is acquired. Each edit is first stored at low rank as a cheap sketch. We then check whether this sketch still satisfies the rewrite, generalization, and locality contract on probe prompts. Edits that pass keep the sketch; those that fail are promoted to a higher rank along a ladder until the contract is met. Because every edit retains some representation, coverage is maintained, and only hard edits consume more rank. Across ZsRE, CounterFact, and WikiBigEdit benchmarks on LLaMA-3-8B, Mistral-7B, and Qwen2.5-7B, LadderEdit tracks exact LoRA storage at 5.2x less memory and remains effective at 50,000 sequential edits.
108. 【2610.11159】Local Prototype Reconstruction for Text-Compatible Speech-to-LLM Bridge Pretraining
链接:https://arxiv.org/abs/2610.11159
作者:Xinnian Zhao,Chia-Hua Wu,Pu Wang,Hugo Van Hamme
类目:Computation and Language (cs.CL); Sound (cs.SD)
关键词:large language model, frozen large language, systems often connect, language model, large language
备注:
点击查看摘要
Abstract:Speech-to-LLM systems often connect a frozen speech encoder to a frozen large language model (LLM) through a small trainable bridge. The bridge is usually treated as plumbing, but it in fact defines the geometry of the speech-to-LLM interface, and the pretraining objective decides whether that interface provides a reusable initialization for downstream tasks. We study a transferable bridge through two complementary properties: global alignment with the text side, and local lexical manifold compatibility, where bridge embeddings remain close to the frozen LLM's input-embedding neighbourhoods. We make this property measurable with a fixed, head-free, timestamp-free diagnostic that applies to any objective, and show that next-word prediction (NWP) and sentence-level contrastive pretraining do not fully capture token-level lexical compatibility. We then introduce Local Prototype Reconstruction (LPR), a lightweight training-only regularizer that requires each aligned bridge token to be reconstructable from a small neighbourhood of frozen LLM token embeddings, with a hard single-prototype anchor as its limiting case. On multilingual ASR and speech translation, LPR improves transfer, with the largest gains on translation and low-resource adaptation. Crucially, our independent diagnostic correlates with downstream gains across objectives, suggesting that lexical manifold compatibility is predictive of reusability for speech-to-LLM bridges.
109. 【2610.11152】Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning
链接:https://arxiv.org/abs/2610.11152
作者:Minchan Kwon,Seunghee Koh,Sunghyun Baek,Minsung Bae,Junmo Kim
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:LLM agents increasingly, agents increasingly improve, updating parameters, LLM agents, increasingly improve
备注: NeurIPS 2026 Spotlight (Negative Results Track)
点击查看摘要
Abstract:LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters. This process is often described as in-context reinforcement learning (ICRL). Whether in-context learning (ICL) can actually play the role of RL, however, has not been tested. We study this question in its simplest form, direct ICRL, where the model conditions directly on raw trajectory-reward pairs, and ask whether the reward acts as a learning signal. Through controlled experiments on four benchmarks across six models, we find that the reward is read, but its effect is small: flipping, randomizing, or removing the reward leaves the improvement curve almost unchanged, and this holds even under meta-prompts that explicitly instruct the model to explore, exploit, or reason over rewards. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for agent memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.
110. 【2610.11140】ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis
链接:https://arxiv.org/abs/2610.11140
作者:Weiwei Ma,Xiaobing Yu,Peijie Qiu,Jin Yang,Zhaoqi An,Xuanzhao Dong,Xiaoqi Zhao,Xiaofeng Liu
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:resolve diagnostic uncertainty, clinicians escalate, escalate from cheap, cheap to costly, costly tests
备注: EMNLP 2026
点击查看摘要
Abstract:Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.
111. 【2610.11136】he "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
链接:https://arxiv.org/abs/2610.11136
作者:Yuchen Miao,Zijun Wang,Chang Han,Yurui Shi,Mingtai Zhang,Siyang Xu
类目:Computation and Language (cs.CL)
关键词:Presupposing the boundaries, study closed-loop bias, closed-loop bias governance, Presupposing, Dutch government documents
备注: 19 pages, 6 figures. Accepted at EMNLP 2026
点击查看摘要
Abstract:Presupposing the boundaries of bias is itself a form of bias. We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning. Existing methods face three challenges: (i) discriminative classifiers capture surface regularities but lack normative grounding; (ii) zero-shot LLMs often adopt generic viewpoints and over-flag ambiguous administrative language; and (iii) fixed taxonomies inherit the Closed-World Assumption, missing emerging local targets. We propose MARS-Gov, a standpoint-aware multi-agent framework that combines legal retrieval, open-set target screening, specialized jurors, conservative routing, and rewrite verification. When screening finds an uncovered group, MARS-Gov instantiates a dynamic "10th juror" to deliberate outside the fixed panel. On DGDB, MARS-Gov sets a new SOTA with 0.880 F1, outperforming the strongest zero-shot LLM detector by 20.2 points (29.8% relative) and the best supervised Dutch encoder by 6.8 points, while reducing unnecessary interventions to 2.5%. Leave-One-Category-Out (LOCO) evaluation recovers held-out categories with 85.1% Correct@1 and 93.8% Correct@3.
112. 【2610.11135】Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged?
链接:https://arxiv.org/abs/2610.11135
作者:Unggi Lee,Haeun Park
类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
关键词:Knowledge tracing, learners, Knowledge, usable model, platform starts
备注: 41 pages, 7 figures. Code and result summaries: [this https URL](https://anonymous.4open.science/r/jevkt-coldstart)
点击查看摘要
Abstract:Knowledge tracing (KT) models need many logged learners, so a new course or platform starts without a usable model. In LLM-based KT the LLM generates the answer, which we call System-Two; it is either fine-tuned on the target data or reasons and votes over ten samples, which is slow and gives coarse probabilities. We ask whether an off-the-shelf System-One LLM, which returns a probability for a typed question directly in a single pass, can perform KT when few or no learners are logged. On seven datasets, Jev without any data from the target platform reaches a mean AUC of .706, above the best of 28 deep KT models trained on 8 learners (.689) and above System-Two Thinking-KT on all seven datasets (.650) at about 1/100 of its API cost. Adding examples and a similar-learner statistic from the logged learners (JevKT) raises this to .722; JevKT stays significantly ahead of deep KT up to 16 learners and ahead on average up to 64, and supervised KT catches up between 64 and 128 learners. Among the readers we tested, the gain is specific to Jev, since three other LLMs queried with the byte-identical typed request through the official System-One adapter fall below it on all seven datasets, and reader swaps and contamination checks find no evidence that the input format or memorised data explain the gain. For new learners the advantage holds from their first interactions, whereas on unseen items with all learners logged, deep KT remains ahead.
113. 【2610.11132】SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning
链接:https://arxiv.org/abs/2610.11132
作者:Kenan Tang,Andong Hua,Chengxuan Qian,Saket Tiwari,Yao Qin
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Supervised fine-tuning, SFT model, equips large language, SFT, parent model
备注: 37 pages
点击查看摘要
Abstract:Supervised fine-tuning (SFT) equips large language models (LLMs) with specialized capabilities, but often comes at the cost of forgetting the general capabilities of their parent models (i.e., the pretrained models before fine-tuning). This trade-off is especially limiting for queries that require both specialized and general capabilities. We introduce SFT-as-context, a training-free method in which the parent model uses the SFT model's response as context to answer the query. This allows the parent model to acquire fine-tuned capabilities from the SFT response through in-context learning while preserving its own general capabilities. Across 19 parent-SFT model pairs and 11 benchmarks, SFT-as-context remains close to the SFT models on fine-tuned capabilities, with gaps of only 2.2 and 2.1 percentage points on AIME 2024 and LiveCodeBench and 2.0 macro MAE on NutriBench-English, while staying within 2.2 percentage points of the parent models on general capabilities on average. Remarkably, it can solve queries requiring both fine-tuned and general capabilities, even when neither the parent nor SFT model succeeds alone. This approach also extends beyond parent-SFT pairs: responses from a small open-source SFT model can improve a strong closed-source LLM, outperforming either model alone. Furthermore, we use a Bayesian framework to derive theoretical guarantees that bound the error of SFT-as-context relative to the SFT model on fine-tuned capabilities and to the parent model on general capabilities. In addition, we visualize the attention weights and find that the parent model attends more to useful SFT responses and less to irrelevant ones, suggesting that selective attention helps the parent model use the SFT response through in-context learning.
114. 【2610.11129】GameCommBench: A Unified Benchmark and Type-Aware Evaluation for AI-Generated Game Commentary
链接:https://arxiv.org/abs/2610.11129
作者:Qirui Zheng,Zhengteng Lin,Yunyi Xiao,Junhao Li,Keyuan Cheng,Xingbo Wang,Yongyi Wang,Lingfeng Li,Yunlong Lu,Wenxin Li
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:requiring multimodal perception, open-ended generation task, generation task requiring, task requiring multimodal, multimodal perception
备注:
点击查看摘要
Abstract:Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary. We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type. We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary. We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators. Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks. Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.
115. 【2610.11111】Lapras: Latent Reasoning for Time Series Language Models
链接:https://arxiv.org/abs/2610.11111
作者:Yuliang Chen,Yu Yvonne Wu,Patrick Langer,Arvind Pillai,Sudarshan Regmi,Martin Maritsch,Juncheng Liu,Robert Jakob,Thomas Kaar,Tess Z. Griffin,Lisa Marsch,Michael V. Heinz,Nicholas C. Jacobson,Andrew Campbell
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Time Series, time series understanding, Series, Lapras, reasoning
备注:
点击查看摘要
Abstract:Time Series Language Models (TSLMs) offer a promising path toward time series understanding by reasoning over temporal signals and producing natural language answers and explanations. A common approach is Chain-of-Thought (CoT), which generates step-by-step rationales linking relevant signal patterns to final answers. Although these models learn from reference CoT traces during post-training, generating faithful descriptions of input time series at inference remains challenging. Expressing high-dimensional, continuous temporal representations in discrete language tokens may cause the model to neglect task-relevant patterns or describe them inaccurately. Because later reasoning steps build on these descriptions, early errors propagate, leading to incorrect answers with plausible explanations that are inconsistent with the input signal. We propose Lapras (Latent Post-trained Reasoning Across Series), a post-training framework that equips TSLMs with latent reasoning. A model trained with Lapras reasons through a sequence of continuous thoughts in the joint time series-language space, producing text only for the final answer. It learns this through teacher-student self-distillation, where a teacher trained on CoT reference traces reasons explicitly through text. The student aligns its hidden states with the teacher's at the answer stage, transferring the teacher's reasoning ability into its latent computation. We evaluate Lapras across four TSLM backbones on five time series question answering benchmarks. Lapras improves average F1 by up to 10.79% over explicit CoT while generating 23.9x fewer tokens. Lapras's continuous thoughts can also be decoded into readable reasoning traces via standard language decoding, preserving textual explanations. Together, these results highlight Lapras as a promising post-training paradigm for efficient, effective, and interpretable TSLM reasoning.
116. 【2610.11069】Clinician use of language models diverges from how the models are evaluated
链接:https://arxiv.org/abs/2610.11069
作者:Krithik Vishwanath,Haitong Lin,Anton Alyakin,Jin Vivian Lee,D. Brock Hewitt,Jie J. Yao,William Robert Small,Hammad A. Khan,Cordelia Orillac,Aakaash Varma,Brandon Ye,Daniel Alexander Alber,Gustavo Stolovitzky,Batia Wiesenfeld,Oded Nov,Wei Wu,Kang Zhang,Yindalon Aphinyanaphongs,Tim Requarth,Eric Karl Oermann, TheInternational Digital Twin Consortium in Healthcare,Medicine
类目:Computation and Language (cs.CL)
关键词:Large language model, readiness rest largely, Large language, curated cases, deployed to clinicians
备注:
点击查看摘要
Abstract:Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts performance in deployment only to the extent that its items resemble real use, yet whether benchmarks reflect the work these systems receive has rarely been measured. Here we analyze 127,833 queries sent by 6,342 physicians, advanced practice providers and nurses in 35 specialties to an institutional assistant during an eight-month roll-out. We characterize each query with RCQ-Map, a clinician-validated framework grounded in taxonomies of clinical questions and of LLM evaluation, which records its task, intent, answerability, missing information and potential harm. Documentation and administration (36.2%) and knowledge retrieval (28.9%) made up nearly two-thirds of use, and diagnosis 3.7%; more than a third of queries could not be answered well as posed. Applying RCQ-Map to 58 public benchmarks drawn from major evaluation suites and frontier model reports, which we assemble into the Clinical AI Benchmark Atlas, showed that the median benchmark contained no documentation requests and shared 31% of the task mix of real use, less than an even spread across task categories would. Benchmarks in suites designed to resemble clinical practice were individually no closer to real use than those used in frontier model reports. Benchmark scores therefore say little about how clinical AI performs on most of the work it is actually given, and evaluation should be matched to real clinical use.
117. 【2610.11064】Measuring and Mitigating Solution Mode Collapse in RLVR
链接:https://arxiv.org/abs/2610.11064
作者:Liv G. d'Aliberti,Marwa Abdulhai,Sofiia Druchyna,Peter Henderson,Manoel Horta Ribeiro
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:reinforcement learning, learning with verifiable, verifiable rewards, model, model produces
备注:
点击查看摘要
Abstract:A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution diversity changes under RLVR post-training. We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated. We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly. A solution found once is, therefore, practiced as often as one found repeatedly. Across three model scales, two RL objectives, and harder task constructions, replay improves both how often a policy succeeds and how many different ways it can succeed.
118. 【2610.11033】FedAlphaEdit: Null-Space-Aligned Merging for Collaborative Knowledge Editing
链接:https://arxiv.org/abs/2610.11033
作者:Sota Sugawara,Yukihiko Okada
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:single large language, large language model, large language, private knowledge edits, raw edit requests
备注: 12 pages, 2 figures, 8 tables. Includes appendices (proofs, implementation reconciliation, experimental details). Code: [this https URL](https://github.com/soutasuga/FedAlphaEdit)
点击查看摘要
Abstract:Multiple institutions may each hold their own private knowledge edits and wish to integrate them into a single large language model without sharing raw edit requests. Null-space-constrained editing methods such as AlphaEdit mathematically guarantee that each update leaves unrelated knowledge intact, while collaborative frameworks such as CollabEdit aggregate edits from multiple clients without data sharing. Combining the two appears trivial. However, we show that this naive combination fails structurally, and we identify its cause. Guided by this analysis, we propose FedAlphaEdit. To our knowledge, this is the first collaborative knowledge editing framework that aligns both local editing and the server-side merging rule under a single null-space principle for preserving existing knowledge. FedAlphaEdit builds on null-space-aligned merging, in which clients share projected statistics and the server provably recovers the result of editing everything in one place under a one-shot idealization. Empirically, the proposed method repairs the collapse and brings edit success and preservation simultaneously close to the level of centralized editing across two architecture families. FedAlphaEdit thus lets institutions that cannot share raw edit data, such as hospitals and financial firms, jointly maintain a shared model that closely approximates editing all facts in one place.
119. 【2610.11015】Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators
链接:https://arxiv.org/abs/2610.11015
作者:Riqiang Wang,Elena Khasanova,Harsh Saini,Lex Konnelly,Parsa Kavehzadeh,Matthias Lee,Mohamed Attia
类目:Computation and Language (cs.CL)
关键词:voice agents gain, popularity commercially, include more realistic, naturalness behaviors, realized naturalness behaviors
备注: Accepted to the UserSim @ NeurIPS 2026 workshop (non-archival)
点击查看摘要
Abstract:As voice agents gain more popularity commercially, the user simulators used to evaluate the deployed agents are also being developed to include more realistic, variable, and diverse speech naturalness behaviors -- disfluency, interruption and backchanneling. The quality of the user simulator directly affects the validity of agent evaluation results. However, we find that most studies so far have not examined in detail whether the intended configuration for these behaviors is realized in the simulation. In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling. We find that prompting for these behaviors is unreliable and produces speech inconsistent with the instructions, placed and distributed less naturally than the instruction implies. In contrast, our rule-based, model-free algorithm produces controllable and diverse naturalness behaviors more aligned with natural speech. Our results suggest that LLMs not purpose-trained for user simulation are not sufficient on their own to represent authentic user behavior, and that linguistically informed deterministic approaches or specialized models are needed to close the gap; auditing and reporting realized naturalness behaviors, rather than configured settings, is what makes that gap visible.
120. 【2610.10988】Back in Style: A Sociolinguistic Approach to Authoring and Measuring Persona Fidelity in User Simulation
链接:https://arxiv.org/abs/2610.10988
作者:Lex Konnelly,Elena Khasanova,Riqiang Wang,Matthias Lee,Harsh Saini,Parsa Kavehzadeh
类目:Computation and Language (cs.CL)
关键词:gain commercial popularity, agentic systems gain, systems gain commercial, simulators increasingly serve, user simulators increasingly
备注: Accepted to the UserSim @ NeurIPS 2026 workshop (non-archival)
点击查看摘要
Abstract:As agentic systems gain commercial popularity, user simulators increasingly serve as measurement instrument for their evaluation. However, the fidelity of simulated users in comparison to real human users is generally low, and typically assessed by costly, subjective LLM judges. In this pilot study, we ask whether fidelity can instead be measured deterministically by treating a user persona sociolinguistically: as a social type that emerges from observable linguistic style, rather than one predicted by labels or descriptions a model must extrapolate into behaviour. We author personas as concrete stylistic rates, which lets us transfer two established, model-free instruments -- authorship-verification stylometry and lexicon-based content analysis -- as fidelity diagnostics. We A/B-test the sociolinguistic schema against a flat descriptive baseline across five task-oriented customer-service agents. Results show that the sociolinguistic schema improves both stylistic adherence and stylometric distinguishability for most of the tested models, with a caveat that persona style fidelity does not necessarily equal persona "naturalness". We argue that a sociolinguistic approach to persona design is a promising path towards more diverse and representative user personas, and that these metrics are most valuable in an error-attribution analysis, localizing where fidelity breaks down. This is a first step towards interventions that move user simulations closer to faithful renderings of diverse and variable linguistic outputs.
121. 【2610.10971】When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection
链接:https://arxiv.org/abs/2610.10971
作者:M. Mikail Demir,M. Abdullah Canbaz
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large language models, Large language, research and drafting, cited source, York State Court
备注: 9 pages, 2 figures, 7 tables. Published in the 21st International Conference on Artificial Intelligence and Law (ICAIL 2026), Singapore. Dataset: [this https URL](https://github.com/mmikaildemir/PARCEL)
点击查看摘要
Abstract:Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even when the full opinion text is provided. Across models, missing support is harder to detect than direct contradiction, and fabricated but plausible citations cause the largest drop in performance. Overall, PARCEL provides a practical benchmark for testing claim-level groundedness in legal RAG systems.
122. 【2610.10946】AI4Fire: Evaluating Large Language Models on Wildfire Tasks
链接:https://arxiv.org/abs/2610.10946
作者:Yue Zhao,Xiyang Hu,Zuobin Xiong,Zhangyu Wang,Ruolin Li
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large language models, entering wildfire management, Large language, property and lives, overstated evaluations
备注: 51 pages. Code: [this https URL](https://github.com/yzhao062/AI4Fire)
点击查看摘要
Abstract:Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a smoke-free reference frame from the same camera. AI4Fire runs six core models bare and grounded on five wildfire tasks, zero-shot; a sweep adds 29 more. Our literature search on fire tasks found 138 works; none combines this roster, task coverage, and paired bare and grounded runs. We report three findings. (1) Grounding helped most where the addition carried the answer: a read-only SQL tool lifted every core model's database accuracy from at most 16 to at least 88 percent. (2) Simple rules were hard to beat: no core model outperformed repeating today's staffing count, and two open-weight models mostly copied the median of similar earlier fire-days, a worse forecast. (3) Public releases carry hazards: a fire-danger column separates the holdout perfectly, and 67 aerial fire frames carry smoldering or fire-free labels read from a clipped thermal maximum. We release prompts, responses, scores, code, and the survey record.
123. 【2610.10942】StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
链接:https://arxiv.org/abs/2610.10942
作者:Daksh Raghuvanshi,Ved Vedere,Yifan Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Reinforcement learning environments, improving large language, Reinforcement learning, large language model, benchmarks remain static
备注: 22 pages, 4 figures, 8 tables
点击查看摘要
Abstract:Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.
124. 【2610.10920】Language Models for Page-Level Layout Decisions in E-commerce Search
链接:https://arxiv.org/abs/2610.10920
作者:Varun Joshi,Eva C. Song,ChengXiang Zhai
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:E-commerce search pages, E-commerce search, modern E-commerce search, critical touchpoints, touchpoints for millions
备注: Accepted at the OARS Workshop, ACM RecSys 2026
点击查看摘要
Abstract:E-commerce search pages are critical touchpoints for millions of online shoppers. While traditional search engines return a ranked list of results, modern E-commerce search pages increasingly incorporate recommender system modules -- for example, secondary stacks that surface alternative product groupings at specific positions. When introduced appropriately, secondary stacks can improve user engagement; however, suboptimal placement may disrupt browsing flow and degrade the primary results. Unlike traditional search ranking, where evaluation techniques such as interleaving are well established, evaluating page-level layout changes e.g., when and where to insert a secondary stack remains challenging without costly online A/B testing. To address this, we study offline methods for evaluating whether a given layout decision -- specifically, the inclusion of a secondary stack at a particular position -- is beneficial to users. We investigate language models as scalable evaluators by comparing direct prompt-based, prompt-derived feature, and representation-based methods. Our results show that representation-based approaches consistently outperform prompt-based judging in predicting user engagement, suggesting they provide a reliable foundation for offline layout evaluation in E-commerce search.
125. 【2610.10918】Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
链接:https://arxiv.org/abs/2610.10918
作者:Hala Almaghout,Christian Federmann,Qin Gao
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Machine Translation, Error Span Annotation, judgment for Machine, language pairs
备注: Accepted at EMNLP 2026
点击查看摘要
Abstract:Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation. With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation. In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators. We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains. Our results show that while LLM agreement with human annotators exceeds agreement between human annotators for some evaluation tasks, both vary substantially across annotation schemes, language pairs and domains and remain unreliable for most settings. Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation. Our results highlight the potential for improvement for both human and LLM annotation performance, possibly through human-LLM collaborative annotation pipelines that address the reliability issues identified in this work.
126. 【2610.10878】Stochastic Teacher Intervention for Agentic On-Policy Distillation
链接:https://arxiv.org/abs/2610.10878
作者:Junnan Liu,Linhao Luo,Zhijun Chen,Qianren Mao,Thuy-Trang Vu,Gholamreza Haffari
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:efficiently transfers capabilities, On-policy distillation, student language model, dense token-level supervision, efficiently transfers
备注: Work in progress
点击查看摘要
Abstract:On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
127. 【2610.10871】Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values
链接:https://arxiv.org/abs/2610.10871
作者:Fang Wan,Xufeng Liu,Fan Li,Yi Liu
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Language Models, achieve strong performance, Large Language, Sparse attention
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.
128. 【2610.10865】Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
链接:https://arxiv.org/abs/2610.10865
作者:Beimnet Bekele Guta,Xiaoyu Yang,Guangzhi Sun,Philip C. Woodland
类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:entangled representation space, Self-supervised speech encoders, Self-supervised speech, entangled representation, representation space
备注: In submission
点击查看摘要
Abstract:Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.
129. 【2610.10845】Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
链接:https://arxiv.org/abs/2610.10845
作者:Sietse Schelpe
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
关键词:large language model, internal key-value, large language, prompt, model
备注:
点击查看摘要
Abstract:A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream. Asked about facts planted millions of tokens earlier, the 12B model gave the right answer 82 times out of 100 and the 31B model 98 times out of 100. Neither model made up an answer. The limits are as follows. This is reuse of stored state, not a wider attention window: one block is loaded at a time, and how well a question is answered depends on the model. Writing the memory is a one-time cost, and the store takes terabytes of local NVMe disk. We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.
130. 【2610.10827】Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models
链接:https://arxiv.org/abs/2610.10827
作者:Marjan Celikik,Ana Peleteiro Ramallo,Javier Morales
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Corrective feedback, lessons rarely accumulate, second-language acquisition, best-evidenced drivers, drivers of second-language
备注:
点击查看摘要
Abstract:Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16$\times$. A feature-level online experiment shows significant gains in learner engagement ($+15.8\%$) and key business metrics, including scheduled hours ($+2.1\%$) and GMV from new lessons ($+13.2\%$).
131. 【2610.10787】NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
链接:https://arxiv.org/abs/2610.10787
作者:Gengze Zhou,Yicong Hong,Jiazhao Zhang,Xunyi Zhao,Jian Zhou,Zixing Lei,Zun Wang,Chongyang Zhao,Xionghui Chen,Stephen Gould,Anton van den Hengel,Qi Wu
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:long-horizon agentic reinforcement, agentic reinforcement learning, Language models trained, express precise actions, Language models
备注: 36 pages, 14 figures. Project page: [this https URL](https://metacognitionai.github.io/NavGPT3/)
点击查看摘要
Abstract:Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching. Beneath it, our action policy NavGPT VLA, trained on 19.28M examples, allocates visual tokens using codec allocation, in proportion to scene change; its 8B model alone reaches 74.51 SR on R2R-CE and leads RxR-CE with 78.19 SR. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs. 90.4 SR) and path fidelity (78.47 vs. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human. We comprehensively ablate the harness design and the interaction between the two models, showing how tools and the action policy shape the path from language-model reasoning to physical control: when NavGPT VLA executes the route, the reasoning loop shortens and the system's minimum reaction time falls from 3-19 s per language-model decision to 0.5-1 s per action-policy step (1-2 Hz). These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. We will release all models, code, and evaluation records.
132. 【2610.10786】Plan-and-Patch: Diffusion Language Models for Agentic Planning
链接:https://arxiv.org/abs/2610.10786
作者:Syamantak Kumar,Jiang Guo,Hassan Hamad,Hideo Kobayashi,Yi Xiang,Yezhou Yang,Yanjun Qi,Daniele Bonadiman,Jiarong Jiang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:requires coordinating subgoals, successful execution requires, execution requires coordinating, coordinating subgoals, increasingly important
备注:
点击查看摘要
Abstract:Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agents must therefore not only generate plans, but also revise them. Such revisions often affect only part of a plan, leaving the preceding and subsequent structure intact. Rather than regenerate the entire plan and risk unnecessary changes, repair can regenerate the affected region conditioned on the preserved prefix and suffix. We introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed. We compare DreamReasoner-8B and Qwen3-8B as diffusion and autoregressive (AR) planners. On Natural Plan without task-specific training, diffusion (53.7%) achieves nearly twice the plan repair success rate of AR (27.0%). After task-specific training on agentic benchmarks, ALFWorld and TextCraft, the planners achieve similar observed success in plan generation, while diffusion reduces mean plan-generation latency by 39-46% relative to AR. Our results show that Plan-and-Patch provides a framework for faster plan generation and effective plan repair in long-horizon agents.
133. 【2610.10758】Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
链接:https://arxiv.org/abs/2610.10758
作者:Mikhail L. Arbuzov(Independent researcher),Karan Dave(Independent researcher),Evgeniya Dontsova(Independent researcher),Yaodong Hu(Independent researcher),Vincent Lao(Independent researcher),Navita Jain(Independent researcher),Sisong Bei(Independent researcher),Dmitry Dimov(Independent researcher)
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Enterprise conversation analytics, Enterprise conversation, Enterprise, millions of interactions, Abstract
备注: 20 pages, 1 figure
点击查看摘要
Abstract:Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.
134. 【2610.10740】Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training
链接:https://arxiv.org/abs/2610.10740
作者:Nafiseh Ghoroghchian,Luis Scoccola,Tina Sedaghat,Omid Vaheb,Hannah Chen,Dino D'Agostino,Keyvan Golestan
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Conversational task disambiguation, resolve missing information, Conversational task, user intended task, tables or databases
备注: 39 pages (9 main, 30 appendix), 12 figures (5 main, 7 appendix)
点击查看摘要
Abstract:Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.
135. 【2610.10738】Lossy Compressive Text Autoencoders
链接:https://arxiv.org/abs/2610.10738
作者:Vinko Sabolčec,Angelos Katharopoulos,David Grangier
类目:Computation and Language (cs.CL)
关键词:work explores learning, explores learning, work explores, compressed latent representation, representation learning
备注:
点击查看摘要
Abstract:Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning. We propose an autoencoder architecture that performs residual downscaling and upscaling of hidden representations along the time axis, with a residual low-dimension discrete bottleneck. We analyze our approach for different quantization methods, training objectives, and datasets. For different levels of compression, we evaluate the similarity between the original and reconstructed text both at the surface-level (BLEU) and at the semantic-level (LLM-based judge). Additionally, we evaluate our models on downstream question-answering and semantic text similarity benchmarks. Our approach results in compressed representations which are on par with lossless text compression algorithms at 2.24 bits per byte on web text data, while having good reconstruction and downstream task performance.
136. 【2610.10724】Cognitive Thermometers: Machine Learning and Logical Complexity
链接:https://arxiv.org/abs/2610.10724
作者:Shane Steinert-Threlkeld,Jakub Szymanik
类目:Computation and Language (cs.CL)
关键词:human mind represent, represent semantic categories, mind represent semantic, human mind, mind represent
备注:
点击查看摘要
Abstract:How does the human mind represent semantic categories? Why do natural languages favor certain meanings over others? Prior explanations have relied on logical definability and complexity, but these are highly sensitive to the choice of logical language, rendering some design choices unmotivated. In this article, we propose that machine learning provides a somewhat more agnostic approach to measuring semantic complexity. We review emerging evidence that logic and machine learning often yield converging results on relative complexity and its resulting effects in semantic typology. Where they diverge, learning appears to be a better explanation than logical complexity. We argue that treating machine learning models as ``cognitive thermometers'' enables a unified approach to complexity that bridges symbolic logic and connectionist AI.
137. 【2610.10650】Large Language Model-Assisted Preparation of Transportation Management Plans: A Case Study with WisDOT WisTMP System
链接:https://arxiv.org/abs/2610.10650
作者:Zihao Sheng,Pei Li,Zilin Huang,Yen-Jung Chen,Yuhao Luo,Zhengyang Wan,Steven T. Parker,David A. Noyce,Sikai Chen
类目:Computation and Language (cs.CL)
关键词:Transportation Management Plans, designed Transportation Management, carefully designed Transportation, Management Plans, requiring carefully designed
备注:
点击查看摘要
Abstract:Work zones are critical yet hazardous components of transportation infrastructure, requiring carefully designed Transportation Management Plans (TMPs) to ensure safety and mobility. However, TMP preparation remains labor-intensive and heavily dependent on practitioner expertise. This paper proposes a Large Language Model (LLM)-assisted framework to automate TMP content generation, leveraging the WisDOT WisTMP system as the application context. The framework fine-tunes multiple open-source LLMs across different model scales and deploys them locally to ensure data security. To support model training, we construct a domain-specific dataset from historical WisTMP documents by converting PDF files into structured question-answer pairs in JSON format. Experimental results show that fine-tuning significantly improves performance across standard text generation metrics. Further section-wise and strategy-level analyses reveal that, while LLMs achieve strong overall performance, they tend to over-generate strategies and struggle to produce project-specific justifications and accurate cost estimates. In addition, scaling from 7B/8B to 14B yields limited gains. These findings demonstrate the potential of LLMs to improve TMP preparation efficiency while highlighting remaining challenges in LLM-assisted TMP development. The source code and demo videos will be publicly available at this https URL.
138. 【2610.10623】Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
链接:https://arxiv.org/abs/2610.10623
作者:Yi Wang,Rui Qian,Yu Li,Haoyang Yao,Wenjie Wang
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Looped Language Models, Looped Language, Language Models, parameter efficient approach, efficient approach
备注: 28 pages, 8 figures. Submitted to ICLR 2027
点击查看摘要
Abstract:Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approaches either provide reward based supervision that is sparse or costly to extend across loops, or rely on external teachers or privileged information, leading to limited teacher availability or teacher-student context mismatch. To address these limitations, we introduce LoopOPD, a cross-loop on-policy distillation framework that uses additional recurrent computation within a LoopLM as its own source of supervision. LoopOPD uses a frozen terminal loop policy as a compute privileged teacher for an intermediate loop student on student generated rollouts, providing dense supervision without an external teacher or privileged information. We further propose Dynamic LoopOPD (D-LoopOPD), which continually refreshes the terminal loop teacher as the shared model parameters are updated, enabling recurrent self-improvement. We characterize how distillation updates propagate across loop depths and derive sufficient conditions under which a single update yields simultaneous local improvement at both loop depths. Experiments on Ouro-Thinking models show that LoopOPD improves mathematical reasoning, while D-LoopOPD yields further gains through dynamic teacher updates. Despite being trained only on mathematical data, the resulting models also improve on general reasoning and code generation benchmarks, demonstrating that recurrent computation can serve as an effective source of supervision for LoopLMs. Our code and model checkpoints will be released upon acceptance.
139. 【2610.10622】WorldBench: Evaluating LLMs on Three.js Voxel World Generation
链接:https://arxiv.org/abs/2610.10622
作者:Krish Bakshi
类目:Graphics (cs.GR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Software Engineering (cs.SE)
关键词:Large language models, Large language, write complete, automatically is unreliable, language model reads
备注:
点击查看摘要
Abstract:Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents. We present WorldBench, a benchmark and judge for open-ended, LLM-generated this http URL worlds. From one prompt describing a floating voxel island with ten biomes, physics, and day/night and seasonal cycles, the judge explores the running world, controlling its clock, orbiting it, and sending a navigator agent to frame each biome, and reads the code for what it sees. Neither channel is trusted on its own: a code quote counts only if it is text the source contains, and visual claims are checked against measured pixels where the property is measurable. A mutation test, in which we remove features by construction, shows that code-only judging gives full credit to four of five removed features, because their code remains in the file. Our judge cuts the points kept on removed features by a third (5.44 to 3.55 of 7.11), and what it still credits is mostly code that exists but never runs. We evaluate five frontier models: Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7 and Gemini 3.1 Pro. Code, prompt, tests and judge configuration are available at this https URL
140. 【2610.10592】Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch
链接:https://arxiv.org/abs/2610.10592
作者:Szymon Kocur
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL)
关键词:Historical Polish, optical character recognition, paper covers, variable quality, machine-readable form
备注: 35 pages, 2 figures, 11 tables. Corpus: [this https URL](https://huggingface.co/datasets/SKocur/polish-pre1918-corpus) ; code: [this https URL](https://github.com/SKocur/wieszcz-xix) ; weights: [this https URL](https://doi.org/10.5281/zenodo.22099358)
点击查看摘要
Abstract:Historical Polish is well documented as a language but annotated in machine-readable form only to about a million words for the period this paper covers; the rest sits behind optical character recognition of variable quality. We present Wieszcz-XIX, a corpus of 6.75 billion tokens (about 3.1 billion words) in 294,369 documents, most of them periodical issues, of Polish published from 1800 to 1918, assembled from Wolne Lektury and the Internet Archive by a pipeline that filters, deduplicates, audits for post-1918 leakage and splits at the document level. It is over three orders of magnitude larger than the annotated corpus of the same period, and we quantify its defects: recognition corruption against a false-positive floor, near-identical duplication, which is removed, and post-1918 leakage, which is excluded from the training corpus itself down to a known residue of 0.04 to 0.38% of its bytes, found in the transcribed source, so the published corpus is the trained one document for document. On a hand-corrected sample the character error rate is 0.68% where the text is legible, and 45% of the sampled passages cannot be corrected. On it we train a ladder of decoder-only models from 47M to 349M parameters from scratch, and measure their temporal boundedness. Against two modern Polish base models, one far larger, the 349M shows a crossover, as does the 107M against the comparator of its size: post-1918 vocabulary costs them about 3.1 bits per byte more than period vocabulary, a gap the comparators do not show, and period vocabulary costs them fewer bits than it costs the comparators. Shown period text, the models keep its spelling and the comparators only partly. Adding parameters gains about twice as much as a second pass over the data. We release the corpus, code and weights. Content warning: the models reproduce period prejudice, including antisemitic statements.
141. 【2610.10550】Diffu-LoRA: A Novel Low-Rank Adaptation for Personalized Diffusion Models
链接:https://arxiv.org/abs/2610.10550
作者:Tianjing Li,Wei Zhu
类目:Computation and Language (cs.CL)
关键词:reference images requires, images requires preserving, requires preserving subject, preserving subject identity, describe new contexts
备注:
点击查看摘要
Abstract:Personalizing text-to-image diffusion models from a few reference images requires preserving subject identity while following prompts that describe new contexts. Full-model fine-tuning is parameter-intensive, whereas low-rank adaptation (LoRA) reduces the number of trainable parameters but leaves open how adaptation capacity should be distributed across layers. We introduce Diffu-LoRA, a parameter-efficient method that learns this allocation through gated low-rank adaptation. Diffu-LoRA inserts trainable low-rank components into the linear layers of Transformer blocks and assigns a learnable gate to each component. Bilevel optimization updates the adaptation weights and gate parameters on separate data splits, while progressive pruning removes components with the lowest gate values to meet a prescribed rank budget. This procedure allocates adaptation capacity nonuniformly across layers while keeping the pretrained backbone frozen. Experiments with Stable Diffusion on subjects from DreamBooth and additional collected datasets show improved overall subject fidelity and prompt alignment relative to the evaluated fine-tuning baselines. Ablation studies examine the contributions of bilevel optimization, progressive pruning, and adapter placement. These results support learned rank allocation as a practical approach to parameter-efficient diffusion model personalization.
142. 【2610.10541】An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment
链接:https://arxiv.org/abs/2610.10541
作者:Marcelo Valentim Silva,Hannes Herrmann,Valerie Maxville
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Semantic Table Interpretation, Table Interpretation, quality, Data Quality Assessment, downstream graph validation
备注: 18 pages, 4 figures, Workshop on Quality of Knowledge Graphs at ESWC 2026, May 11, 2026, Dubrovnik, Croatia
点击查看摘要
Abstract:Knowledge Graph (KG) quality depends not only on downstream graph validation, but also on the quality of tabular metadata used before integration. In metadata-only Semantic Table Interpretation (STI), where cell values are unavailable, noisy, or unsuitable, column headers become a critical source of semantic evidence for traceable KG preparation. We present an explainable, header-centric framework for metadata-only Column Type Annotation (CTA) and Data Quality Assessment (DQA). The framework maps headers to 39 interpretable FinalFormat types using curated lexical resources and preserves token-level traceability through SourceKeywords. Each assigned type activates validation rules based on a taxonomy of Data Quality Issues (DQIs), producing detections such as missing data, duplicates, domain violations, wrong data type, and temporal mismatch. These detections are aggregated into HeadersIQ, a lightweight, unweighted data source-level quality metric. The framework was evaluated across heterogeneous benchmarks, including UCI, Prague, Kaggle, VizNet/Sato, SOTAB, T2Dv2, and the SemTab 2024 Metadata-to-KG track, comprising around 120,000 header columns. The results show broad practical coverage across noisy real-world metadata, while a parallel KG-mapping pathway supports alignment to DBpedia and this http URL. On the SemTab 2024 Metadata-to-KG track, the official GT-strict evaluation was modest. However, a blinded diagnostic audit indicates that many mismatches reflect benchmark granularity, aliasing, and ontology-selection effects rather than wholly implausible header-centric predictions. We report this audit as diagnostic evidence on disagreement patterns, not as revised benchmark performance. Overall, the paper presents a reusable workflow for metadata-driven semantic annotation, data source-level quality monitoring, and KG-oriented benchmark diagnosis.
Comments:
18 pages, 4 figures, Workshop on Quality of Knowledge Graphs at ESWC 2026, May 11, 2026, Dubrovnik, Croatia
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:
arXiv:2610.10541 [cs.AI]
(or
arXiv:2610.10541v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2610.10541
Focus to learn more
arXiv-issued DOI via DataCite</p>
143. 【2610.12201】SteerablePlex: Can We Steer Full-Duplex Models?
链接:https://arxiv.org/abs/2610.12201
作者:Haolong Zheng,Maike Züfle,Dominik Macháček,Peter Polák,Xulin Fan,Xavier Sumba,Siyin Wang,Ondřej Klejch,Mark Hasegawa-Johnson
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
关键词:enabling natural interaction, conversation history grows, speak simultaneously, enabling natural, natural interaction
备注: 5 pages, 2 figures
点击查看摘要
Abstract:Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows. When used as user simulators, this lack of control can cause them to deviate from prescribed scenarios and produce unreliable evaluation outcomes. We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order. The benchmark reveals that current open-source full-duplex models struggle to follow such constraints. We then introduce a Group Reward-Decoupled Normalization Policy Optimization (GDPO)-based training recipe that enables a full-duplex model to follow textual instructions during an ongoing conversation while maintaining its turn-taking ability. By connecting the resulting SteerablePlex to an asynchronous backend language model that monitors the conversation and provides instructions when needed, we build a more controllable full-duplex user simulator that follows multi-stage constraints more reliably than existing open-source models and GPT-Realtime.
144. 【2610.10868】Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
链接:https://arxiv.org/abs/2610.10868
作者:Xilin Jiang,Shun Zhang,Tejas Jayashankar,Yinghao Aaron Li,Osama Hanna
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD)
关键词:large language model, introduce Conversational Voice, natural conversational contexts, Conversational Voice Aesthetic, Voice Aesthetic Model
备注:
点击查看摘要
Abstract:We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.
信息检索
1. 【2610.12300】Compact and Efficient Indexes for Learned Sparse Retrieval
链接:https://arxiv.org/abs/2610.12300
作者:Franco Maria Nardini,Luca Rizzo,Cosimo Rulli,Rossano Venturini
类目:Information Retrieval (cs.IR); Databases (cs.DB)
关键词:retrieval data structures, learned sparse retrieval, data structures, paper investigates, substantially reduce
备注: 15 pages, 2 figures. Accepted at IEEE International Conference on Data Engineering 2027 (IEEE ICDE 2027)
点击查看摘要
Abstract:This paper investigates how to substantially reduce the memory footprint of learned sparse retrieval indexes without sacrificing the efficiency of state-of-the-art retrieval data structures. Building on SEISMIC, we revisit both levels of its design: the inverted index used to select candidates and the forward index used to score them. For the inverted index, we replace costly per-block summaries with medoids, namely existing documents elected as block representatives, collapsing the per-block metadata from a sparse vector to a single document identifier. For the forward index, we compress both components and values. We reorder the vocabulary to place co-occurring components closer together and encode the resulting $\Delta$-gaps with DOTPACKING8, a SIMD-friendly bit-packing scheme that fuses decompression with dot-product evaluation; values are quantized with compact per-component 4-bit codebooks fitted to each component's distribution. We further introduce JUMPDOT, a blocked dot-product kernel tailored for queries that contain only a few non-zero entries. Our forward-index compression is independent of SEISMIC and can be plugged into any system relying on forward-index-based scoring, as we demonstrate by integrating it into KANNOLO. A comprehensive evaluation on MS MARCO with three state-of-the-art learned sparse encoders shows that our solutions markedly improve the speed-space trade-off of learned sparse retrieval: at equal accuracy, our indexes answer queries up to 5.3x faster than the best competitor while using about 3x less memory, and in the most memory-constrained regime, they remain up to 1.9x faster while using up to 3.9x less memory.
2. 【2610.12256】Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
链接:https://arxiv.org/abs/2610.12256
作者:Youngtaek Oh,Qiyu Wu,Hiromi Wakaki,Junmo Kim,Yuki Mitsufuji
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:embeddings naturally involve, Omnimodal embeddings naturally, heterogeneous inputs, naturally involve, features across heterogeneous
备注: Accepted to EMNLP 2026 (Long, Findings). Code: [this https URL](https://github.com/sony/syn-omni)
点击查看摘要
Abstract:Omnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space over mixed-modality data, limiting structural separation between universal and modality-specific representations. To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration. Specifically, we introduce Orthogonal Modality-Expert LoRA (OME-LoRA), which decomposes adaptation into a shared LoRA path for universal semantics and modality-expert LoRA paths for modality-aware specialization. Furthermore, Progressive Synergy Routing (PSR) enables experts to first establish modality-specific priors, then gradually interact with other modality-experts for cross-modal synergy. Evaluated across 81 diverse tasks spanning image, video, audio, and audiovisual modalities, Syn-Omni consistently outperforms omnimodal baselines, demonstrating the effectiveness of structured specialization and cross-modal progressive collaboration.
3. 【2610.12243】NativeScope: Relation-Localized Retrieval over Native Topology with a Correct Anchor
链接:https://arxiv.org/abs/2610.12243
作者:Long Wang
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:semantic similarity, ranks text chunks, Dense retrieval, percent, text chunks
备注: 14 pages, 4 figures, and 6 tables. Includes an appendix with reproduction information and an evidence inventory
点击查看摘要
Abstract:Dense retrieval usually ranks text chunks by their semantic similarity to a question. This ignores structure that many data systems already store, including section membership, session boundaries, and native order. We propose NativeScope, a scope-then-rank method for queries with a known anchor and relation. It represents a query as q - (A, r, B). The anchor A and relation r select native units through belonging, before, or after operators, and the target term B ranks only chunks that overlap the selected scope. An internal variant, NS-FullQ, ranks the same candidates with the full question. We evaluate both methods on 200 controlled document and memory records derived from QASPER and LongMemEval under a 1,024-token budget. NativeScope attains native-unit recall of 89.28 percent for documents and 72.50 percent for memories, improving over instance-wide Dense RAG by 42.75 and 22.00 percentage points. NS-FullQ reaches 87.78 percent and 68.50 percent; its differences from NativeScope are inconclusive, locating the primary gain in relational scoping rather than the shorter ranking query. With automatic Top-1 anchors, memory recall falls to 35.50 percent. NativeScope is therefore effective when anchor coordinates and native relations are reliable, but hard scoping inherits errors from the localization interface.
4. 【2610.11922】Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
链接:https://arxiv.org/abs/2610.11922
作者:Jimmy Lin,Sahel Sharifymoghaddam,Lingwei Gu,Nour Jedidi
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:Project Greenhouse represents, modest computational resources, Project Greenhouse, Greenhouse represents, build fully open
备注:
点击查看摘要
Abstract:Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.
5. 【2610.11816】Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
链接:https://arxiv.org/abs/2610.11816
作者:Yubo Sun,Chunyi Peng,Yukun Yan,Zhenghao Liu,Zhipeng Xu,Sen Mei,Linlin Xin,Zheni Zeng,Maosong Sun
类目:Information Retrieval (cs.IR)
关键词:made significant progress, capabilities extend reliably, Dense retrievers, documents remains unclear, remains unclear
备注: Code: [this https URL](https://github.com/OpenBMB/Trident)
点击查看摘要
Abstract:Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear. In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition. As image documents are progressively replaced with semantically corresponding text representations, retrieval performance follows a pronounced V-shaped curve, remaining strong on single-modality corpora but degrading substantially when modalities coexist. In particular, irrelevant text causes more severe degradation than an equal number of irrelevant images, a phenomenon we term Chaos in the Text. Further analysis reveals modality preference, whereby text representations receive systematically higher similarity scores, allowing irrelevant text to outrank relevant images. To mitigate this bias, we introduce Trident, which constructs text, image, and fused text-image views of each document as co-equal positives and jointly optimizes relevance discrimination and positive-view balance through Multi-Positive View InfoNCE. Experiments across visual document and natural image benchmarks show that trident improves mixed-modality retrieval on both CLIP-based and VLM-based architectures, reduces sensitivity to modality composition and text distractors, and increases average single-modality retrieval performance.
6. 【2610.11666】Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
链接:https://arxiv.org/abs/2610.11666
作者:Jianfei Zhao,Yifan Wang,Feng Zhang,Xin Sun,Chong Feng,Zhixing Tan,Yang Luo,Boyuan Pan,Xu Kai,Yao Hu
类目:Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
关键词:Universal multimodal retrieval, Universal multimodal, independently indexed items, ranks independently indexed, retrieval typically encodes
备注: Under Review
点击查看摘要
Abstract:Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.
7. 【2610.11650】SkillContrast: Difference-Guided Text Selection for Agent Skill Reranking
链接:https://arxiv.org/abs/2610.11650
作者:Jiandong Ding,Honglei Ji,Ming Liu,Tao Duan
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:Similar agent skills, Similar agent, Similar, share instructions, Abstract
备注: 5 pages, 2 figures, 3 tables
点击查看摘要
Abstract:Similar agent skills can share instructions but differ in their conditions of use. Query-based text selection may retain shared instructions and omit these distinctions. We introduce SkillContrast, a training-free selector that compares retrieved skills and retains their differing text with local context for a pretrained reranker. On 1,235 requests from SameCapRisk-Bench, it yields 54-72 more clean hits (requests that retrieve a helpful skill without its marked risky sibling) than TF-IDF query selection at identical per-candidate input lengths, across 2 retrievers and 2 reranker sizes. Length-matched component replacements identify differing text as the main contributor in the primary setting, with smaller, mixed context effects. Relative to full skill bodies, SkillContrast uses 51.1-58.8% fewer model-input tokens, with 10-18 fewer clean hits at 0.6B and matching or higher observed clean-hit counts at 4B. Candidate-relative differences thus complement query relevance in selecting compact reranking inputs.
8. 【2610.11598】Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task
链接:https://arxiv.org/abs/2610.11598
作者:Junjie Chen,Yuxi Dong,Haitao Li,Yiqun Liu,Qingyao Ai
类目:Information Retrieval (cs.IR)
关键词:Large Language Models, Automatic Evaluation, automatic evaluation methods, investigate automatic evaluation, Deep Research Evaluation
备注: NTCIR-19
点击查看摘要
Abstract:In this paper, we provide an overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) task. Building on the success of the NTCIR-18 core task AEOLLM, we proposed AEOLLM-2 for NTCIR-19 to further investigate automatic evaluation methods for Large Language Models (LLMs), particularly in long-form text generation scenarios. In AEOLLM-2, we introduced a new subtask, Deep Research Evaluation, which focuses on the automatic evaluation of long-form deep research reports generated by LLMs. Participants developed evaluation methods to automatically assess the quality of these reports, and the performance of each method was measured by comparing its scores against human-annotated ground-truth labels. This year, we received 91 runs from 10 teams in total. This paper describes the background of the task, the dataset construction, the evaluation measures, the participants' methods, and the final evaluation results.
9. 【2610.11553】EVIE: Evidence-Vector-Informed Embeddings for Visual Document Retrieval
链接:https://arxiv.org/abs/2610.11553
作者:Zifei Wang,Wei Wen,Qiang Ji,Qian-Wen Zhang,Ruizhi Qiao,Xing Sun
类目:Information Retrieval (cs.IR)
关键词:existing approaches struggle, Accurate and scalable, fine-grained page understanding, efficient indexing, scalable visual document
备注: 22 pages, 8 figures
点击查看摘要
Abstract:Accurate and scalable visual document retrieval (VDR) requires both fine-grained page understanding and efficient indexing, yet existing approaches struggle to achieve both. OCR-based text retrieval adds preprocessing latency and can lose visual and structural cues needed to understand complex pages. Single-vector vision-language models bypass OCR, but compressing an entire page into one vector limits the granularity of query--document matching. Multi-vector retrievers with MaxSim provide finer interactions, yet demand large indexes and still leave room for accuracy improvements. We argue that overcoming these limitations requires preserving query-relevant page evidence throughout representation learning and index construction. To this end, we introduce \textbf{\textit{EVIE}} (Evidence-Vector-Informed Embeddings), a family of native visual document retrievers integrating three key innovations: (1) Evidence-judged data governance, which uses a multimodal judge to identify answer-bearing positives and filter unreliable negatives. (2) Bidirectional teacher--student learning with symmetric listwise distillation and prefix-based Matryoshka representation learning (Prefix-MRL), enabling one student checkpoint to serve six nested embedding dimensions without re-encoding. (3) Hierarchical agglomerative index compression (HAC), which clusters page tokens with spatial regularization and stores semantic centroids for single-stage MaxSim retrieval. Extensive experiments across 138 tasks from ViDoRe V1, V2, V3, and JinaVDR validate EVIE. EVIE-8B achieves 66.75 nDCG@10 on V3, exceeding the best external baseline by 1.43 points, with a four-suite average of 79.51. EVIE-4.5B with HAC retains 59.58 nDCG@10 at only 3.81 GiB per million pages, reducing vector payload by $128\times$. Together, these results improve the accuracy--storage trade-off for visual document retrieval.
10. 【2610.11506】Compactness and Consistency: A Conjoint Framework for Deep Graph Clustering
链接:https://arxiv.org/abs/2610.11506
作者:Wei Ju,Siyu Yi,Kangjie Zheng,Yifan Wang,Ziyue Qiao,Li Shen,Yongdao Zhou,Xiaochun Cao,Jiancheng Lv
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Social and Information Networks (cs.SI)
关键词:aiming at grouping, fundamental task, similar characteristics, Graph, data analysis
备注: Accepted by Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026 Oral)
点击查看摘要
Abstract:Graph clustering is a fundamental task in data analysis, aiming at grouping nodes with similar characteristics in the graph into clusters. This problem has been widely explored using graph neural networks (GNNs) due to their ability to leverage node attributes and graph topology for effective cluster assignments. However, representations learned through GNNs typically struggle to capture global relationships between nodes via local message-passing mechanisms. Moreover, the redundancy and noise inherently present in graph data may easily result in node representations lacking compactness and robustness. To address these issues, we propose a conjoint framework CoCo, which captures compactness and consistency in the learned node representations for deep graph clustering. Technically, our CoCo leverages graph convolutional filters to learn robust node representations from both local and global views, and then encodes them into low-rank compact embeddings, thus effectively removing the redundancy and noise as well as uncovering the intrinsic underlying structure. To further enrich the node semantics, we develop a consistency learning strategy based on compact embeddings to facilitate knowledge transfer from the two perspectives. Our experimental results indicate that our CoCo outperforms state-of-the-art counterparts on various datasets.
11. 【2610.11489】Beyond Resolution: Object-to-Image Ratio Mismatch in Instance Retrieval
链接:https://arxiv.org/abs/2610.11489
作者:Boaz Meivar,Ofir Kedem,Amit Edenzon,Gal Chechik,Shai Avidan
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:Visual instance retrieval, Visual instance, instance retrieval, retrieval often fails, apparent sizes
备注: 24 pages. Preprint, under review
点击查看摘要
Abstract:Visual instance retrieval often fails when the same object appears at different apparent sizes in the query and gallery. We show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images. On a controlled benchmark of 3,021 Objaverse objects rendered at five camera distances, more than 80% of the cross-distance degradation is attributable to O2I mismatch rather than resolution for 9 of 12 pretrained backbones; multi-scale architectures cut the resolution-only effect to single digits yet remain equally susceptible. The failure is also asymmetric: tight queries retrieve more reliably against wide gallery images than the reverse. Guided by this analysis, query-side scale augmentation and an OWLv2 crop reranker reach state of the art on ILIAS 100M (29.2 mAP@1000 before reranking, 42.0 after) without training or modifying the precomputed gallery index, and a LoRA fine-tune matches the query-side gains at a single forward pass, showing that O2I robustness is learnable.
12. 【2610.11451】SAIL: Scientific Agentic Intelligence via a Science-Aware Loop
链接:https://arxiv.org/abs/2610.11451
作者:SAIL Model Team:Boyuan Sun,Bryan Dai,Che Liu,Chi Liu,Derek Li,Hongming Piao,Mengzhuo Chen,Xidong Wang,Yan Shu,Yinda Chen,Ziyang Zeng
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
关键词:multi-step research workflows, introduce SAIL, open model, SAIL, research workflows
备注: 16 pages, technical report
点击查看摘要
Abstract:We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows. SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze its task failures and construct training tasks that address the underlying capability gaps. The diagnosis examines search and evidence selection in literature tasks, scientific assumptions and reasoning in coding, and planning and revision in longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools. We repeat this loop over multiple development cycles and train SAIL through supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning. SAIL achieves competitive performance across scientific research tasks with substantially fewer parameters than leading open-weight models.
13. 【2610.11439】On-Chain Archaeology of Bitcoin Oracles: Evidence of Use under Limited Observability
链接:https://arxiv.org/abs/2610.11439
作者:Giulio Caldarelli
类目:Cryptography and Security (cs.CR); Computers and Society (cs.CY); Information Retrieval (cs.IR)
关键词:Ethereum made, key-release services, federated signers, household term, serving as feeds
备注:
点击查看摘要
Abstract:Before Ethereum made the "oracle problem" a household term, Bitcoin already had oracles serving as feeds, key-release services, federated signers, and arbiters that carried real value on the main chain. This study traces their use and the changing evidence of oracle activity from early days through July 2026. We combine a complete census of Counterparty betting (1,149 bets), analysis of the full Bitcoin chain through block 958,628, and searches for documented keys from Reality Keys, Orisi, Bitrated, and Oraclize in an 854-million-row public-key index. We also recover DLC oracle records from an archived explorer and live Nostr relays. Two results emerge. First, early contracts remain on-chain, but many event descriptions have disappeared, and protocol encoding and API limitations complicate access to the surviving record. However, for modern DLCs, public oracle announcements can survive even when the contracts using them cannot be identified on-chain. In the script classes examined, the share of spends that reveal no script peaks at 81.9% in 2024 after excluding spends containing inscription data. Second, public registries can give a misleading picture of oracle use. In Counterparty, 95% of pre-2018 sources declaring an oracle fee were never bet on. In Bitrated, 0.1% of archived keys appear on-chain overall, compared with 10 of 19 keys captured in 2014. Sport dominates Counterparty's matched volume, while a daily price series dominates the archived DLC announcements. These findings show how protocol design and data preservation shape the historical record of Bitcoin oracle use.
14. 【2610.11370】RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees
链接:https://arxiv.org/abs/2610.11370
作者:Meghanadh Pulivarthi,Swaraj Kumar Biswal,Kushagra Bhushan,Yatin Nandwani,Sachindra Joshi,Dinesh Raghu
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:Retrieval-augmented generation, grounds language models, grounds language, RAG enables iterative, RAG
备注: 24 pages (main text through Limitations ends on page 9, followed by references and appendix), 9 figures, 16 tables
点击查看摘要
Abstract:Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
15. 【2610.11277】H2CE: Modeling Geo-Semantic Interactions for POI Reranking with Heterogeneous Two-Stage Cross-Encoders
链接:https://arxiv.org/abs/2610.11277
作者:Zhengwei Bai,Moreno D'Incà,Danielle Class,Alessandro Moschitti
类目:Information Retrieval (cs.IR)
关键词:real-time serving constraints, geospatial proximity, review count, serving constraints, rating and review
备注:
点击查看摘要
Abstract:Point-of-Interest (POI) reranking in local search must model query-conditioned tradeoffs among lexical semantics, geospatial proximity, and numerical quality signals such as rating and review count, while remaining practical under real-time serving constraints. A close POI may only partially satisfy the query intent, while a farther one may offer stronger semantic and quality evidence. We present H2CE, a Heterogeneous Two-stage Cross-Encoder for latency-bounded POI reranking. H2CE represents numerical attributes in two complementary ways: bucketized natural-language descriptors are inserted into the cross-encoder input to support semantic--numeric attention, while exact scalar values are processed by dedicated MLPs to preserve magnitude information. The resulting semantic and numerical embeddings are fused through latent-space aggregation, enabling nonlinear interactions beyond scalar weighted sums. H2CE then applies a two-stage architecture: Stage 1 scores all candidates pointwise for scalable filtering, and Stage 2 performs head-to-head pairwise comparison among the top-$K$ candidates with Copeland aggregation, making fine-grained relative tradeoffs explicit while reducing pairwise cost from O(N^2) to O(N+K(K-1)). On a 5,743-query local search test set, H2CE achieves 67.48% NDCG@5, improving over XGBoost LTR by +22.82% absolute and over a zero-shot LLM reranker by +35.89%. The pairwise stage adds +1.98% NDCG@5 over the pointwise model alone. Ablations confirm the value of numerical features, latent aggregation, top-K pairwise reranking, and aligned training.
16. 【2610.11270】Gated Memory: Admission-Controlled Memory Formation for Conversational AI
链接:https://arxiv.org/abs/2610.11270
作者:Preeti Saraswat,Divya Neelagiri,Ajay Manoj
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:persistent vector stores, Personalized conversational, long-term memory systems, vector stores, conversational AI relies
备注:
点击查看摘要
Abstract:Personalized conversational AI relies on long-term memory systems that extract facts from user utterances and store them in persistent vector stores. Despite progress in retrieval, deduplication, and lifecycle management, the formation stage, the moment a fact is first written to storage has received almost no principled attention. We identify this as the binding constraint on memory quality in production systems. Critical contextual signals, such as the distinction between a permanent user attribute and a transient situation, exist only in the original utterance and are irreversibly lost the moment extraction produces a subject-relation-object triple. No downstream process can recover them. We propose Gated Memory, a lightweight, modular formation framework that interposes two decision checkpoints between conversation and storage: an admission gate that evaluates every candidate fact against the full utterance context before extraction runs, and a conditional enrichment stage that grounds admitted facts through an entity scope taxonomy with privacy constraints. The gate evaluates only the current exchange while using prior turns as read-only reference context, and produces a structured formation record. Admitted content is decomposed into atomic facts, each categorized, tagged with provenance (directly stated versus inferred), scoped to its condition of applicability, and grounded in resolved time and place, subject to a constraint that no entity absent from the context may be asserted. On the LoCoMo-10 benchmark with atypical emotional density in utterance data, Gated Memory achieves an overall +2.6% relative improvement in LLM-judge accuracy over a strong baseline with identical retrieval and generation, establishing formation quality as a measurable constraint on memory performance.
17. 【2610.10955】Learning Multi-Step Query Rewriting via Corpus Feedback for Conversational Search
链接:https://arxiv.org/abs/2610.10955
作者:João Coelho,Hong Wang,Jie Yuan,Zhuoer Wang,Samson Koelle,Wei Niu
类目:Information Retrieval (cs.IR)
关键词:context dependent user, dependent user turn, Conversational Query Rewriting, context dependent, dependent user
备注:
点击查看摘要
Abstract:Conversational Query Rewriting (CQR) turns a context dependent user turn into a standalone query for a retriever, and most methods do this in a single step from the dialogue history before retrieving once. The rewrite is therefore fixed before any corpus evidence is available to correct its reference resolution or its vocabulary. We recast CQR as a sequential retrieval problem: an agent rewrites the current turn, retrieves, and conditions its next rewrite on the returned passages. The agent acts in a typed space of three rewriting operations, resolving conversational intent into a standalone query, generating lexical reformulations, or synthesizing pseudo-documents for document-to-document matching, together with a stop action that ends the episode. We train the policy with supervised fine-tuning followed by reinforcement learning against a single retrieval-quality reward, using no human rewrite annotations. Across TopiOCQA and QReCC, the agent outperforms several retrieval-aligned baselines, while remaining effective across retrieval backends and generalizing to the CAsT benchmarks without additional training. Further analysis shows that, through retrieval-reward optimization alone, the learned policy develops a behavior of grounding pseudo-documents in passages retrieved by earlier steps, substantially improving retrieval.
18. 【2610.10920】Language Models for Page-Level Layout Decisions in E-commerce Search
链接:https://arxiv.org/abs/2610.10920
作者:Varun Joshi,Eva C. Song,ChengXiang Zhai
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:E-commerce search pages, E-commerce search, modern E-commerce search, critical touchpoints, touchpoints for millions
备注: Accepted at the OARS Workshop, ACM RecSys 2026
点击查看摘要
Abstract:E-commerce search pages are critical touchpoints for millions of online shoppers. While traditional search engines return a ranked list of results, modern E-commerce search pages increasingly incorporate recommender system modules -- for example, secondary stacks that surface alternative product groupings at specific positions. When introduced appropriately, secondary stacks can improve user engagement; however, suboptimal placement may disrupt browsing flow and degrade the primary results. Unlike traditional search ranking, where evaluation techniques such as interleaving are well established, evaluating page-level layout changes e.g., when and where to insert a secondary stack remains challenging without costly online A/B testing. To address this, we study offline methods for evaluating whether a given layout decision -- specifically, the inclusion of a secondary stack at a particular position -- is beneficial to users. We investigate language models as scalable evaluators by comparing direct prompt-based, prompt-derived feature, and representation-based methods. Our results show that representation-based approaches consistently outperform prompt-based judging in predicting user engagement, suggesting they provide a reliable foundation for offline layout evaluation in E-commerce search.
19. 【2610.10556】LIFT: A Lifecycle-aware Interaction Factorization Transformer for Unified Retrieval and Ranking
链接:https://arxiv.org/abs/2610.10556
作者:Keji Miao,Enhao Cheng,Yan Li,Qun Li,Qi Zhang,Jie Yuan,Xiaoyong Li
类目:Information Retrieval (cs.IR)
关键词:Cascaded recommender systems, Interaction Factorization Transformer, Lifecycle-aware Interaction Factorization, Cascaded recommender, recommender systems
备注: 16 pages, 5 figures, 7 tables
点击查看摘要
Abstract:Cascaded recommender systems use the same user history for retrieval and ranking, while the two stages have access to different information at different points of an interaction. We propose the Lifecycle-aware Interaction Factorization Transformer (LIFT), which decomposes each interaction into ordered Request, Item, Context, and Action states and models them as a causal sequence. Retrieval reads the Request state, while ranking reads the Context state, allowing both tasks to share history modeling while preserving stage-specific information. LIFT instantiates this representation with Role-Conditioned Attention and a lightweight Pre-LN Bias. On ML-20M and Taobao, LIFT achieves the highest Joint Score among the evaluated joint models, improving over the strongest baselines by 4.9% and 3.6%, respectively. Loss-weight sweeps show favorable retrieval--ranking trade-offs, while ablations and scaling analyses further examine lifecycle sequence construction, model components, and capacity settings.
计算机视觉
1. 【2610.12470】Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration
链接:https://arxiv.org/abs/2610.12470
作者:Jusuk Lee,Sungha Kim,Yeonsoo Park,Jonguk Cheon,Yoonkyo Jung,Yongjun You,H. Jin Kim,Jia-Bin Huang,Furong Huang,Youngseok Jang,Seungjae Lee
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:costly robot demonstrations, recent methods predominantly, methods predominantly imitate, predominantly imitate demonstrated, imitate demonstrated motions
备注: Project page: [this https URL](https://dex-one2many.github.io/)
点击查看摘要
Abstract:While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video. Alternatively, discovering a policy via reinforcement learning (RL) allows for broad generalization, but without prior guidance, it struggles with high-dimensional exploration in complex, multi-stage tasks. To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability. The graphs serve as generative constraints for sampling diverse reset states and provide dense rewards for each stage. Because the graphs constrain relations rather than exact poses, these reset states cover object poses and grasps beyond the video, while initializing each stage from them with dense rewards keeps exploration short and guided. Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand. Across five tool-use and manipulation tasks, Dex-One2Many exceeds baselines by 6.5% in seen configurations, while its robust generalization widens this gap to 71% in unseen scenarios.
2. 【2610.12469】Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
链接:https://arxiv.org/abs/2610.12469
作者:Ritesh Thawkar,Shubham Patle,Shravan Venkatraman,Rao Muhammad Anwer
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Instruction-guided image editors, Instruction-guided image, highly capable, external reward models, human-edited training pairs
备注: Project Page: $\href{ [this https URL](https://riteshthawkar.github.io/Rubric-CEPR/) }{\text{this URL}}$
点击查看摘要
Abstract:Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models. Such supervision is costly to obtain and can reward plausible failures: a realistic output may leave the requested change undone or alter content that should be preserved. In this work, we strive to improve a pretrained image editor using only its own generations, without human-edited targets or an external training-time reward model. To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR). A Planner proposes structured edit instructions from unlabeled images, the Editor samples multiple candidate edits, and a frozen Critic scores each candidate with decomposed rubric checks for edit realization, removal of the old state, and content preservation, using features already exposed by the editor. Non-compensatory gates reject infeasible candidates, and the best verified candidate is distilled into the editor through lightweight adapter training. On Qwen-Image-Edit, Rubric-CEPR improves ImgEdit from 4.36 to 4.60 (+5.5%), with a +24.9% gain on object isolation, and transfers to GEdit-Bench and Complex-Edit. The same procedure also improves Step1X-Edit by +7.8% on ImgEdit. We hope our approach will serve as a solid baseline for image editors that improve themselves from their own verified samples. Our code is publicly available at $\href{this https URL}{\text{this URL}}$
3. 【2610.12468】DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
链接:https://arxiv.org/abs/2610.12468
作者:Junyan Li,Ruizhi Li,Yu Liu,Xiangshuo Liu,Mingchao Sun,Hongyu Pan,Mu Xu,Lue Fan,Zhaoxiang Zhang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:cross-embodiment robot world, cross-embodiment robot, action, model, robot
备注: project page: [this https URL](https://brave-eai.github.io/DreamTrue;) code: [this https URL](https://github.com/brave-eai/DreamTrue)
点击查看摘要
Abstract:We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at this https URL.
4. 【2610.12464】What 30,000 Hours of Ego-centric Video Does Not Teach
链接:https://arxiv.org/abs/2610.12464
作者:Jiahua Dong,Anurag Bagchi,Yash Jangir,Muhammad Zubair Irshad,Sergey Zakharov,Martial Hebert,Homanga Bharadhwaj,Yu-Xiong Wang,Vitor Campagnolo Guizilini,Pavel Tokmakov
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:physics-based simulators, practical deployment, offer a promising, promising alternative, alternative to physics-based
备注:
点击查看摘要
Abstract:World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains. Finally, our conclusions transfer to downstream humanoid modeling. Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.
5. 【2610.12461】OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
链接:https://arxiv.org/abs/2610.12461
作者:You-Zhe Xie,Ting-Wei Chou,Yu-Hsuan Li,Kaipeng Zhang,Zhixiang Wang,Yu-Lun Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:models generate photorealistic, world models generate, Gaussian Splatting scene, generate photorealistic, frozen in time
备注: Project page: [this https URL](https://ouroworld.userwei.com)
点击查看摘要
Abstract:Recent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint. A vision-language model infers plausible dynamics and guides a video model to synthesize a reference video, which we lift and complete into multi-view videos. To learn from this imperfect supervision, we propose Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field guarantees looping by construction, while a Grounded Drift Field anchored at the reference view absorbs cross-view inconsistency. Unlike prior Eulerian methods limited to fluid-like motion, we capture general deformation, object motion, and illumination change. We introduce a ground-truth-free evaluation covering vividness, naturalness, loop seam coherence, and scene quality. On 39 reconstructed and generated scenes, OuroWorld outperforms all baselines and wins 70.8%-99.0% of user-study comparisons. Project page: this https URL
6. 【2610.12459】WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
链接:https://arxiv.org/abs/2610.12459
作者:Ankan Deria,Komal Kumar,Hisham Cholakkal,Fahad Shahbaz Khan,Salman Khan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:plausible visual trajectories, synthesize plausible visual, video-based world models, generators and video-based, synthesize plausible
备注: 34 pages, 14 figures, 15 Tables
点击查看摘要
Abstract:Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce \textbf{WorldGuide Bench}: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33\% Task Success on \textbf{WorldGuide-Bench}, compared with 29.90\% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69\% on \textbf{VideoCraft-Bench} compared with 32.73\% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.
7. 【2610.12458】OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
链接:https://arxiv.org/abs/2610.12458
作者:Zhongyu Yang,Jiale Tao,Ruitao Chen,Zuhao Yang,Yingfang Yuan,Xueliang Zhao,Auden,Kai Wang,Shuai Shao,Biao Wang,Steve Yves,Qinglin Lu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multimodal large language, large language models, Multimodal large, language models, creating an urgent
备注: Accepted by NeurIPS 2026. Code and benchmark can be found at [this https URL](https://01yzzyu.github.io/OmniCapBench/)
点击查看摘要
Abstract:Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio--visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.
8. 【2610.12455】Hybrid Cinematography: Previsualizing and Managing Hallucination Risk in Generative Video Reshooting
链接:https://arxiv.org/abs/2610.12455
作者:Nhan(Nathan)Tran,Neal Wadhwa,Abe Davis,Stefan Stojanov
类目:Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
关键词:film set, present Hybrid Cinematography, Hybrid Cinematography, set, hallucination
备注: Project page: [this https URL](https://hybridcinematography.github.io/)
点击查看摘要
Abstract:On a film set, the camera move is committed during a take. Generative video reshooting lets filmmakers change it afterward, but may require hallucinating unrecorded content, a gap sometimes discovered only after leaving the set. We present Hybrid Cinematography, a workflow that bridges physical capture and generative reshooting to manage hallucination risk while filmmakers can still act on it. Using an editable 3D shot plan and a proxy of the take, our previsualization evaluates hallucination risk in real time. Seeing where the take lacks support, filmmakers can iteratively adjust the plan, explore moves that balance capture and generation, shoot guided pickups, or knowingly accept hallucination. We demonstrate the workflow through a mobile augmented reality application for on-set planning, capture, and review, and an offline pipeline for existing video. A study with experienced filmmakers reveals how previsualizing risk informs camera decisions and exposes tensions between creative intent and generative hallucination.
9. 【2610.12452】BrickBench: Evaluating Agentic Brick Design
链接:https://arxiv.org/abs/2610.12452
作者:Peter Kulits,Yiqing Xu,R. Kenny Jones,Cordelia Schmid,Jiajun Wu
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:agentic text-conditioned LEGO-set, text-conditioned LEGO-set design, propose BrickBench, agentic text-conditioned, text-conditioned LEGO-set
备注: Project page: [this http URL](http://www.brickben.ch)
点击查看摘要
Abstract:We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at this http URL
10. 【2610.12451】VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
链接:https://arxiv.org/abs/2610.12451
作者:Boyao Han,Chen Shi,Jingjing Qian,ZhuoTan Tian,Li Jiang
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:fixed camera configurations, models have emerged, robotic manipulation, emerged as powerful, powerful foundations
备注: Accepted at NeurIPS 2026. Project page: [this https URL](https://boyaohan.github.io/VersaCamVLA.github.io/)
点击查看摘要
Abstract:Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.
11. 【2610.12448】One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
链接:https://arxiv.org/abs/2610.12448
作者:Adrian Bulat,Yassine Ouali,Georgios Tzimiropoulos
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:full-depth vision encoder, comparable inference FLOPs, single Transformer block, single Transformer, applied recurrently
备注:
点击查看摘要
Abstract:In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
12. 【2610.12442】LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
链接:https://arxiv.org/abs/2610.12442
作者:Suhwan Cho,Yonwoo Choi,Soongjin Kim,Jicheol Park,Taegyu Lim
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:single exocentric recording, video diffusion model, exocentric recording, challenging case, share little overlap
备注:
点击查看摘要
Abstract:Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.
13. 【2610.12427】FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
链接:https://arxiv.org/abs/2610.12427
作者:Yuxuan Hu,Weikang Shi,Yang Bo,Xudong Lu,Xintong Guo,Shuhan Li,Yuyang He,Huankang Guan,Peiwen Sun,Yunqiao Yang,Wenbo Li,Rui Liu,Hongsheng Li
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Large Language Models, Video Large Language, Large Language, existing benchmarks focus, Streaming Video Large
备注:
点击查看摘要
Abstract:Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios. Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events. We introduce FastBench to evaluate high-dynamic perception in real-world video streams. Its trajectory-grounded pipeline combines QA generation from high-FPS clips, filtering of questions answerable at 2 FPS, answer verification using SAM3 and CoTracker3 trajectories, and three rounds of human inspection. FastBench contains 306 QA pairs across eight domains, six capabilities, and forward, instant, and backward temporal scopes, with human-annotated evidence intervals. We also present ProactiveFrame, a training-free baseline that adjusts incoming frame rates through text tokens. A dual-tier sliding window retains recent high-FPS observations while downsampling older ones into sparse history. Experiments reveal substantial limitations: the strongest model, Gemini-3.5-Flash, scores only 50.7%. Denser sampling improves Qwen3-VL-8B from 32.9% at 2 FPS to 44.6% at 24 FPS, but gains saturate as history is compressed. ProactiveFrame outperforms sparse uniform sampling by 5.4 and 1.5 percentage points, yet remains well below oracle-guided focusing, showing that current VLMs struggle to determine from the stream alone when finer temporal perception is needed. FastBench provides a testbed for high-dynamic streaming video understanding. Code and data: this https URL.
14. 【2610.12423】Pumpire: Unified Benchmark for Metric Distance Estimation
链接:https://arxiv.org/abs/2610.12423
作者:Siyu Chen,Zehan Wang,Jiayang Xu,Yihan Wu,Jialei Wang,Junming Chen,Ziang Zhang,Yutong Ying,Zhou Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:distance estimation capability, evaluating metric point-pair, point-pair distance estimation, Pumpire directly assesses, present Pumpire
备注: Project page: [this https URL](https://pumpire.github.io/)
点击查看摘要
Abstract:We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors. In contrast to previous approaches that normally evaluate depth and camera intrinsics separately or evaluate point-clouds with geometric similarity metrics, which cannot directly reflect models' point-to-point distance estimation capability, Pumpire directly assesses point-to-point distances from the reconstructed geometry. To this end, we collect a large-scale and diverse dataset (pumpire-6k) comprising 100 real-world scenes, each annotated with physically measured point-pair distances and containing 64 frames, for a total of 6,400 frames. Building on this dataset, we establish a holistic evaluation protocol that covers both image- and video-level 3D foundation models and enables direct assessment of point-pair distance errors and cross-setting comparison. We conduct extensive experiments across 29 baseline configurations of representative 3D foundation models and provide a comprehensive analysis of the results. By offering this benchmark, we target the more fundamental ability to perceive and estimate physical scale in the reconstructed 3D space, which prior evaluation protocols have largely overlooked. The project page can be found at this https URL
15. 【2610.12421】Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
链接:https://arxiv.org/abs/2610.12421
作者:Luping Liu,Bingyi Kang,Yifan Wang,Dong Xu
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:simplifying spatio-temporal priors, spatio-temporal priors, rigid geometry, Dense correspondence matching, matching has historically
备注: Accepted at NeurIPS 2026. 24 pages, 7 figures, including appendices
点击查看摘要
Abstract:Dense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these assumptions break down in image editing and reference-guided generation (IEG), where transformations can preserve visual identity while breaking physical continuity. To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes. Teacher-guided iterative refinement further improves correspondence in IEG without dense correspondence annotations. Experimentally, a single FreeMatching model substantially improves correspondence quality on challenging IEG image pairs while retaining competitive performance on classical benchmarks. Furthermore, we demonstrate its utility as a quantitative metric for evaluating identity preservation, with scores that correlate with human judgment. The code is available at this https URL.
16. 【2610.12419】OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
链接:https://arxiv.org/abs/2610.12419
作者:Hongyu Li,Manyuan Zhang,Kaituo Feng,Shu Chen,Dian Zheng,Hao Li,Hao Yu,Zhangquan Chen,Zoey Guo,Ray Zhang,Shaofei Huang,Tianrui Hui,Linjiang Huang,Si Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:deep research require, external retrieval, share a workflow, Grounded Evidence Graph, video deep research
备注:
点击查看摘要
Abstract:Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: this https URL
17. 【2610.12417】WOVEN: Weaving Visual World Modeling into Multimodal LLMs
链接:https://arxiv.org/abs/2610.12417
作者:Zheyu Fan,Yue Zhang,Mingkai Deng,Kangrui Wang,Qineng Wang,Canyu Chen,Jie Hao,Xing Fan,Chenlei Guo,Eric P. Xing,Mohit Bansal,Manling Li
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Multimodal large language, Multimodal large, large language models, visual transition reasoning, struggle with spatial
备注:
点击查看摘要
Abstract:Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
18. 【2610.12416】MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances
链接:https://arxiv.org/abs/2610.12416
作者:Mingyuan Lei,Yoonchang Sung,Tat-Jen Cham
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Generating realistic human-object, Generating realistic, Generating, synthesizing realistic human-object, human-object
备注: Project page: [this https URL](https://leimingyuan.github.io/MAMHOI-project-page/)
点击查看摘要
Abstract:Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion. However, supervision for these capabilities is rarely available jointly at scale. Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce. We present MAMHOI, an affordance-mediated factorization for scene-aware human-object interaction generation. MAMHOI factorizes scene-aware HOI generation through an explicit motion-affordance interface between scene understanding and motion synthesis: a scene-conditioned model first predicts where and how an interaction can be feasibly executed, and an affordance-conditioned HOI model then generates the corresponding human-object motion. This factorization allows scene understanding and interaction dynamics to be learned from complementary sources of supervision without requiring paired human-object-scene data. Experiments in complex indoor environments show that MAMHOI reduces object--scene penetration while better preserving human--object interaction quality, yielding more realistic and physically feasible scene-aware interactions. Project page: this https URL
19. 【2610.12412】WorldCast: Distributed Multiplayer World Models
链接:https://arxiv.org/abs/2610.12412
作者:Ziyang Ye,Junchao Huang,Evelyn Zhang,Zhihao Xie,Ruicheng Zhang,Boyao Han,Litao Ban,Ziye Wang,Xinting Hu,Shaoshuai Shi,Zhuotao Tian,Li Jiang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:generate independently controlled, independently controlled views, player, generate independently, independently controlled
备注: 31 pages, 19 figures, 20 tables. Project page: [this https URL](https://ziyang-ye.github.io/WorldCast-Page)
点击查看摘要
Abstract:Multiplayer world models must generate independently controlled views with consistent representations of both players and their shared environment. Most existing approaches coordinate multiple players through joint multi-view generation, whose cost grows with each additional player. We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model. Using recorded player positions and map geometry during training, the state model estimates the player's position from generated video and control inputs. Clients exchange player states and project them into camera-aligned player state fields that guide where and how other players are rendered. Shared scene state enables clients to reuse one another's generated observations to maintain consistent scene appearance across views. Experiments on Counter-Strike 2 demonstrate WorldCast's consistency, real-time performance, and distributed scalability. The camera-aligned player state field improves player rendering rates by over an order of magnitude over joint-generation methods, while shared scene state improves visual consistency over whole rounds. Each client runs in real time and exchanges only player and scene states, enabling scalable multiplayer generation without a centralized computational bottleneck. Image quality remains stable over hour-long rollouts.
20. 【2610.12407】LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
链接:https://arxiv.org/abs/2610.12407
作者:Shashank Hegde,Alexander Popov,Elie Aljalbout,Nikolai Smolyanskiy
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:complicate downstream predictions, JEPA world model, World Model, future observations, carries noisy
备注: 14 pages, 6 figures
点击查看摘要
Abstract:World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.
21. 【2610.12403】ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
链接:https://arxiv.org/abs/2610.12403
作者:Hongxing Li,Dingming Li,Yixin Li,Yong Du,Wenqi Zhang,Weiming Lu,Jun Xiao,Yueting Zhuang,Yongliang Shen
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Skill-augmented agents improve, improve sample efficiency, agents improve sample, Skill-augmented agents, reusable strategies
备注: Code: [this https URL](https://github.com/ZJU-REAL/ViSkill)
点击查看摘要
Abstract:Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at this https URL.
22. 【2610.12402】SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
链接:https://arxiv.org/abs/2610.12402
作者:Hongxing Li,Jinyue Su,Dingming Li,Wenqi Zhang,Weiming Lu,Jun Xiao,Yueting Zhuang,Yongliang Shen
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Existing spatial reasoning, Existing spatial, reading off relations, relations already visible, test spatial perception
备注: Code: [this https URL](https://github.com/ZJU-REAL/SpaceCast-Bench) Dataset: [this https URL](https://huggingface.co/datasets/hongxingli/SpaceCast-Bench)
点击查看摘要
Abstract:Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
23. 【2610.12399】SpaceFlow: Locally Controllable 3D Generation
链接:https://arxiv.org/abs/2610.12399
作者:Neil De La Fuente,Joan Lafuente,Mukhammadali Sayfiddinov,Felicia Scharitzer,Marc Pollefeys,Ata Celen,Sayan Deb Sarkar,Elisabetta Fedele
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)
关键词:methods lack explicit, global control strength, lack explicit local, generation methods lack, explicit local control
备注:
点击查看摘要
Abstract:Current 3D generation methods lack explicit local control: geometric adherence is often defined by a global control strength, and appearance cannot be specified locally. We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives. Each primitive serves as a proxy for an object part and is assigned a local control level, enabling users to specify whether regions should strictly follow the input shape or allow generative completion. During structure generation, we enforce these spatial constraints within the generative flow process. For appearance synthesis, the generated structure is segmented and matched to the primitives. Each generated part is conditioned only on its assigned text or image cue, thereby limiting cross-part leakage. Regional geometry metrics demonstrate that SpaceFlow preserves the specified geometry in high-control regions and enables plausible shape variation in low-control areas. A user study further indicates that the resulting balance between geometric fidelity and generative freedom remains competitive in overall quality. When evaluating appearance on fixed geometry, text-conditioned routing achieves state-of-the-art prompt faithfulness and color/material accuracy. Qualitative results additionally show localized routing of image cues. The project page is available at this http URL.
24. 【2610.12388】GenIA: Generative Reconstruction with Test-Time Input Alignment
链接:https://arxiv.org/abs/2610.12388
作者:Stefano Esposito,Naama Pearl,Polina Karpikova,Samuel Rota Bulò,Lorenzo Porzi,Peter Kontschieder,Andreas Geiger,Jonathon Luiten
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:observations remains challenging, Reconstructing complete, remains challenging, sparse multi-view observations, multi-view observations remains
备注:
点击查看摘要
Abstract:Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at this https URL.
25. 【2610.12382】WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation
链接:https://arxiv.org/abs/2610.12382
作者:Jing He,Kaixin Ding,Xingye Tian,Guibao Shen,Wenhang Ge,Xin Tao,Pengfei Wan,Ying-Cong Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:requires generated videos, videos to maintain, consistency, dynamic, generated videos
备注: Includes an appendix with implementation details, human evaluation, and additional ablation analysis. Project page: [this https URL](https://worldalign.github.io/)
点击查看摘要
Abstract:Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency. Static consistency requires coherent 3D structure in static environments across viewpoints, while dynamic consistency requires plausible subject motion and consistent appearance over time. Geometry-aware post-training offers a promising way to improve world consistency. However, existing methods often rely on a static-scene assumption. Even those that accommodate dynamic scenes struggle to provide reliable static-consistency feedback, while dynamic consistency is often overlooked or inadequately assessed. To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions. For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions. For dynamic subjects, WorldAlign uses a strong vision-language model (VLM) as a dynamic world prior and constructs a VLM-as-a-judge reward based on sample-specific checklists that assess dynamicity, physical plausibility, shape, and texture consistency. This decoupled design enables more effective online post-training without requiring human preference annotations. Across two pretrained image-to-video generators, Wan2.1 and Wan2.2, WorldAlign jointly improves static and dynamic consistency over existing methods without suppressing overall or subject motion. These results support decoupled world-prior alignment for more faithful visual world simulation. Project page: this https URL.
26. 【2610.12374】AgentGarten: Code Worlds for Evolving Agents
链接:https://arxiv.org/abs/2610.12374
作者:Jiawei Chi,Shangchen Miao,Zhiyuan Shi,Kailu Wu,Hanyang Wang,Weiliang Chen,Qiyu Dai,Jinshan Ren,Jun Gao,Mingsheng Long,Yueqi Duan,Jiangran Lyu,Jialong Wu,Fangfu Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Interactive virtual worlds, agents, Adversarial Forcing, Interactive virtual, virtual worlds
备注: Project page: [this https URL](https://mirros-lab.github.io/agent-garten)
点击查看摘要
Abstract:Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.
27. 【2610.12369】Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
链接:https://arxiv.org/abs/2610.12369
作者:Kairui Hu,Siyuan Hu,Fangzhou Hong,Zhaoxi Chen,Ziwei Liu
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Harness VLA queries, VLA maps observations, Agent Harness, Harness VLA, Embodied Turing Machine
备注: 31 pages, 19 figures
点击查看摘要
Abstract:Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.
28. 【2610.12363】HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
链接:https://arxiv.org/abs/2610.12363
作者:Xiazhen Wu,Wansong Qin,Yangbin Zheng,Liangda Fang,Zhan Li,Xiujie Huang,Liushen Zhou,Quanlong Guan
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:automated scoring technologies, scoring technologies constitute, technologies constitute critical, constitute critical infrastructure, Intelligent grading
备注:
点击查看摘要
Abstract:Intelligent grading and automated scoring technologies constitute critical infrastructure for smart education. However, existing document parsing and handwriting recognition benchmarks are predominantly designed for well-structured printed documents or isolated mathematical expressions, lacking datasets that capture the complex characteristics inherent to student answer sheets, including multi-line derivation processes, heterogeneous mixtures of text and mathematical formulae, and noise artifacts such as strikethroughs. To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research. Building upon HANS, we propose NA-GOT, an end-to-end framework that achieves two-stage noise suppression through a lightweight noise suppression module operating at the feature level, complemented by a noiseaware attention mechanism incorporated into the decoding stage. Experimental results demonstrate that HANS poses substantial challenges to existing methods, while NA-GOT achieves significant improvements in both accuracy and stability for answer process recognition. The dataset will be made publicly available upon publication.
29. 【2610.12355】Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
链接:https://arxiv.org/abs/2610.12355
作者:Hongxing Li,Yixin Li,Dingming Li,Zixuan Wang,Yuchen Yan,Wenqi Zhang,Weiming Lu,Yongliang Shen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:RGB inputs, directly provide geometric, persistent weakness, weakness of vision-language, provide geometric evidence
备注: Code available at [this https URL](https://github.com/ZJU-REAL/GPD)
点击查看摘要
Abstract:Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.
30. 【2610.12343】Reasoning-Informed Visual Editing
链接:https://arxiv.org/abs/2610.12343
作者:Xue Yang,Peiyuan Zhang,Yilun Zhu,Qihao Yang,Mingxin Liu,Xiangyu Zhao,Ziqian Fan,Zhaokai Wang,Yan Li,Yifan Yang,Xu Yang,Xiaosong Jia,Yue Zhou,Zhihang Zhong,Junchi Yan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large Multi-modality Models, Large Multi-modality, made significant progress, preserving appearance consistency, supporting flexible input
备注:
点击查看摘要
Abstract:Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task. RISEBench++ extends the taxonomy into a hierarchical scheme spanning six reasoning dimensions: Temporal, Causal, Spatial, Logical, and Counterfactual Reasoning, together with Hybrid Reasoning integrating multiple reasoning types across multi-turn edits. These dimensions are further decomposed into 12 subcategories and 65 fine-grained task types. We expand input formats to include multi-image conditioning and scale the benchmark to 1000 human-annotated test cases, released in English and Chinese. We also improve our evaluation framework, assessing Instruction Reasoning, Appearance Consistency, and Visual Plausibility with human judges and an LMM-as-a-judge approach for more reliable and calibrated judgements. Beyond benchmarking, we introduce RISE-Agent, a training-free agentic framework integrating reasoning-driven planning, tool-augmented execution, and verifier-guided refinement, outperforming most strong existing approaches across diverse RISE tasks. We evaluate 58 visual editing approaches, including 34 open-source models, 19 closed-source models, and 5 agentic methods. The results reveal substantial challenges in reasoning-based visual editing, with even the strongest evaluated approach, GPT-Image-2.5 Sunburst, achieving only 56.6% accuracy. RISEBench++ highlights the limitations of contemporary editing models, provides insights, and indicates future directions for reasoning-aware visual editing.
31. 【2610.12337】ContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image Restoration
链接:https://arxiv.org/abs/2610.12337
作者:Jialei He,Enhe Liu,Sifan Song,Pengfei Jin,Jionglong Su,Hongbin Wang,Zhixiang Lu,Yanhao Huang,Anteng Cai,Zhengyong Jiang,Jiaman Ding,S. Kevin Zhou,Jinfeng Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:medical image restoration, continuity provides complementary, complementary information, information for medical, medical image
备注:
点击查看摘要
Abstract:Anatomical continuity provides complementary information for medical image restoration, but its use requires accounting for local anatomy and variations in slice sampling. We introduce ContiLNN, which augments two-dimensional restoration backbones with bidirectional closed-form continuous-time (Bi-CfC) modules for cross-slice modeling while retaining in-plane feature extraction. Slice-index intervals modulate gates determined by local features and hidden states, enabling propagation to respond to sampling variations without numerical ODE integration. Reference-guided consistency aligns first- and second-order cross-slice intensity differences to preserve anatomical variation, while distillation from a frozen backbone helps retain in-plane fidelity. Across five training seeds, ContiLNN improves mean PSNR over Restore-RWKV by 0.1907, 1.0176, and 1.2482 dB for CT denoising, MRI super-resolution, and reduced-count PET restoration, respectively, with lower RMSE in all three tasks. CT results are descriptive for one held-out patient. PET ablations support ordered propagation beyond additional pointwise capacity. Under contiguous training, Bi-CfC achieves higher fidelity than a Bi-GRU with similar parameter counts and arithmetic costs across all tested sampling conditions. Matched seven-slice profiling shows 52.8% lower latency and 57.0% lower peak GPU memory use than Bi-GRU. Mixed-gap training improves sparse and irregular-context performance for both operators, without a uniform ranking across metrics and contexts. Experiments with fewer training patients and a second backbone further support data efficiency and backbone compatibility.
32. 【2610.12333】RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact Reasoning
链接:https://arxiv.org/abs/2610.12333
作者:Ruixiang Ouyang,Guanren Qiao,Fansen Meng,Yueci Deng,Ruixing Jin,Kui Jia,Guiliang Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Machine Learning (cs.LG)
关键词:Accurate simulation, essential for predictive, Accurate, predictive physical world, contact
备注:
点击查看摘要
Abstract:Accurate simulation of rigid-body interactions is essential for predictive physical world models. Despite recent progress in modeling object dynamics, capturing how local contacts between surfaces shape object motion remains challenging. While end-to-end world models predict interactions across entire scenes or objects, in practice, rigid-body contact is inherently local, and only nearby surfaces can directly exchange contact forces. Motivated by this observation, we introduce Rigid-body Contact Reasoning (RiCo), which represents interactions between objects through sparse neighborhoods of contact surface points. RiCo combines each point's state with the relative geometry, motion, and physical properties of nearby surfaces, then reasons across the object's points to determine how these local contacts jointly affect its motion. By confining cross-object reasoning to nearby surfaces while propagating contact information within each rigid body, RiCo retains fine-grained interaction details without the cost of modeling every pair of scene points. Such properties enable RiCo a higher accuracy and contact fidelity. Experiments on MOVi-benchmark demonstrate that RiCo reduces 100-frame position and orientation errors by 31-35% and approximately 38%, respectively, compared with baselines. Moreover, RiCo achieves high contact fidelity, with ground-truth-relative penetration-time and mean-depth differences of 11.0% and 2.22 mm, respectively. RiCo further generalizes zero-shot from small-scale training scenarios to scenes containing 270 objects. Our real-world multi-ball collision experiments further provide preliminary evidence of sim-to-real transfer.
33. 【2610.12316】Controllable Exaggeration for Generative Motion Models via Training-Time Adaptation and Inference-Time Guidance
链接:https://arxiv.org/abs/2610.12316
作者:Amirhossein Zamani,Arianna Rampini,Bruno Roy
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent motion generative, overlook established animation, motion generative pipelines, synthesizing physically plausible, Recent motion
备注:
点击查看摘要
Abstract:Recent motion generative models have demonstrated strong capabilities in synthesizing physically plausible character motion, but often overlook established animation principles used by professional animators to ground and design their animation work. Understanding and incorporating these principles into motion generative pipelines is essential for producing motions that serve not only physically grounded applications but also the needs of the character animation community. This enables the creation of characters that not only move in physically plausible ways but also feel alive, expressive, and engaging. To close this gap, we focus on the Exaggeration principle of animation and investigate how it can be incorporated into modern motion generative pipelines to produce more expressive character motions. To this end, we introduce a framework that operates at two stages of existing motion generative pipelines. The first stage introduces exaggeration during training, where we perform supervised fine-tuning of pre-trained text-to-motion models on our curated exaggeration dataset. The second stage operates at inference time, where we: (i) introduce a mathematical formulation of exaggeration based on dynamic movement primitives (DMPs); and (ii) leverage this formulation as an exaggeration guidance signal to guide existing diffusion and flow-matching text-to-motion generation models toward exaggerated motion without additional training. Through qualitative and quantitative evaluations against three strong motion generation models, we show that our methods generate more exaggerated and expressive motions while preserving neutral reference motion intent and physical plausibility.
34. 【2610.12310】Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?
链接:https://arxiv.org/abs/2610.12310
作者:Xingwu Zhang,Duanyang Du,Huiling Zhu,Jiayue Dai,Yixiao Liu,Guozhi Liu,Zhihan Zhang,Zijun Long
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:large language model, single multimodal large, multimodal large language, struggles to excel, simultaneously at detection
备注: 21 pages, 5 figures, 8 tables
点击查看摘要
Abstract:A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model. Meanwhile, different MLLMs offer complementary strengths but share this weakness in fine-grained perception, so combining them alone cannot remove it. We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert. A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality. SiGMA reaches 85.2% average accuracy on MMAD, 4.0% above the strongest trained specialist and Gemini-2.5-Pro and within 1.5% of human experts. Even with three agents of at most 9B parameters, it reaches 84.4%, and new MLLMs join without retraining.
35. 【2610.12307】BudgetPix: Compute-Adaptive Tokenization for Pixel-Space Image Diffusion
链接:https://arxiv.org/abs/2610.12307
作者:Ozgur Kara,Yujia Chen,Daniel Watson,David Forsyth,James Matthew Rehg,Wen-Sheng Chu,Du Tran
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:equally-sized image patches, generation models rely, allocating the exact, models rely, rely on uniform
备注: More details are available at our project page: [this https URL](https://karaozgur.com/BudgetPix)
点击查看摘要
Abstract:Most image generation models rely on uniform tokenization, allocating the exact same computational budget to equally-sized image patches. This static paradigm cannot adapt to different resource constraints at inference time, and yields suboptimal quality-cost tradeoff by devoting the same effort to both plain backgrounds and intricate details. We propose BudgetPix, an adaptive tokenization framework that dynamically allocates compute based on visual complexity and spatial layout, enabling flexible computational budgeting at inference time. BudgetPix comprises three key components: (1) an adaptive encoder that maps a fixed-size image to a variable-length token sequence using an entropy-guided quadtree alongside a multi-scale patch embedder; (2) a scale-aware decoder reconstructs fixed-resolution images from multi-scale token sets; and (3) a flexible training and sampling schedule that enables pixel-space denoisers to operate across variable token counts. BudgetPix seamlessly integrates with existing pixel-space diffusion architectures, enabling a single checkpoint to be operated at a wide range of compute budgets. Evaluated on text-to-image generation, BudgetPix matches the fidelity of MiniT2I-L at $512^2$ and PixelDiT at $1024^2$ using just 25% of the original compute budget. In class-conditional generation using a MeanFlow backbone, BudgetPix requires merely 60% of the full compute budget to produce images with near-zero quality degradation, observing a marginal 0.8-point increase in FID. Comprehensive assessments by human and VLM judges confirm that BudgetPix establishes a significantly improved quality-efficiency tradeoff over prior budget-adaptive baselines. More details are available at our project page: this https URL
36. 【2610.12305】From What to Which: Decoding Modifier Grounding in Frozen MLLMs
链接:https://arxiv.org/abs/2610.12305
作者:Barbara Toniella Corradini(AI for Good (AIGO), Istituto Italiano di Tecnologia, Italy),Caterina Gallegati(University of Siena, Italy),Ludovica Genovese(AI for Good (AIGO), Istituto Italiano di Tecnologia, Italy, University of Genoa, Italy),Vittorio Murino(AI for Good (AIGO), Istituto Italiano di Tecnologia, Italy, University of Verona, Italy)
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large Language Models, Multimodal Large Language, Language Models, Multimodal Large, Large Language
备注:
点击查看摘要
Abstract:As Multimodal Large Language Models (MLLMs) can describe increasingly complex visual scenes, token-level grounding becomes crucial. Yet, when an MLLM generates "the yellow banana on the left", established grounding approaches focus on what is in the image ("banana"), overlooking tokens that help describe which instance is meant ("yellow", "left"). In this work, we ask whether frozen MLLM representations contain decodable grounding information about the referred instance across generated tokens, extending to modifiers such as attributes, spatial expressions, and relational/action terms. To address this question, we introduce OTTER, a lightweight supervised probe over frozen MLLM representations that uses Optimal Transport (OT) to align generated tokens with visual regions and produce compact grounding maps. Our results show that (i) instance-discriminative visual information can be decoded from modifier tokens, with the clearest evidence for spatial terms, but (ii) is not confined to them, as contextualized object nouns also carry referential information; (iii) the recovered grounding remains informative under context perturbations, while selected regions remain relevant to generation; and (iv) the learned OT-based grounding extends beyond the controlled setting to free generation and cross-dataset transfer.
37. 【2610.12299】Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
链接:https://arxiv.org/abs/2610.12299
作者:Dahyun Chung,Siyoon Jin,Hyunwook Choi,Honggyu An,Junyoung Seo,Hyunsung Kim,Seung Wook Kim,Seungryong Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:predict first-person observations, first-person observations conditioned, models predict first-person, multi-agent egocentric world, world models predict
备注:
点击查看摘要
Abstract:Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.
38. 【2610.12282】Slot3R: Set-Associative Spatial Memory for Streaming 3D Reconstruction
链接:https://arxiv.org/abs/2610.12282
作者:Xiyuan Zhang,Yanming Yang,Kaiyuan Xu,Ruibo Li,Chi Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:expanding scene online, reconstruction must preserve, scene online, processing an expanding, expanding scene
备注: Project Page: [this https URL](https://ashleyxyz.github.io/Slot-3R/)
点击查看摘要
Abstract:Streaming 3D reconstruction must preserve evidence from each frame while processing an expanding scene online. Spatial memory is a natural fit because it organizes history by reconstructed 3D location. Yet Point3R uses spatial proximity both to associate a new observation with an existing memory entry and to decide whether to fuse it, conflating co-location with state identity. Because pointers summarize image patches, nearby pointers may encode distinct surfaces, viewpoints, or visibility conditions; averaging them can destroy complementary evidence before later frames disambiguate it. We argue that location should determine address, not whether observations must merge. Slot3R realizes this principle as a training-free, set-associative retrofit that lets multiple states coexist at a shared address while keeping the pretrained Point3R backbone frozen. A bounded sparse readout further decouples persistent storage from per-frame decoder access. At 300-500 sampled frames, Slot3R reduces Point3R's point-cloud accuracy error (Acc) by 57.1%-63.1% on 7Scenes and 64.0%-72.0% on NeuralRGBD, lowers Sim(3)-aligned absolute trajectory error (ATE) on all three pose benchmarks, and remains competitive on video-depth estimation. It completes all evaluated settings from 600 to 1000 sampled frames at about 19 FPS under the same protocol, whereas Point3R and InfiniteVGGT run out of memory at 800 frames and beyond.
39. 【2610.12266】DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
链接:https://arxiv.org/abs/2610.12266
作者:Jinghua Hou,Zhe Liu,Hengshuang Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Multimodal large language, large language models, made remarkable progress, Multimodal large, perception tasks essential
备注:
点击查看摘要
Abstract:Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.
40. 【2610.12256】Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
链接:https://arxiv.org/abs/2610.12256
作者:Youngtaek Oh,Qiyu Wu,Hiromi Wakaki,Junmo Kim,Yuki Mitsufuji
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:embeddings naturally involve, Omnimodal embeddings naturally, heterogeneous inputs, naturally involve, features across heterogeneous
备注: Accepted to EMNLP 2026 (Long, Findings). Code: [this https URL](https://github.com/sony/syn-omni)
点击查看摘要
Abstract:Omnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space over mixed-modality data, limiting structural separation between universal and modality-specific representations. To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration. Specifically, we introduce Orthogonal Modality-Expert LoRA (OME-LoRA), which decomposes adaptation into a shared LoRA path for universal semantics and modality-expert LoRA paths for modality-aware specialization. Furthermore, Progressive Synergy Routing (PSR) enables experts to first establish modality-specific priors, then gradually interact with other modality-experts for cross-modal synergy. Evaluated across 81 diverse tasks spanning image, video, audio, and audiovisual modalities, Syn-Omni consistently outperforms omnimodal baselines, demonstrating the effectiveness of structured specialization and cross-modal progressive collaboration.
41. 【2610.12248】EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
链接:https://arxiv.org/abs/2610.12248
作者:Heeseung Kim
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
关键词:Wearable augmented reality, continuous real-world interaction, Wearable augmented, timely spoken guidance, provide timely spoken
备注: Accepted to EMNLP 2026 (Main Conference). 25 pages, 12 figures, 11 tables. Project page: [this https URL](https://egocentricvoice.github.io/)
点击查看摘要
Abstract:Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. While proactive video assistants, spoken dialog systems, and egocentric task understanding have each advanced rapidly, existing systems do not address the joint problem of deciding when to speak and what to say from continuous first-person streams. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation and speech resynthesis, and convert each video session into a format where the model must decide at each moment whether to remain silent or provide spoken guidance. We fine-tune an omni-modal LLM with our data, and further improve its proactive intervention behavior with direct preference optimization. Experiments across closed and open-source models show that existing systems rarely produce well-timed, meaningful proactive interventions, while EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.
42. 【2610.12230】From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
链接:https://arxiv.org/abs/2610.12230
作者:Yitong Wang,Fangyun Wei,Jinjing Zhao,Sirui Zhang,Hongyang Zhang,Dong Chen,Bo Dai,Yan Lu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:encode inherently two-dimensional, Spatial Canvas Interface, Spatial Canvas, Text Specifications, inherently two-dimensional composition
备注: Project page: [this https URL](https://snowflakewang.github.io/Compo-Page/) GitHub: [this https URL](https://github.com/snowflakewang/Compo)
点击查看摘要
Abstract:Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
43. 【2610.12229】VibeEdit: Image Editing with Canvas Instructions
链接:https://arxiv.org/abs/2610.12229
作者:Jinjing Zhao,Fangyun Wei,Yitong Wang,Xiuyu Wu,Yunuo Chen,Yang Yue,Sirui Zhang,Wenbo Wang,Hongyang Zhang,Dong Chen,Yan Lu,Chang Xu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:text-guided image editing, image editing, describing the desired, identifying the intended, image editing interface
备注: Project page: [this https URL](https://zhaojingjing713.github.io/VibeEdit/)
点击查看摘要
Abstract:In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.
44. 【2610.12216】Stride Independent Patching for Deep Learning
链接:https://arxiv.org/abs/2610.12216
作者:Olivier Rukundo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:paper presents semi-automatic, presents semi-automatic stride-independent, stride-dependent patching techniques, semi-automatic stride-independent patching, automatic stride-dependent patching
备注: 8 pages, 6 figures, 2 tables
点击查看摘要
Abstract:This paper presents semi-automatic stride-independent patching (SSP) as an alternative to automatic stride-dependent patching techniques. SSP uses user or expert input to position predefined patches over one or more objects of interest. To evaluate its effectiveness, three patch-based datasets were created using SSP, overlapping patching (Overlap), and non-overlapping patching (Noverlap). DeepLabV3+ models with ResNet50, ResNet18, and MobileNetV2 backbones were trained sepa-rately on each dataset. Quantitative evaluations were performed on the respective test splits and a common external test set. SSP generally achieved higher segmentation scores on the test splits and required the shortest model training time across all three backbones. On the external test set, SSP achieved the highest average precision and F1-score across backbones, whereas Noverlap achieved the highest average recall. These preliminary results demonstrate that the potentially greater spatial coverage of Noverlap and Overlap does not generally translate into better segmentation perfor-mance and that SSP offers a favorable balance between segmentation performance and model training time.
45. 【2610.12189】Just Weather Scoring: Efficient End-to-end Nowcasting with Distributional Diffusion
链接:https://arxiv.org/abs/2610.12189
作者:Jannik Wiese,Johannes Schusterbauer,Tommaso Martorella,Björn Ommer
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:Generative diffusion models, separately trained compression, probabilistic precipitation nowcasting, deterministic forecasting components, Generative diffusion
备注: Project Page: [this https URL](https://compvis.github.io/jws)
点击查看摘要
Abstract:Generative diffusion models are well-suited for probabilistic precipitation nowcasting, but existing approaches often rely on separately trained compression or deterministic forecasting components and remain costly at inference due to iterative denoising. We introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation. Radar-space modeling greatly simplifies training and inference and eliminates uncertainty arising from lossy compression. JWS combines Masked Asynchronous Diffusion, a timestep-sampling scheme that preserves clean context while adapting diffusion training to high-dimensional spatio-temporal data, with a simple scoring-rule objective that aligns training with probabilistic forecasting and unlocks few-step generation. On the SEVIR and MeteoNet benchmarks, JWS achieves state-of-the-art probabilistic forecasting performance at reduced training and inference cost. Even our smallest model remains competitive using substantially fewer parameters and more than 17x faster inference.
46. 【2610.12182】AI-Based On-Board Maritime Object Detection for Earth Observation Payload Data Reduction on Versal Embedded Hardware
链接:https://arxiv.org/abs/2610.12182
作者:Thomas Goudemant,Aurélien Bobey,Omar Hlimi,Marjorie Bellizzi
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Earth-observation satellites acquire, Earth-observation satellites, satellites acquire, maritime surveillance, cover a tiny
备注: 8 pages. Accepted at the 10th On-Board Payload Data Compression Workshop (OBPDC 2026), Barcelona, Spain, 12-14 October 2026
点击查看摘要
Abstract:Very-high-resolution Earth-observation satellites acquire more data than they can store and downlink, while in maritime surveillance the vessels cover a tiny fraction of each scene. We study onboard vessel detection as a way to select what is downlinked, which reduces the data according to its content rather than coding every pixel; it is complementary to conventional onboard compression. The work follows three axes. (i) Data and algorithm: a controlled dataset is generated from 68 annotated Maxar scenes with 43 vessel classes, and a YOLOX-S detector is trained on it. (ii) Embedded deployment: the detector is quantized and deployed on the DPU of a Versal VC1902, with a limited loss of detection quality and a processing time of a few seconds per scene. (iii) Data reduction: we propose several downlink modes, from metadata only (box, class and score of each detection) to image crops around vessels, tiles holding detections, or the whole scene with a degraded background, and estimate from the measured detection errors the trade-off each offers between the vessels kept and the volume downlinked. On our dense harbor and coastal scenes, tiles keep 98% of the vessels with 29% of the scene volume, and crops 83% with 3%.
47. 【2610.12156】Connected Self Forcing: Beyond Local Learning in Video Autoregression
链接:https://arxiv.org/abs/2610.12156
作者:Dongbin Zhang,Chaoda Zheng,Kangjie Chen,Xiangyu Li,Shijia Chen,Jinhao Deng,Yuqi Zhang,Guangfeng Jiang,Hongbin Lin,Choo Sin Wai,Minqi Wang,Puyi Wang,Jingye Zhang,Yu Zhang,Xianming Liu,Boyang Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:mitigates exposure bias, Forcing mitigates exposure, stream long videos, histories with key-value, stream long
备注: Project Page: [this https URL](https://eastbeanzhang.github.io/CSF/)
点击查看摘要
Abstract:To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching. To keep memory manageable, it detaches historical caches, preserving forward dependencies between chunks but severing the backward gradient paths. We introduce Connected Self Forcing, a training framework that reconnects gradient paths across autoregressive chunks, allowing feedback from later predictions to guide how earlier context is generated. These connections go beyond historical KV-writing: gradients pass through generated latents into the computations that produced them, linking the generation of earlier context to its use in later predictions. To make this connected training memory-efficient, we develop shortcut gradient replay, which recovers cross-chunk gradients without retaining the full rollout computation graph. Integrated with distribution matching distillation, Connected Self Forcing trains historical chunks according to both their direct supervision and their contribution to subsequent generation. Experiments on autoregressive video generation show improvements in long-horizon visual quality and temporal consistency, without changing the inference procedure.
48. 【2610.12147】Healthy Counterfactual Generation via Diffusion Inpainting for Mammography Classification
链接:https://arxiv.org/abs/2610.12147
作者:Inês Cruchinho Garcia,Mariana Mourão,Francisco Maria Calisto,Carlos Santiago,Jacinto Nascimento
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:False negatives remain, False negatives, cancer screening due, Denoising Diffusion Probabilistic, computer-aided diagnosis
备注: Accepted at MICCAI Workshop Deep-Brea3th 2026
点击查看摘要
Abstract:False negatives remain a critical limitation of computer-aided diagnosis (CAD) systems for breast cancer screening due to delayed detection and treatment. To address this issue, we propose a counterfactual data augmentation strategy that generates healthy mammograms by "erasing" lesions from anomalous images, thereby enriching the training distribution. We train a Denoising Diffusion Probabilistic Model on BI-RADS 1 (healthy) mammograms and use a RePaint-based sampling strategy to inpaint realistic normal tissue within annotated lesion bounding boxes. The resulting healthy counterfactuals replace annotated lesion regions with realistic healthy tissue while preserving patient-specific anatomical structure, as supported by similarity metrics between real and generated images. Image realism was further assessed by radiologists and found to be consistent with the original dataset quality. We evaluate counterfactual augmentation across four representative classifier architectures: a convolutional neural network (ConvNeXt), a vision transformer (ViT), a vision-language model pre-trained on mammogram-report pairs (Mammo-CLIP) and a multi-scale attention-based multiple-instance learning framework (FPN-MIL). Experiments conducted on the VinDr-Mammo dataset show improvements in sensitivity across all architectures, particularly at 80\% fixed specificity, contributing towards more reliable CAD systems for breast cancer. Code is available at: this https URL.
49. 【2610.12127】LVS: Local View Synthesis from Relative Camera Pose by Reusing Previous Views
链接:https://arxiv.org/abs/2610.12127
作者:Qizhou Huo,Xuan Sun,Yongfei Guo,Zhipeng Wang,Yuanhao Gong
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Multimedia (cs.MM); Image and Video Processing (eess.IV)
关键词:small camera motions, camera motions preserve, requires frequent view, exploration requires frequent, Interactive scene exploration
备注:
点击查看摘要
Abstract:Interactive scene exploration requires frequent view updates, although small camera motions preserve much of the visible content. Conventional 3D Gaussian Splatting nevertheless renders each target view, leaving this image overlap unexploited. Reusing rendered images offers an alternative. Geometric warping alone cannot recover newly exposed content and remains sensitive to depth errors. We propose a per-scene framework that replaces repeated scene rendering for nearby views with relative-pose-guided RGB-D image reuse. Geometric warping uses depth and relative pose to transport source content, while a lightweight multiscale network predicts RGB residuals to correct artifacts and infer missing appearance. Cached source features further reduce repeated computation. On GS-render, residual refinement improves PSNR by 0.72~dB over pure warping; evaluations on captured and rendered scenes demonstrate low query latency. This separation of scene rendering from local view updates supports responsive scene exploration, with potential applications in augmented and virtual reality.
50. 【2610.12126】SuperNav: An Agentic Navigation System for Any Task in Any Scene
链接:https://arxiv.org/abs/2610.12126
作者:Jinkai Zhang,Jingyi Xu,Yuanhong Yu,Jiarui Guo,Ruizhen Hu,Hujun Bao,Xiaowei Zhou,Sida Peng
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:handle diverse human, diverse human requests, combining task generality, General-purpose service robots, handle diverse
备注: 20 pages, 7 figures. Project page: [this https URL](https://zju3dv.github.io/SuperNav/)
点击查看摘要
Abstract:General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments. Our key insight is to let the MLLM focus on interpreting requests, understanding scenes, and making decisions while preserving its general-purpose capabilities and delegating motion execution to navigation tools. To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM. Our harness supports these decisions with Navigation Skills, agent-oriented Tools for physical interaction, and task-progress and context management. A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback. Together, these components support sustained navigation across different task requirements and environments. SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks. Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments. Project Page: this https URL
51. 【2610.12107】ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation
链接:https://arxiv.org/abs/2610.12107
作者:Ruicheng Zhang,Kaiwen Shen,Jiaqi Hou,Shuhan Yang,Junchao Huang,Kewei Zhang,Jun Zhou,Li Jiang,Shen Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generalized referring expression, referring expression segmentation, requires dynamically balancing, dynamically balancing high-level, fine-grained visual evidence
备注:
点击查看摘要
Abstract:Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation. This requirement challenges existing cascaded vision-language architectures, which typically rely on static feature interfaces and single-pass mask prediction, limiting adaptive perception and geometric correction. We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks. Evolution-Aware Semantic Scheduling (EASS) couples contour-guided bidirectional boundary sampling with state-conditioned routing of multilevel multimodal features, adapting perception to each contour state. Following supervised initialization, Dustbin-Augmented Entropic Credit Transport GRPO (DECT-GRPO) jointly optimizes discrete grounding and continuous contour actions with instance-level credits. Its rollout rewards and credits are derived from soft prediction-target correspondences that account for false positives and missed targets. ContourVLA improves gIoU over the strongest evaluated baselines by 8.7, 2.8, and 2.7 points on gRefCOCO val, testA, and testB, respectively, and achieves the highest mIoU across all eight RefCOCO, RefCOCO+, and RefCOCOg splits.
52. 【2610.12104】VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling
链接:https://arxiv.org/abs/2610.12104
作者:Leigang Qu,Feng Cheng,Ziyan Yang,Bangbang Yang,Zhaoyang Huang,Wei Chow,Yicong Li,Wenjie Wang,Tat-Seng Chua,Yan Zeng
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
关键词:editor remains significantly, remains significantly harder, Building a capable, capable video editor, video editor remains
备注: Accepted to NeurIPS'26. Project page: [this https URL](https://vincie-next.github.io/)
点击查看摘要
Abstract:Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video - Image - Image - Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
53. 【2610.12102】Few-Step Generation via Data-Space Iteration
链接:https://arxiv.org/abs/2610.12102
作者:Shanchuan Lin,Yansong Peng,Fu-Yun Wang,Haoqi Fan
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:high-quality generative models, training high-quality generative, learned probability flow, generative models, scalable paradigm
备注:
点击查看摘要
Abstract:Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations. Distillation can reduce this cost to one or a few evaluations; however, one-step generation often sacrifices quality, making few-step generation the practical operating regime. Existing few-step methods perform their iterative computation along the probability flow and therefore require a fixed, manually chosen timestep discretization. This discretization is often chosen heuristically and is expensive to tune; it may also be restrictive when refinement difficulty differs across samples or spatial locations. We introduce data-space iteration, a few-step generation framework that removes flow discretization altogether. Starting from noise, a shared generator directly refines its prediction in data space, with every iteration trained to produce the best sample permitted by its capacity. Our formulation integrates with distribution matching distillation (DMD) with minimal changes, enabling a controlled comparison between iteration methods under matched training settings. On class-conditional ImageNet 256x256, data-space iteration outperforms standard discretization baselines and matches or improves upon variants selected through schedule search, without requiring schedule-specific training. These results show that data-space iteration provides a simple and effective alternative to discretized flow-space iteration for fast generation.
54. 【2610.12095】DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
链接:https://arxiv.org/abs/2610.12095
作者:Wenhao Li,Xianjing Meng,Qiangchang Wang,Zhongyi Han,Yilong Yin,Liqiang Nie
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Few-shot learning aims, Few-shot learning, aims to recognize, recognize novel categories, Few-shot
备注: This work has been submitted to the IEEE TPAMI for possible publication
点击查看摘要
Abstract:Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at this https URL.
55. 【2610.12081】Perception Test 2026: Challenge Summary and Extension to City-scale Audio-Visual Reasoning
链接:https://arxiv.org/abs/2610.12081
作者:Fedor Kitashov,João Carreira,Shiry Ginosar,Dima Damen,Andrew Zisserman,Viorica Pătrăucean
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Perception Test challenge, Test challenge series, Computer Vision, European Conference, Conference on Computer
备注:
点击查看摘要
Abstract:Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks used for the city-scale tracks and summarise the winning solutions across all tracks, including a generalist model that competed across all tracks with satisfactory performance. The winning solutions in the newly added city-scale tracks demonstrated that complex spatial and multimodal reasoning can be solved by expensive agentic pipelines, but remains difficult for multimodal models used standalone.
56. 【2610.12073】A Minimal Optical-Flow Representation for Vision-Based Tactile Rotation Classification in Robotic Manipulation Across Gravity Domains
链接:https://arxiv.org/abs/2610.12073
作者:Oscar Martinez-Bernal,Mario Cavero-Vidal,Francesco Grella,Carol Martinez
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision-based tactile sensors, rich contact information, sensors provide rich, provide rich contact, processing high-resolution images
备注:
点击查看摘要
Abstract:Vision-based tactile sensors provide rich contact information, but processing high-resolution images can be costly for resource-constrained platforms such as space robots. This work investigates whether a compact representation of tactile motion can classify object rotation across different gravity conditions. Dense optical flow from a simulated GelSight Mini is aggregated over a 7x9 grid into 126 features and used to classify the direction of load-induced rotation under Earth, Mars, Moon, and orbital gravity. Gravity causes a small but significant shift in these features, accounting for 1.6% of their variance (R2 = 0.016). Despite its small magnitude, this shift affects models trained only on Earth data: XGBoost accuracy decreases from 94.4% on Earth to 75.9% in orbit. In contrast, a single model trained across all four gravity domains achieves 96.3% overall accuracy and 95.1%-97.0% across individual domains, without using gravity as an input. The representation can also be reduced to 40 features while retaining 95.7% accuracy, with XGBoost requiring only 0.14 ms per inference. These findings show that Earth-gravity performance alone is insufficient to establish the transferability of tactile perception for space robotic manipulation, highlighting the need to account for gravity-induced domain shifts during training and validation.
57. 【2610.12069】LIVIN: Benchmarking Spatial and Embodied Intelligence in Digital Twins of Lived-In Homes
链接:https://arxiv.org/abs/2610.12069
作者:Peijun Xu,Chuansen Nie,Yiyang He,Yinuo Bai,Jingyang Liu,Kuixiang Shao,Yuyang Jiao,Kuanhao Xia,Jiayi Zhu,Zitian Yang,Yanqi Zhang,Tianye Tan,Shuwei Di,Junyi Xu,Jingyi Yu,Jiayuan Gu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:shape robot motion, Realistic household simulation, household simulation, simulation must capture, constraints that shape
备注:
点击查看摘要
Abstract:Realistic household simulation must capture not only diverse environments but also the lived-in object arrangements and spatial constraints that shape robot motion and interaction. Existing resources often trade off scale, real-world correspondence, and interaction readiness, leaving a gap in faithful, interactive replicas of how real homes are actually arranged. To this end, we introduce LIVIN, a benchmark for spatial and embodied intelligence built on digital twins of 30 diverse lived-in homes. These replicas preserve observed room layouts, furniture configurations, and everyday belongings. To construct them, we design a human-in-the-loop workflow comprising instance recognition, architectural reconstruction, and object generation and placement, with intermediate results reviewed and corrected by humans against the source observations at each stage. We evaluate four tasks in LIVIN: 3D detection, 3D reconstruction, navigation, and loco-manipulation. Our evaluations show that current methods remain challenged by the dense object arrangements, occlusions, limited free space, and constrained interaction regions found in realistic lived-in homes. We hope LIVIN will help advance embodied AI in real-world homes, from spatial understanding to robotic interaction, and ultimately bring embodied intelligence into everyday home environments.
58. 【2610.12060】Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning
链接:https://arxiv.org/abs/2610.12060
作者:Yicheng Xue,Han Wu,Jufeng Yang,Minjing Dong,Xinghao Chen,Hanting Chen,Jianyuan Guo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large Language Models, multimodal Large Language, Language Models, Large Language, Processing long visual
备注: 27 pages, 8 figures
点击查看摘要
Abstract:Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as reasoning unfolds, making it difficult for a fixed compressed context to retain all the details needed across stages. To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve. Deformable Aggregation of Region-wise Tokens (DART) learns content-adaptive groups and aggregation capacities, constructing compact Coarse representations linked to recoverable original Fine tokens. Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups. Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens. On Qwen3-VL-4B, ViMoD outperforms all evaluated baselines on all eight reasoning benchmarks at a 20% target visual token budget, improving the mean normalized score by 39.0% over the strongest evaluated one-shot baseline. These gains are achieved with only 0.0546% additional trainable parameters relative to the frozen backbone.
59. 【2610.12048】Do Not Train Away Uncertainty: Early Uncertainty Anchored Calibration
链接:https://arxiv.org/abs/2610.12048
作者:Yutong Xie,Jiawei Tang,Zhenglin Hua,Yuxiang Ma,Si Qin,Yaxin Hou,Hui Liu,Junhui Hou,Yuheng Jia
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Deep neural networks, including large language, achieved remarkable performance, Deep neural, large language models
备注: 18 pages, 9 figures, 11 tables
点击查看摘要
Abstract:Deep neural networks, including large language models, have achieved remarkable performance across various tasks. However, they are prone to overconfidence during training or fine-tuning. In this work, we observe a consistent phenomenon across different models that the early model is better calibrated, while later training or fine-tuning yields marginal accuracy gains but substantially increases calibration errors. Our analysis suggests that the early model retains uncertainty awareness in both its predictions and features, which is gradually lost with continued training. To avoid training away this uncertainty awareness, we propose \textbf{EUA-Cal}, a novel method that exploits the \textbf{E}arly model as an \textbf{U}ncertainty \textbf{A}nchor for \textbf{Cal}ibration. EUA-Cal introduces early prediction regularization to preserve early predictive uncertainty and prototype structure regularization to exploit uncertainty reflected in the early feature space, jointly mitigating overconfidence. Extensive experiments on image classification and multiple-choice question answering across eight diverse models demonstrate that EUA-Cal outperforms state-of-the-art calibration methods.
60. 【2610.11993】DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
链接:https://arxiv.org/abs/2610.11993
作者:Yupeng Xie,Zhenyang Wang,Jiayi Zhu,Yinghao Tang,Zhouan Shen,Yiyu Chen,Yuyu Luo
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:integrates data visualization, data video understanding, Data video, video understanding, Data
备注: 46 pages, 22 figures, 14 tables
点击查看摘要
Abstract:Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at this https URL.
61. 【2610.11986】FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment
链接:https://arxiv.org/abs/2610.11986
作者:Xiaoshan Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:show great potential, Vision-language models, declare a hazard, great potential, potential for damage
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate. This study examines that deficiency by using signal detection theory to decompose the decision behavior into perceptual capability and decision-criterion placement. We then propose a novel method for correcting the over-conservative decision policy, inspired by the finding that fear makes humans risk-averse, and ask whether an affective representation associated with fear can be causally manipulated to similarly alter a VLM's decision tendency. Using mechanistic interpretability, we localize a causally implicated affective circuit in the model and use activation steering to manipulate it while observing the effect on downstream prediction. The method is tested on a two-stage SeisMLLM pipeline built on Qwen2.5-VL-7B-Instruct, which flags only 27.0% of genuinely unsafe buildings on the SeisMLLM-1K test split and never issues a false Red, an SDT criterion of c = +1.354, despite adequate evidence quality (d' = 1.521). An affective direction is localized on emotion-rich natural scenes, causally validated by sparse-neuron knockout and distributed steering on held-out emotion data, and then injected into the building task. Fear-direction injection raises Red recall to 75.7% (p0.001), and subtracting the same direction suppresses Red predictions entirely, whereas norm-matched random and matched happiness directions show no significant effect. The mechanism is a shift in criterion (c=-1.515) while discrimination is not improved (d'=-0.493). These results show that VLM decisions can be adjusted at inference time without retraining and demonstrate how mechanistic interpretability can be used to diagnose and control VLM decision behaviors in engineering applications.
62. 【2610.11967】Learning Which Correspondences to Trust: Confidence-Weighted Event-Camera Localization in LiDAR Maps
链接:https://arxiv.org/abs/2610.11967
作者:Panagiotis Kiousis,Kuangyi Chen,Jun Zhang,Friedrich Fraundorfer
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:pre-built LiDAR map, dense optical-flow estimation, Localizing an event, rendered depth view, event image
备注: 8 pages, 8 figures/tables. Submitted to IEEE ICRA 2027 (under review). Code: [this https URL](https://github.com/panagiotisq/CELL)
点击查看摘要
Abstract:Localizing an event camera against a pre-built LiDAR map can be cast as dense optical-flow estimation between a rendered depth view and an event image, followed by a Perspective-n-Point (PnP) solver over the induced 3D-2D correspondences. Existing pipelines rely on geometric consensus during pose estimation, but do not explicitly model the reliability or pose informativeness, i.e., how strongly a correspondence constrains the camera pose, of individual correspondences. We show that the natural way to learn it -- using the per-correspondence error to constrain the learning of confidence -- suffers from a depth-dependent bias: small pixel errors reside predominantly at large depths and do not lead to high pose informativeness. Instead, in our method (CELL), we learn a per-correspondence confidence end-to-end through the pose, using a differentiable probabilistic PnP whose log-partition term encourages weight configurations that yield a better-constrained pose distribution. The learned confidence is used in three ways: (i) it reweights the flow supervision in a decoupled training scheme that keeps pose gradients out of the flow/edge backbone; (ii) it drives a probabilistic correspondence selection at test time; and (iii) together with the network's edge-probability it weights a final edge-matching refinement. We further design a partial-completion depth representation that adds signal without hallucinating across large gaps. On M3ED and DSEC our full system improves over the LEAR baseline on the majority of the evaluated sequences: it reduces the median translation error by up to 26.9% and the median rotation error by up to 15.8%.
63. 【2610.11954】Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion
链接:https://arxiv.org/abs/2610.11954
作者:Kartik Bali,Mahish Guru,Yiderigun Borjigin,Alexandra Starostina,Christian J. Cyron,Roland Aydin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:reality rarely grants, unrestricted visual access, luxury reality rarely, rarely grants, unrestricted visual
备注:
点击查看摘要
Abstract:CAD reconstruction methods assume a luxury reality rarely grants: unrestricted visual access to the object, photographed from any desired angle. Real objects, however, are scene-embedded, bolted against walls, wedged into corners, resting on floors, where the scene renders much of the view sphere unreachable and the remaining views unequally informative. We introduce \textbf{SightCAD}, a framework for parametric CAD reconstruction that treats view feasibility as a first-class constraint. In this work we consider objects from standard CAD benchmarks embedded in realistic indoor scenes with physically derived visibility constraints over a discrete view sphere. A learned view selector must choose $K$ feasible views for a vision--language model (VLM) that generates executable CadQuery code, scored by geometric fidelity of the executed solid. Because reward arrives only after discrete view selection, autoregressive generation, and CAD-kernel execution, we propose a joint training paradigm in which the view selector and the CAD-generation VLM are trained together against this reward. The learned selection policy departs sharply from random, uniform, and coverage-greedy alternatives, outperforming surface-area maximization (SA-max) by up to $6.4$ Intersection-over-Union (IoU) points across budgets $K\in\{1,\dots,5\}$. The full system surpasses strong external baselines on scene-embedded, occluded multi-view renders of DeepCAD and Fusion360 objects ($+21$ and $+17$ effective-mIoU points over the best baseline, respectively), as well as on test-time domain-canonicalized real images from the industrial T-LESS benchmark and on both synthetic and real images from the MP6D industrial metal-parts benchmark, while producing the highest rate of executable programs of any method compared (invalid-code rate ${\leq}1.5\%$).
64. 【2610.11942】Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
链接:https://arxiv.org/abs/2610.11942
作者:Jiaming Zhang,Xuan Wang,Fuyao Zhang,Yang Cao,Lingjuan Lyu,Wei Yang Bryan Lim
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:tap on Sign, tap on View, GUI world models, View order, GUI world
备注:
点击查看摘要
Abstract:A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack. For GUI agents, safety is therefore a property of the transition rather than of the screen, and a monitor that inspects only screens can be defeated by reusing a legitimate one. Judging a transition requires an expectation of what should have followed the action. Existing GUI world models provide one, but they output it as text, code, or images, so checking it against the observed screen requires a second model to judge the two. We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions. Verification reduces to a vector comparison, and the same signal reveals whether a mismatch is harmful. We evaluate on RSWT-BENCH, a diagnostic where each credential screen appears under both a legitimate and a hijacked transition, so detectors that see only the screen are at chance by construction. The training-free score reaches 0.987 AUC at 17 ms per decision, on par with the strongest closed-source VLMs and about ten AUC points above generative GUI world models at over three orders of magnitude lower latency. The residual direction reaches 0.953 AUC at separating harmful from benign violations, where prompted VLMs are near chance. Further analyses show that the prediction is a usable future state rather than an anomaly score. World models have mostly served as simulators or planners; our results point to a third role, verification, for which predicting in representation space is the natural design.
65. 【2610.11938】Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders
链接:https://arxiv.org/abs/2610.11938
作者:Tao Yang,Jianying Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:vision-language models optimize, model similarity score, attacked model similarity, Adversarial attacks, optimize an image
备注: 15 pages
点击查看摘要
Abstract:Adversarial attacks on vision-language models optimize an image toward a text target, then cite the attacked model's similarity score as evidence of success. We ask whether that score - victim-space target alignment (VTS) - predicts recovery of the target by an independent model. We first build a measurement instrument: supervised judges outside the attacked geometry, real-target blend controls, shuffled-target negatives, and a reference level derived from a 50% target-image blend. Two preregistered studies then compare six contrastive encoders under a matched attack at three perturbation budgets. Robustly trained encoders (FARE, TeCoA, PMG, TRADES) transfer substantially more independent evidence than vanilla CLIP or SigLIP; all eight contrasts reject at the bootstrap floor. However, no cell reaches the blend-derived reference level. The three best cells fall within its replication band, leaving practical recovery undecided. Within robust encoders, per-sample alignment gain correlates with evidence gain ($\rho = 0.24-0.51$); within vanilla CLIP the correlation is consistent with zero. Across encoders we find no monotone alignment-evidence relation. VTS is therefore informative only within a fixed robust encoder, and we provide a reporting protocol in its place.
66. 【2610.11924】Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows
链接:https://arxiv.org/abs/2610.11924
作者:Jierui Lei,Wenjian Zhang,Qingyi Yang,Yuyang Hong,Fangzheng Chen,Zhengbo Zhang,Haina Tang,Shiming Xiang
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Media-bridged time series, time series forecasting, time series, Media-bridged time, Existing Time Series
备注:
点击查看摘要
Abstract:Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at \href{this https URL}{this https URL}.
67. 【2610.11907】Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations
链接:https://arxiv.org/abs/2610.11907
作者:Shuran Ma,JiaLe Li,Yuxin Dong,Shan Zheng,Qingyun Jiang,Xiang Chen,Qi Zhu,Deyi Ji,Yifan Yang,Jianfeng Pan,Yu Tian,Xue Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Large Vision-Language Models, challenge in Large, Large Vision-Language, remains a significant, significant challenge
备注:
点击查看摘要
Abstract:Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations? In this work, we investigate and quantify how LVLMs coordinate multiple context sources during decoding and examine how this intrinsic behavior can guide hallucination mitigation. We find that LVLMs exhibit an intrinsic vision-attending tendency that can guide adaptive visual steering, while textual contexts can also contribute to hallucination mitigation. Motivated by these findings, we propose AIMS (Adaptive Information Multi-source Steering), a lightweight training-free framework that adaptively coordinates visual, prefilled textual, and generated contexts during decoding. Specifically, AIMS constructs compact prototypes for the three context domains and estimates their affinities with the current query to determine head-wise steering weights. The resulting multi-source steering direction is applied to the query representation, enabling adaptive context integration without additional model training or auxiliary forward passes. Extensive experiments across multiple LVLMs and decoding strategies demonstrate that AIMS effectively mitigates object hallucination while maintaining competitive general-purpose multimodal capabilities.
68. 【2610.11876】Relative Patch Response Learning for Generalizable AI-Generated Image Detection
链接:https://arxiv.org/abs/2610.11876
作者:Tianyu Wang,Ouxiang Li,Yanbin Hao,Zhenhua Tang,Shuo Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:synthesize highly realistic, Generative models, highly realistic images, simultaneously increasing, visual forgery
备注: 19 pages, 6 figures, 13 tables
点击查看摘要
Abstract:Generative models can now synthesize highly realistic images, simultaneously increasing the risks of misinformation and visual forgery. Therefore, detecting AI-generated images becomes more essential, and a reliable detector must generalize to unseen generators and stay robust to unseen perturbations in the wild. Existing detectors are typically trained on either independently collected real and generated images or aligned real-generated pairs designed to mitigate content bias. Building on aligned pairs, recent methods form a mixed view by replacing some patches of the real image with their generated counterparts. However, we find that self-attention lets real and generated patches interact, so the feature of each patch no longer reflects its own source alone. This contextual shift makes a per-patch source label an imprecise target. To this end, we propose Relative Patch Response Learning (PRL). Instead of labeling each patch, PRL compares the same patch across two mixed views of an aligned pair and learns from its patch response, the change of its score between the views. (i) To give precise supervision under the contextual shift, a relative response objective measures the responses of source-changed patches against those of source-unchanged patches, which respond to the shift alone. (ii) To provide a reliable reference for the shift, a reference coherence objective keeps each group of source-unchanged patches moving as a whole. (iii) Since the two views contain different amounts of generated content, an area ranking objective asks the view with the larger generated area to have a higher mean patch score. Extensive experiments demonstrate the superior performance of PRL, which surpasses the best prior methods by 4.3% and 5.9% in average balanced accuracy across eight standard and three in-the-wild benchmarks, respectively.
69. 【2610.11857】Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement
链接:https://arxiv.org/abs/2610.11857
作者:Jingyi Pan,Dan Xu,Qiong Luo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:textural consistency, scene inpainting aims, missing or occluded, ensuring geometric, geometric and textural
备注: Accepted to NeurIPS 2026 (poster). Project page: [this https URL](https://rorisis.github.io/FreeInpaint/)
点击查看摘要
Abstract:3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency. Existing approaches, however, typically require accurately calibrated camera poses, which restricts their applicability in casual, in-the-wild scenarios and introduces additional preprocessing overhead. To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions. At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry. Our method addresses two key challenges in adapting feed-forward 3D foundation models to masked inputs. First, masked regions can corrupt cross-view correspondence reasoning, degrading pose estimation and geometry recovery. To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers. Second, under severe occlusions, a single forward pass often lacks sufficient appearance evidence for high-fidelity completion. Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor. Extensive experiments across diverse datasets demonstrate that FreeInpaint achieves superior inpainting quality, eliminating the reliance on pre-computed camera poses while keeping a fast inference speed. The project page is this https URL.
70. 【2610.11850】VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval
链接:https://arxiv.org/abs/2610.11850
作者:Shahaf Wagner,Gabriele Serussi,Dan Ben Ami,Tomer Galanti,Chaim Baskin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:requires distinguishing similar, distinguishing similar scenes, events occur, requires distinguishing, distinguishing similar
备注:
点击查看摘要
Abstract:Finding the right video often requires distinguishing similar scenes in which different events occur. Joint matching improves retrieval, but processing rich video representations for each query is costly. VEDJE compresses features within sampled frames while keeping their representations separate in a reusable cache. Feature-change prediction supplies an auxiliary training signal that improves retrieval from the compressed cache without adding work at query time. On MSR-VTT, MSVD, DiDeMo, and ActivityNet, VEDJE improves R@1 over matched first-stage retrievers in both retrieval directions. On MSR-VTT, it reaches 59.8 text-to-video R@1 with a fine-tuned VideoCLIP-XL first stage. In the VideoPrism configuration, shrinking the per-video cache fourfold to 12 KiB preserves text-to-video recall within 0.2 points. These results show that accurate video search can operate on compact evidence, encoded once and reused as new queries arrive.
71. 【2610.11846】Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
链接:https://arxiv.org/abs/2610.11846
作者:Anirudh Praveen,Koteswar Rao Jerripothula,Pratik Joshi,Aveen Dayal,Neela Sawant
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Image and Video Processing (eess.IV)
关键词:Audio-Visual Event Localization, Open-Vocabulary Audio-Visual Event, Event Localization, Audio-Visual Event, event class
备注: Accepted to British Machine Vision Conference (BMVC) 2026
点击查看摘要
Abstract:Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training. The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text. Existing methods then collapse this pair into a single scalar score with a fixed rule (geometric mean, weighted average) before taking the argmax. Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN). Each modality's standard representation becomes the real part of our pipeline, and a paired companion stream supplies the imaginary part. We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams. This results in two complex similarities, which are then fused. While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained. The four-stream complex architecture sets a new state of the art on both OV-AVEL benchmarks. On the open (unseen-class) split of OV-AVEBench we reach 66.5/59.1/54.1% Acc/Seg-F1/Event-F1 (+1.6/+4.1/+6.6 over the previously reported fine-tuned baseline), with consistent gains for seen classes as well. We also modify AVE dataset for this task and observe that our architecture reaches 60.7/51.9/50.4% Acc/Seg-F1/Event-F1, achieving state-of-the-art OV-AVEL results on it as well. We also propose a two-stream alternative, which also sees great improvements over the baseline.
72. 【2610.11826】From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment
链接:https://arxiv.org/abs/2610.11826
作者:Chen Zhao,Xingping Dong,Jiachun Shi,Liang Peng,Chong Wang,Zhen Lei,Ran He,Bo Du
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:generate reliable content, large vision-language models, reliable content, remains a major, major obstacle
备注:
点击查看摘要
Abstract:Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken the model's multimodal capabilities. In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment. Specifically, ResOT projects dominant hallucinated directions away from the faithful subspace, forming a low-dimensional residual subspace for intervention. Within this subspace, ResOT uses Gaussian optimal transport (OT) to align the hallucinated distribution with the faithful one. The resulting map defines repair targets with minimal changes to the original representations. At inference, ResOT adaptively controls how far each token state moves toward its OT target. Experiments on three representative LVLMs show that ResOT substantially reduces object hallucination while improving image caption quality and multimodal performance across multiple benchmarks. Code will be released.
73. 【2610.11818】From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
链接:https://arxiv.org/abs/2610.11818
作者:Uddipan Basu Bir,Vincent Christlein,Andreas Maier,Mathias Zinnen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:set strong benchmarks, data autonomy concerns, APIs limits adoption, institutional archives due, closed-source Vision-Language Models
备注: 17 pages. Published in Document Analysis and Recognition - ICDAR 2026, LNCS vol. 16974, Springer. Code: [this https URL](https://github.com/uddipan77/Analysis-of-Lightweight-Vision-Language-Models-for-Document-OCR-and-Structured-Output-Generation)
点击查看摘要
Abstract:While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image, models must extract text and generate schema-compliant JSON, enabling automatic validation and downstream use. We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS*), and mean Average Precision F1 (mAP-F1). Against a fine-tuning baseline, we further test the independent impact of (i) hyperparameter optimization, (ii) classical image preprocessing (illumination flattening, denoising, and CLAHE), and (iii) multi-stage training. Finally, we analyze the trade-off between dataset-specific fine-tuning and a single multi-dataset checkpoint, where joint training enables one model to operate across collections but can shift performance between datasets. Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.
74. 【2610.11815】Fast Pose Tracking of Rigid Objects with Compact Pose Graph Optimization
链接:https://arxiv.org/abs/2610.11815
作者:Xiaojie Zhang,Tom Fischer,Viktor Larsson,Eddy Ilg
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:long horizons, horizons currently requires, requires either expensive, expensive onboarding, reconstruction maintained
备注:
点击查看摘要
Abstract:Tracking a novel object's 6D pose over long horizons currently requires either expensive onboarding or a reconstruction maintained throughout the sequence. This makes current trackers impractical for robotic manipulation and augmented reality, which need trackers that are ready to use and run in real time. We show that a lightweight tracking module can be applied on top of a wide range of correspondence estimation methods to keep drifts bounded while maintaining fast runtime. Our key idea is to avoid point-based optimization in the pose graph and operate only on relative pose constraints, which we weight by a derived uncertainty from the geometric alignment. This makes optimization independent of the number of correspondences while avoiding the direct inclusion of noisy point measurements, leading to fast and robust long-term tracking. Across four real-world benchmarks, our approach achieves tracking accuracy comparable to reconstruction-based trackers with a fraction of the optimization cost. Overall, these results suggest that a compact and reliable pose graph optimization can provide long-horizon consistency at substantially lower computational cost.
75. 【2610.11810】Seek-and-View Reasoning for Multi-View Spatial Understanding
链接:https://arxiv.org/abs/2610.11810
作者:Qixiang Chen,Cheng Zhang,Fucai Ke,Chi-Wing Fu,Jianfei Cai,Jingwen Ye
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Existing approaches, sparse input views, sparse input, spatial reasoning operate, spatial
备注: Project page: [this https URL](https://seekandview2026.github.io) ; Code: [this https URL](https://github.com/q1xiangchen/Vantage)
点击查看摘要
Abstract:Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, we formulate a novel Seek-and-View reasoning approach to find implicit cross-view spatial evidence by locating a question-relevant view to support the spatial reasoning. To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning. Comprehensive experiments on six VLMs demonstrate consistent improvements on five benchmarks without fine-tuning. Overall, by revealing spatial evidence through view-grounded reasoning, Vantage can largely reduce reliance on language-based cross-view alignment and improve multi-view spatial understanding. Our code is available at this https URL.
76. 【2610.11794】Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
链接:https://arxiv.org/abs/2610.11794
作者:Haoyu Zhao,Zhengxu Yu,Zhiyuan He,Meng Fang,Rasul Tutunov,Haitham Bou-Ammar,Weilin Luo,Jun Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:unfamiliar environments requires, act in unfamiliar, multiple world models, world models, environments requires agents
备注:
点击查看摘要
Abstract:Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
77. 【2610.11791】Phase-aware video generation for physics-grounded dynamics and interactions
链接:https://arxiv.org/abs/2610.11791
作者:Jingfeng Ou,Kun Wang,Rui Zhao,Jingwei Guan,Limin Wang,Chao Dong,Xingyu Zeng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generating physically plausible, Generating physically, remain coupled, Generating, exhibit distinct dynamics
备注: 31 pages, 5 figures, 9 tables, including appendix
点击查看摘要
Abstract:Generating physically plausible videos for solid-gas dynamics is challenging as different phases exhibit distinct dynamics yet remain coupled through physical interactions. We present PAVG, a Phase-Aware Video Generator for solid-gas dynamics and interactions. It employs a dual-branch architecture to explicitly model the distinct dynamics of solids and gases, while spatiotemporal cross-attention captures their physical interactions. This design enables PAVG to preserve phasespecific motion characteristics while producing physically consistent responses across phases. To facilitate this task, we further construct a simulation corpus comprising over 700K physical trajectories across diverse solid, gas, and solid-gas interaction scenarios. Extensive evaluations demonstrate that our PAVG produces videos with improved motion adherence, physical plausibility, and visual quality compared with existing approaches.
78. 【2610.11781】Skill-V: Verifiable Self-Evolving Skill Library for Interactive Agents
链接:https://arxiv.org/abs/2610.11781
作者:Jie Ma,Zhipeng Qian,Yufei Ma,Zihan Liang,Jiayi Ji,Qingpeng Cai,Ben Chen,Peng Jiang,Xiaoshuai Sun
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:libraries primarily improve, Interactive agents, skill libraries primarily, skill, agents can turn
备注:
点击查看摘要
Abstract:Interactive agents can turn experience into reusable skills, yet existing self-evolving skill libraries primarily improve by accumulating new knowledge. Failures may lead to new skills, while previously stored skills are less often revisited as new evidence arrives. However, growth alone does not ensure reliability, as a retrieved skill may be inapplicable under the current task conditions, and an existing skill may encode a mis-specified operational boundary. Reliable skill evolution therefore requires not only adding knowledge, but also testing and revising what is already stored. We introduce Skill-V, a verifiable self-evolving skill library. To make stored knowledge testable, we propose representing skills as versioned, falsifiable contracts that link semantic intent to observable behavioral criteria. We use environment outcomes to drive library evolution. Specifically, task failures motivate skill addition, while disagreements between contract evaluations and task outcomes guide revisions to existing skill boundaries. To validate these revisions, we require them to preserve protected semantic constraints and satisfy non-regression criteria for rubric-outcome metrics on historical replay evidence. Finally, we employ an applicability-aware filter to exclude candidates judged confidently inapplicable to the current task. Across ALFWorld and WebShop, Skill-V achieves success rates of 95.3% and 85.9%, respectively, while maintaining a more compact skill library than growth-oriented baselines. Applicability-aware filtering reduces incorrect skill invocations, and outcome-grounded revisions correct mis-specified skill boundaries without degrading performance on previously observed evidence. These results show that reliable skill evolution requires more than accumulating experience: the library must learn which knowledge to retain, when to revise it, and when it should be applied.
79. 【2610.11770】From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation
链接:https://arxiv.org/abs/2610.11770
作者:Sicong Yang,Ruihuan Yang,Jian Lu,Jianfei Yuan,Xiaodong Cun,Xiuli Bi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:video creation, AI-Native video creation, video creation workflow, video, iterative video creation
备注: 18 pages, 17 figures
点击查看摘要
Abstract:AI-Native video creation is shifting from isolated video clips toward iterative video creation workflows. However, existing datasets remain largely video clips, representing video generation and editing as separate tasks rather than connected stages of a video creation workflow. In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory. Specifically, we first identify video creation trajectories and decompose them into three subsets according to their structural roles: text-to-video generation records as roots, single-turn video editing records as editing edges, and multi-turn video editing records as complete trajectories. Then, we use a VLM to assign semantic annotations for generation roots and editing-operation annotations for editing edges. A strict construction pipeline further reconstructs source-to-edit lineage, editing order, and intermediate video states while ensuring data quality. Finally, we perform lightweight adaptation on LTX-2 models to assess the supervision value of Sora100K. The results show improvements in visual quality, multi-shot generation, and cross-shot consistency, while successive-turn evaluation reveals that following multi-turn editing instructions remains challenging. Sora100K establishes a new data foundation for AI-Native video creation beyond isolated video clips and toward structured video creation trajectory. The dataset and supplementary materials are publicly available at this https URL.
80. 【2610.11756】Memory Forcing: Attendable Mid-Horizon History for Streaming Video Generation
链接:https://arxiv.org/abs/2610.11756
作者:Jiaming Zhang,Xinyu Wang,Huafeng Shi,Gangshan Wu,Limin Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:diffusion enables causal, Autoregressive video diffusion, existing few-step systems, video diffusion enables, enables causal video
备注: 10 pages, 6 figures
点击查看摘要
Abstract:Autoregressive video diffusion enables causal video streaming without a bidirectional pass over the full clip, but existing few-step systems usually retain only the opening and most recent frames in a fixed-size KV cache. Once an event leaves this window, later frames can no longer attend to it, a failure we term mid-horizon forgetting. We present Memory Forcing, a few-step streaming method that preserves this missing history without increasing the cache size. Its Archive \ Working Banks partition the cache into sink, archive, and working regions, retaining diverse intermediate events alongside recent motion under fixed memory. Because absolute temporal indices drift outside the training range, Bank-aware RoPE reassigns indices at attention time so each bank remains distinguishable. At 1.3B, Memory Forcing leads on longer clips, shows the smallest drop from 5s to 60s among methods reporting all four lengths, and preserves subjects and scenes through leave-and-return. The same design scales to Wan2.2 5B, producing more physically plausible, realistic, and dynamic videos and, to our knowledge, the first public 5B model on this forcing line.
81. 【2610.11751】Dino Forcing Flow Models: Do not denoise what you can predict
链接:https://arxiv.org/abs/2610.11751
作者:Arijit Ghosh,Lucas Degeorge,Paul Couairon,Alexei A Efros,Vicky Kalogeiton,David Picard
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Co-denoising pretrained representations, requires carefully designed, carefully designed schedules, flow matching models, Co-denoising pretrained
备注:
点击查看摘要
Abstract:Co-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules. We propose a simpler alternative: predict the pretrained representation directly, then condition the model on its own prediction. This removes the need for a second ODE and any representation-specific denoising schedules, while retaining the benefits of representation guidance. Our approach converges substantially faster and achieves better generation quality as measured by FID score. On ImageNet, it outperforms the state of the art in latent space at 2x fewer epochs than prior methods; in pixel space, it improves FID over comparable prior methods by more than 20%. These results support a simple principle: do not denoise what you can predict. Our code is openly available at this https URL.
82. 【2610.11746】Streaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step Attention
链接:https://arxiv.org/abs/2610.11746
作者:Harris Partaourides,Sotirios Chatzis
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:strict latency constraints, requires high spatio-temporal, high spatio-temporal fidelity, iterative sampling cost, super-resolution requires high
备注:
点击查看摘要
Abstract:Real-time video super-resolution requires high spatio-temporal fidelity under strict latency constraints, challenging diffusion models due to their iterative sampling cost and limited temporal coordination. We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams. Our Cross-Step Attention mechanism reuses intermediate denoising features across adjacent frames and diffusion steps, enabling temporal information exchange without explicit temporal modeling. We further introduce Trajectory-Coupled Diffusion Scheduling, which aligns adjacent diffusion states and provides cleaner intermediate representations for cross-step conditioning, improving temporal coherence. These components are integrated into a streaming inference pipeline that incrementally propagates latent states across frames, reducing the effective computational complexity from $O(N \cdot S)$ to $O(N + S)$ for $N$ frames and $S$ diffusion steps. Experiments on REDS4 and YouHQ40-Test demonstrate improved perceptual quality and temporal realism while maintaining frame-wise stability. Our method achieves over 40 FPS at $512 \times 512$ resolution after cold start, enabling real-time VSR without explicit temporal modeling.
83. 【2610.11736】owards Unified Evaluation of Prompt Enhancers for Video Generation
链接:https://arxiv.org/abs/2610.11736
作者:Yawen Shao,Yubo Zhu,Ziyun Dai,Zixun Fang,Kai Zhu,Zeyinzi Jiang,Yufeng Ai,Siyang Sun,Haolan Xue,Yu Shang,Yuxiang Bao,Zoubin Bi,Jingming Luo,Jie Xiao,Chaojie Mao,Zhehan Kan,Hongchen Luo,Yu Liu,Sheng Zhong,Wei Tong,Xueyang Fu,Yang Cao,Wei Zhai,Zheng-Jun Zha
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Modern video generators, realize increasingly complex, concise user instructions, complex visual narratives, structured cinematic plans
备注: Project page: [this https URL](https://github.com/yawen-shao/PEBench)
点击查看摘要
Abstract:Modern video generators can realize increasingly complex visual narratives, positioning the prompt enhancer (PE) as a critical bridge from concise user instructions and multimodal references to structured cinematic plans. However, existing PE evaluation relies on rendered videos, imposing substantial computational and human costs, slowing PE training and iteration, and conflating PE quality with downstream generator behavior. To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement. It comprises 1,100 expert-verified cases and 1,005 visual assets, spanning 35 fine-grained tasks with diverse temporal, cinematic, audiovisual, and multi-reference requirements. In addition, we develop PEBench evaluation, an evidence-grounded framework that combines modality-aware fact extraction with rubric-based assessment across 24 criteria. Our systematic evaluation of representative open- and closed-source PE methods reveals an emerging shift from fine-grained descriptive expansion toward intent-preserving cinematic planning, while the caption-reconstruction and forward-refinement methods show complementary strengths in cinematic coverage and semantic fidelity or internal coherence, respectively. Human validation shows that PEBench scores align closely with expert judgments of enhanced prompts and downstream videos from Wan3.0 and MiniMax-H3, indicating that prompt-level evaluation reliably reflects downstream utility.
84. 【2610.11735】Onboard Marine Anomaly Detection on $Φ$sat-2: From Simulation-Based Development to In-Orbit Demonstration
链接:https://arxiv.org/abs/2610.11735
作者:Clotilde Szywala,Thomas Goudemant,Marjorie Bellizzi,Benjamin Francesconi,Adrien Girard
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Earth Observation systems, Onboard Artificial Intelligence, Artificial Intelligence, Earth Observation, European Space Agency
备注:
点击查看摘要
Abstract:Onboard Artificial Intelligence can improve responsiveness and bandwidth efficiency of Earth Observation systems by processing data directly on the satellite. This paper presents the experience gained from the development, onboard integration, and post-launch adaptation of a lightweight marine anomaly detection pipeline deployed on the European Space Agency's $\Phi$sat-2 mission. The application combines sea segmentation, self-supervised feature encoding of marine regions, generic anomaly detection based on deviations from a normal sea state, and optional characterization of selected anomaly types. Before launch, the pipeline was trained and validated on simulated $\Phi$sat-2 imagery to assess algorithmic performance and compatibility with resource-constrained onboard hardware. After integration and functional validation in the mission environment, early experiments on real $\Phi$sat-2 acquisitions revealed a significant mismatch between simulated and in-orbit data. The pipeline was therefore retrained on real Level-1 imagery using an improved annotation strategy to better handle ambiguous marine regions, substantially enhancing performance. Beyond demonstrating the onboard feasibility of the application, the $\Phi$sat-2 experience highlights the importance of robust annotation strategies and sensor-aware design, and shows that simulation-based development is valuable for pre-flight risk reduction, while reliable scientific validation requires representative in-orbit data and should be clearly distinguished from functional validation.
85. 【2610.11723】MultiWorldBench: Do Independently Controlled Views Describe One Shared World?
链接:https://arxiv.org/abs/2610.11723
作者:Zhangbo Xu,Ruoxi Zhang,Rui Hu,Yisong Wang
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:independently controlled views, controlled views remain, views remain consistent, models must ensure, ensure that independently
备注: 31 pages, 17 figures, 3 tables
点击查看摘要
Abstract:Multiplayer world models must ensure that independently controlled views remain consistent with one shared and persistent world. We introduce MultiWorldBench, a diagnostic Minecraft benchmark containing 495 case configurations across seven task suites and ten capabilities, including independent control, cross-view motion, shared-state synchronization, persistence, structural reasoning, concurrent interaction, and delayed revisit. We evaluate Solaris, Gamma-World, and MineWorld, using Engine GT as a reference. Gamma-World achieves the highest ten-capability average among the generated systems at 21.39, followed by Solaris at 20.88 and MineWorld at 1.89, while Engine GT reaches 91.69. Gamma-World performs better on several control, shared-state, and revisit capabilities, whereas Solaris leads in cross-view motion and race-condition consistency. Nevertheless, all generated systems score at most 8.00 on state persistence and 1.33 on structural consistency, and none succeeds in spatial reasoning or building-identity preservation. Human preferences produce the same overall ranking and show strong alignment with the automatic evaluation, with a mean dimension-level Spearman correlation of 0.96. These results show that plausible individual views do not yet constitute a coherent multiplayer world.
86. 【2610.11685】HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution
链接:https://arxiv.org/abs/2610.11685
作者:Ziying Li,Shengchu Zhao,Huiang He,Yiyang Chen,Jianwen Huang,Bailin Li,Changhao Li,Jianhui Li,Jie Li,Ruiyang Liu,Yibo Luo,Tengjiao Sun,Pei Tang,Shiwen Wang,Jiaqi Wu,Kang Wu,Kaiqiao Yang,Zherui Yang,Hu Zhang,Xuezhi Zhao,Xinhe Zheng,Yukun Li,Heliang Zheng,Rongfei Jia
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:input image, including the specific, increasingly capable, capable of producing, closely resemble
备注: Hi3D 3.0 (Twinkle3D) Technical Report
点击查看摘要
Abstract:Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.
87. 【2610.11674】VESSI - VLM-Enhanced Support for Surveillance and Investigations
链接:https://arxiv.org/abs/2610.11674
作者:Saverio Cavasin,Pietro Tedeschi,Mattia Tamiazzo,Alessandro Brighente,Simone Milani,Mauro Conti
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Law Enforcement agencies, Law Enforcement, Enforcement agencies, infrastructures and Law, Automated video surveillance
备注: 12-page main manuscript, 3 main figures; supplementary material included
点击查看摘要
Abstract:Automated video surveillance analysis has become a critical component of intelligence infrastructures and Law Enforcement agencies. Traditional systems lack the semantic module for comprehensive situational awareness and forensic tasks, limiting their ability to interpret events meaningfully or support post-incident investigations. This slows operational insight and increases the burden on human analysts. Recent advances in Vision-Language Models (VLMs) offer promising pathways to bridge this gap. To address this, we propose VLM-Enhanced Support for Surveillance and Investigations (VESSI), a VLM-based framework designed to enhance automated video surveillance analysis through prompt-driven interrogation of video sequences where salient visual features are converted into textual descriptions. We test our framework with four state-of-the-art models. Since most datasets for this task are unlabeled, we also propose the Composite Model Utility Score (CMUS) to assess VLM performance. Experimental results show that our solution substantially improves the analysis capabilities of human operators and enhances the flexibility of automated surveillance systems. In our evaluation, the most reliable model flagged potentially relevant activity in more than 66% of the videos while reducing review time by more than 85%, offering a practical balance between selectivity and efficiency. The model ordering produced by the reference-free CMUS evaluation was reproduced by the normal-video CMUS evaluation and matched the false-positive-rate ordering obtained from 5,909 manually referenced frames. This agreement supports the operational use of the score within the evaluated setting.
88. 【2610.11669】Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
链接:https://arxiv.org/abs/2610.11669
作者:Nick Milkin,Lanmiao Liu,Esam Ghaleb,Asli Ozyurek,Zerrin Yumak
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Holistic and semantics-aware, appropriateness remains difficult, co-speech gesture generation, semantic appropriateness remains, reflect human perception
备注:
点击查看摘要
Abstract:Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify. We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation. We first curate a list of 13 objective metrics covering different aspects, including distributional similarity, geometric fidelity, kinematic quality, cross-modal synchrony, and semantic appropriateness. For the semantic-appropriateness category, we propose a new metric, Semantic Gesture Preservation (SGP), which measures how far semantic gestures in the ground truth are preserved in the generated gestures. For this, we augment the BEAT2 dataset's annotations using a multi-modal LLM. We then conduct a perceptual study where 101 participants score generated gestures among five dimensions, including human-likeness, motion diversity, absence of animation errors, speech timing and content match. We systematically analyze objective metric--subjective score correlations. Unlike Semantic Score (SC), which shows no significant association with the evaluated perceptual dimensions, SGP is selectively aligned with speech-aware human judgments. We construct five target-specific composite metrics aligned with the subjective dimensions. These composites improve perceptual alignment across all five dimensions, with the largest gains for absence of animation errors and content match, indicating that complementary objective signals can better approximate human judgments than individual metrics alone. Overall, our results show that objective metrics require validation against subjective evaluations.
89. 【2610.11666】Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
链接:https://arxiv.org/abs/2610.11666
作者:Jianfei Zhao,Yifan Wang,Feng Zhang,Xin Sun,Chong Feng,Zhixing Tan,Yang Luo,Boyuan Pan,Xu Kai,Yao Hu
类目:Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
关键词:Universal multimodal retrieval, Universal multimodal, independently indexed items, ranks independently indexed, retrieval typically encodes
备注: Under Review
点击查看摘要
Abstract:Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.
90. 【2610.11656】DisFace3DNet: Explainable Facial Attractiveness Prediction via 3D Component Disentanglement
链接:https://arxiv.org/abs/2610.11656
作者:Fenggui Rao,Yan Luximon,Jie Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:viewing conditions implicit, leaving the roles, conditions implicit, viewing conditions, Facial attractiveness prediction
备注: Includes supplemental materials
点击查看摘要
Abstract:Facial attractiveness prediction usually assigns one overall rating, leaving the roles of shape, appearance, and viewing conditions implicit. We propose DisFace3DNet, which uses 3D component disentanglement to learn seven component reference scores from overall ratings with auxiliary weak semantic supervision, without human-labeled component targets. Designated 3D representations and image cues feed jointly learned routes for identity, skin, hair, light, background, expression, and pose. A constrained fit then combines five static and two signed dynamic scores into the overall rating, exposing each component's numerical contribution and supporting component-specific comparisons across images. On SCUT-FBP5500, DisFace3DNet achieves a Pearson correlation of $0.8904\pm0.0063$ (mean $\pm$ standard deviation across five folds) with average human ratings; its component terms reconstruct every held-out prediction to numerical precision. Skin, hair, and facial shape account for the largest component-wise prediction variation. Human evaluation supports the score directions for facial shape, skin, and hair; expression agreement is weaker. DisFace3DNet thus connects overall prediction to quantitative analysis of the facial and contextual cues entering each estimate.
91. 【2610.11653】abula Rasa: Monte Carlo estimation of unit-variance noise with controlled spatio-temporal correlation
链接:https://arxiv.org/abs/2610.11653
作者:Tobias Ritschel,Yang Zhou,Nick Milef,Mikhail Dereviannykh,Chen Liu,Christophe Hery,Carl Marshall
类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
关键词:generate time-varying Gaussian, time-varying Gaussian noise, time-varying Gaussian, controlled temporal correlation, Gaussian noise
备注: SIGGRAPH Asia 2026 Conference Papers. Code: [this https URL](https://github.com/facebookresearch/Tabula-Rasa)
点击查看摘要
Abstract:We suggest a method to generate time-varying Gaussian noise with controlled variance and controlled temporal correlation. This noise is used in several downstream tasks for temporal control and temporal coherence. The core technical idea is to phrase this problem as joint Monte-Carlo estimation of both a classic pixel reconstruction and estimation of variance using the concept of "sketching" from the database literature. We demonstrate that our method allows temporal control for downstream tasks with simpler and faster code than previous methods.
92. 【2610.11641】Revisiting Handcrafted Minutiae Detection: A Simple and Effective Open Source Baseline for Modern Fingerprint Workflows
链接:https://arxiv.org/abs/2610.11641
作者:Raffaele Cappelli
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:algorithms remain fundamental, forensic practice due, Handcrafted minutiae detection, GPU hardware, detection algorithms remain
备注:
点击查看摘要
Abstract:Handcrafted minutiae detection algorithms remain fundamental to biometric science and forensic practice due to their full auditability, adherence to international standards, and operational independence from training datasets or GPU hardware. However, current open-source traditional baselines are severely outdated, relying almost exclusively on legacy C/C++ codebases that lack seamless integration with modern scientific software ecosystems. To bridge this gap, the present work introduces SBMEX (Skeleton-Based Minutiae EXtraction), a fast and deterministic minutiae detection method integrated into the open source \texttt{pyfing} package. SBMEX achieves high computational throughput by employing a dual Look-Up Table architecture that replaces runtime neighborhood scanning during Crossing Number computation and skeleton tracking. Additionally, it incorporates a continuous quality scoring framework driven by tracking path length, dual ridge-valley skeleton fusion, and spatial density decay. Rigorous evaluation on NIST SD302 datasets demonstrates that SBMEX delivers feature extraction accuracy comparable to or outperforming traditional open-source baselines without fine-tuning, while achieving a drastic reduction in minutiae detection latency relative to classical Crossing Number Python implementations.
93. 【2610.11631】Neural Networks for Temporal Pattern Recognition and Dynamic Arm Gesture Speed Estimation for Robot Control
链接:https://arxiv.org/abs/2610.11631
作者:Milán Zsolt Bagladi,László Gulyás
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Deploying intelligent robotic, Deploying intelligent, requires neural networks, neural networks capable, diverse temporal patterns
备注: 10 pages, 6 figures, 4 tables. Published in Proceedings of the Intelligent Robotics FAIR 2026 (IntRob '26), June 18-19, 2026, Budapest, Hungary, ACM
点击查看摘要
Abstract:Deploying intelligent robotic systems that interact with humans through gestures requires neural networks capable of recognizing diverse temporal patterns. We present a systematic benchmark of ten abstract sequential tasks--five permutation-invariant (set) and five order-dependent (sequence) problems--evaluated across eighteen neural network architectures spanning recurrent, convolutional, attention-based, and set-function families. Beyond the core architecture-task grid, we explore numerous preprocessing and target-variable transformations, yielding more than 250 distinct experimental configurations. All variants are trained and tested under strictly identical conditions (fixed random seeds, shared hyperparameters, shared data splits) to ensure fair and reproducible comparison. Ranking across all ten tasks reveals four consistently top-performing architectures--BiGRU, TCN, Conv1D, and GRUReLU--all compact enough for real-time deployment (under 2,000 parameters in the benchmark setting). Based on this ranking, we apply three architecturally diverse top models (BiGRU, TCN, and GRUReLU) to a practical robotics problem: estimating the execution speed of dynamic arm gestures from skeletal keypoint sequences. Three speed interpretations (peak count, period time, and mean spike spacing) are evaluated on a custom dataset of eight traffic-related gesture classes comprising 256,710 frames recorded via OpenPose. The best configuration achieves a mean absolute error of 0.198 on the peak-count interpretation, corresponding to roughly 5% relative error, while the period-time interpretation reaches approximately 4% relative error, and the mean spike spacing interpretation approximately 8% relative error. These results demonstrate that neural networks can reliably estimate gesture speed from skeletal data, opening a path toward speed-aware gesture-controlled robotic systems.
94. 【2610.11617】AM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction
链接:https://arxiv.org/abs/2610.11617
作者:Yuqi Li,Xiaoqin Feng,Fan Xu,Weilun Feng,Chuanguang Yang,Yingli Tian,Hao Wu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:enables efficient spatiotemporal, distillation enables efficient, efficient spatiotemporal prediction, enables efficient, efficient spatiotemporal
备注: 19 pages
点击查看摘要
Abstract:Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student. However, matching outputs or features independently for each sample leaves cross-sample predictive structure underused. Exploiting this structure requires representations and historical references that reflect the dynamics of each task. We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history. Memory entries encode latent features, forecast changes, or flow residuals, while task-specific selection rules identify relevant historical references. The student either matches the teacher's similarity distribution over shared references or regresses observation-conditioned residual prototypes. These objectives complement supervised prediction and conventional distillation. The teacher, memory, and auxiliary adapters are used only during training, leaving student inference unchanged. We evaluate TAM on video prediction, weather forecasting, and traffic flow prediction across multiple teacher-student configurations. Averaged over four paired runs, adding TAM improves SSIM on all six video datasets and reduces MSE on five relative to the corresponding KD baselines. Mean paired MSE reductions reach 1.86% on KittiCaltech, 1.93% on WeatherBench with a gSTA teacher, and 1.01% on TaxiBJ. These results demonstrate the utility of historical teacher supervision across distinct forecasting tasks without additional student inference cost.
95. 【2610.11612】PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors
链接:https://arxiv.org/abs/2610.11612
作者:Haobo Jiang,Liang Yu,Jianmin Zheng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:RGB-D point cloud, unordered RGB-D scans, consistent coordinate frame, point cloud registration, estimate global rigid
备注: 19 Pages, 6 figures
点击查看摘要
Abstract:This paper addresses multiview RGB-D point cloud registration, aiming to estimate global rigid poses for unordered RGB-D scans and align them in a metrically consistent coordinate frame. The conventional pairwise-then-global paradigm suffers from locally optimized pairwise registration, severe error propagation and high computational burden. In particular, existing methods typically treat RGB data as a mere auxiliary matching cue and overlook the holistic geometric priors (e.g., camera poses and 3D models) encoded across image sequences. This paper introduces PointVGGT, a zero-shot framework built upon a novel \emph{foundation-then-refinement} paradigm that systematically leverages visual geometry foundation models (e.g., VGGT) as the computational backbone for robust, training-free multiview RGB-D registration. In the foundation stage, we directly recover metrically consistent global poses (without any pairwise estimation) by grounding the scale-ambiguous pose predictions of the foundation model against metric depth observations. In the refinement stage, we introduce an efficient voxelized spatial hashing mechanism that exploits the globally coherent 3D reconstruction (induced by the foundation model) as a shared spatial anchor, enabling dense multiview correspondences in near-linear time. On top of this, an IRLS-based robust motion-only bundle adjustment is performed using a conjugate gradient solver to jointly minimize the correspondence and reprojection residuals for multiview pose refinement. Extensive experiments on indoor/object-centric/outdoor datasets verify the outstanding zero-shot registration accuracy and computational efficiency of our proposed method.
96. 【2610.11610】Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
链接:https://arxiv.org/abs/2610.11610
作者:Yuchen Yang,Xin Wang,Lufan Wang,Yinghong Pan,Yujuan Feng,Yuqing Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:archived key frames, key frames capture, multiple images requires, images requires aggregating, requires aggregating clinical
备注: Accepted to BMVC 2026. Code: [this https URL](https://github.com/NiHaoWoJiaoYYC/CAMEO)
点击查看摘要
Abstract:Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiable from the images available to a model. This gap creates a clinical behavior alignment problem. A model must preserve visible findings, avoid diagnostic reversals and unsupported completion, and not collapse into conservative templates. We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation. Stage I learns ultrasound visual-language primitives; Stage II performs Cross-View Evidence Grounding by distilling trusted visible report points into multi-image QA and report-style supervision; and Stage III performs Clinically Aware Preference Alignment using clinical-error-oriented preference pairs. From USReport, we construct USReport-Distilled with 17,670 evidence-grounded paired-image training instances and USReport-Pref with 21,869 preference pairs; we additionally use 25,631 PubMedVision-US ultrasound instruction samples for domain adaptation and multi-image instruction tuning. On the primary USReport-Distilled benchmark, CAMEO improves over EchoVLM from 0.25 to 0.40 BLEU-1, 0.28 to 0.45 ROUGE-1, and 0.27 to 0.43 METEOR, while raising ClinicalScore from 55.02 to 74.20. These results underscore the value of evidence-grounded supervision, clinically aware alignment, and clinically structured evaluation for reliable ultrasound report generation.
97. 【2610.11608】S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization
链接:https://arxiv.org/abs/2610.11608
作者:Ziqian Mo,Hill Zhang,Haosheng Tan,Ling Li,Jiaheng Wei
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:estimate geographic locations, matching images captured, aims to estimate, satellite views, estimate geographic
备注:
点击查看摘要
Abstract:Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views. Existing methods mainly rely on visual representations, but often fail to jointly model fine-grained structural correspondences and semantic priors, making them prone to confusion between visually similar but semantically different regions, and thus limiting robustness under large viewpoint variations. To address these challenges, we propose \textbf{S$^3$Geo}, a structure-semantic synergistic learning framework for cross-view matching. Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns. We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment. Furthermore, we incorporate a Semantic Knowledge Distillation (SKD) strategy from a frozen CLIP teacher to transfer semantic priors and relational structures, thereby improving discrimination on hard negatives. By operating synergistically, the semantic priors provide robust contextual filtering, which guides the structural module to establish precise spatial alignments. Experiments on the University-1652 and SUES-200 datasets demonstrate that \textbf{S$^3$Geo} consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.
98. 【2610.11579】SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection
链接:https://arxiv.org/abs/2610.11579
作者:Ricardo Pizarro,Roberto Valle,José M. Buenaposada,Luis M. Bergasa,Luis Baumela
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:billion-parameter Vision Transformers, adapt billion-parameter Vision, recent methods freeze, Vision Transformers, lightweight convolutional modules
备注:
点击查看摘要
Abstract:To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules. While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed. Token selection can reduce attention cost by pruning redundant tokens, but it breaks the spatial grid structure required by convolutional adapters. This forces an expensive dense reconstruction, nullifying much of the potential speedup. We address this by introducing native sparse 2D convolutions, a primitive that allows these adapters, for the first time, to operate directly and efficiently on dynamically pruned token sets. We integrate this primitive into SV-TAD, an adapter framework for temporal action detection, reducing VideoMAEv2-L computation by up to 64% and achieving 2.2x faster inference, while maintaining state-of-the-art accuracy on THUMOS-14 and ActivityNet-1.3. When scaled to InternVideoNext-L, our approach surpasses the previous state of the art at roughly half its computational cost. Moreover, the sparse formulation naturally supports auxiliary task tokens, which improves fine-grained assembly detection on ATTACH.
99. 【2610.11577】CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving
链接:https://arxiv.org/abs/2610.11577
作者:Soham Pahari,Sudip Das,Arindam Das,Ujjwal Bhattacharya
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:blind spots, due to occlusions, surrounding environments, complex nature, nature of surrounding
备注:
点击查看摘要
Abstract:Autonomous vehicles often suffer from limited perception due to occlusions, blind spots, limited sensor range, and the complex nature of surrounding environments. Multi-agent collaborative perception (CP) addresses these challenges by allowing vehicles to share sensory information and reconstruct the scene cooperatively. However, camera-only perception remains fundamentally limited by the uncertainty of distance-dependent monocular depth estimation. We propose CoCam4D, a Bayesian framework for collaborative perception that explicitly models geometric uncertainty. It uses a VGGT-based feedforward network to generate 3D Gaussian scene representations with associated uncertainty estimates, enabling multiple vehicles or agents to efficiently combine their observations. By sharing compact Gaussian primitives, reliable observations from one agent can reduce the depth uncertainty of another without requiring LiDAR sensors. To support real-world deployment, we introduce Dynamic Object Primitives (DOPs), a compact 35-byte representation designed for efficient C-V2X communication. Extensive experiments show that our proposed method consistently outperforms recent vision-only methods, achieving improvements of 11.48% on OPV2V+ and 10.62% on DAIR-V2X-C, demonstrating the potential of geometrically grounded collaborative perception for LiDAR-free autonomous driving.
100. 【2610.11572】PAM-ToD: Plug-and-Play Appearance Modeling for Cross-Time-of-Day 3D Gaussian Splatting
链接:https://arxiv.org/abs/2610.11572
作者:Kota Shimomura,Sungho Moon,Tsubasa Hirakawa,Takayoshi Yamashita,Sunghoon Im,Hironobu Fujiyoshi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian Splatting, real-time rendering, preserving consistent, Adapting a pre-trained, anchor images
备注: 21 pages, 9 figures
点击查看摘要
Abstract:Adapting a pre-trained 3D Gaussian Splatting (3DGS) road scene to a new time of day requires learning appearance changes from a few anchor images while preserving consistent, real-time rendering. We propose PAM-ToD, a lightweight plug-in that learns color corrections while keeping the pre-trained 3DGS parameters fixed. PAM-ToD scales each Gaussian's existing color to model illumination changes and uses an additive term for additional brightness, such as when street lamps turn on at night. Under a simplified image formation model, unchanged surface albedo can be eliminated from the relation between source and target appearances, allowing us to learn these corrections without separately estimating albedo and illumination. The model corrects colors across the scene while allowing the corrections to vary by location and by Gaussian. To guide learning from a few anchor images, it discourages abrupt spatial changes in these corrections. We also introduce CARLA-ToD, a benchmark with matching geometry, camera poses, and moving-object trajectories across three times of day. A few target-time anchor images are used to train each plug-in, while separate views are used for evaluation. Across the static and dynamic settings, PAM-ToD achieves higher PSNR and lower LPIPS than the baselines, even when the anchor images come from a single synchronized capture across multiple cameras.
101. 【2610.11560】Hankel Subspace Self-Supervised Learning for Parallel MRI Reconstruction
链接:https://arxiv.org/abs/2610.11560
作者:Mingyu Hu,Siquan Zhu,Xijun Zhong,Qiegen Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:ill-posed inverse problem, Hankel, Parallel magnetic resonance, problem under undersampling, ill-posed inverse
备注:
点击查看摘要
Abstract:Parallel magnetic resonance imaging reconstruction is an ill-posed inverse problem under undersampling. Multi-coil acquisition and Hankel lifting expose complementary repeated information: observations of the same anatomy across coils and repeated local k-space neighborhoods in overlapping windows. These dependencies guide recovery of missing k-space data. However, splitting lifted Hankel entries for self-supervision can place the original sample in both input and target, causing data leakage. We propose Hankel Subspace Self-Supervised Reconstruction (HSSRecon), a scan-specific reconstruction framework for parallel magnetic resonance imaging. HSSRecon partitions data by physical acquisition units before Hankel lifting and applies multiplicity normalization to repeated Hankel copies in overlapping windows. Rather than learning a mapping that directly predicts missing data, the network learns a compact complex-valued Hankel subspace operator. Reconstruction is performed over the original k-space variables using a conjugategradient solver with hard data consistency. This design separates structural learning in the Hankel domain from data consistency in the physical domain: the former exploits multi-coil and local Hankel correlations, while the latter solves over unacquired degrees of freedom. We provide theoretical analyses of physicalgroup splitting and multiplicity normalization, and establish positive definiteness, uniqueness, hard data consistency, and a finite-step conjugate-gradient error bound for the system. On fastMRI brain data with three contrasts and three sampling masks, HSSRecon achieves competitive peak signal-to-noise ratio, structural similarity, and normalized mean squared error across six aggregated conditions.
102. 【2610.11547】OX-NeRF: 3D X-ray Tomography Reconstruction from Sparse Views Using Implicit Neural Representation
链接:https://arxiv.org/abs/2610.11547
作者:Thomas Welsch,Min-Hsin Tu,David J. Chapman,Daniel E. Eakins
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian splatting methods, NeRF and Gaussian, Gaussian splatting, successfully applied, X-ray Neural Radiance
备注:
点击查看摘要
Abstract:NeRF and Gaussian splatting methods have been successfully applied on X-ray scenes where the views are too sparse for 3D reconstruction via classical methods. Ultra-sparse scenes with 10 or fewer views such as those with high-rate or low-dose acquisition still, however, present a significant challenge. To address this problem we present a new framework, Optimised X-ray Neural Radiance Fields (OX-NeRF), that combines cross-scene feature learning with scene-specific optimisation to reconstruct sets of related scenes. OX-NeRF employs a convolutional neural network (CNN) to identify cross-scene features while maintaining scene-specific multi-resolution hash grids of spatial features. The paired representations are fused and passed to a multilayer perceptron (MLP); the CNN, hash grids and MLP are then jointly optimised end-to-end. Benchmarking on parallel-beam and cone-beam X-ray datasets shows OX-NeRF provides significantly higher reconstruction accuracy on ultra-sparse scenes compared to existing radiance field methods.
103. 【2610.11534】HAND: A Biologically-Inspired Activation Function that Improves Generalisation and Sample Efficiency in Image Classification
链接:https://arxiv.org/abs/2610.11534
作者:Michael W. Spratling,Heiko H. Schütt
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:DNNs exhibit robustness, DNNs exhibit, exhibit robustness, training, Inductive bias
备注:
点击查看摘要
Abstract:DNNs exhibit robustness and generalisation issues not seen in humans. They are also far less data-efficient learners, requiring considerably more training samples to accurately classify novel exemplars. Inductive bias could help with these issues by providing in-built mechanisms to improve generalisation, and hence, reduce reliance on learning from data. We incorporate a biologically-inspired inductive bias into a new activation function, HAND (Homeostasis, Accelerating Nonlinearity, and Divisive-nomalisation), and show its effectiveness with CNNs trained on image classification. Using HAND a ConvNeXt-tiny required 25 training epochs to reach the same accuracy on ImageNet1k as the unmodified model achieved after 200 epochs. Consistent with the effects of an inductive bias, the performance gap reduced with training time and increased data augmentation. When the volume of training data was reduced and unevenly distributed between classes (Long-tailed ImageNet) the improvements in accuracy were even larger and did not reduce with increased training time. Generalisation performance with the common-corruptions data, and the ability to reject samples from unknown classes, were unaffected or improved by HAND. Results generalised across CNN architectures and training data-sets. HAND can, therefore, reduce the required training time and/or the required volume and variety of training data, helping to improve sample efficiency.
104. 【2610.11526】MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text Retrieval
链接:https://arxiv.org/abs/2610.11526
作者:Xintao Zong,Wenxuan Liu,Jianhao Ding,Zhaofei Yu,Tiejun Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:showing great potential, energy-efficient computing paradigm, efficient multimodal learning, sparse event-driven computation, Spiking neural networks
备注:
点击查看摘要
Abstract:Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning. However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment. Existing spiking ITR methods rely on local alignment and additional soft-label supervision during training, while lacking awareness of structural and multi-granularity relationships. To address these issues, we propose a Multi-head Spiking Graph Attention Network (\textbf{MSGAT}) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation. However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations. Therefore, we introduce \textbf{Sim-Fuse}, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations. Experiments on Flickr30K and MSCOCO show our method outperforms ANN methods under matched settings and existing SNN retrieval baselines. Moreover, with only two time steps, our SNN achieves comparable or superior performance to its ANN counterpart while reducing theoretical module-level energy by 55\%. The code is provided in the Supplementary Materials.
105. 【2610.11508】WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models
链接:https://arxiv.org/abs/2610.11508
作者:Junmyeong Lee,Dongmin Shin,Min-Gyu Park,Wooseok Jeon,Inho Chang,Hae-Gon Jeon
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:performance remains sensitive, robotic manipulation, recent advances, performance remains, remains sensitive
备注: 9 pages, 6 figures
点击查看摘要
Abstract:Despite recent advances in Vision-Language-Action models (VLAs) for robotic manipulation, their performance remains sensitive to changes in camera configuration. The problem becomes more evident in cross-setup deployment, as reproducing the exact camera pose used for training is nearly impossible. Unlike fixed external views, wrist views are more challenging because the camera moves with the robot, causing even small mounting variations to alter fine-grained geometric cues. To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations. WARP-VLA adopts a Mixture-of-Experts (MoE) architecture where individual experts learn view-specific feature transformations, and a router combines them based on implicit view information. This allows the policy to be deployed without requiring camera extrinsic parameters as additional input. Through experiments on the LIBERO benchmark, WARP-VLA improves the average success rate of pi-0.5 from 39.2% to 78.3% under wrist-view perturbations. The real-robot experiments further show that the feature-level adaptation learned in simulation successfully transfers to diverse deployment settings. To facilitate reproducibility and future research, we release our wrist viewpoint robustness benchmark and a plug-and-play implementation.
106. 【2610.11498】Parametric Trajectory Distillation for Few-Step Video Generation
链接:https://arxiv.org/abs/2610.11498
作者:Lan Feng,Peter Karkus,Maximilian Igl,Julius Berner,Yuxiao Chen,Shuhan Tan,Alexandre Alahi,Boris Ivanovic,Marco Pavone
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:making generation computationally, generation computationally expensive, Video diffusion, flow models require, computationally expensive
备注:
点击查看摘要
Abstract:Video diffusion and flow models require many sequential evaluations, making generation computationally expensive. Few-step distillation reduces this cost but poses a capacity allocation problem: a student must match the teacher's iterative generation with far less sequential computation. Existing trajectory methods ask the student to reproduce teacher transitions that are highly curved at high noise, which can exceed its capacity and degrade fine detail. We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path. PTD is designed to let the learned curvature adapt to the backbone's predictive capacity, preserving motion and diversity. The curvature head is used only in training; inference keeps the original backbone architecture. On Wan2.1-14B, four-step PTD sets a new state of the art for trajectory distillation, significantly improving dynamic quality and naturalness over PDD, the best-performing trajectory-only method on this model, under the same training setting. On the 33B audio-video MiniMax-H3, LoRA-trained PTD significantly improves diversity and naturalness over the state-of-the-art LightX2V Turbo. Blinded human votes give PTD 55.1% and 63.4% preference shares against PDD and LightX2V Turbo. Project page: this https URL.
107. 【2610.11496】Stop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake Videos
链接:https://arxiv.org/abs/2610.11496
作者:Fazhong Liu,Yan Meng,Tian Dong,Guoxing Chen,Haojin Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
关键词:Motion Aware Deepfake, Pose-guided diffusion models, Aware Deepfake, synthesize entire human, entire human figures
备注:
点击查看摘要
Abstract:Pose-guided diffusion models can now synthesize entire human figures in motion, spawning a new class of deepfakes: Motion Aware Deepfake (MAD) that have already reached hundreds of millions of viewers. To better understand this emerging threat, we construct the first MAD-specific benchmark and measurement framework, containing over 1.5 million frames that mix 1,363 real and 30,122 synthetic videos from six controllable generators, with realistic perturbations and open-world evaluation splits. Then, we dissect MAD and discover that, despite their global coherence, these videos betray faint yet reliable cues: because the model relies on limited input frames for motion synthesis, it must predict and simulate coherent movement at motion boundaries, thereby producing high-frequency artifacts along with model-specific spectral fingerprints. Based on the observations obtained from analysis on dataset, we propose MoDA, the first defense framework tailored to detect and attribute MAD videos. MoDA couples spatial semantics with steganalysis-rich frequency features via cross-domain alignment and multi-scale aggregation, achieving 94.8% in-distribution and 89.1% cross-dataset detection accuracy gains of 10% to 25% over prior work and 91.5% model attribution accuracy. MoDA achieves 81.94% accuracy on 200 clips produced by two unseen commercial MAD platforms, indicating promising zero-shot transfer, and 78.13% detection accuracy on 1,200 unseen MAD video clips (55k frames in total) collected from the open Internet. Under white-box, gray-box, and black-box adaptive attacks, MoDA maintains relatively stable detection and attribution performance while the accuracies of the baselines drop rapidly.
108. 【2610.11489】Beyond Resolution: Object-to-Image Ratio Mismatch in Instance Retrieval
链接:https://arxiv.org/abs/2610.11489
作者:Boaz Meivar,Ofir Kedem,Amit Edenzon,Gal Chechik,Shai Avidan
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:Visual instance retrieval, Visual instance, instance retrieval, retrieval often fails, apparent sizes
备注: 24 pages. Preprint, under review
点击查看摘要
Abstract:Visual instance retrieval often fails when the same object appears at different apparent sizes in the query and gallery. We show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images. On a controlled benchmark of 3,021 Objaverse objects rendered at five camera distances, more than 80% of the cross-distance degradation is attributable to O2I mismatch rather than resolution for 9 of 12 pretrained backbones; multi-scale architectures cut the resolution-only effect to single digits yet remain equally susceptible. The failure is also asymmetric: tight queries retrieve more reliably against wide gallery images than the reverse. Guided by this analysis, query-side scale augmentation and an OWLv2 crop reranker reach state of the art on ILIAS 100M (29.2 mAP@1000 before reranking, 42.0 after) without training or modifying the precomputed gallery index, and a LoRA fine-tune matches the query-side gains at a single forward pass, showing that O2I robustness is learnable.
109. 【2610.11479】Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
链接:https://arxiv.org/abs/2610.11479
作者:Bowen Zheng,Zhiguang Liu,Jiarong Ou,Rui Chen,Tianyang Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Causal video diffusion, suits streaming, model, video diffusion models, long-video generation
备注: 26 pages, 6 figures, 4 tables
点击查看摘要
Abstract:Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation. Under standard training, however, they often yield lower generation quality than bidirectional models of the same size. Many existing approaches address this gap by initializing from or distilling a pretrained bidirectional teacher. We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage. Because this path requires neither a large bidirectional teacher nor a complex distillation pipeline, it is simpler and more scalable. On this path, we find that a causal model trained on ground-truth history becomes strongly dependent on it, so that at inference errors in its own generated history propagate forward. We hypothesize that much of this dependence is unnecessary, because the current input already determines much of what the history provides. We propose Conditional Residual Prediction (CRP), a simple recipe for reducing a model's reliance on a condition: the model first predicts the target without the condition, and the condition may only add a residual on top of this prediction. Applied to history, CRP makes the model predict each chunk from the present as far as it can and use the past only for what the present cannot supply. In controlled experiments, CRP nearly closes the 6.14-point gap to a bidirectional model trained under the same setup. Scaling this recipe, we train Optica, a 2B-parameter causal video model that autoregressively generates 5-second 480p videos and reaches 82.78 on VBench with only about 15M training videos.
110. 【2610.11469】SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
链接:https://arxiv.org/abs/2610.11469
作者:Jeonghyo Song,YoungJoon Yoo
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:large language model, Recent large vision-language, large vision-language models, Recent large, diverse image-text tasks
备注: Accepted to EMNLP 2026 Findings
点击查看摘要
Abstract:Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.
111. 【2610.11460】ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification
链接:https://arxiv.org/abs/2610.11460
作者:Mohammad Zare,Pirooz Shamsinejadbabaki
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:readable form, prototype images, image, vector prototypes, document image
备注:
点击查看摘要
Abstract:Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.
112. 【2610.11444】Learning to Retrieve: Internalizing Memory Retrieval for Video World Models
链接:https://arxiv.org/abs/2610.11444
作者:JiaKui Hu,Tailai Chen,Yuqi Pan,Xuerui Qiu,Jialun Liu,Xiao Cao,Zhenxin Zhu,Guang Chen,Hangjun Ye,Bing Wang,Yanye Lu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:scene videos conditioned, generate explorable, camera trajectories, world models aim, aim to generate
备注:
点击查看摘要
Abstract:Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories. Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation. However, these auxiliary memory pathways operate outside the model's internal generative dynamics, preventing the model from intrinsically learning when and what historical information should be retrieved. We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system. Based on this principle, we introduce \textbf{Learning-to-Retrieve (L2R)}, which repurposes the model's persistent internal state as a memory for historical context. A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining \textit{what to retrieve}, while a retrieval trigger determines \textit{when to retrieve}. We further supervise the trigger with a 3D re-visibility signal, activating retrieval when previously observed content re-enters the current view while otherwise preserving the existing context. Together, these components enable the model to intrinsically acquire memory retrieval behavior and incorporate relevant historical observations into generation without a separate retrieval pathway. Across multiple base models and camera-revisit benchmarks, L2R improves long-term scene consistency while eliminating the need for an external memory bank or 3D conditions. this https URL
113. 【2610.11433】DLC: A Metric-Guided Dynamic Loss Controller for Multi-Objective Training
链接:https://arxiv.org/abs/2610.11433
作者:Jaewan Ko,Janghoon Choi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:metric-guided dynamic loss, DLC, metric-guided dynamic, restoration, image restoration
备注: Accepted to ACCV 2026
点击查看摘要
Abstract:In this paper, we introduce a metric-guided dynamic loss controller (DLC) for multi-objective image restoration. Conventional image restoration pipelines usually train with a fixed weighted combination of multiple losses, without changing the relative importance of fidelity, perceptual similarity, and no-reference quality during optimization. DLC is an architecture- and loss-term-agnostic training-time controller: it does not modify the restoration architecture or introduce new differentiable loss terms, but dynamically reweights the existing training losses. During training, DLC periodically evaluates the current model on a small fixed feedback subset and uses the resulting quality metrics to update the loss-weight vector through an LLM-based controller. Because DLC operates on existing loss terms rather than task-specific architectures, the same controller formulation can be instantiated across diverse image restoration training pipelines. We evaluate DLC on three restoration domains: low-light image enhancement, deraining, and real-world super-resolution, using both reference-based and no-reference quality metrics. Across these settings, DLC considers metric-dependent trade-offs during optimization and guides training toward balanced operating points across fidelity and perceptual quality. The results show that DLC can move models toward more favorable operating points across different restoration domains, supporting its role as a practical plug-in controller for multi-objective image restoration.
114. 【2610.11431】EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation
链接:https://arxiv.org/abs/2610.11431
作者:Feiyue Qi,Xingyue Wei,Jianwen Luo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remains challenging due, Motion estimation, myocardial motion estimation, limited image information, image artifacts
备注: 19 pages, 14 figures
点击查看摘要
Abstract:Motion estimation in echocardiography is essential for quantitative assessment of cardiac function and myocardial mechanics, but remains challenging due to image artifacts, limited image information, speckle decorrelation, and the scarcity of ground-truth displacement fields. Anatomy-guided approaches can provide structural information, yet often rely on expert-labeled myocardial segmentations. We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network. Here, unsupervised motion estimation refers to learning without ground-truth displacement fields. The self-distillation strategy jointly optimizes anatomical segmentation and myocardial motion estimation under limited anatomical annotations. Diffusion-based conditioning is used during training with stochastic perturbations, while inference requires only a single deterministic forward pass without iterative reverse-diffusion sampling. EchoDiST was evaluated on three echocardiographic datasets, including two external test datasets under cross-view and cross-dataset settings. Compared with seven representative learning-based methods, EchoDiST consistently improved anatomical alignment, myocardial strain assessment, and motion-derived functional and cardiac-phase assessment. These gains were statistically significant across the evaluated tasks and datasets. Overall, EchoDiST provides an effective approach for reliable myocardial motion estimation under limited anatomical supervision and supports downstream quantitative assessment of cardiac function.
115. 【2610.11419】Missing Modality-Aware Calibration for Trustworthy Brain Tumor Segmentation
链接:https://arxiv.org/abs/2610.11419
作者:Sol Lee,Hyunji Kim,Sungrae Hong,Donghee Han,Mun Yi
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:leverages multiple MRI, multiple MRI modalities, typically leverages multiple, Multimodal brain tumor, multiple MRI
备注: MICCAI2026 poster
点击查看摘要
Abstract:Multimodal brain tumor segmentation typically leverages multiple MRI modalities, yet incomplete modality acquisition is common in clinical practice due to protocol heterogeneity and scan failures. Although recent methods maintain segmentation accuracy under missing modality conditions, they frequently overlook prediction reliability, leading to miscalibrated confidence estimates that hinder clinical adoption. Existing calibration techniques are largely modality-agnostic or assume that prediction difficulty decreases monotonically as additional modalities become available. However, in brain tumor segmentation, prediction difficulty depends primarily on which modalities are absent rather than how many, leading to combination-specific and spatially heterogeneous calibration errors. To address this, we propose Missing Modality-Aware Local Temperature Scaling (MMA-LTS), a post-hoc voxel-wise confidence calibration method. It estimates a spatially adaptive temperature field conditioned on a modality-availability learnable token and a voxel-wise difficulty score. Experiments on BraTS 2020 and FeTS 2024 show that MMA-LTS improves calibration while preserving the segmentation accuracy of state-of-the-art models across diverse missing-modality scenarios, thereby enhancing trustworthiness toward clinical deployment.
116. 【2610.11412】Fresco++: Frequency-Guided and Canonical-Consistent Optimization for Fine-Grained Head Avatar Modeling
链接:https://arxiv.org/abs/2610.11412
作者:Shikun Zhang,Yong Li,Yiqun Wang,Qiuhong Ke,Cunjian Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:view-consistent head avatar, unified optimization framework, Head avatar optimization, framework for fine-grained, fine-grained and view-consistent
备注: 14 pages, 10 figures
点击查看摘要
Abstract:We propose Fresco++, a unified optimization framework for fine-grained and view-consistent head avatar reconstruction. Head avatar optimization is typically driven by per-view image supervision, which can lead to premature fitting of unstable high-frequency details and inconsistent local appearance across viewpoints. Fresco++ addresses these challenges by regulating both the progression of visual detail and the formation of cross-view supervision during optimization. For frequency-aware optimization, a progressive curriculum first stabilizes low-frequency structures and then introduces high-frequency constraints to recover fine facial and hair details without amplifying spurious responses at early stages. For cross-view optimization, we introduce Canonical Group Consensus, which associates local observations through shared canonical surface regions and establishes correspondence across different viewpoints. Geometric and visibility-aware screening removes unreliable observations, while the remaining multi-view evidence is aggregated in feature space to form a consensus target for supervising the current rendering. This design enforces local consistency without relying on a specific image-space parameterization and avoids additional rendering of the auxiliary view. Together, the frequency curriculum and canonical consensus provide stable optimization from coarse structures to fine details while maintaining coherent appearance across viewpoints. Extensive experiments on NeRSemble demonstrate improved reconstruction quality and cross-view consistency, while evaluations across diverse avatar representations further confirm the generality and transferability of Fresco++.
117. 【2610.11402】GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
链接:https://arxiv.org/abs/2610.11402
作者:Kun Wang,Yupeng Hu,Ruping Cao,Hao Liu,Zhiran Li,Qianlong Xiang,Harry Cheng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:jointly answer driving-scene, making precise evidence, GoldenViewVQA requires models, supporting visual evidence, answer driving-scene questions
备注:
点击查看摘要
Abstract:GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75\% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92\% Answer Accuracy and 86.44\% View Accuracy. The final submitted run achieves 88.14\% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.
118. 【2610.11401】WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models
链接:https://arxiv.org/abs/2610.11401
作者:Kai Ding,Yang He,Ruijie Quan,Yi Yang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:World Action Models, video Diffusion Transformer, pretrained video Diffusion, Diffusion Transformer, enable generalist robot
备注: 19 pages, 5 figures, 7 tables. Project page: [this https URL](https://dingkai0302.github.io/wam-cache/)
点击查看摘要
Abstract:World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.
119. 【2610.11394】CoPoE: Multimodal Fusion via Decomposable Disease-Coordinate Product-of-Experts for Missing-Modality Alzheimer's Diagnosis
链接:https://arxiv.org/abs/2610.11394
作者:Chihun An,Ikbeom Jang
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:Multimodal Alzheimer disease, cohorts frequently suffer, integrating heterogeneous clinical, clinical cohorts frequently, Alzheimer disease
备注: Accepted at IEEE BIBM 2026
点击查看摘要
Abstract:Multimodal Alzheimer's disease (AD) diagnosis benefits from integrating heterogeneous clinical, imaging, genomic, and biomarker evidence, but clinical cohorts frequently suffer from irregular modality missingness. Existing fusion methods often synthesize absent inputs, risking the introduction of artificial surrogates, or pool available signals into uninterpretable latent spaces. We present CoPoE (Disease-Coordinate Product-of-Experts), a disease-coordinate framework that maps multimodal evidence into a structured latent space partitioned into four distinct biological and clinical axes: genetic Risk, molecular Pathology, Neurodegeneration, and clinical Stage (R/P/N/S). Each observed modality parameterizes a diagonal Gaussian expert over the full RPNS vector, and a masked Product-of-Experts architecture fuses only the available modalities. Consequently, absent modalities add no factor to the fusion path, allowing the network to preserve a robust, decomposable posterior for any non-empty modality subset without synthetic imputation in the RPNS path. Through extensive missing-modality experiments on the ADNI dataset, CoPoE achieves the best all-modality performance and the highest mean AUROC across all 15 observed-subset evaluations among standardized missing-modality fusion baselines under a shared non-PET ADNI embedding benchmark, while substantially improving raw-probability ECE, Brier score, and NLL. Furthermore, PET-supervised probing shows evidence enrichment within the pathology (P) block under full modalities, with tau-related signal retained even when direct fluid biospecimen inputs are withheld. Our code is available at this https URL.
120. 【2610.11381】EvoKnow: Continual Knowledge Evolution for AI-Generated Image Detection
链接:https://arxiv.org/abs/2610.11381
作者:Zhiheng Peng,Wenwei Jin,Yangshi Ge,Siyu Xia,Jiawei Li,Xu Tang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:generative models emerge, models emerge, AI-generated image detectors, detectors are commonly, commonly trained
备注: 19 pages, 6 figures
点击查看摘要
Abstract:AI-generated image detectors are commonly trained on fixed generator domains and become difficult to maintain as new generative models emerge. Continual adaptation is challenging because replaying historical generated images is costly, whereas updating shared parameters with limited current-domain data can overwrite prior forensic knowledge. We propose EvoKnow, a replay-free framework that formulates continual AI-generated image detection as forensic knowledge evolution. EvoKnow preserves a shared forensic basis learned from base domains, incrementally adds isolated residual experts for complementary generator-relevant evidence, and retrieves expertise through an Analytical Incremental Router (AIR) updated in closed form from current-stage generated images and accumulated sufficient statistics. Experiments demonstrate effective cross-generator generalization, few-shot expansion, and long-horizon continual adaptation. With ten generated images per arriving generator, EvoKnow achieves 96.70% average accuracy on non-base GenImage generators and 94.48% accuracy on Chameleon without target-benchmark adaptation. Under a strict replay-free continual learning protocol, EvoKnow achieves state-of-the-art continual learning performance, attaining 96.32% mean stage-wise accuracy and 4.32% average forgetting.
121. 【2610.11379】FastJEV: Understanding Redundancy for Compact JEV Inference
链接:https://arxiv.org/abs/2610.11379
作者:Jie Ma,Jie Gao,Yihang Liu,Zhike Qiu,Junle Li,Chongyi Zhuang,Jiayi Ji,Xiaoshuai Sun
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:make multimodal decisions, directly scoring candidates, models make multimodal, make multimodal, directly scoring
备注:
点击查看摘要
Abstract:JEV models make multimodal decisions by directly scoring candidates. Although the common context is encoded once, candidate evaluation can still repeat matching token histories, duplicate inference states, and execute the full backbone. In this paper, we study these sources of redundancy and present FastJEV for compact candidate evaluation. We jointly organize history reuse and state storage, since sharing computation requires preserving states for later branches. We first introduce shared context anchoring to reuse recurrent initial states and omit unused final recurrent caches. We extend this reuse through candidate prefix sharing, retaining the intermediate states needed by subsequent branches. To further reduce the depth of these paths, we apply decision guided pruning based on relative score changes measured on a small unlabeled set. Our method retains full context encoding and all candidates without additional training. We evaluate FastJEV across three OmniJev model sizes on five public benchmarks and reconstructed LIBERO-10 offline questions. At the selected pruning budgets, the complete method reduces candidate depth by 43.75% to 45.83%, while retaining 93.66% to 97.52% of the original task scores on average across the six evaluation sets. Through controlled experiments, we show how candidate overlap and branching structure affect the execution cost of history reuse. In our implementation, candidate prefix sharing can reduce repeated computation while increasing latency. These findings motivate designing sharing granularity and execution schedules together for efficient JEV inference.
122. 【2610.11376】CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding
链接:https://arxiv.org/abs/2610.11376
作者:Andrea Ceron,Michael Schmidt,Alvaro Marcos-Ramiro,Sebastian Schmidt,Benjamin Busam
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:LiDAR pipelines suffer, blur sharp radial, Latent LiDAR pipelines, yielding edge depths, convolutional VAEs blur
备注: Accepted at NeurIPS 2026 (poster). 41 pages, 10 figures, 18 tables. Project page: [this https URL](https://andrea25512.github.io/CRISP/)
点击查看摘要
Abstract:Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.
123. 【2610.11374】Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
链接:https://arxiv.org/abs/2610.11374
作者:Zidan Wang,Yaqian Li,Xiaokai Zhang,Kaiwen Long,Kun He,Hanpeng Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:foundational vision-language model, facto vision encoder, vision encoder, CLIP serves, foundational vision-language
备注:
点击查看摘要
Abstract:CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature $\tau$: with $\tau$ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing ($48.99$ vs.\ $42.28$ on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine ($24.20$ vs.\ $19.01$); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across $8$ VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at this https URL.
124. 【2610.11371】SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
链接:https://arxiv.org/abs/2610.11371
作者:Zhi Rao,Yucheng Zhou,Qianran Sun,Yiqing Huang,Longcan Yuan,Jiayi Hou,Chengwen Yao,Lin Cheng,Donghui Sun,Xiaoxin Chen,Jun Wan
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:Contemporary decoder-only large, demonstrated strong capabilities, decoder-only large language, Contemporary decoder-only, large language models
备注:
点击查看摘要
Abstract:Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{this https URL}{GitHub}, together with models of different sizes to support future academic research.
125. 【2610.11364】FlyMark: Training-Free Invisible Watermarking of 3D Gaussian Splatting via a Fruit Fly Connectome
链接:https://arxiv.org/abs/2610.11364
作者:Ziyuan Luo,Haoliang Li,Renjie Wan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian Splatting, remains checkable long, portable parameter array, training pipeline, remains checkable
备注:
点击查看摘要
Abstract:A trained 3D Gaussian Splatting (3DGS) scene ships as a portable parameter array that can be copied, pruned, requantized, or repackaged outside its training pipeline, so ownership evidence is most useful when it lives in the released parameters and remains checkable long after the embedding tooling is gone. Existing 3DGS watermarks typically tie embedding or extraction to scene optimization, a learned decoder, or rendered views, so the evidence survives only as long as a second trained artifact does. FlyMark instead writes a keyed message into the parameters a 3DGS file already stores. Its carrier directions are derived from the photoreceptors of a published connectome, a citable versioned artifact that fixes the geometry exhaustively and leaves nothing to tune per scene. A virtual observer reads cone-wise apparent luminance along a scene-normalized orbit from stored centers, colors, and opacities; a keyed dithered quantization-index-modulation code replicates each message bit across these observations; and one sparse bounded least-squares solve realizes the targets through achromatic shifts of existing degree-zero colors under a hard per-channel linear-RGB bound. All geometry and higher-order appearance parameters are preserved bit-identically, and extraction needs only cone queries, rounding, and majority voting. Under a model-domain threat model on synthetic and real scenes, FlyMark attains high clean bit accuracy and visual fidelity while cleanly separating matched from wrong keys.
126. 【2610.11362】Bernoulli Flow Models: Self-Consistent Generative Modeling for Binary Data
链接:https://arxiv.org/abs/2610.11362
作者:Hao Mo,Liying Yang,Shumin Yao,Xinxing Yu,Ajian Liu,Xudong Mao,Yanyan Liang
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:generate high-quality samples, making practical inference, inference computationally expensive, Binary diffusion models, practical inference computationally
备注:
点击查看摘要
Abstract:Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffusion models define a discrete one-step forward path and then derive the reverse posterior. In low-NFE settings requiring cross-step sampling, they approximate the true multi-step likelihood with a single-step likelihood transition, which severely degrades sample quality. To address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than relying on sequential one-step Markov diffusion chains, BFM defines a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE is no longer an approximation based on skipping discrete steps; it only requires re-evaluating the analytical posterior over a new time grid. This eliminates the structural training-inference mismatch inherent to discrete chains and yields self-consistent low-NFE sampling. Experiments show that BFM is highly robust to aggressive NFE reduction. On LSUN Churches 256x256, a BFM trained with 256 steps achieves an FID of 9.22 using only 16 sampling steps, whereas the state-of-the-art discrete baseline degrades to 204.10. BFM also remains competitive with continuous and discrete generative baselines under standard full-step inference. These results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.
127. 【2610.11361】SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model
链接:https://arxiv.org/abs/2610.11361
作者:Aviad Dahan,Rajaei Khatib,Yonatan Bitton,Idan Szpektor,Lior Wolf,Raja Giryes
类目:ound (cs.SD); Computer Vision and Pattern Recognition (cs.CV)
关键词:audio-visual scene comprises, geometry it depicts, position and trajectory, dynamic geometry, audio-visual scene
备注: 24 pages, 8 figures, 18 tables. Project page: [this https URL](https://sepgen.github.io/)
点击查看摘要
Abstract:A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory. Rendering such a scene from a novel viewpoint requires that every source be available as an individual waveform, so that it can be localized in the scene and propagated to the observer before the signals are mixed. Joint audio-video generators can synthesize the video and its soundtrack, but the soundtrack is emitted as a single audio-mix in which the sources are not individually accessible. We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run. SepGen supports two complementary modes: generation and separation. In generation mode, each source caption specifies what its stem contains. A two-speaker dialogue, for example, comes out as one stem per speaker in the original turn order. In separation mode, the input audio-mix remains clean while the captions specify what to extract, so the model can decompose a recording from a free-text description. We evaluate generation on scenes synthesized from text, and separation on scenes rendered by other generators and on real recordings of speech, music, and sound effects. Given an audio-mix and captions that carry the spoken lines, SepGen outperforms language-conditioned separators, most clearly on speech, and it keeps the lead when the lines are removed from the captions. Code, checkpoints, and datasets are available at this https URL
128. 【2610.11356】Adaptive Adversarial Augmentation for Controllable Face Synthesis
链接:https://arxiv.org/abs/2610.11356
作者:Saransh Suri,Shivang Agarwal,Mayank Vatsa,Richa Singh
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:low resolution, Ensemble Feedback Controllable, Feedback Controllable Synthesis, scalable alternative, face recognition models
备注:
点击查看摘要
Abstract:Synthetic data provides a scalable alternative to real-world datasets for training face recognition models, particularly under challenging conditions such as low resolution, occlusion, and masks. Yet, most approaches lack diversity and fail to generalize effectively. We propose Ensemble Feedback Controllable Synthesis (EFCS), a guided framework that generates diverse and challenging samples while preserving visual realism. EFCS expands distributional variability, often reflected in higher FID and KID scores compared to single-feedback and random synthesis, while maintaining high precision. Recognition models trained on EFCS data consistently outperform baselines across multiple benchmarks, showing improved generalization to real-world scenarios. Furthermore, we introduce an analytically motivated formulation linking perturbation-induced difficulty, sample utility, and performance degradation, offering principled insights into balancing synthetic data complexity for optimal training. Together, these contributions establish EFCS as an effective and analytically grounded approach for bridging the gap between synthetic and real datasets.
129. 【2610.11347】EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video
链接:https://arxiv.org/abs/2610.11347
作者:Zhuo Dong,Jianhua Yang,Haohao Li,Yumeng Zhao,Keji He,Yan Huang,Liang Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Physically grounded manipulation, articulated objects requires, objects requires understanding, Physically grounded, maximum forces encountered
备注:
点击查看摘要
Abstract:Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of \(5.205 \pm 0.584\) $N$ and $0.894 \pm 0.081$ $J$, respectively.
130. 【2610.11342】Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud Representation
链接:https://arxiv.org/abs/2610.11342
作者:Kanglin Qu,Pan Gao,Qun Dai,Yuanhao Sun
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Synergistically capturing intricate, global contextual dependencies, cloud representation learning, Synergistically capturing, point cloud representation
备注: Accepted by ICLR'26
点击查看摘要
Abstract:Synergistically capturing intricate local structures and global contextual dependencies has become a critical challenge in point cloud representation learning. To address this, we introduce PointLearner, a point cloud representation learning network that closely aligns with biological vision which employs an active, foveation-inspired processing strategy, thus enabling local geometric modeling and long-range dependency interactions simultaneously. Specifically, we first design a point-focused attention, which simulates foveal vision at the visual focus through a competitive normalized attention mechanism between local neighbors and spatially downsampled features. The spatially downsampled features are extracted by a pooling method based on learnable inducing points, which can flexibly adapt to the non-uniform distribution of point clouds as the number of inducing points is controlled and they interact directly with point clouds. Second, we propose a context-scan state space that mimics eye's saccade inference, which infers the overall semantic structure and spatial content in the scene through a scan path guided by the Hilbert curve for the bidirectional S6. With this focus-then-context biomimetic design, PointLearner demonstrates remarkable robustness and achieves state-of-the-art performance across multiple point cloud tasks.
131. 【2610.11329】PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology
链接:https://arxiv.org/abs/2610.11329
作者:Fanqi Cheng,Kuo Gong,Shangke Liu,Beidi Zhao,Junchao Zhu,Zheyu Zhu,Leiyue Zhao,Fengbei Liu,John Cannon,Gang Wang,Zu-hua Gao,Kenji Ikemura,Yihe Yang,Yaohong Wang,Yuankai Huo,Xiaoxiao Li,Mert R. Sabuncu,Ruining Deng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:visual perception ability, remains poorly characterized, shown strong visual, strong visual perception, domain remains poorly
备注:
点击查看摘要
Abstract:Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized. Existing pathology VLM benchmarks largely rely on canonical closed-set prompts or perturb only generic templates, treating language as a fixed evaluation component rather than a variable axis of model behavior. In clinical practice, however, diagnostic language varies across reports, institutions, and candidate diagnoses. We introduce PathLang, a language-centered and clinically grounded zero-shot benchmark. PathLang holds the underlying slides, ground-truth labels, and image-text evaluation direction fixed while systematically varying only the diagnostic language, so that performance differences reflect how a diagnosis is phrased rather than what is imaged. The language variation follows how pathologists actually rephrase diagnoses (terminology, specificity, and reporting style), and all prompts and candidate pools are validated by six board-certified pathologists. PathLang covers four task families: (1) zero-shot classification with image-text alignment analysis, (2) cross-modal retrieval, (3) paraphrase robustness, including semantic-equivalence paraphrases, length and reporting-style variation, and prompt ensembling, and (4) open-vocabulary diagnosis retrieval over four candidate pools with distinct forms of semantic competition. Across nine VLMs and five public datasets spanning four organs, we find that performance is highly sensitive to clinically equivalent paraphrases, varies substantially across forms of semantic competition, and that image-text alignment quality does not necessarily translate into inter-class separability. We release the prompt corpus, candidate pools, pre-computed text embeddings, and evaluation code at this https URL.
132. 【2610.11320】It's Always 10:10: Reference Images Break a Bias That Prompts Only Dent
链接:https://arxiv.org/abs/2610.11320
作者:Luca Cazzaniga
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
关键词:reproduce the habits, photographs they learned, time, images, clocks
备注: 16 pages, 5 figures, 6 tables. Data and code: doi: [https://doi.org/10.5281/zenodo.23224681](https://doi.org/10.5281/zenodo.23224681)
点击查看摘要
Abstract:Text-to-image models appear to reproduce the habits of the photographs they learned from. Analog clocks are an extreme case: in advertising, watches almost always show 10:10, and generated clocks return to 10:10 even when another time is requested. We measure this bias and test three ways of overcoming it on 52 models available on the Magnific platform, with a replication on Higgsfield. Every image shows three identical clocks that must show 2:35, 6:50 and 11:20. The description of the object is fixed and only the request about the time changes: no time (A), the time in digits (B), the hand positions described by construction relative to the dial numerals (C), or the same description plus a drawn reference dial (D). Two AI readers read all 1,799 images blind from coded copies, with a third reader and the author settling disagreements (dial-level agreement 96.0% and 97.3%). With no time requested, 67% of the images have all three clocks at 10:10. On the 20 current models, all three clocks are correct in 34% of the images with digits, 30% with the hands described in words and 75% with the reference dial (D-B: +37 points, 95% CI +28 to +45); we found no evidence that describing the hands in words beats the digits (C-B: -4 points, CI -10 to +1). The replication on the 12 models shared by both platforms gives the same picture (B 54%, C 50%, D 81%). Writing the time reduces the bias but leaves two thirds of the images of the 20 current models with at least one wrong clock; adding a drawn reference raises full accuracy to three quarters and almost eliminates images entirely at 10:10. We release all images, prompts, raw readings and a script that recomputes every result.
133. 【2610.11315】Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
链接:https://arxiv.org/abs/2610.11315
作者:Shiwen Wang,Pengxiang Zhao,Xiaoming Yuan
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Optimization and Control (math.OC)
关键词:high-precision low-rank component, branches can mitigate, loss by decomposing, Existing low-rank PTQ, low-rank PTQ approaches
备注:
点击查看摘要
Abstract:In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating the output-side low-rank factor yields a \emph{deflated Hessian} that discounts residual errors already captured by the low-rank component, while an activation-noise surrogate is incorporated to suppress activation quantization error. Across five diffusion backbones, rank-4 \method{} consistently outperforms rank-4 SVDQuant in PSNR and LPIPS. It further surpasses rank-32 SVDQuant on SANA-1.6B, FLUX.1-schnell, and FLUX.1-dev with an $8\times$ smaller rank and up to $6.25\times$ faster quantization. Furthermore, on the Qwen3-8B LLM, rank-4 \method{} improves MMLU accuracy from 61.50\% to 68.17\% over rank-32 SVDQuant. Overall, \method{} achieves better W4A4 performance with substantially lower rank and quantization cost.
134. 【2610.11310】FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
链接:https://arxiv.org/abs/2610.11310
作者:Kyeong-Rae Kim,Sungnyun Kim,Tae-Hyun Oh
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:lack explicit mechanisms, internalize global geometry, global geometry directly, large language models, audio-visual large language
备注: Project page: [this https URL](https://byulharang.github.io/FloorSAV/)
点击查看摘要
Abstract:While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
135. 【2610.11306】Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples
链接:https://arxiv.org/abs/2610.11306
作者:Zhi Li,Haowei Liu,Hongchen Yang,Xiaoxuan Wang,Song Gao,Shaowen Yao,Wei Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Adversarial distillation, Adversarial, Collaboratively Guided Adversarial, distillation, compact students
备注:
点击查看摘要
Abstract:Adversarial distillation transfers robustness from high-capacity teachers to compact students. Existing adversarial distillation methods mainly use teacher predictions on clean or adversarial examples to supervise student learning. However, teacher-favorable supervision within the perturbation neighborhood remains underexplored in adversarial distillation. We therefore propose Collaboratively Guided Adversarial Robust Distillation (CGARD), which jointly optimizes distinct student-adversarial and teacher-collaborative examples within the same perturbation neighborhood. The teacher-collaborative example is constrained to incur no greater cross-entropy loss under the teacher than the clean input. CGARD combines collaborative teacher guidance with adversarial teacher supervision to improve robust knowledge transfer. Experiments on CIFAR-10 and CIFAR-100, including white-box evaluation and additional black-box transfer evaluation, demonstrate consistent robustness improvements over strong adversarial distillation baselines.
136. 【2610.11303】Efficient Multi-Granularity Knowledge Transfer for Radiology Report Generation
链接:https://arxiv.org/abs/2610.11303
作者:Xubin Zhong,Zheyu Zhang,Wenjian Qin,Ning Wen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Radiology report generation, automatically generate clinical, generate clinical descriptions, X-ray images, Radiology report
备注:
点击查看摘要
Abstract:Radiology report generation can automatically generate clinical descriptions from X-ray images, thereby significantly improving the efficiency of radiologists. This task is challenging because it requires medical knowledge to accurately identify diseases and describe them in a professional manner. However, existing methods often overlook the importance of enhancing medical knowledge in describing pivotal areas, a capability that requires models to effectively extract and aggregate knowledge at multiple levels of granularity. Accordingly, we herein propose a novel and compact Efficient Multi-Granularity Knowledge Transfer (\textbf{EMGKT}) method to address the above issues. First, we encode global knowledge embeddings using a medical vision-language model, which provides contextual medical knowledge. Moreover, we devise a novel Fine-Grained Knowledge Distillation (FGKD) training task which efficiently extract fine-grained knowledge. Specifically, the FGKD training task contains teacher embeddings and student embeddings. Teacher embeddings are encoded using extra priors; while student embeddings are learned from the teacher embeddings through knowledge distillation. During inference, the student embeddings are used to enhance fine-grained knowledge while the teacher embeddings are discarded, resulting in negligible computational costs and no need for extra priors. Finally, we further develop a mixture of disease diagnosis expert classifiers to enhance knowledge extraction. The classifiers are initialized using disease embeddings and are modeled as different experts to address various granularity features. Notably, \textbf{EMGKT} can be efficiently applied to most existing methods. Extensive experiments are conducted on two widely-used public datasets and various baselines, which demonstrates the effectiveness and transferability of \textbf{EMGKT}.
137. 【2610.11302】CATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity
链接:https://arxiv.org/abs/2610.11302
作者:Chengfeng Han,Baole Ai,Xianlu Bian,Jie Yao,Zilong Huang,Ang Wang,Dandan Ding
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Diffusion Transformers, Training-free sparse attention, solution to Diffusion, practical acceleration solution, Training-free sparse
备注:
点击查看摘要
Abstract:Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning. It typically involves estimating the importance of query-key regions and deriving sparse masks to compute only the important candidates, which inevitably introduces approximation errors that may degrade generation quality. To better balance the efficiency-quality trade-off, we propose iCATS, integrating improved importance estimation and sparse mask construction with an efficient hardware execution strategy. Specifically, for importance estimation, unlike previous works that perform independent clustering over query and key tokens based on feature similarity to estimate attention scores, iCATS demonstrates that clustering based on query-key dot-product interactions is more accurate and further reformulates this objective as a simple quadratic form for low-cost computation. For sparse mask construction, instead of using a fixed top-p rule, we observe that tolerance to sparse approximation errors varies across denoising timesteps and therefore introduce an SNR-guided sparsity schedule to adjust sparsity dynamically, leading to higher accuracy. Finally, for hardware execution, we devise a tail-merging strategy to reduce padding overhead caused by irregular cluster sizes, improving GPU kernel utilization. Extensive experiments show that iCATS achieves $2.03\times$ acceleration with 31.017 dB PSNR on HunyuanVideo-T2V-13B and $1.55\times$ acceleration with 29.301 dB PSNR on Wan2.1-T2V-14B, delivering a state-of-the-art efficiency-quality trade-off.
138. 【2610.11296】Spatial-Frequency-Aware Implicit Neural Representation of Multidimensional Signals via MLP-KAN Fusion
链接:https://arxiv.org/abs/2610.11296
作者:Wen Yan,Ligen Shi,Jun Qiu,Haimiao Zhang,Lina Wu,Chang Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Implicit Neural Representations, Implicit Neural, MLP branch, Neural Representations, compelling paradigm
备注:
点击查看摘要
Abstract:Implicit Neural Representations (INRs) have emerged as a compelling paradigm for modeling multidimensional signals by mapping continuous coordinates to signal values. However, Multi-Layer Perceptrons (MLP)-based INRs inherently suffer from spectral bias, which favors low-frequency components and suppresses the reconstruction of essential high-frequency details. While existing techniques, such as Fourier feature mappings, mitigate this issue, they often rely on sensitive manual tuning and are prone to spectral artifacts. In this paper, we propose a spatial-frequency-aware INR framework that combines an MLP branch with a Kolmogorov-Arnold network (KAN) branch for complementary frequency-oriented modeling. The MLP branch provides a low-frequency-oriented representation of smooth structures, whereas the KAN branch complements localized variations and fine details. To coordinate the two branches, we integrate the discrete wavelet transform (DWT) and inverse discrete wavelet transform (IDWT) into the output fusion stage. The outputs of the two branches are decomposed into wavelet coefficients, and the corresponding coefficients are additively fused before inverse wavelet reconstruction. A wavelet-domain band-separation regularization further penalizes high-frequency responses in the MLP branch and low-frequency responses in the KAN branch, thereby encouraging complementary frequency-oriented behavior. Experiments on 1D signals, 2D images, 3D volumes and signed distance functions, videos, and 4D light-fields demonstrate the applicability of the proposed representation across the evaluated signal modalities. Results demonstrate improved reconstruction fidelity across the evaluated signal modalities.
139. 【2610.11286】When Scene Text Hijacks the Scene: Uncovering, Exploiting, and Mitigating Rendered-Text Semantic Leakage in Image Generation Models
链接:https://arxiv.org/abs/2610.11286
作者:Feifei Li,Runjie Wang,Xiaohan Zhang,Zhenxing Qian,Mi Wen,Mi Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:trustworthy AI systems, reliability and accountability, essential for building, building responsible, responsible and trustworthy
备注: To appear in the 2027 IEEE Symposium on Security and Privacy (IEEE SP 2027)
点击查看摘要
Abstract:The reliability and accountability of image generative models (IGMs) are essential for building responsible and trustworthy AI systems. Recent IGMs, such as Nano Banana and GPT-Image, now support complex instruction following, realistic image synthesis, and controllable scene-text rendering. As these capabilities expand, safety analysis must also account for new control channels introduced by complex prompts. In this work, we study rendered-text semantic leakage, a largely overlooked phenomenon in open-domain text rendering. Although rendered text is intended to serve as a local visual constraint that should be reproduced verbatim in the generated image, it also carries linguistic semantics that may be interpreted by the model as part of the input instruction. This makes rendered text a potential semantic control channel whose safety implications remain insufficiently understood. We systematically characterize this phenomenon by decoupling the main visual prompt from the rendered text and measuring their individual and compositional effects on generated images. We quantify semantic leakage and rendering fidelity, and further analyze how leakage emerges from intermediate model evidence. We then show that harmful semantics embedded in scene text can persist through LLM-based prompt enhancement pipelines and steer non-text image regions, even when the main visual prompt remains benign. Finally, we propose a preliminary mitigation approach that reduces unsafe semantic transfer from rendered text to non-text regions while preserving the intended text-rendering behavior on FLUX-2-dev. Our findings reveal rendered text as a dual-use carrier of visible data and latent semantics, exposing a text-centric cross-modal attack surface in modern IGMs.
140. 【2610.11283】Being-M0.7: A Latent World-Action Model for Humanoid Robots
链接:https://arxiv.org/abs/2610.11283
作者:Junpeng Yue,Boyuan Li,Yuxuan Wang,Zepeng Wang,Yuhui Fu,Feiyang Xie,Yu Zhang,Jing Zhang,Xianqi Zhang,Weibo Li,Xiaofei Zheng,Yuming Fang,Jiangxing Wang,Zongqing Lu
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:requires coordinated locomotion, scarce robot demonstrations, future scene evolution, loco-manipulation requires coordinated, requires coordinated
备注:
点击查看摘要
Abstract:Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training. We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure. Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics. Robot mid-training adapts this coarse-grained prior to robot viewpoints and body dynamics. During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands. Being-M0.7 achieves the highest aggregate success rate among the compared baselines on SIMPLE and matches the strongest baseline on real-world Unitree G1 loco-manipulation tasks.
141. 【2610.11264】LR-V2X: Loss-Resilient Collaborative Perception under Low-Bandwidth Communication
链接:https://arxiv.org/abs/2610.11264
作者:Kang Yang,Tianci Bu,Peng Wang,Deying Li,Yongcai Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:yield practical benefits, vehicular wireless communications, packet loss, BEV feature, BEV feature fusion
备注:
点击查看摘要
Abstract:Given the inherent unpredictability of packet loss in vehicular wireless communications, V2X collaborative perception can yield practical benefits only if agents can achieve reliable collaboration under lossy and low-bandwidth communication conditions. Existing dense BEV feature fusion methods depend on redundant BEV feature exchange, which is infeasible in low-bandwidth scenarios, while compact-communication methods aggressively compress messages but can hardly recover the missing feature content after packet loss. In this paper, we present LR-V2X, a loss-resilient, latent-space reconstruction framework that converts corrupted received latents (even under severe 90% packet loss) into a spatial prior and then reconstructs the missing BEV information from this informative prior and using ego context as condition. Notably, the model can be trained under complete communication conditions and can be directly applied to lossy conditions at test time, eliminating the need for training under numerous lossy conditions. Experiments on DAIR-V2X and V2XREAL show that LR-V2X delivers the strongest robustness under severe packet loss and preserves reliable collaboration as communication quality degrades. And it reduces communication overhead by $64\times$ compared to dense BEV feature fusion baselines. Code will be released at this https URL.
142. 【2610.11251】V-CoLA: Vision Token Compression with Linear Attention
链接:https://arxiv.org/abs/2610.11251
作者:Hao Jiang,Yiru Mao,Tianpeng Bu,Hao Zhou,Hongtao Duan,Wang Jing,Bowen Xu,Xin Chen,Lulu Hu,Bin Yang,Yongliang Tao,Minying Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:substantial computational overhead, demonstrated impressive capabilities, Vision-language models, computational overhead, input sequence
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attention. V-CoLA introduces a novel \textit{uniqueness-aware importance criterion} for identifying critical vision tokens, coupled with an \textit{adaptive token merging strategy} that performs compression. All components are optimized at the implementation level to remain compatible with the chunk-wise parallelism of linear attention, ensuring strong practical value. Extensive experiments across multiple benchmarks demonstrate the superiority of V-CoLA: it achieves 99.5\% of the original performance with only 50.0\% of vision tokens, and over 88.0\% with as few as 12.5\%, while delivering a 1.86$\times$ to 6.15$\times$ prefill speedup.
143. 【2610.11241】AP3D: Thermal-Assisted 3D Human Point Clouds
链接:https://arxiv.org/abs/2610.11241
作者:Xie Zhang,Chengxiao Li,Xuan Liu,Chenshu Wu
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
关键词:versatile representation, representation for AI-enabled, AI-enabled human sensing, point clouds, Human
备注:
点击查看摘要
Abstract:Human body point clouds are a versatile representation for AI-enabled human sensing. However, existing methods using LiDAR, radar, and depth cameras suffer from inherent drawbacks in high cost, sparse reconstruction, and privacy concerns, etc. In this paper, we exploit low-cost thermal arrays and present TAP3D, the first system to reconstruct 3D human point clouds from body heat signatures, offering significant advantages in cost, density, human sensitivity, and privacy. To overcome major challenges in depth estimation, thermal interference, and multi-person separation, we propose a novel physics-informed design, which integrates a forward thermal physics model with two distinct modules: multi-primitive estimation for self-supervised joint recovery of depth and other thermal properties, and geometric perspective fusion for suppressing interference and disentangling multiple people. We implement TAP3D using a single commodity thermal array sensor and build a large-scale dataset (160K samples, 8 environments, 11 users) for evaluation. TAP3D achieves remarkable accuracy for dense point cloud generation, enabling downstream tasks like fall detection (91.46%), indoor tracking (21.86 cm MAE), and human mesh recovery (4.87 cm error). By transforming body heat into point clouds for the first time, TAP3D pioneers a new paradigm for privacy-first, fully passive human sensing for many applications. TAP3D is open-sourced at this https URL.
144. 【2610.11238】Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models
链接:https://arxiv.org/abs/2610.11238
作者:Haoqian Zhang,Ziyuan Yang,Zerui Shao,Yi Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:achieved remarkable success, remains challenging, achieved remarkable, remarkable success, success in image
备注:
点击查看摘要
Abstract:Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging. Existing attribution methods often compress factor-specific effects into scalar responses, making distinct internal changes indistinguishable. This is particularly limiting for diffusion models, where semantic factors emerge through evolving representation dynamics during denoising. We therefore reformulate diffusion data attribution as attributing factor-induced internal response trajectories. In this paper, we propose a novel Concept Attribution method through Dynamic Trajectories(CADT). We argue that attribution should therefore ask not only \emph{which} examples matter, but also \emph{how} their influence unfolds during generation. Specifically, we construct matched counterfactual pairs at identical noisy states to isolate factor-specific representation displacements, and model their directional and magnitude evolution across denoising as dynamic attribution signatures. For each training example and generated query, CADT extracts stage-wise feature vectors and integrates them along the denoising process to form a trajectory descriptor. Applying the same construction across the training set yields a bank of factor-specific trajectory descriptors. The covariance statistics of this bank are then used to construct . CADT uses this covariance-aware positive-semidefinite kernel to calibrate the query and training representations, and compares the calibrated query trajectory with each training trajectory to produce the final training-sample attribution scores. Experiments on multiple public datasets show consistent improvements over existing diffusion attribution baselines across hierarchical, compositional, and style attribution.
145. 【2610.11237】Breaking the Group Size Barrier: Parameter-Efficient Group Dance Generation with Chain-of-Dancers
链接:https://arxiv.org/abs/2610.11237
作者:Jing Xu,Cunjian Chen,Qiuhong Ke
类目:Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
关键词:interactive content creation, synthesize coordinated multi-dancer, coordinated multi-dancer choreography, Group dance generation, dance generation aims
备注: Accepted at NeurIPS 2026
点击查看摘要
Abstract:Group dance generation aims to synthesize coordinated multi-dancer choreography from music, with broad applications in animation and interactive content creation. This task requires modeling dense inter-person dependencies to ensure spatial coordination, while naturally preserving individual dancer identities. Existing approaches model all dancers jointly with end-to-end transformers, which tie the architecture to a fixed group size and entangle per-dancer identities across frames. We propose ChainDance, a scalable framework that reformulates group dance generation as a Chain-of-Dancers: a sequential decomposition over per-dancer conditional distributions, allowing a single model to scale across variable group sizes without retraining and naturally preserving per-dancer identity. Built on a frozen single-dancer diffusion backbone, ChainDance introduces two lightweight modules: a Role-Aware Text Encoder (RATE) for per-dancer semantic conditioning, and a Group-Aware Motion Encoder (GAME) that aggregates previously generated dancers via a distance-weighted graph convolutional network, and incorporates a training-free noise optimization procedure at inference time to enforce global spatial coherence. Experiments on AIOZ-GDance demonstrate that ChainDance achieves state-of-the-art motion quality and group coordination while structurally preserving per-dancer identity, with $3$-$4\times$ fewer parameters and requiring $3$-$6\times$ less training time compared to prior approaches.
146. 【2610.11233】MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
链接:https://arxiv.org/abs/2610.11233
作者:Leran Chen,Lingnan Kong,Zile Cai
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Multimedia (cs.MM)
关键词:evidence, core challenge, challenge in short-video, short-video fact-checking, fact-checking is identifying
备注: 33 pages, 2 figures, 20 tables
点击查看摘要
Abstract:A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.
147. 【2610.11231】Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?
链接:https://arxiv.org/abs/2610.11231
作者:Minhao Fan,Yinyi Liu,Jiayu Zhao,Zihan Teng,Song Chen,Weichen Liu
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:frozen small VLM, Small vision-language models, read external evidence, Small vision-language, struggle to obtain
备注: 58 pages, 10 figures, including appendices
点击查看摘要
Abstract:Small vision-language models may be able to read external evidence yet struggle to obtain it. We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness. A large teacher uses student execution traces to revise reusable content and control, while a separate validation set selects the deployed harness. Deployment requires neither weight updates nor teacher calls. Across seven visual question-answering settings with students of at most 9B parameters, HC improves scores over bare students by 9.9-23.9 points, averaged over three independent builds per setting. Interventions on five runtime decision types (invocation, selection, argument generation, evidence integration and abstention) show why this allocation matters: requesting evidence and generating open queries can be costly, whereas bounded choices and reading supplied text can remain useful student work. Fact cards benefit all ten evaluated students, but decision policies transfer unevenly. Recompilation for a new student model helps when the transferred interface no longer fits the student. With 100 practice items, HC exceeds answer-only LoRA on three tasks. Larger training budgets can match or surpass a fixed harness, while combining the two improves SlideVQA beyond either alone. These findings support allocating work from measured student behavior rather than uniformly removing decisions.
148. 【2610.11221】How Firm Should a Grasp Be?
链接:https://arxiv.org/abs/2610.11221
作者:Matthew Beveridge,Shree K. Nayar
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:avoid damaging, properties, ideal robot grasp, securely handle, ideal robot
备注: CoRL 2026
点击查看摘要
Abstract:An ideal robot grasp is firm enough to securely handle an object, yet gentle enough to avoid damaging it. Achieving this balance requires knowledge of the object's material properties, such as its mass, elasticity, and surface friction. These properties, however, are seldom precisely known a priori. In this work, we propose a visuotactile approach to estimating material properties in real time, during the process of grasping. Our method uses these estimated properties to determine the minimum grasp force required to handle the object. We contribute a new dataset of real-world objects (fruits and vegetables) with measured physical properties (shape, mass, elasticity, and friction), which we use to construct our force estimation model via simulations. We experimentally validate our approach to grasp force control using a robot with a parallel-jaw gripper. We demonstrate our system's ability to gently grasp a wide variety of objects, in each case adapting to their unique physical properties.
149. 【2610.11215】GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images
链接:https://arxiv.org/abs/2610.11215
作者:Qirui Wu,Stan Birchfield,Hesam Rabeti,Angel X. Chang,Bowen Wen
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:inferring surfaces hidden, casual images requires, images requires integrating, Reconstructing complete, requires integrating sparse
备注:
点击查看摘要
Abstract:Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: this https URL
150. 【2610.11205】Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study
链接:https://arxiv.org/abs/2610.11205
作者:Kim-Cuc Nguyen,Ngai-Man Cheung
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:understanding Vision Transformer, Vision Transformer, understanding Vision, structure is crucial, crucial for understanding
备注: Accepted in IEEE VCIP 2026
点击查看摘要
Abstract:Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior. However, prior studies neither isolate nor analyze module-level features nor investigate how their interactions contribute to performance estimation. In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design. Our contributions are fivefold: Across diverse architectural scales, 1) We identify feature collapse at initialization, which leads to redundancy, and propose a reduction scheme to mitigate this issue. 2) We quantify feature information using entropy and the minimum eigenvalue, demonstrating that these metrics serve as reliable indicators for generalization prediction. 3) We show that feature in the token space provides a more faithful representation than those in embedding space. 4) We discover an unexpected finding: features produced by linear submodules within ViT layers are critical for the prediction of generalization performance. 5) Our proposed proxy improves the correlation ranking by 18-48% over prior baselines and can effectively identify ViT architectures that achieve higher accuracy at lower or comparable computational cost.
151. 【2610.11198】3DTexMOR: 3D Gaussian Multi-Object Removal via Texture-Space Inpainting
链接:https://arxiv.org/abs/2610.11198
作者:Kunxin Guang,Yonghao Zhao,Jian Yang,Beibei Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remove target objects, object removal aims, target objects, aims to remove, remove target
备注: 17 pages, including appendix. Kunxin Guang and Yonghao Zhao contributed equally
点击查看摘要
Abstract:3D object removal aims to remove target objects from reconstructed scenes and complete the geometry and appearance of occluded regions. Existing NeRF- and 3DGS-based methods typically inpaint 2D images to guide 3D completion. However, complex multi-object layouts limit the surrounding context visible in each view, making 2D inpainting prone to artifacts. Inconsistent completions across views also introduce conflicting supervision and blurry reconstructions. We propose 3D Gaussian Multi-Object Removal via Texture-Space Inpainting (3DTexMOR). Our key idea is to perform inpainting in a unified texture space shared by all views. By combining complementary observations, this space provides richer context for recovering missing regions and promotes cross-view appearance consistency. We aggregate multi-view observations into texture maps, inpaint the missing regions, and reproject the completed maps into camera views to supervise Gaussian scene completion. To avoid the influence of view-dependent highlights and reflections, we decompose appearance and aggregate view-independent intrinsic attributes instead of RGB colors. We further introduce geometrically regularized Gaussian completion to constrain the geometry of the completed regions. Extensive experiments demonstrate visually plausible completions and state-of-the-art multi-object removal performance, improving PSNR by 5.8 dB and reducing LPIPS by at least 22% compared with existing methods.
152. 【2610.11194】OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes
链接:https://arxiv.org/abs/2610.11194
作者:Naiyu Fang,Zhongjin Luo,Yuxin Mo,Siyuan Huang,Jianbo Liu,Yufei Liu,Zheyuan Zhou,Chenkai Jin,Xiaogang Wang,Hongsheng Li
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:demanding massive data, demanding massive, foundational primitive, primitive in embodied, massive data
备注:
点击查看摘要
Abstract:Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.
153. 【2610.11187】RGBD-to-3D Object Mesh Refinement via Depth Matching and Symmetry Propagation
链接:https://arxiv.org/abs/2610.11187
作者:Ahyun Seo,Minsu Cho
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:produce plausible meshes, input view, discontinuities and self-occlusions, produce plausible, plausible meshes
备注: To be appear in ACCV2026
点击查看摘要
Abstract:Single-view 3D reconstructors often produce plausible meshes that disagree with the input view, especially near depth discontinuities and self-occlusions. We present a lightweight, plug-and-play RGBD-to-3D refinement that improves any RGB-to-3D reconstructor without retraining. Given a depth map, we correct the visible surface by bipartite matching to back-projected depth points, mirror these corrections onto the occluded side across a detected symmetry plane, and propagate them with a smoothness solver. Every stage is closed-form, making the method orders of magnitude faster than optimization-heavy test-time refinement. On GSO and OmniObject3D with five backbones, it yields consistent gains, also with monocular pseudo-depth, benefits more from symmetry on symmetric objects, and compares favorably with prior refinement in accuracy and runtime. It further improves an RGB-D-to-mesh reconstructor and transfers to real captures with noisy sensor depth.
154. 【2610.11184】WorldFact-Bench: Beyond Image-Internal Plausibility to Image-World Consistency
链接:https://arxiv.org/abs/2610.11184
作者:Zhuohong Chen,Zhengxian Wu,Yunyao Yu,Hangrui Xu,Zijian Yu,Hao Tan,Zhifang Liu,Peng Jiao,Jun Lan,Haoqian Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:authenticity increasingly difficult, made visual authenticity, visual authenticity increasingly, difficult to assess, authenticity increasingly
备注:
点击查看摘要
Abstract:Advances in image generation have made visual authenticity increasingly difficult to assess. Although image forensics now examines both generation artifacts and higher-level visual inconsistencies, a plausible image can still contradict real-world facts or rules. We introduce WorldFact-Bench to evaluate image-world consistency from a single image, without a predefined claim or verification target. The benchmark contains 1,274 source-aligned real-fake pairs across four verification regimes and ten semantic domains. Each pair introduces a specific, evidence-supported factual conflict while seeking to preserve non-target content and visual plausibility. Images are evaluated independently, and pair accuracy requires both members of a pair to be classified correctly. We further propose PERSIST-Agent, which organizes iterative verification around a persistent state linking candidate facts, visual observations, evidence, and verification statuses. This state guides subsequent inspection and retrieval while retaining unresolved candidates. With backbone weights fixed, harness self-optimization refines the agent's prompts and execution rules through validation feedback. Experiments reveal strong label biases in several detectors and uneven gains from retrieval. On the evaluated 8B backbones, PERSIST-Agent improves pair accuracy over both direct judgment and retrieval-augmented baselines, while ablations support the role of persistent verification state. These findings highlight the value of state-guided verification and the remaining gap between visual plausibility and factual correctness.
155. 【2610.11181】MATE4D: Matrix-Guided Editable 4D Generation from a Single Image
链接:https://arxiv.org/abs/2610.11181
作者:Xiaotian Chen,Dongfu Yin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:rapidly pushed content, Generative models, content creation be-yond, pushed content creation, models have rapidly
备注:
点击查看摘要
Abstract:Generative models have rapidly pushed content creation be-yond 2D imagery toward dynamic 3D and 4D scene synthesis. Yet pro-ducing realistic and temporally stable 4D content from a single image is still difficult because one view provides limited structural cues and weak motion evidence. We introduce MATE4D, a framework that converts one input image into editable dynamic 4D content. Our method constructs a spatio-temporal multi-view image matrix with text-guided background manipulation, delivering coherent supervision over viewpoint, appear-ance, and motion. These synthesized observations are used to optimize 3D Gaussian primitives, which are then animated through a lightweight deformation module to form a 4D representation. The resulting scenes preserve geometry more faithfully, maintain smoother temporal behavior, and keep background edits more consistent, reducing context ambiguity and motion artifacts. Experiments on Objaverse-XL and Diffusion4D show that MATE4D outperforms strong baselines in visual quality, effi-ciency, and controllability, supporting practical AR/VR content creation.
156. 【2610.11176】Multimodal Remote Sensing Image Registration: A Comprehensive Review, Challenges and Prospects
链接:https://arxiv.org/abs/2610.11176
作者:Zhiqiang Han,Yuanxin Ye,Qiuyun Wu,Jinhao Chen,Bai Zhu,Siyuan Hao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remote sensing data, Multimodal remote sensing, change detection, remote sensing image, target recognition
备注: 12 figures, 8 tables, 135 references. Review article accepted for publication in Photogrammetric Engineering and Remote Sensing (ASPS), manuscript number PERS-26-00034
点击查看摘要
Abstract:Multimodal remote sensing image registration is a crucial prerequisite for the collaborative processing and downstream application of remote sensing data, such as image fusion, change detection, and target recognition. However, significant variations in radiometry, geometry, scale, viewpoint, and time often exist between multimodal images. These differences, driven by varying sensor geometries, physical radiation mechanisms, imaging platforms, and environmental disturbances, pose severe challenges to achieving high-precision, robust registration. This paper systematically reviews the progress of mainstream multimodal remote sensing image registration methods. Based on their registration pipelines, existing approaches are categorized into three main types: region-based, feature-based, and deep learning-based methods. We detail the core principles, representative algorithms, advantages, and limitations of each category. Additionally, we summarize publicly available multimodal image datasets in the remote sensing domain, analyzing their specific characteristics and applicable scenarios. Finally, we highlight current bottlenecks in high-precision registration research and outline future development trends. This review aims to provide a comprehensive reference and valuable insights for researchers in related fields.
157. 【2610.11174】IntactWorld: Joint World Modeling with Intact Features
链接:https://arxiv.org/abs/2610.11174
作者:Boming Tan,Xiangdong Zhang,Yan Xia,Qi Zhu,Deyi Ji,Xue Yang,Shaofeng Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:intrinsic real-world logic, recent video generation, video generation models, generation models synthesize, highly realistic visuals
备注:
点击查看摘要
Abstract:While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity $v$ within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature $x_0$ at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4\% and cutting inference latency by 43.8\%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.
158. 【2610.11171】VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
链接:https://arxiv.org/abs/2610.11171
作者:Runquan Gui,Hanzhu Chen,Zehao Wang,Hanxin Zhu,Xin Li,Zhibo Chen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Long-form video understanding, textbf, Long-form video, involves multiple questions, involves multiple
备注:
点击查看摘要
Abstract:Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1\% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9\%} and \textbf{61.4\%}.
159. 【2610.11162】AutoAdapt: Reliable Few-Shot Adaptation under Clinical Distribution Shifts
链接:https://arxiv.org/abs/2610.11162
作者:Song Wang,Jie Peng,Davis Hobley,Zachary Plotkin,Tianlong Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:reuse learned prior, learned prior knowledge, Large pretrained clinical, Large pretrained, clinical models provide
备注:
点击查看摘要
Abstract:Large pretrained clinical models provide a practical way to reuse learned prior knowledge across hospitals by adapting models to them. In practice, a target hospital may only have a small labeled patient cohort, a setting commonly referred to as few-shot adaptation. This requires making multiple decisions, such as which pretrained model to adapt, how much of the model to update, and which patients to use. Nevertheless, this process faces two primary challenges. First, the best adaptation strategy varies across clinical tasks. Second, evaluating and comparing candidate strategies becomes unreliable due to the small patient cohort. In this work, we introduce AutoAdapt with two core designs to deal with these challenges. The Adapter defines an extensible space of adaptation recipes, and the Automator forms a weighted recipe combination from evidence within the adaptation patients. We propose a reliability rule to ensure that only the most effective strategy on most available patients will be selected. These selected strategies then form a combination for effective few-shot adaptation. We conduct extensive experiments across critical care, emergency care, and diagnostic datasets, and the results show that AutoAdapt consistently achieves state-of-the-art performance using only a few patients for adaptation.
160. 【2610.11161】VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving
链接:https://arxiv.org/abs/2610.11161
作者:Zhaoyang Liu,Kun Jiang,Ziying Song,Diange Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:recovering unified, visual observations, strong foundation, foundation for geometry-centric, VGGT
备注: 20 pages, 9 figures
点击查看摘要
Abstract:VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations. Although recent extensions enable temporal 3D prediction, their future evolution remains weakly conditioned on driving intentions and actions, limiting their ability to model alternative action-dependent futures. We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving. First, we introduce an action--semantic conditioning mechanism that injects complementary driving semantics and ego-motion representations into the future-token stream, enabling different future geometry predictions for the same observed scene under alternative ego actions. Second, we develop a geometry--language--action bridge that adapts historical geometry, VLA semantic features, and maneuver and trajectory representations for joint conditioning of future geometry prediction. We evaluate future geometry prediction on NAVSIM, while conditioning ablations further examine the contributions of semantic and action information. Compared with the baseline, our method demonstrates competitive geometry prediction performance. Ablation studies further support the effectiveness of semantic and action conditioning. These results demonstrate the potential of semantic and action conditioning for controllable VGGT-based world prediction in autonomous driving.
161. 【2610.11160】LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs
链接:https://arxiv.org/abs/2610.11160
作者:Xiaobing Yu,Peijie Qiu,Jin Yang,Xuanzhao Dong,Weiwei Ma,Zhaoqi An,Xiaoqi Zhao,Xiaofeng Liu
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM)
关键词:LLMs requires storing, requires storing thousands, Lifelong editing, editing of LLMs, LLMs requires
备注: EMNLP 2026 Main Conference Long Paper
点击查看摘要
Abstract:Lifelong editing of LLMs requires storing thousands of edits after acquisition. A widely used family of approaches attaches one LoRA adapter per edit, which preserves behavior but grows linearly in storage. To address this challenge, we propose LadderEdit, a method that compresses each LoRA adapter after it is acquired. Each edit is first stored at low rank as a cheap sketch. We then check whether this sketch still satisfies the rewrite, generalization, and locality contract on probe prompts. Edits that pass keep the sketch; those that fail are promoted to a higher rank along a ladder until the contract is met. Because every edit retains some representation, coverage is maintained, and only hard edits consume more rank. Across ZsRE, CounterFact, and WikiBigEdit benchmarks on LLaMA-3-8B, Mistral-7B, and Qwen2.5-7B, LadderEdit tracks exact LoRA storage at 5.2x less memory and remains effective at 50,000 sequential edits.
162. 【2610.11153】CARE: Constrained Attention Refinement for Fine-Grained Visual Classification via Teacher-Student Distillation
链接:https://arxiv.org/abs/2610.11153
作者:Ruibo Wen,Hang Shao,Yiming Lei
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:recognize subtle local, visual classification requires, subtle local traits, classification requires models, Fine-grained visual classification
备注: 15 pages
点击查看摘要
Abstract:Fine-grained visual classification requires models to recognize subtle local traits while exposing the visual evidence behind their predictions. Class-specific attention pathways provide a natural basis for interpretable recognition, but their constrained prediction structure limits discriminative capacity and underuses intermediate representations from strong pretrained backbones. To address this problem, we propose CARE, a constrained attention refinement framework for interpretable fine-grained recognition via teacher-student distillation. CARE keeps the final prediction and explanation within a class-specific attention student, while introducing a training-only auxiliary query teacher that reads selected intermediate DINOv2 layers with learnable queries. The teacher fuses multi-level representations and transfers logit-standardized class-discriminative knowledge to the student. To further refine the explanation pathway, we design diversity and sparsity terms to regularize student attention heads, reducing redundancy and encouraging compact trait localization. Experiments on CUB, Oxford-IIIT Pet, Stanford Dogs, and Stanford Cars show that CARE achieves strong classification performance under an interpretable frozen-backbone setting, reaching 78.5% Top-1 accuracy on CUB. Faithfulness analysis with insertion and deletion metrics further indicates that the top-ranked attention regions retain class-relevant evidence for explanation.
163. 【2610.11149】A Unified Score Matching Paradigm for Video Anomaly Detection and Anticipation
链接:https://arxiv.org/abs/2610.11149
作者:Congqi Cao,Zhenhe Liang,Hanwen Zhang,Yifan Zhao,Qinyi Lv,Lingtong Min,Yanning Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:computer vision, Video anomaly, fundamental and safety-critical, safety-critical task, task in computer
备注:
点击查看摘要
Abstract:Video anomaly detection (VAD) is a fundamental and safety-critical task in computer vision. Recent generative approaches detect anomalies from a distributional perspective, but remain limited by local anomaly modes. Meanwhile, video anomaly anticipation (VAA), as a proactive extension beyond post-hoc detection, introduces additional challenges. In particular, the contrastive inference paradigm in VAD, which relies on ground-truth frames, is not applicable to VAA, hindering its development. To address these challenges, we propose a unified score-driven framework, termed Uni-DSM, based on denoising score matching (DSM), which models anomaly patterns through likelihood estimation and score functions over the learned data distribution. Within this unified framework, we adopt a shared noise-conditioned score transformer backbone with scene-dependent embeddings and motion-aware weighting for distribution-level modeling. Instead of introducing separate architectures, Uni-DSM unifies VAD and VAA through different inference and supervision paradigms built upon the same score-based formulation. For VAD, we instantiate an autoregressive denoising score matching (ADSM) mechanism, which progressively accumulates anomalous evidence via autoregressive denoising, enabling enhanced perception of local modes beyond visual cues. For VAA, we extend the same architecture by incorporating a lightweight auxiliary decoder and a novel self-distilled denoising score matching (SDSM) mechanism. By constructing supervision from output discrepancies instead of relying on unavailable future ground truth, our method achieves efficient training suitable or early anomaly anticipation. Extensive experiments on multiple benchmark datasets demonstrate state-of-the-art performance in both VAD and VAA while maintaining high efficiency, establishing a unified and scalable pipeline from anomaly detection to anticipation.
164. 【2610.11148】SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
链接:https://arxiv.org/abs/2610.11148
作者:Wenjie Liao,Xiaohui Song,Liangjie Zhao,Haonan Lu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Accurate page transcription, Accurate page, vision language models, transcription remains difficult, page transcription remains
备注:
点击查看摘要
Abstract:Accurate page transcription remains difficult for vision language models under limited input and training budgets. We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning. Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes. Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions. Only the OCR module is trained, while the backbone remains frozen. We derive the combined gradient to distinguish relative score optimization from direct supervision. Compared with SFT-2, SP-DR-3 reduces Vary-600K character error rate on both backbones. On Qwen3-VL-4B, it reduces character error rate by approximately 54 percent and improves DocVQA Average Normalized Levenshtein Similarity (ANLS) by 3.7 points. These results show the value of focusing self-play training on the discrepancies that remain after supervised fine-tuning.
165. 【2610.11144】Improving Image-Based Nutrition Estimation Through Multimodal Food-Item Verification and Recovery
链接:https://arxiv.org/abs/2610.11144
作者:Jingbo Yue,Bruce Coburn,Jinge Ma,Jui-Feng Chi,Fengqing Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
关键词:Single-image nutrition estimation, Single-image nutrition, fail silently, Single-image, portion estimation
备注: 5 pages, 2 figures, 3 tables. Submitted to IEEE ICASSP 2027
点击查看摘要
Abstract:Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity and whether each proposed 2D region supports portion estimation. One whole-image review uses these verification results to identify unresolved gaps and omitted foods, triggering at most one targeted recovery pass. Recovered regions are re-verified without access to the recovery prompt, then reconciled into a final item set for nutrition estimation. The framework requires no task-specific fine-tuning. Matched evaluation on common valid-output samples shows that item-level grounding improves mass accuracy across all tested settings and energy accuracy relative to an adapted retrieval baseline, with item-identity precision and recall also improving, while post-recovery visual coverage is assessed separately at inference time without ground-truth annotations.
166. 【2610.11140】ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis
链接:https://arxiv.org/abs/2610.11140
作者:Weiwei Ma,Xiaobing Yu,Peijie Qiu,Jin Yang,Zhaoqi An,Xuanzhao Dong,Xiaoqi Zhao,Xiaofeng Liu
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:resolve diagnostic uncertainty, clinicians escalate, escalate from cheap, cheap to costly, costly tests
备注: EMNLP 2026
点击查看摘要
Abstract:Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.
167. 【2610.11138】IntrinSync: Joint Intrinsic Decomposition and Reciprocal Rendering
链接:https://arxiv.org/abs/2610.11138
作者:Zheng Gu,Rui Huang,Xilu Zhang,Jingbo Zhang,Min Lu,Zhida Sun,Dani Lischinski,Daniel Cohen-Or,Hui Huang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:inherently interdependent, Inverse rendering decomposes, intrinsic, rendering, image
备注: 21 pages
点击查看摘要
Abstract:Inverse rendering decomposes an image into intrinsic properties such as appearance, illumination, geometry, and material, yet these properties are inherently interdependent. A reliable decomposition should produce intrinsic maps that are not only individually plausible, but also mutually compatible in explaining the image. However, existing methods either model intrinsic channels in isolation or treat inverse and forward rendering as separate processes, leaving the interdependence underexploited. In this paper, we introduce IntrinSync, a unified framework that captures this interdependence through joint-channel modeling and reciprocal inverse-forward rendering. At the channel level, we jointly decompose an input RGB into albedo, shading, surface normal, roughness, and metallic maps through a 1-to-N mapping, enabling information exchange across channels throughout generation. At the process level, we establish inverse-forward reciprocity through a dual cycle-consistent objective that aligns corresponding predictions across a closed loop. Experiments on three datasets demonstrate that our method achieves competitive intrinsic estimation and forward rendering performance, improving coherence and physical consistency. Beyond decomposition, IntrinSync provides a physically grounded interface for image editing, allowing intrinsic properties to be explicitly manipulated and rendered back into RGB images.
168. 【2610.11117】MCL: Meta Convolution Layer
链接:https://arxiv.org/abs/2610.11117
作者:Naim Reza,Md Al Amin,Ho Yub Jung
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:enhances convolutional neural, Dynamic convolution enhances, mixture size grows, Meta Convolution Layer, convolutional neural networks
备注:
点击查看摘要
Abstract:Dynamic convolution enhances convolutional neural networks (CNNs) by adapting kernels to input content, but it expresses the effective kernel as a linear mixture of a small number of basis kernels, which limits expressivity and complicates optimization as the mixture size grows. In this work, we revisit dynamic convolution from a functional perspective and propose the Meta Convolution Layer (MCL), which directly models the convolutional kernel as an input-conditioned function W(x) realized via a high-order polynomial expansion. Leveraging nested residual blocks inspired by deep polynomial networks, MCL implements a structured polynomial meta-network that generates a single input-adaptive kernel, thereby decoupling representational power from the explicit number of mixture kernels and alleviating training instability. MCL is a plug-in addition with standard convolutions and can be seamlessly integrated into both CNN and transformer backbones. Experimental evaluation shows that adding MCL improves the Top-1 accuracy of Resnet- 18, Resnet-50 and ResNet-101 by 6.61%, 3.42% and 3.05% on the ImageNet dataset. Moreover, the proposed method significantly boosts the accuracy of Resnet and Wide-Resnet variants on CIFAR-10 and CIFAR-100 datasets. Additionally, the proposed method outperforms previous methods on fine-grained visual classification tasks using Swin and ViT backbones. These results demonstrate that high-order polynomial kernel generation is a powerful and scalable alternative to linear mixture based dynamic convolution.
169. 【2610.11113】DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models
链接:https://arxiv.org/abs/2610.11113
作者:Mengping Dong,Jinbao Li,Fei Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:varied perception tasks, Pre-trained vision-language models, Pre-trained vision-language, generalization remains non-trivial, sacrificing generalization remains
备注: 20 pages, 6 figures, 11 tables. Accepted to ECCV 2026
点击查看摘要
Abstract:Pre-trained vision-language models excel across varied perception tasks, but adapting them to novel downstream settings without sacrificing generalization remains non-trivial. Existing parameter-efficient prompt learning method often yields inconsistent representations and fails to account for semantic distribution shifts. In this work, we present DiscoVL, a disentangled cross-modal representation learning framework that couples orthogonal adversarial regularization with structured cross-modal alignment for vision-language models. To address the insufficient cross-modal interaction, our DiscoVL designs a multi-branch low-rank residual aligner that decomposes representations into subspaces and enables bidirectional cross-modal feedback between visual and textual streams at each layer. Furthermore, while conventional triplet constraints overfit features to class centroids, we design an orthogonal regularization for adversarial triplet loss, which prevents centroid collapse and substantially boosts generalization. Evaluations on 15 benchmarks demonstrate that DiscoVL delivers consistent improvements over state-of-the-art methods for base-to-novel generalization, cross-dataset evaluation, and few-shot learning
170. 【2610.11112】False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
链接:https://arxiv.org/abs/2610.11112
作者:Zeyu Ye,Yanchun Li,Sibei He,Meng Xie,Hangtao Zhang,Xianlong Wang,Li Zeng,Jiahao Chen,Yichen Wang,Junhui Wang,Ziqi Zhou
类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:Image-generation models, natural-looking visual artifacts, textbook pages, reports and textbook, models
备注: 29 pages, 19 figures, 6 tables. Project website: [this https URL](https://github.com/Ye-ze-yu/EpiReal-Bench)
点击查看摘要
Abstract:Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages. Yet, the same capability introduces a new risk: these models can just as easily fabricate visual misinformation. Even commercial models (e.g., GPT-Image-2) readily produce it. Curiously, we find that these models can recognize a claim as false when asked, yet still render that very claim as credible visual evidence. This discrepancy points to a blind spot in current alignment: safeguards judge what an image shows, not what it asserts; however, existing red-teaming benchmarks target conventional harmful content, such as violent or explicit imagery, and say little about where the alignment boundaries lie for visual misinformation, especially in commercial models. To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats. We further introduce EpiReal-Attack, a skill-guided black-box optimization framework that uses Pareto-based selection and multimodal feedback to identify commands that bypass alignment safeguards while preserving visual realism, textual legibility, and semantic fidelity. Experiments on four commercial models reveal that more than 70% of false-claim prompts elicit images that faithfully depict the corresponding misinformation, and EpiReal-Attack pushes this rate to 95%. Most worryingly, these models are only a click away, and their outputs are cheap to spread yet hard to disbelieve, leaving this dimension of alignment largely unguarded.
171. 【2610.11105】KCAM: Text and Keyframe to Camera Trajectory Generation
链接:https://arxiv.org/abs/2610.11105
作者:Haozhe Yang,Zhiyang Dou,Zekai Gu,Cheng Lin,Wenping Wang,Yuan Liu,Taku Komura
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:controllable camera motion, scene understanding, Residual Vector Quantizer, Keyframe-conditioned CAMera-motion synthesis, AI-assisted cinematography
备注: 20 pages, 7 figures. Paper accepted to NeurIPS2026 (submission number 15461)
点击查看摘要
Abstract:Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at this https URL.
172. 【2610.11104】Skeleton-Guided Progressive Test-Time Adaptation for Thin Curvilinear Structures
链接:https://arxiv.org/abs/2610.11104
作者:Boa Jang,JunGyu Lee,Gwanho Lee,Jinwook Choi,Young-Gon Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:Accurate segmentation, thin curvilinear structures, real-world applications, road extraction, segmentation of thin
备注: 9 pages, 6 figures
点击查看摘要
Abstract:Accurate segmentation of thin curvilinear structures is vital for various real-world applications, from vessel analysis to road extraction. Yet their intricate geometry makes even minor pixel-wise errors enough to break the global topology, and this structural fragility turns severe domain shifts into catastrophic failures. The difficulty is most acute under cross-modality gaps, where the imaging process itself differs fundamentally between source and target. While test-time adaptation (TTA) offers a practical source-free remedy, existing methods adapt feature statistics and confidence, neither of which constrains connectivity, and thus degrade under such extreme gaps. To address this, we propose Skeleton-Guided Progressive Test-Time Adaptation (SGP-TTA). Progressive Batch Normalization (ProgBN) shifts normalization from frozen source statistics toward current target estimates under a sample-count schedule, so that the source-target balance follows the stage of adaptation rather than a fixed coefficient. Consensus Skeleton Recall (CSR) then derives a structural target from geometrically aligned multi-view predictions and updates only the BN affine parameters to preserve connected structures. Extensive experiments show that SGP-TTA consistently outperforms existing TTA methods in topological connectivity, with the largest margins under cross-modality shift. The project page is available at this https URL.
173. 【2610.11096】Continuous Ground-Truth Construction and a Recovery Policy for Air--Water Robotic Tracking
链接:https://arxiv.org/abs/2610.11096
作者:Jiangong Xiao,Zhe Sun,Kanzhong Yao,Yuanbo Bi,Haofei Zhao,Ruixuan Hu,Guan Huang,Xuelong Li
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:temporarily invalidate observations, challenged by splashes, air-water interface, interface is challenged, temporarily invalidate
备注: 8 pages, 6 figures. Submitted to IEEE International Conference on Robotics and Automation (ICRA 2027)
点击查看摘要
Abstract:Visual tracking across the air-water interface is challenged by splashes, bubbles, refraction, reflections, and abrupt appearance changes that can temporarily invalidate observations. This setting poses two coupled difficulties: first, for evaluation, image-only annotation cannot reliably describe the target's physical location during visual blindness; second, for online tracking, corrupted observations can contaminate motion estimates and appearance templates. We address the first difficulty with a construction pipeline that synchronizes camera frames with motion-capture poses, projects known target geometry, corrects underwater projection with a medium-gated residual, and subjects the annotations to manual review. This yields an evaluation-only cross-medium test set of 22,346 frames. We further introduce a Cross-Medium Recovery Policy (CMRP) centered on confidence-triggered template selection. It supplies MixFormerV2 with the fixed initial template, a window-best pre-trigger template, and a trigger-frame Kalman-guided image crop, together with their associated weights, without retraining the visual backbone. In the accuracy evaluation, CMRP achieves 49.90 Macro Success AUC, 2.95 points above MixFormerV2 Official. On selected cross-medium transition and occlusion-recovery intervals, CMRP increases MixFormerV2 tracking coverage from 47.91\% to 50.43\% relative to Official updating, while mean loss-to-recovery latency over successfully recovered videos decreases from 55.3 to 49.3 frames.
174. 【2610.11086】Contrast Enhancement or Noise Reduction? On Improving Cervical Cancer Classification
链接:https://arxiv.org/abs/2610.11086
作者:Ach Khozaimi,Ulfatun Nahdhiyah
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:image preprocessing algorithms, mortality worldwide, preprocessing algorithms, Purpose, image preprocessing
备注:
点击查看摘要
Abstract:Purpose: Cervical cancer is one of the leading causes of mortality worldwide. Deep learning has shown promising performance in medical image classification. The influence of image preprocessing algorithms on classification performance remains insufficiently investigated in the literature. This research aims to evaluate the impact of image preprocessing algorithms on the performance of CNNs for Pap smear image classification. Methods: Three CNN architectures (ResNet-34, MobileNet-V2, and DenseNet-121) were trained and evaluated using the SIPaKMeD dataset. Two preprocessing algorithms were applied: the PMD filter for noise reduction and CLAHE for contrast enhancement. The model performance was assessed using a confusion matrix. Results: Preprocessing improved the classification performance of all models. CLAHE significantly increased the accuracy of ResNet-34 from 76.73% to 84.16% and DenseNet-121 from 76.73% to 84.16%. The PMD filter yielded limited improvement and slightly reduced the MobileNet-V2 performance. Novelty: This research provides a systematic comparison of contrast enhancement and noise reduction techniques across CNN architectures. This research demonstrates that contrast enhancement is more effective than noise reduction in improving CNN performance. The research provides new pipelines for improving cervical cancer classification.
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2610.11086 [cs.CV]
(or
arXiv:2610.11086v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2610.11086
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Related DOI:
https://doi.org/10.15294/sji.v13i3.48099
Focus to learn more
DOI(s) linking to related resources
Submission history From: Ach Khozaimi [view email] [v1]
Thu, 8 Oct 2026 01:57:18 UTC (774 KB)
175. 【2610.11070】No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
链接:https://arxiv.org/abs/2610.11070
作者:Yu Han,Dejan Markovic,Alexander Richard,Wojciech Zielonka,Akshay Venkatesh,Cheng-hsin Wuu,Michael Zollhoefer
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:underpins real-time avatars, embodied virtual agents, Audio-driven facial animation, animation underpins real-time, facial animation underpins
备注: Project website: [this https URL](https://wojciechzielonka.com/facegan/)
点击查看摘要
Abstract:Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents. And it must run online: each frame emitted from audio observed up to the current time, at interactive rates. Recent progress is dominated by diffusion models, which need many network evaluations per sample and are therefore a poor fit for streaming. We argue the cost is unnecessary in this domain. Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices. The obstacle is not capacity but stochastic structure. We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation. We proposed FaceGAN, which dissolved the limitation by shaping the noise pathway acausally. Because the driving noise is synthetic, its future can be sampled now, so the audio-to-expression path stays causal, and the model supports fully causal operation. FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality. Being feed-forward with bounded attention windows, it generates indefinitely without drift.
176. 【2610.11067】Diffusion Meta-Prompting and Steering for Generalizable Foundation Model Adaptation
链接:https://arxiv.org/abs/2610.11067
作者:Deepak Sridhar,Yi Li,Kartikeya Bhardwaj,Shuangjun Liu,Taotao Jing,Yuan Li,Shuai Zhang,Jiancheng Lyu,Dashan Gao,Nuno Vasconcelos
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:adapting foundation models, DMP, adapting foundation, typically task-specific, task-specific and fail
备注: Accepted to NeurIPS 2026. Project page: [this https URL](https://deepaksridhar.github.io/dmp.github.io/) . Code: [this https URL](https://github.com/DeepakSridhar/dmp)
点击查看摘要
Abstract:Prompt learning is a popular method for adapting foundation models, but learned prompts are typically task-specific and fail to generalize to new classes, domains, or compositions of tasks. In this paper, we introduce a Diffusion Meta-Prompt (DMP) model , a framework that models the distribution of learned prompts using diffusion models. Given a repository of previously learned prompts, DMP is trained and sampled without access to the original task examples or task losses, and synthesizes new prompts conditioned on natural language task descriptions. To improve the sampling stability, we introduce a test-time steering strategy for DMP, which uses the best training-selected prompt in the repository as a latent anchor during diffusion sampling, without retraining the DMP or accessing test classes. DMP improves generalization across classification, retrieval and text-to-image generation tasks, supports concept composition and negative prompting without explicit training. It reduces storage and inference costs by over 90% compared to prompt retrieval methods. For composite classification, DMP achieves upto 2.0% average gain over prior meta-learning methods across 55 pairs of datasets with gains as high as 8.5% on specific pairs such as Eurosat and Flowers. DMP also enhances cross-task generalization with ~2-9% improvement for hierarchical classification task. We further provide a theoretical guarantee bounding the expected task loss of prompts sampled from a DMP. Code is available: this https URL
177. 【2610.11060】AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding
链接:https://arxiv.org/abs/2610.11060
作者:Tianhui Cai,Xinglong Sun,Chao Fang,Zhenxin Li,Rui Song,Jose M. Alvarez,Yunxiang Mao,Jiaqi Ma,Langechuan Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:recently improved autonomous, improved autonomous driving, recently improved, improved autonomous, learning future scene
备注:
点击查看摘要
Abstract:World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation. Most existing approaches model the future primarily through RGB appearance, and recent works have begun to incorporate geometric prediction to improve spatial understanding. However, dense geometry describes the spatial layout of the entire scene without indicating which parts are most relevant to the ego vehicle's action. For driving, the model must also identify and anticipate where it can safely move and which regions may pose collision risks. Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure. We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure. In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents. On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.
178. 【2610.11058】Refine Connections, Close the Gap: A Reliable Enhancement Framework for Driving Scene Topology
链接:https://arxiv.org/abs/2610.11058
作者:Xiaoqi Wang,Dingyi Zhaung,David Paz,Wenbin He,Yucai Bai,Peng Zhou,Rui Zhang,Jinhua Zhao,Liu Ren
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:safe path planning, understanding scene topology, autonomous driving, understanding scene, motion control
备注:
点击查看摘要
Abstract:In autonomous driving, understanding scene topology - the connectivity between lanes and traffic elements - is critical for safe path planning and motion control. While current methods excel at detecting individual map elements, their connectivity reasoning often falls short of its theoretical potential, leaving a significant performance gap relative to the theoretical upper-bound achievable given the underlying detections. Furthermore, the decision-ready topology graphs passed to downstream tasks often remain unreliable. Current approaches typically derive connectivity by thresholding continuous topology scores; however, these scores often fail to reflect the true logical likelihood of connectivity, resulting in false positives or missing connections. Existing benchmarks further overlook this issue by primarily evaluating continuous metrics, rather than assessing the discrete connectivity required for decision-making. To bridge these gaps, we propose TopoEnhance, a novel topology enhancement framework designed to unlock the latent potential of existing methods and improve the reliability of decision-ready topology. We formulate topology enhancement as a denoising-based reconstruction process, where the model learns to recover structural consistency from stochastically corrupted ground-truth graphs. This formulation enables the model to resolve logical inconsistencies and rectify unreliable connections, producing robust discrete topology graphs that closely approach theoretical maximum performance. Extensive experiments across different baselines show that TopoEnhance consistently improves both continuous topology metrics (TOP score), and discrete connectivity measured by our adapted Topology Jaccard Similarity (TJS) metric. As a flexible, source-agnostic framework, TopoEnhance delivers substantial gains across diverse state-of-the-art baselines without requiring retraining.
179. 【2610.11057】Learning What to Trust in Multimodal Learning under Noisy Supervision
链接:https://arxiv.org/abs/2610.11057
作者:Jiashuo Zou,Xiaobo Xia
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Multimodal classification processes, accurate predictions, classification processes, processes and relates, multiple modalities
备注:
点击查看摘要
Abstract:Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions. However, existing methods typically rely on high-quality ground-truth labels, which are difficult to obtain in real-world scenarios. While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully. This motivates us to build a more reliable noise detector in multimodal learning. To this end, we theoretically analyze the relationship between representation structure and noise detection capability. Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection. Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class. Within each trusted space, REFINE measures the alignment between each instance representation and the discriminative eigenvectors. It then combines the subsets selected from these spaces. The combined set provides cleaner supervision for updating the multimodal classifier, thereby reducing the influence of mislabeled examples during training and improving model generalization. Extensive experiments across diverse tasks demonstrate REFINE's superiority compared to baseline methods. The source code will be publicly available.
180. 【2610.11049】SatFix: Absolute Visual Localization of UAVs in Satellite Maps from a Single Oblique Image
链接:https://arxiv.org/abs/2610.11049
作者:Jiarui Zeng,Kun Shi,Chiman Vong,Zhedong Zheng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:provided geo-referenced satellite, short multi-view clip, study absolute metric, absolute metric UAV, UAV
备注:
点击查看摘要
Abstract:We study absolute metric UAV localization within a provided geo-referenced satellite region, recovering continuous map position and viewing heading from a single oblique image or a short multi-view clip. Existing cross-view geo-localization methods retrieve the most similar satellite tile from a gallery and report Recall@K, but retrieval depends on gallery sampling, provides no heading estimate, and returns a tile index rather than a continuous coordinate. We propose SatFix, a feed-forward UAV--satellite localization framework built on VGGT-$\Omega$. Satellite-grid features act as queries that aggregate UAV visual evidence, and two lightweight heads regress a 3-DoF pose in the satellite-map frame: continuous 2D position and heading. SatFix requires no explicit 3D map, rendered bird's-eye image, auxiliary sensor, or test-time pose alignment. A single model supports both single- and multi-view inputs, with trajectory constraints used during multi-view training. For metric evaluation, we introduce University-Metric, where satellite imagery is re-collected over a region up to 10.7$\times$ longer on a side (about 114$\times$ the ground area) than the original University-1652 tiles, with continuous position and heading labels for the original UAV tours. With one UAV view, SatFix localizes 52.08% of test frames within 50 m and 17.34% within 10 m, with median position and heading errors of 45.66 m and $20.81^\circ$, respectively. Inference takes under 0.1 s per single-view query on an NVIDIA RTX 4090. With nine UAV views, the median position error falls to 21.96 m and the median heading error to $8.73^\circ$. Compared with a fine-tuned VGGT-$\Omega$ baseline, SatFix reduces median position error by 34.0% and nine-view median heading error from $25.43^\circ$ to $8.73^\circ$.
181. 【2610.11039】Rendering-Free Lookahead for Question-Guided Active Vision
链接:https://arxiv.org/abs/2610.11039
作者:Koya Sakamoto,Daichi Azuma,Shuhei Kurita,Naoya Chiba,Yusuke Iwasawa,Yutaka Matsuo,Taiki Miyanishi
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:Active robot vision, reveal task-relevant information, Active robot, robot vision requires, vision requires controlling
备注: Project page: [this https URL](https://k0uya.github.io/rfl-proj/)
点击查看摘要
Abstract:Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions that expose the visual evidence needed to answer the question. Although vision-language models (VLMs) can interpret observed images, selecting such motions requires anticipating the usefulness of unseen views. We quantify this usefulness as answerability, a VLM's estimate that a view suffices to answer the question, and present Rendering-Free Lookahead (RFL), a viewpoint-selection policy that ranks candidate camera motions by predicted future answerability. RFL transfers visual lookahead from deployment to offline training. At training, a privileged teacher renders candidate future views in 3D Gaussian Splatting (3DGS) scenes and uses a frozen VLM to compute one- and two-step answerability targets. Through two-stage distillation, a student learns to predict these action values from the question, recent visual observations, and a candidate camera motion. At deployment, RFL uses these predicted values to select camera motions without rendering future views. On 377 E3VS-Bench test episodes in unseen environments, RFL improves the mean judge score by 43\% over a direct-action baseline using the same VLM. These results support learning camera-control policies from privileged visual lookahead for viewpoint-dependent question answering.
182. 【2610.11037】ransforming Image Editors into Video Editors
链接:https://arxiv.org/abs/2610.11037
作者:Feng Wang,Zijie Li,Ceyuan Yang,Alan Yuille,Peng Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:impressive semantic understanding, Recent image editing, achieved impressive semantic, video editing, editing remains substantially
备注: In NeurIPS 2026
点击查看摘要
Abstract:Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video. Code is available at this https URL.
183. 【2610.11023】Expression-Diverse References for Identity-Preserving Video Generation
链接:https://arxiv.org/abs/2610.11023
作者:Tianwen Fu,Wenbin Teng,Gonglin Chen,Junyi Ouyang,Haolin Xiong,Yajie Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Identity-preserving video generation, Identity-preserving video, aims to maintain, Identity-preserving, reference
备注: 9 pages, 8 figures
点击查看摘要
Abstract:Identity-preserving video generation aims to maintain a subject's identity while synthesizing realistic videos. Yet a single reference portrait captures the subject's appearance under only one facial configuration. As expressions change, facial appearance can vary in highly identity-specific ways, leaving the subject's appearance under unseen expressions underdetermined by the reference alone. This expression-dependent variation also complicates evaluation: similarity to a neutral reference may decrease under strong expressions even for real images of the same person. We investigate this limitation from both generation and evaluation perspectives. First, we quantify how face-recognition similarity varies with expression intensity using controlled photographs and MEAD videos. We then construct a compact yet expressive reference gallery that captures diverse expression-dependent facial configurations. Matching against this gallery provides a more robust measure of identity similarity under expressive motion. To further expose performance degradation with expression intensity, we report identity similarity separately for mild, intense, and extreme expressions. For generation, we extend Stand-In to condition on our expression-diverse reference sets and develop a data-curation pipeline that extracts consistent yet diverse face crops from training videos. In practical settings where only a single portrait is available, we construct the reference set by synthesizing additional expressions with a pretrained facial reenactment model. On our controlled benchmark, both real and synthesized reference sets outperform the evaluated baselines in identity similarity across all three expression-intensity regimes, with the largest improvements for extreme expressions.
184. 【2610.11019】Mid-Training Language Models on Raw Video
链接:https://arxiv.org/abs/2610.11019
作者:Jaedong Hwang,Xiaoqian Shen,Ernie Chang,Changsheng Zhao,Chong Zhou,Saksham Suri,Qi Qian,Zechun Liu,Lemeng Wu,Qinsi Wang,Raghuraman Krishnamoorthi,Wei Wen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Multimodal large language, Multimodal large, large language models, existing language model, language model
备注:
点击查看摘要
Abstract:Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded into continuous visual tokens, and the language model learns to predict the next visual token. We mid-train Qwen3-1.7B on raw clips from YT-Temporal-1B and then apply the same image-text instruction tuning to it and to the model without mid-training, so that the two differ only in mid-training. The mid-trained model scores 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks, spanning perception, document, and chart tasks. Text performance is preserved even though mid-training includes no text, with an average of 48.9 across 14 text benchmarks compared with 48.0 for the model without mid-training. Analyses across training show that the image and video gains emerge within 30% of training and plateau thereafter, varying by less than 0.5 points. Predicting captions fails to outperform next-visual-token prediction, demonstrating that video mid-training can remain purely self-supervised without the computational overhead or labeling noise of automated captioning.
185. 【2610.11011】PCAsplat: Gaussian Splatting with Local PCA Regularization
链接:https://arxiv.org/abs/2610.11011
作者:Vitor Matias,Filipe Nascimento,Kiyohiro Nakayama,João Paulo Lima,Márcus Lobo,Gordon Wetzstein,Leonidas Guibas,Afonso Paiva,Tiago Novello
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:posed images, flexible representation, Gaussian splatting, Gaussian, Gaussians
备注:
点击查看摘要
Abstract:Gaussian splatting has emerged as a flexible representation for 3D reconstruction from posed images. However, existing methods are optimized primarily using rasterization-based losses, which supervise a splat only when it contributes to sampled camera rays. Gaussians that are occluded or contribute little to the sampled view therefore receive weak or no geometric gradients and may drift away from the underlying surface, producing undesired floaters. We introduce PCAsplat, a geometry-aware regularization framework for Gaussian splatting based on differentiable local principal component analysis (PCA). Our PCA regularizer acts directly on neighborhoods of Gaussian centers and can therefore update Gaussians that do not contribute to the current training view. We regularize the PCA eigenvalues to encourage Gaussians to move to the underlying surface with isotropic tangent-plane coverage. We also align each Gaussian normal with the PCA-estimated neighborhood normal to enforce consistent orientation. Experiments on DTU, Tanks and Temples, and NeRF Synthetic show that the splats produced by PCAsplat better approximate samples of the reference surface while substantially reducing undesired floaters. These surface-aligned splats enable downstream geometry-processing tasks, including point cloud segmentation, and direct Poisson reconstruction. Additionally, PCAsplat remains competitive under conventional novel view synthesis and mesh extraction tasks. Code will be released.
186. 【2610.10991】Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
链接:https://arxiv.org/abs/2610.10991
作者:Ian de Holanda Cavalcanti Bezerra,Vivek Trivedy,Lucas Pascotti Valem,Longin Jan Latecki
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:semantic descriptor extracted, Image retrieval methods, tokens, CLS, single global semantic
备注:
点击查看摘要
Abstract:Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at this https URL.
187. 【2610.10990】Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models
链接:https://arxiv.org/abs/2610.10990
作者:Hong Huang,Chenhongyi Yang,Junzhe Sun,Animesh Sinha,Wuyang Chen,Yifan Jiang
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:large language models, diffusion large language, multimodal diffusion large, Unified multimodal diffusion, iterative decoding requires
备注:
点击查看摘要
Abstract:Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.
188. 【2610.10989】Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models
链接:https://arxiv.org/abs/2610.10989
作者:Jialin Zhu,Xing Liu,Feixiang He,He Wang
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:Distribution Matching Distillation, one-step generation recently, one-step generative model, exploring effective one-step, generative model continuously
备注:
点击查看摘要
Abstract:Researchers are exploring effective one-step generative model continuously, and, Drifting Models (Deng et al., 2026), demonstrate great potential in one-step generation recently. There are works that reveal the connection between Diffusion Flow Style Generative Models (DFSGMs) (Ho et al., 2020; Song et al., 2020a;b; Lipman et al., 2022; Liu et al., 2022) and Drifting Models (Li Zhu, 2026; Lai et al., 2026; Turan et al., 2026). But no one has yet established a precise correspondence between the Drifting Model and the widely used distillation method- Distribution Matching Distillation (DMD/DMD2) (Yin et al., 2024b;a) to the best of our knowledge, even though their optimization objective formulas are virtually identical. In this paper, we prove that by converting the velocity-field / noise-field from the pre-trained DFSGMs into the attraction force field in Drifting Models and estimating the repulsion force field from the generative distribution, training the Drifting Model is naturally equivalent to the Distribution Matching Distillation. With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
189. 【2610.10984】Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without Training
链接:https://arxiv.org/abs/2610.10984
作者:Hong Huang,Yuqiu Liu,Chenyu You,Daniel Martin,Chuhang Zou,Wuyang Chen
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:decouples physical reasoning, training-free framework, framework for physics-aware, decouples physical, physical reasoning
备注:
点击查看摘要
Abstract:We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis. Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators. We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping. This plug-and-play design is compatible with current video foundation models. We further introduce a benchmark for fluid-object interaction video generation. Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality. In a human preference study, raters favor Fluid-Gen-Zero in 55.1%-74.4% of same-backbone comparisons across three backbones, and in 90.4%-94.2% of comparisons against simulation-based methods. Code and data will be released upon acceptance.
190. 【2610.10960】LVSPM: Long Sequence View Synthesis and Pose Estimation Model
链接:https://arxiv.org/abs/2610.10960
作者:Xi Chen,Yachi Zhang,Linghao Chen,Minghua Liu,Hao Su,Zexiang Xu,Xiaoshuai Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:jointly estimates camera, uncalibrated image collections, estimates camera poses, jointly estimates, estimates camera
备注: ECCV 2026. Project Page: [this https URL](https://burningdust21.github.io/Projects/LVSPM/)
点击查看摘要
Abstract:We present LVSPM, a generalizable model that jointly estimates camera poses and synthesizes novel views from uncalibrated image collections. Trained with only RGB images and pose supervision, LVSPM avoids dense 3D ground truth and employs test-time training (TTT) layers to scale seamlessly to hundreds of input views. On RealEstate10k, Co3Dv2, and DL3DV, LVSPM surpasses VGGT in pose estimation across 16-256 views, with especially large margins at strict thresholds. For novel view synthesis under a practical protocol where more views cover larger scenes, LVSPM achieves state-of-the-art pose-free quality---surpassing even pose-dependent models in PSNR---and still maintains high quality as scene scale grows, while baselines collapse. The code is available at this https URL .
191. 【2610.10959】GPU-Accelerated Computation of Persistent Homology for Topological Analysis of Image Data
链接:https://arxiv.org/abs/2610.10959
作者:Fan Wang,Hubert Wagner,Rezaul Chowdhury,Chao Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:recent years, rapid adoption, remains a major, major bottleneck, persistent homology
备注: Accepted for publication in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Includes supplementary material (8 pages) after the main paper. Code: [this https URL](https://github.com/seravee08/GPU-Computation-of-Persistent-Homology-for-Image-Data)
点击查看摘要
Abstract:In recent years, persistent homology has seen rapid adoption in deep learning, yet its computation remains a major bottleneck in network training. This paper introduces TopoGPU, a GPU streaming pipeline that computes persistence diagrams of cubical complexes induced by 2D and 3D images. TopoGPU streams the input image chunk by chunk, processing each chunk with massively parallel GPU kernels on a grid of GPU blocks; the resulting boundary relations are accumulated in host memory, where the CPU performs the boundary matrix reduction. TopoGPU introduces a stratification-aware discrete Morse matching that provably preserves persistent homology under streaming, together with a parallel topological sorting algorithm and a parallel V-path parity algorithm for deriving Morse boundaries on the GPU. TopoGPU outperforms Cubical Ripser, a state-of-the-art method for persistent homology computation, on every benchmark evaluated, achieving an average end-to-end speedup of 53.24x and a maximum of 198.01x. We further integrate TopoGPU into a topology-preserving deep network, demonstrating that it substantially reduces the cost of persistent homology computation during network training. TopoGPU is open source, with pre-built binaries, Google Colab notebooks, and Docker images available at the project's GitHub page: this https URL.
192. 【2610.10945】GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior
链接:https://arxiv.org/abs/2610.10945
作者:Ali Benlalah,Sepehr Johari,Patricia Vitoria,Armin Kappeler,Artem Sevastopolsky,Alexander Jung,Gabriele Fanelli,Kevin Mader,Manuel Breitenstein,Claudia Plüss,Jan Rüegg,Simon Biland,Thomas Etterlin,Dmitry Kostiaev,Mathias Deschler,Brian Amberg,Sebastian Martin
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Real-time Gaussian Head, Gaussian Head Animation, Large-scale Reconstruction Prior, present GHARP, human heads
备注: Accepted to ACCV 2026. 35 pages (14 main + references + 15 pages supplementary), 14 figures, 17 tables
点击查看摘要
Abstract:We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
193. 【2610.10912】When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs
链接:https://arxiv.org/abs/2610.10912
作者:Jasper Gerigk,Kenzo Aspuru-Takata,Chin-Hsuan Wu,Mohammad Mohammadi,Shuhong Zheng,Igor Gilitschenski
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:prevalent issue, Shortcut learning, visual shortcut learning, visual, visual shortcut
备注:
点击查看摘要
Abstract:Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.
194. 【2610.10893】Less from More: Reinforcing Sparse Video Reasoning from Dense References
链接:https://arxiv.org/abs/2610.10893
作者:Wenfang Sun,Yingjun Du,Cees G. M. Snoek
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Video-language models commonly, models commonly assume, Video-language models, temporal observations lead, models commonly
备注:
点击查看摘要
Abstract:Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Notably, SAVER is trained only on 1,250 randomly sampled temporal grounding examples, without using any video question answering annotations. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.
195. 【2610.10889】SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
链接:https://arxiv.org/abs/2610.10889
作者:Raja Kumar,Rajat Koner,Ritwick Chaudhry,Zhuowei Li,Nishant Sankaran,Yifan Xing
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:language reasoning, reasoning requires, extract relevant, relevant and accurate, accurate information
备注:
点击查看摘要
Abstract:Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation. Although trained in two phases, trained policy is evaluated like GRPO model, with a single CoT call at inference time. Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B. Evaluating each capability using an oracle based diagnostic shows that answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both VR and LR.
196. 【2610.10859】Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
链接:https://arxiv.org/abs/2610.10859
作者:Gaurav Patel,Jun Fang,Greg Ver Steeg,Qiang Qiu,Sravan Sripada
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:enable fast inference, fast inference, FSD models, deployed to enable, enable fast
备注: Accepted at NeurIPS 2026
点击查看摘要
Abstract:Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.
197. 【2610.10823】Velocity Scaling in Flow Matching
链接:https://arxiv.org/abs/2610.10823
作者:Youssef Saied,François Fleuret
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:learned flow-matching velocity, flow-matching velocity field, learned flow-matching, recently shown, time
备注: 41 pages, including appendix
点击查看摘要
Abstract:Scaling a learned flow-matching velocity field $v_\theta$ by a gain $\gamma(t)$ was recently shown to greatly improve generation quality. Prior work argued that velocity fields trained with mean-squared error (MSE) systematically underestimate velocity magnitude and that scaling corrects this error. We show that MSE training does not create a velocity-magnitude deficit. We find instead that velocity scaling reduces population time lag: sampled states at model time $t$ resemble training states from an earlier time. Velocity scaling and moving model time back are two ways to address this population time lag. Across architectures and model sizes, measuring population time lag and using it to select a gain greatly improves generation quality, reducing FID from 28.0 to 12.2 (estimated by linear interpolation between FID measurements at neighboring gains) on ImageNet-256 at NFE 25 without guidance.
198. 【2610.10787】NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
链接:https://arxiv.org/abs/2610.10787
作者:Gengze Zhou,Yicong Hong,Jiazhao Zhang,Xunyi Zhao,Jian Zhou,Zixing Lei,Zun Wang,Chongyang Zhao,Xionghui Chen,Stephen Gould,Anton van den Hengel,Qi Wu
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:long-horizon agentic reinforcement, agentic reinforcement learning, Language models trained, express precise actions, Language models
备注: 36 pages, 14 figures. Project page: [this https URL](https://metacognitionai.github.io/NavGPT3/)
点击查看摘要
Abstract:Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching. Beneath it, our action policy NavGPT VLA, trained on 19.28M examples, allocates visual tokens using codec allocation, in proportion to scene change; its 8B model alone reaches 74.51 SR on R2R-CE and leads RxR-CE with 78.19 SR. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs. 90.4 SR) and path fidelity (78.47 vs. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human. We comprehensively ablate the harness design and the interaction between the two models, showing how tools and the action policy shape the path from language-model reasoning to physical control: when NavGPT VLA executes the route, the reasoning loop shortens and the system's minimum reaction time falls from 3-19 s per language-model decision to 0.5-1 s per action-policy step (1-2 Hz). These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. We will release all models, code, and evaluation records.
199. 【2610.10782】VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
链接:https://arxiv.org/abs/2610.10782
作者:Meng Lu,Ligeng Zhu,Olivia Xiao,Yuchen Zhuang,Zihan Wang,Kuncheng Wu,Bangya Liu,Yu Wang,Charles Fleming,Wenqi Shi,Xuan Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Reinforcement learning, Toggle, Toggle Hugging Face, Bibliographic Explorer Toggle, Explorer Toggle Bibliographic
备注:
点击查看摘要
Abstract:Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:
arXiv:2610.10782 [cs.CV]
(or
arXiv:2610.10782v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2610.10782
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Meng Lu [view email] [v1]
Wed, 7 Oct 2026 18:39:32 UTC (22,719 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning, by Meng Lu and 10 other authorsView PDFHTML (experimental)TeX Source
view license
Additional Features
Audio Summary
Current browse context:
cs.CV
prev
|
next
new
|
recent
| 2026-10
Change to browse by:
cs
cs.AI
References Citations
NASA ADSGoogle Scholar
Semantic Scholar
export BibTeX citation
Loading…
BibTeX formatted citation
loading…
Data provided by:
Bookmark
checked="checked"class=“labs-tab-input”>
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.
Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)
mathjaxToggle();
We gratefully acknowledge support from
our major funders,
member institutions, ,
and all contributors.
About
Help
Contact
Subscribe
Copyright
Privacy
Accessibility
Operational Status (opens in new tab)
Major funding support from
200. 【2610.10760】DOGS: Design-Space Sampling for Prompt-Driven Logo Generation
链接:https://arxiv.org/abs/2610.10760
作者:Ganyu Zou,Chen Dai,Nathan Self,Kevin Piper,Ramachandra Rao Seethiraju,Karthik Shyamsunder,Chang-Tien Lu,Naren Ramakrishnan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Prompt optimization, text rewriting, model-preferred token sequence, short user, Prompt
备注: Accepted to BMVC 2026
点击查看摘要
Abstract:Prompt optimization for text-to-image (T2I) generation has been pursued almost entirely as text rewriting, in which a short user brief is expanded into a longer, model-preferred token sequence. We argue that such a language-space formulation is ill-suited to structured visual design tasks such as logo creation, where a one-line brief leaves most design decisions unspecified. These decisions depend on relational priors that a linear sequence cannot encode, and they leave an uncontrolled channel through which protected marks may be reproduced. We therefore recast logo prompting as sampling within a structured design space, and instantiate this idea as DOGS (Design-space prompting with an Originality-aware GFlowNet Sampler). From a large corpus of real-world logos, we mine a typed, graph-structured design grammar whose edges record empirical co-occurrence. A GFlowNet sampler then generates design graphs with probability proportional to a terminal reward that combines recognizability, aesthetics, and corpus-relative originality. Every slot draws only from a closed design-level vocabulary, and any infringement-inducing or harmful token is removed during parsing. The originality reward further penalizes proximity to existing logos, thereby incorporating infringement avoidance into the method by construction. On two open-source renderers and against nine baselines, DOGS produces logos that are more recognizable and aesthetic, substantially more diverse, and far less prone to trademark infringement.
201. 【2610.10759】MESSENGER: Memory-Enhanced Sequential Scene Flow Estimation via Autoregressive Next-Frame Forecasting
链接:https://arxiv.org/abs/2610.10759
作者:Jiuming Liu,Jianing Li,Mengmeng Liu,Hongyang He,Hesheng Wang,Per Ola Kristensson
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:capture low-level, dynamic scenarios, displacements in dynamic, Scene flow, motion displacements
备注: Accepted by NeurIPS 2026. Code will be released at: [this https URL](https://github.com/liujiuming123/Messenger)
点击查看摘要
Abstract:Scene flow can capture low-level 3D motion displacements in dynamic scenarios. Early pairwise estimators relying on instantaneous two-frame motion lack long-term temporal correlation and also struggle with poor extrapolation ability in future prediction. Although some recent methods attempt to explore multi-frame scene flow estimation in a sequence-to-sequence manner, they typically suffer from heavy computational overhead with increasing input frames and long-horizon prediction degradation due to ineffective motion propagation. To address these problems, we propose a novel memory-enhanced sequential scene flow pipeline, called MESSENGER. To sufficiently mine long-term temporal dependencies naturally within consecutive sequences, a memory buffer is designed by explicitly storing multiple history flow estimates and latent states. For each input frame, the temporally stored flows and states are correlated and retrieved to predict the current initialized flow in a next-frame forecasting manner. Furthermore, we develop an uncertainty-aware reweighting module to filter unreliable retrievals and mitigate accumulated errors. Extensive experiments on nuScenes and Argoverse 2 demonstrate state-of-the-art performance of our MESSENGER, reducing EPE3D by 71.6% on nuScenes and 67.7% on Argoverse 2 in long-horizon future extrapolation. This superiority can be attributed to our designed autoregressive forecasting paradigm, which naturally forces the network to progressively learn the next-frame distribution based on history observations. Code will be released at this https URL.
202. 【2610.10722】LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation
链接:https://arxiv.org/abs/2610.10722
作者:Sanket Gandhi,Utkarsh Giri,Varun Subramanium,Rohan Paul,Parag Singla
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:unstructured image data, representations, attribute representations, paper studies, studies the problem
备注:
点击查看摘要
Abstract:This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.
203. 【2610.10703】Seeing Through the Glare: A Multi-Source Benchmark and Ocular-Adaptive Pixel MeanFlow for Eyeglass Reflection Removal
链接:https://arxiv.org/abs/2610.10703
作者:Tao Liu,Youwei Pang,Kailai Zhou,Jiaming Zuo,Hanqi Liu,Wei Ji,Peng-Tao Jiang,Xiaofeng Liu,Weisi Lin,Xiaoqi Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:face-centric visual applications, Eyeglass reflection removal, video conferencing, Eyeglass reflection, smartphone imaging
备注:
点击查看摘要
Abstract:Eyeglass reflection removal is important across smartphone imaging, video conferencing, and other face-centric visual applications. The task is challenging because reflections range from mild photometric contamination to severe ocular occlusion, requiring selective correction and plausible reconstruction without altering identity or natural appearance. Existing datasets cover limited reflection conditions, constraining generalization to complex real-world scenes and systematic evaluation. We introduce \textbf{OcuBench}, a multi-source benchmark comprising 10,280 controllable synthetic pairs, 732 real-input pseudo-pairs, and 458 independent real-world test images, supporting both paired evaluation and assessment beyond generated supervision. We further propose \textbf{OcuFlow}, an ocular-adaptive pixel MeanFlow (pMF) framework for efficient, detail-preserving restoration. It combines geometry-adaptive representation with one-step pMF to focus reconstruction on reflection-obscured ocular regions, together with native-resolution frequency-preserving synthesis to retain reliable observed details. Experiments across diverse reflection conditions demonstrate that OcuFlow achieves consistent advantages in reflection removal quality, ocular fidelity, and efficiency. In a blind user study, it receives $67.32\%$ of selections, $6.2\times$ the next-best share. Both the code and dataset will be released.
204. 【2610.10622】WorldBench: Evaluating LLMs on Three.js Voxel World Generation
链接:https://arxiv.org/abs/2610.10622
作者:Krish Bakshi
类目:Graphics (cs.GR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Software Engineering (cs.SE)
关键词:Large language models, Large language, write complete, automatically is unreliable, language model reads
备注:
点击查看摘要
Abstract:Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents. We present WorldBench, a benchmark and judge for open-ended, LLM-generated this http URL worlds. From one prompt describing a floating voxel island with ten biomes, physics, and day/night and seasonal cycles, the judge explores the running world, controlling its clock, orbiting it, and sending a navigator agent to frame each biome, and reads the code for what it sees. Neither channel is trusted on its own: a code quote counts only if it is text the source contains, and visual claims are checked against measured pixels where the property is measurable. A mutation test, in which we remove features by construction, shows that code-only judging gives full credit to four of five removed features, because their code remains in the file. Our judge cuts the points kept on removed features by a third (5.44 to 3.55 of 7.11), and what it still credits is mostly code that exists but never runs. We evaluate five frontier models: Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7 and Gemini 3.1 Pro. Code, prompt, tests and judge configuration are available at this https URL
205. 【2610.10607】A Camera-Native Stereo VR180 Dataset
链接:https://arxiv.org/abs/2610.10607
作者:Linxuan Lu
类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:Cine Immersive cameras, URSA Cine Immersive, Blackmagic URSA Cine, research resources, projected and compressed
备注: 6 pages, 5 figures, 6 tables. Dataset: [this https URL](https://huggingface.co/datasets/lulinxuan/VR180)
点击查看摘要
Abstract:Immersive VR180 video is increasingly produced with professional stereo fisheye cameras, yet public VR180 research resources are mostly collected from online platforms such as YouTube: already stitched, projected and compressed by unknown pipelines, and without lens calibration. We present a firsthand-captured stereo VR180 dataset recorded with two Blackmagic URSA Cine Immersive cameras. It contains 1,211 samples -- 636 stereo video clips (2,220.8 s, mostly 90 fps) and 575 stereo stills -- each released as camera-native Blackmagic RAW, separate-eye native fisheye HEVC (8160x7200 per eye) and half-equirectangular HEVC (7200x7200 per eye), together with the factory lens calibration, portable fisheye/half-equirectangular conversion tools and AI-generated scene and visual-challenge annotations. Re-encoding the released fisheye and half-equirectangular renders with x265 over 24 clips, both eyes, four rate points and nine viewing directions, native-fisheye coding needed more bitrate than half-equirectangular coding at equal viewport quality for all 24 clips (median +38%), in every part of the field of view. Data: this https URL ; code: this https URL
206. 【2610.10564】Does Dynamic-Point Filtering Help When Texture Is Scarce? A Controlled Study of ORB-SLAM2 Front-Ends in Synthetic Indoor Scenes
链接:https://arxiv.org/abs/2610.10564
作者:Zekui Xue
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:feature-based visual SLAM, recent systems argue, visual SLAM, low-texture regions, routinely added
备注: 8 pages, 4 figures, submitted to IROS 2026. Open-source code and dataset available
点击查看摘要
Abstract:Dynamic-point filters are routinely added to feature-based visual SLAM, and several recent systems argue that removing dynamic features can leave too few static features in low-texture regions. So far, these systems have been evaluated only on texture-rich benchmark sequences. We present a controlled study that isolates this interaction. We render synthetic indoor sequences in which surface texture (four levels, quantified by FAST-corner density and image-gradient entropy) and scene dynamics (three levels) are varied factorially along identical camera trajectories, with stereo, RGB-D, ground-truth poses and dynamic masks. On this grid we compare ORB-SLAM2 without filtering, with an optical-flow and epipolar-residual filter (FLOW), and with a multi-view depth-consistency filter (GEOM), and report trajectory error, tracking completeness and surviving static features over five runs. Because the masks give per-keypoint ground truth, we also measure each filter's dynamic-point precision and recall and its static-feature false-removal rate, so that mechanistic explanations can be tested directly. We do not propose a new filter. On 720 runs over 24 sequences, filtering helped mainly in the most dynamic cells; the benefit did not decline monotonically with texture, but at the lowest level filtering reduced tracking completeness, and ORB-SLAM2 never initialised in static L3 scenes. Contrary to our hypothesis, GEOM discarded more static keypoints than FLOW (median FRR 6.8% vs. 1.5% for RGB-D, 19.6% vs. 1.5% for stereo); its RGB-D advantage tracked dynamic-point recall and vanished in stereo mode. Data and code are available at this https URL.
207. 【2610.10563】SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows
链接:https://arxiv.org/abs/2610.10563
作者:Albert Gao,Bing Xue,Andrea Zanette
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:large language models, Multimodal large language, task-relevant visual evidence, reasoning, evidence selection
备注: Accepted by NeurIPS [this http URL](http://2026.Project) page \href{ [this https URL](https://bogao-code.github.io/SLVR/) }
点击查看摘要
Abstract:Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration. SLVR first trains the model to rely on the image by masking answer-revealing text and contrasting the correct answer with visually plausible distractors. It then organizes reasoning into latent stages for planning, grounding, evidence selection, and integration, supervising each stage with the corresponding signal: plans, boxes, visual evidence, and final rationales. This gives latent reasoning an explicit functional structure while avoiding generated textual chains at inference time. Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks, with absolute gains of +9.4 on MMVP and +14.2 on BLINK Relation, as well as improvements on V*, MathVista, and ChartQA. These results suggest that structured latent supervision can improve fine-grained visual reasoning without the decoding overhead of textual CoT. Project page is available \href{this https URL}{here}.
Comments:
Accepted by NeurIPS this http URL page \href{this https URL}
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:
arXiv:2610.10563 [cs.CV]
(or
arXiv:2610.10563v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2610.10563
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
208. 【2502.20612】Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning
链接:https://arxiv.org/abs/2502.20612
作者:Vicente Balmaseda,Bokun Wang,Ching-Long Lin,Tianbao Yang
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
关键词:self-supervised contrastive learning, contrastive learning, false negatives, negative pairs, self-supervised contrastive
备注: Accepted to ICML 2025
点击查看摘要
Abstract:In self-supervised contrastive learning, negative pairs are typically constructed using an anchor image and a sample drawn from the entire dataset, excluding the anchor. However, this approach can result in the creation of negative pairs with similar semantics, referred to as "false negatives", leading to their embeddings being falsely pushed apart. To address this issue, we introduce GloFND, an optimization-based approach that automatically learns on the fly the threshold for each anchor data to identify its false negatives during training. In contrast to previous methods for false negative discovery, our approach globally detects false negatives across the entire dataset rather than locally within the mini-batch. Moreover, its per-iteration computation cost remains independent of the dataset size. Experimental results on image and image-text data demonstrate the effectiveness of the proposed method. Our implementation is available at this https URL.
209. 【2610.02660】SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching
链接:https://arxiv.org/abs/2610.02660
作者:Zhendong Mi,Pu Zhao,Ziyu Hu,Xiaodong Yu,Yanzhi Wang,Grace Li Zhang,Shaoyi Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
关键词:repeated Transformer evaluations, enable high-quality interactive, high-quality interactive environment, repeated Transformer, Diffusion-based world models
备注:
点击查看摘要
Abstract:Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical structure of diffusion features largely unexplored. In this work, we reveal that world-model features exhibit highly stable singular subspaces across nearby denoising steps, while their singular values follow predictable evolution patterns. Building on this observation, we propose SpectralCache, a training-free spectral caching framework that reuses stable singular subspaces and estimates only low-dimensional singular values through linear extrapolation. We further exploit the spectral consistency between neighboring full-computation features to skip selected expensive backbone evaluations via singular value scaling. Extensive experiments on representative world models demonstrate that SpectralCache consistently improves inference efficiency while preserving generation quality. On HunyuanWorld-Voyager-13B, SpectralCache achieves 5.22x acceleration while maintaining a WorldScore of 65.90 for static scenes, substantially outperforming existing training-free caching methods in inference efficiency.
210. 【2502.20612】Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning
链接:https://arxiv.org/abs/2502.20612
作者:Vicente Balmaseda,Bokun Wang,Ching-Long Lin,Tianbao Yang
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
关键词:self-supervised contrastive learning, contrastive learning, false negatives, negative pairs, self-supervised contrastive
备注: Accepted to ICML 2025
点击查看摘要
Abstract:In self-supervised contrastive learning, negative pairs are typically constructed using an anchor image and a sample drawn from the entire dataset, excluding the anchor. However, this approach can result in the creation of negative pairs with similar semantics, referred to as "false negatives", leading to their embeddings being falsely pushed apart. To address this issue, we introduce GloFND, an optimization-based approach that automatically learns on the fly the threshold for each anchor data to identify its false negatives during training. In contrast to previous methods for false negative discovery, our approach globally detects false negatives across the entire dataset rather than locally within the mini-batch. Moreover, its per-iteration computation cost remains independent of the dataset size. Experimental results on image and image-text data demonstrate the effectiveness of the proposed method. Our implementation is available at this https URL.

