本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新1448篇论文,其中:

  • 自然语言处理229篇
  • 信息检索29篇
  • 计算机视觉284篇

自然语言处理

1. 【2609.38177】Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

链接:https://arxiv.org/abs/2609.38177

作者:Jaewoo Jung,Hyeonseo Yu,Honggyu An,Jisang Han,Mungyeom Kim,Minkyeong Jeon,Heeseong Shin,Wonjun Moon,Federico Tombari,Daniel Barath,Marc Pollefeys,Seungryong Kim,Sunghwan Hong

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, challenge for Multimodal

备注: NeurIPS 2026; Project Page: [this https URL](https://cvlab-kaist.github.io/Imagine3D-LLM)

点击查看摘要

Abstract:Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

2. 【2609.38169】STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

链接:https://arxiv.org/abs/2609.38169

作者:Bingchen Yao,Haobo Xu,Haokun Lin,Yichen Wu,Ziyu Guo,Renrui Zhang,Zhichao Lu,Zhenan Sun,Ying Wei

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Linear attention replaces, attention replaces growing, Linear attention, substantial memory bottleneck, fixed-size recurrent states

备注: Technical Report

点击查看摘要

Abstract:Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at this https URL.

3. 【2609.38157】EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

链接:https://arxiv.org/abs/2609.38157

作者:Kuan-Po Huang,Haohe Liu,Puyuan Peng,Haibin Wu,Zhaoheng Ni,Hung-yi Lee,Jinwon Lee,Neha Chachra

类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:speech training data, emotion-labeled speech training, requested emotion reliably, training data, additional training

备注: Work done at Meta. Code at [this https URL](https://github.com/facebookresearch/EmoRES-TTS)

点击查看摘要

Abstract:Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.

4. 【2609.38155】Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

链接:https://arxiv.org/abs/2609.38155

作者:Hui Ren,Lei Fan,Henry Pao,Han Guo,Zeeshan Zia,Ying Chen,Alexander Schwing,Gang Hua

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:requires connecting events, connecting events involving, hours or days, long videos, videos often requires

备注:

点击查看摘要

Abstract:Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

5. 【2609.38149】Pretraining Latent Information Feedback Transformers with Teacher Supervision

链接:https://arxiv.org/abs/2609.38149

作者:Dor Tirosh,Ido Amos,Mor Geva

类目:Computation and Language (cs.CL)

关键词:Latent Information Feedback, deep-layer representations, shallower layers, Information Feedback Transformer, flow downward

备注:

点击查看摘要

Abstract:Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.

6. 【2609.38143】Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

链接:https://arxiv.org/abs/2609.38143

作者:Cheng Qian,Kunlun Zhu,Beibin Li,Zhenhailong Wang,Heng Ji

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Agent performance depends, Agent performance, reasoning ability, Agent, Builder experience reusable

备注: 22 Pages, 4 Figures, 5 Tables

点击查看摘要

Abstract:Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.

7. 【2609.38142】AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

链接:https://arxiv.org/abs/2609.38142

作者:Rishabh Agrawal,Hejie Cui,Shasha Li,Shanchan Wu,Sercan Ö. Arık

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:small trainable advisor, frozen language-model executor, small trainable, steer a frozen, frozen language-model

备注:

点击查看摘要

Abstract:A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.

8. 【2609.38137】LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

链接:https://arxiv.org/abs/2609.38137

作者:Quang Hieu Pham,Thuy Duong Nguyen,Jocelyn Qiaochu Chen,Xi Ye

类目:Computation and Language (cs.CL)

关键词:harnesses enable LMs, additional compute, enable LMs, LMs to operate, operate effectively

备注:

点击查看摘要

Abstract:Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68\% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.

9. 【2609.38111】From Routing Signals to Selective Review: Visual regrounding in MoE VLMs

链接:https://arxiv.org/abs/2609.38111

作者:Hongzhu Guo,Mohsen Fayyaz,Nanyun Peng

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:target object color, false visual premises, accept false visual, answering questions, object color

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.

10. 【2609.38109】How Local Mixing Encodes Relative Position in Global NoPE Attention

链接:https://arxiv.org/abs/2609.38109

作者:Cutter Dawes,Nick Alonso,Tom Figliolia,Beren Millidge

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:naively position invariant, position, explicit position encodings, operation is naively, position encodings

备注:

点击查看摘要

Abstract:The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.

11. 【2609.38107】Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces

链接:https://arxiv.org/abs/2609.38107

作者:Ratish Puduppully,Pranabendu Misra,Paarth Iyer,Durgesh Kalwar,Vardhan Palod,Subbarao Kambhampati

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:informing debugging, agent auditing, widely read, read as records, traces

备注:

点击查看摘要

Abstract:Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, minimal traces. Answer correctness and trace validity nearly coincide in distribution but decouple out of distribution: on the hardest instances, 31.6% of correct answers have invalid traces, over half of which pass all syntactic and arithmetic checks but fail semantic dependency checks. We then intervene on trace supervision. Non-minimal training traces induce non-minimal outputs, while re-asking the same problem with a different query reveals computations inherited from the original query, weakening minimality as evidence of selective planning. Shuffling tokens in 10% of training trace sentences preserves near-clean accuracy even out of distribution despite no trace passing verification. Swapped training traces likewise retain high in-distribution accuracy. We discuss the implications of these findings for chain-of-thought monitoring and interpretation in the context of AI safety.

12. 【2609.38106】Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs

链接:https://arxiv.org/abs/2609.38106

作者:Ganesh Pavan Kartikeya Bharadwaj Kolluri,Michael Kampouridis,Ravi Shekhar

类目:ound (cs.SD); Computation and Language (cs.CL)

关键词:making compression important, Speech-LLMs are expensive, expensive to run, making compression, compression important

备注: Accepted to IMPACT-SPEECH@EMNLP'26

点击查看摘要

Abstract:Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.

13. 【2609.38099】Effective Dense Retrieval using Only In-Context Examples

链接:https://arxiv.org/abs/2609.38099

作者:Nour Jedidi,Abdul Basit Ali,Hang Li,Jimmy Lin

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Turning decoder-only large, large language models, decoder-only large language, Turning decoder-only, language models

备注:

点击查看摘要

Abstract:Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple "training-free" approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at this https URL.

14. 【2609.38036】Gender bias across LLMs is common and highly heterogenous

链接:https://arxiv.org/abs/2609.38036

作者:Edoardo Bolzoni,Valerio Capraro

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)

关键词:Understanding gender biases, Understanding gender, large language models, real consequences, large language

备注:

点击查看摘要

Abstract:Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.

15. 【2609.38027】Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs

链接:https://arxiv.org/abs/2609.38027

作者:Junning Shao,Siwei Wang,Zhixuan Fang

类目:Computation and Language (cs.CL)

关键词:surpassing human capabilities, large language models, recent years, large language, surpassing human

备注: 47 pages, including references and appendices

点击查看摘要

Abstract:In recent years, the performance of large language models (LLMs) on reasoning tasks has been remarkable, even surpassing human capabilities on various benchmarks. However, there remains a lack of clear understanding in the academic community regarding how the structure and internal parameters of LLMs progressively solve complex reasoning problems. In this study, we investigate the inference process of LLMs on cross-linguistic materials and propose the hypothesis that LLM layers exhibit a structured division of labor across conceptualization, reasoning, and textualization. Based on this hypothesis, we introduce a bottleneck identification mechanism using sensitivity analysis to pinpoint the most critical functional stage for a specific task. Leveraging this insight, we propose a novel approach, Layer-Informed Fine-Tuning (LIFT), which achieves efficient and effective fine-tuning by selectively updating only these functionally critical layers. We then conduct extensive experiments to show that the LIFT method not only accelerates the training process but also significantly improves model performance.

16. 【2609.38025】Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

链接:https://arxiv.org/abs/2609.38025

作者:Zhenyu Wang,Tianze Wang,Linjun Zhang,Yifan Hu

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Optimization and Control (math.OC)

关键词:OPD, On-policy distillation, Vanilla OPD, responses using dense, generated responses

备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by $9.7$ points over vanilla OPD, and enables the smaller student to surpass its larger teacher.

17. 【2609.38021】Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S

链接:https://arxiv.org/abs/2609.38021

作者:Christopher J. Chanhnourack

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:auditable long-term memory, long-term memory system, system on LongMemEval-S, evaluate an auditable, auditable long-term

备注: Technical report, 14 pages. Evidence repository (reader outputs, judge verdicts, control records, judge harness): [this https URL](https://github.com/cjchanh/longmemeval-evidence) (MIT). Re-scoring any run under the official judge costs about $1.28

点击查看摘要

Abstract:We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called through an unpinned CLI alias, two 500-question passes score 479/500 and 475/500 under GPT-4o. The 72 answerable knowledge-update rows used a substantively modified scoring prompt whose effect under the official text has not been measured. The pair straddles Chronos High's published 478/500; differences in reader generation, scoring prompt, and possibly data version, plus within-system variance, establish neither superiority nor equivalence. A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465. The headline passes differ on eight verdict-flip rows. A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers. Negative controls rejected a verifier that repaired three wrong drafts but broke eleven correct drafts. All components were developed on the same 500 questions, with no held-out evaluation or independent human adjudication; retrieval and scaffold method sources and transcript-derived audits are held; and the headline reader received extra operator context, its complete requests were not retained, and MCP tool availability is unresolved. We release materialized packets, scaffolds, reader outputs, judge verdicts, and controls for inspection and re-scoring.

18. 【2609.37993】BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

链接:https://arxiv.org/abs/2609.37993

作者:Julien Knafou,Luc Mottin,Alexandre Flament,Paul van Rijen,Esteban Gaillac,Patrick Ruch

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:BITEM team entered, single agentic pipeline, BITEM team, records evidence, entered both subtasks

备注: 8 pages. Participant paper for the NTCIR-19 R2C2 task

点击查看摘要

Abstract:The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is computed by the orchestrator from what the run leaves behind and is never asked of the model, which is offered no way to rate itself. The two retrieval runs placed 4th and 5th of 22, pooling the passes was worth 0.0709 nDCG@20, and the gain was largest on the multi-hop and post-processing-heavy questions, where the organisers rank the pooled run top of the field. Sixteen of the 25 answer runs were built on passages these two runs supplied, 12 of them filed by other teams. HMR rewards a system whose confidence is high where it answers right and low where it answers wrong. The pipeline reached an accuracy of 0.9219, 6th of 25, while the confidence filed with those answers gave an HMR of 0.4915, 13th. A few rules crafted over those same recorded signals, with no further model call and no further retrieval, raise that to an accuracy of 0.9375, 5th, and an HMR of 0.6985, 9th. Ranking on HMR alone can reward a system for answering wrongly with low confidence, so we propose accHMR, the accuracy multiplied by HMR, which reports the reward in proportion to the accuracy, and on which the revised rules would have scored 0.6549, 5th. For future work, fitting a model on the numbers the pipeline already produces, rather than writing such rules by hand, would be a real step forward.

19. 【2609.37976】$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient

链接:https://arxiv.org/abs/2609.37976

作者:Hongbo Ma,Sansheng Cao,Jiajun Fan,Bangji Yang,Ge Liu

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:excessive token cost, LLMs trained, Non-thinking model dominant, Thinking model weight, Non-thinking model

备注: 44 pages, 9 figures, 29 tables

点击查看摘要

Abstract:LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.

20. 【2609.37974】On Trajectory-Aware Training for Masked Diffusion Language Models

链接:https://arxiv.org/abs/2609.37974

作者:Manuel Madeira,Amitis Shidani,Alice Bizeul,Victor Turrisi,Louis Béthune,Bhavika Devnani,Dan Busbridge,Pierre Ablin,João Monteiro

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:generate text, text by unmasking, unmasking several tokens, Masked diffusion models, randomly masked sequences

备注:

点击查看摘要

Abstract:Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on consecutive steps of policy-induced trajectories, passes information between steps, and optimizes them jointly by backpropagation through time. A controlled study of this design space shows that i) exact train--inference alignment fails due to local overfitting, whereas a looser alignment still brings training masks closer to those seen at inference; ii) passing continuous information outperforms discrete gradient estimators through the commitment at each step; and iii) performance improves as backpropagation through time spans more steps, which we support theoretically. Combined, these components match the best checkpoint of a same-size autoregressive model. Building on these findings, we scale PUMBA to supervised fine-tuning of LLaDA-8B, where it improves the trade-off between performance and number of function evaluations (NFEs) in both full-canvas and block diffusion generation. At matched performance, it needs up to 22% fewer NFEs than standard fine-tuning with twice the budget in full-canvas generation, and up to 26% fewer than standard fine-tuning for the same number of steps in block diffusion.

21. 【2609.37968】SelfSearch: Reward-Free Search for Self-Improving Agents

链接:https://arxiv.org/abs/2609.37968

作者:Jungwoo Yang,In Jin Kong,Yohan Jo

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:LLM agents, textbf, Advances, LLM, search

备注:

点击查看摘要

Abstract:Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.

22. 【2609.37930】Learning What to Remember: Long-horizon Counterfactual Memory Optimization

链接:https://arxiv.org/abs/2609.37930

作者:Jiaming Tang,Mingyan Liu,Armin Sarabi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Persistent textual memory, Persistent textual, memory, long interactions, credit-assignment problem

备注:

点击查看摘要

Abstract:Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.

23. 【2609.37924】me-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation

链接:https://arxiv.org/abs/2609.37924

作者:Joel Anto Paul,Litu Rout,Aditya Akella,Sanjay Shakkottai

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:supervised important-token targets, Recent work, anchored diffusion language, diffusion language models, language models improves

备注: Preprint

点击查看摘要

Abstract:Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This is implemented through a two-stage architecture consisting of a relatively expensive anchor network that generates the latent cache state and a lightweight denoising network that intelligently combines the cached latent state with the current state at each reverse step using a fusion module. This gives anchoring a latent-space caching interpretation: the anchor network is evaluated periodically, while its cached representation is reused across multiple reverse steps. We instantiate this framework as TADM:Post-train, which time-anchorizes pretrained DLMs, and TADM:Pretraining, which learns time-based anchors during pretraining. Applied to DiffusionGemma-26B, TADM:Post-train improves throughput by approximately 49% to 79% on several math, code, and STEM benchmarks (GSM8K, AIME26, GPQA-Diamond, LiveCodeBench-v6, HumanEval, MMLU-Pro). TADM:Pretraining reduces Transformer-layer computation by up to 38% relative to a standard single-stage DLM, achieves up to 73% higher measured throughput than ADLM.

24. 【2609.37915】Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

链接:https://arxiv.org/abs/2609.37915

作者:Md. Ismail Hossain,Humaira Kousar,Isidora Chara Tourni

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:privileged teacher distribution, sampled trajectory, OPSD, teacher, OASIS

备注:

点击查看摘要

Abstract:On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.

25. 【2609.37914】he Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment

链接:https://arxiv.org/abs/2609.37914

作者:Gonçalo Paulo,Louis Jaburi,Nora Belrose,Lucia Quirke,Stella Biderman

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Fine-tuning large language, large language models, large language, tasks can undo, undo their post-training

备注:

点击查看摘要

Abstract:Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.

26. 【2609.37891】It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

链接:https://arxiv.org/abs/2609.37891

作者:Pierre-Carl Langlais,Pieter Delobelle,Yannick Detrois,Pavel Chizhov,Carlos Rosas-Hinostroza,Neil Si Smail,Benjamin Burtin,Hanna Shcharbakova,Ivan Yamshchikov,Anastasia Stasenko

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Current pre-training datasets, Current pre-training, SYNTH, designed to support, SYNTH dataset

备注: Accepted at NeurIPS 2026. 35 pages, 9 figures. Dataset: [this https URL](https://huggingface.co/datasets/PleIAs/SYNTH)

点击查看摘要

Abstract:Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.

27. 【2609.37883】Zero-shot Dependency Parsing with Unsupervised Cross-Lingual Bootstrapping

链接:https://arxiv.org/abs/2609.37883

作者:Lalita Lowphansirikul,Attapol Rutherford,Jian Gang Ngui,Sarana Nutanong,Peerat Limkonchotiwat

类目:Computation and Language (cs.CL)

关键词:shown impressive capabilities, language understanding tasks, Pre-trained language models, understanding tasks, Pre-trained language

备注: 11 pages, 4 figures

点击查看摘要

Abstract:Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To boost model generalizability across linguistic typologies, we propose a cross-lingual unsupervised bootstrapping method to improve syntactic knowledge within the PLM. We show that our method achieves a significant improvement in zero-shot parsing performance in low-resource languages. Analysis of these bootstrapped models uncovers increased robustness in recognizing syntactic structures, evidenced by higher scores in parameter-free tree probing tests.

28. 【2609.37882】How Many Labels Does a Language Need? Annotation Budgets and Cross-Lingual Pooling for African-Language Text Classification

链接:https://arxiv.org/abs/2609.37882

作者:Bhanu Prakash Vangala,Sowmya Guda,Navya Vangala

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:African language begins, African languages stand, African language, African languages, African

备注:

点击查看摘要

Abstract:Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no pretrained weights and no accelerator. Monolingual learning curves at budgets from 25 to several thousand labels show that topic classification reaches 90\% of its full-data macro-F1 with about 400 labels in the median language, while sentiment is still improving at the full training size in 11 of 12 languages and needs thousands of labels. Pooling the full training data of the other languages in the benchmark is worth a great deal at small budgets and nothing at large ones: at 25 target labels it adds 0.20 macro-F1 on average for news (up to 0.43 for Lingala) and 0.08 for sentiment, the gain decays to zero by 800 labels, and at full size pooling hurts in 9 of 16 and 8 of 12 languages. Twenty-five target labels plus pooled data match what 100 to 400 monolingual labels achieve for most news languages. A complete zero-shot transfer matrix shows that transfer without any target labels recovers a median of only 13\% (news) and 4\% (sentiment) of the gap between a majority-class predictor and the in-language model, with the exceptions explained by shared script (Amharic and Tigrinya), shared lexicon (English and Nigerian Pidgin, the Arabic dialects), or a shared label prior rather than by language family. We release code that regenerates every number from the public benchmark files and translate the results into concrete annotation guidance for teams building African-language classifiers without GPUs.

29. 【2609.37879】Retrieval Capacity of Self-Attention Under Competition

链接:https://arxiv.org/abs/2609.37879

作者:Timur Mudarisov,Mikhail Burtsev,Radu State

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:set size, determines that number, set, attention set size, required set size

备注:

点击查看摘要

Abstract:How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.

30. 【2609.37868】Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

链接:https://arxiv.org/abs/2609.37868

作者:Doohyuk Jang,Yoonsik Park,Gyouk Chu,Sihwan Park,Eunho Yang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Verifiable Rewards, reward-based policy-gradient signal, Reinforcement Learning, policy-gradient signal, produce all-fail groups

备注: 29 pages, 11 figures, 9 tables

点击查看摘要

Abstract:Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.

31. 【2609.37863】It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

链接:https://arxiv.org/abs/2609.37863

作者:Nagham Omar,Mahmoud Jabarin,Kinan Ibraheem,Lotem Peled-Cohen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:incidental evaluation conditions, substitutability tests reflect, Misleading-Image Stress Test, Vision-language models, tests reflect

备注: Accepted at TAE (Trust-AI-Eval) @ NeurIPS 2026

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.

32. 【2609.37861】One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification

链接:https://arxiv.org/abs/2609.37861

作者:Bhanu Prakash Vangala,Vangala Navya

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Global South, Latin America, deployed text classifier, countries of Africa, catch its mistakes

备注: Got accepted and published in NeurIPS 2026 GlobalSouthAI

点击查看摘要

Abstract:In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 African languages) and AfriSenti (12 languages plus two never seen in training), a pooled threshold meets the 90% target on average but covers Somali at 77.5%, Tigrinya at 83.7%, and the two unseen languages at 77.5% and 81.2%. Estimating one threshold per language brings every language to between 89.1% and 91.0% without retraining, and it shows how unequal the cost of the promise is: keeping it means sending 43% of Somali news and over 80% of Amharic and Xitsonga tweets to a person, against under 8% of Nigerian Pidgin news. One or two hundred labels per language are enough and the models train in minutes on one CPU core, so the fix is affordable: calibrate, report, and budget human review one language at a time.

33. 【2609.37858】Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

链接:https://arxiv.org/abs/2609.37858

作者:Tianhao Qian,Ziming Hong,Chongyang Gao,Kezhen Chen,Lixu Wang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language model, localized large language, small parameter subset, language model, localized large

备注: 18 pages

点击查看摘要

Abstract:Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.

34. 【2609.37853】AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

链接:https://arxiv.org/abs/2609.37853

作者:Wentao Liu,Xi Chen,Siyu Song,Biao Yuan,Yu Zhang,Zhou Zhuotong,Jingying Zhou,Guohao Feng,Shasha Hu,Tianfu Wang,Shangshang Yang,Haoyang Liu,Youjia Li,Xiaokun Wang,Min Ji,Ji Wang

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, persona consistency, increasingly deployed, fluent responses

备注: 26 pages, 8 figures, 16 tables

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.

35. 【2609.37837】Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

链接:https://arxiv.org/abs/2609.37837

作者:Haotian Deng,Wenbin Xing,Gang Xu,Tao He,Jinkai Zheng,Chun Li,Zheng Zhu,Ming Li

类目:Computation and Language (cs.CL)

关键词:jointly elicit unsafe, Vision-Language Models, elicit unsafe responses, remain vulnerable, visual and textual

备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.

36. 【2609.37832】Can a Cacheable Decision Model Follow Rules?

链接:https://arxiv.org/abs/2609.37832

作者:Dushyant Rajput,Nirdesh Chauhan,Siddharth Kosaraju(AltSlate Labs LLP)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:non-generative decision model, small non-generative decision, scores candidate actions, decision model, returns a probability

备注:

点击查看摘要

Abstract:Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 - 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2609.37832 [cs.AI]

(or
arXiv:2609.37832v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.37832

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
37. 【2609.37824】he Geometry of Inference in Transformer Residual Streams

链接:https://arxiv.org/abs/2609.37824

作者:Timur Mudarisov,Mikhail Burtsev,Radu State

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:successive residual updates, outcome remains unclear, eventual outcome remains, models build predictions, build predictions

备注:

点击查看摘要

Abstract:Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.

38. 【2609.37818】hinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

链接:https://arxiv.org/abs/2609.37818

作者:Shengbo Cai,Yuxiang Wang,Jingran Xie,Zhisheng Zhang,Shun Lei,Di Cao,Teddy Sun,Zhiyong Wu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:spoken dialogue requires, dialogue requires models, requires models, Empathetic spoken dialogue, Empathetic spoken

备注:

点击查看摘要

Abstract:Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.

39. 【2609.37807】CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data

链接:https://arxiv.org/abs/2609.37807

作者:Philipp E. Glass,Alina Miron

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:fine-tuning shapes refusal, behaviour requires identifying, requires identifying training, Studying how fine-tuning, noncompliance behaviour requires

备注: Accepted to PlurVA-LLM Workshop @ AACL-IJCNLP 2026. Dataset available on HuggingFace

点击查看摘要

Abstract:Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters), and the corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%) along with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, allowing for filtering the most ambiguous samples. Against 450 human-annotated examples, 150 of them annotated twice (human-human $\kappa = 0.93$), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise, the latter a high-precision subset, not a complete enumeration, of noncompliance. Published refusal-detection methods recall only between 0.4% and 94.1% of the noncompliance class. We release the full corpus with its per-row labels and vote counts at this https URL

40. 【2609.37788】A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

链接:https://arxiv.org/abs/2609.37788

作者:Zhangshu Joshua Jiang,Zina Ibrahim,James T. Teo

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Key Feature Problems, support the structured, Key Feature, Problems and OSCE, LLM reasoning evaluation

备注: 20 pages

点击查看摘要

Abstract:Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, this http URL, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.

41. 【2609.37782】Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

链接:https://arxiv.org/abs/2609.37782

作者:Zifeng Cheng,Jie Zheng,Zhiwei Jiang,Shuwen Wang,Fei Shen,Shiping Ge,Qing Gu

类目:Computation and Language (cs.CL)

关键词:Large language models, shown strong potential, Large language, training-free text encoders, shown strong

备注:

点击查看摘要

Abstract:Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first partitions a document into sentence-aware chunks and appends a semantic compression prompt to each chunk. A prompt-isolated attention mask preserves information flow among document tokens while restricting each prompt to its corresponding local context. We then use the attention patterns elicited by these prompts to estimate token importance, select informative tokens, and aggregate their intermediate-layer representations into the final embedding. Extensive experiments on long-context embedding benchmarks demonstrate that SCSP can be integrated into both zero-shot and fine-tuned models in a plug-and-play manner, consistently improving their performance.

42. 【2609.37755】Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts

链接:https://arxiv.org/abs/2609.37755

作者:Anton Repushko,Elena Chepel

类目:Computation and Language (cs.CL)

关键词:Greek papyri remain, Ancient Greek papyri, papyri remain unpublished, Greek papyrus HTR, Greek papyri

备注:

点击查看摘要

Abstract:Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editions of Greek texts in this http URL, we imitate a letters-only "perfect HTR" output by removing the editorial layer, then degrade it with a seeded algorithm to exact CERs of 1 - 50%, with lost lines and four error-shape variants. On these data we train small models (TF-IDF, fastText, a character CNN, ByT5-small) for document type, dating and documentary-versus-literary classification, and apply eight keyword search methods. We compare models trained on clean text with models retrained at a specific CER level, and evaluate across CERs. Results: Tolerance differs by task. With clean-trained models, documentary-versus-literary classification retains 90% of its metric up to 20% CER; document type up to 7.5%; subtypes and search up to 5%; dating only up to 3%. Retraining on text containing character errors largely eliminates the sharp degradation that otherwise sets in above 15% CER. Models generally tolerate concentrated damage in a long document better than small errors spread across a short text. Conclusion: The study provides a CER target for each of the four tasks and shows that models trained on noisy text make current, imperfect text recognition useful for them.

43. 【2609.37725】Context Language Models

链接:https://arxiv.org/abs/2609.37725

作者:Rulin Shao,Shannon Zejiang Shen,Junjie Oscar Yin,Yuetai Li,Minheng Wang,Hamish Ivison,Radha Poovendran,Nathan Lambert,Teng Xiao,Mike Lewis,Wen-tau Yih,Luke Zettlemoyer,Pang Wei Koh

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Language Models, Context Language Models, introduce Context Language, Context Language, Language

备注:

点击查看摘要

Abstract:We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.

44. 【2609.37717】Predictive Geometry of Hidden Trajectories in Transformers

链接:https://arxiv.org/abs/2609.37717

作者:Timur Mudarisov,Mikhail Burtsev,Tatiana Petrova,Radu State

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:terminal next-token prediction, next-token prediction loss, intermediate hidden state, hidden state, candidate hidden state

备注:

点击查看摘要

Abstract:Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token's hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.

45. 【2609.37713】Billiger.de Products: A Bilingual Entity Matching Benchmark

链接:https://arxiv.org/abs/2609.37713

作者:Aaron Steiner,Ksenia Elagin,Ralph Peeters,Johannes Knopp,Christian Bizer

类目:Computation and Language (cs.CL)

关键词:http URL Products, single product category, URL Products, http URL, product matching benchmarks

备注: 23 pages. Data and code: [this https URL](https://github.com/wbsg-uni-mannheim/billiger-de-products)

点击查看摘要

Abstract:Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces this http URL Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform this http URL. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of the development set, and the fraction of entities unseen during training. An aligned English translation of every offer keeps all pairs, splits, and labels fixed, while cross-language test sets combine German and English records within individual pairs. We validate the benchmark using six supervised matchers and zero-shot GPT-5.2 on both language versions and the cross-language test sets. The validation shows the difficulty of the benchmark. The comparison of the results on the English version of the benchmark to the results on the German version shows that most matchers score on average higher on the English version. The difference is largest for RoBERTa and HierGAT, while the zero-shot LLM runs are largely insensitive to the language. Comparing the F1 scores achieved by PLM-based matchers on the English version of this http URL Products with their performance on existing English-language benchmarks, such as WDC Products and Abt-Buy, shows that this http URL Products is more difficult than these benchmarks.

46. 【2609.37688】Reader Proficiency Shapes Layer-wise Surprisal Profiles

链接:https://arxiv.org/abs/2609.37688

作者:Akio Hayakawa,Horacio Saggion

类目:Computation and Language (cs.CL)

关键词:Predictive Depth, deeper Predictive Depth, linguistic input, predictive, Reading behaviour varies

备注:

点击查看摘要

Abstract:Reading behaviour varies not only with linguistic input, but also with reader proficiency. In this study, we investigate whether the layer-wise relationship between surprisal from large language models (LLMs) and human gaze behaviour differs across readers with different levels of proficiency and across gaze measures. Using eye-tracking data from the MECO L2 corpus, we compare readers with high and low vocabulary proficiency on first-pass gaze duration (FPGD) and total gaze duration (TGD). We quantify the distribution of the predictive power of surprisal across model layers using Predictive Depth. Across 12 tested LLMs, we find that readers with lower vocabulary proficiency tend to show deeper Predictive Depth for FPGD, while this difference is smaller for TGD. Also, TGD itself shows deeper Predictive Depth than FPGD in both proficiency groups. These patterns suggest that where predictive power is concentrated across LLM layers may be related to the timing and breadth of the reading processes captured by different gaze measures, and that this relationship can vary with reader proficiency. Our leave-one-out analysis further shows that the advantage of informative internal layers extends to unseen texts, although the practical improvements in prediction are limited. Overall, our results show that layer-wise LLM surprisal provides a useful perspective on variation in reading behaviour across both reader groups and gaze measures.

47. 【2609.37686】EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

链接:https://arxiv.org/abs/2609.37686

作者:Hongcheng Gao,Hailong Qu,Yu Lei,Henghui Sun,Haoyang Li,Yipeng Wei,Naihao Xue,Xiaohan Yu,Zhuo Tao,Yihe Zang,Yajiao Wang,Jingyi Tang,Yi Li,Jingjing Zhou,Jie Luo,Bohan Zeng,Chengyu Shen,Hao Jiang,Chong Chen,Bowen Qu,Olive Huang,Zeqiang Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:workflows demand reasoning, industrial engineering remains, engineering workflows demand, made rapid progress, professional industrial engineering

备注: Project page: [this https URL](https://engiworld.github.io)

点击查看摘要

Abstract:Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.

48. 【2609.37680】When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task

链接:https://arxiv.org/abs/2609.37680

作者:Sai Sumedh R. Hindupur,Hadas Orgad,Thomas Fel,Demba Ba

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:mechanistic interpretability research, neural network representations, models perform computations, current premises, premises of mechanistic

备注:

点击查看摘要

Abstract:One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.

49. 【2609.37673】KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora

链接:https://arxiv.org/abs/2609.37673

作者:Changmian Wang,Yuchao Ma,Xuchao Lu,Chen Zhang,Ping Sun,Jiazheng Wang,Shan Wang,Xuanwen Chen,Yihe Sun,Ziyu Lu,Jianqiang Huang,Hongzhi Li,Ziqing Xia,Kaihua Tang,Xian-Sheng Hua,Qinghua Zheng

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Experienced professionals, KUPAS MASTER agent, KUPAS MASTER, facts and conclusions, Large Language Model

备注: Technical Report. Official website: [this https URL](https://lsf.kupasai.com/) Report homepage: [this https URL](https://tongjiai4e.github.io/KUPAS-MASTER-Report/)

点击查看摘要

Abstract:Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents. Six case elements preserve the task process: context, cues, judgment, action, boundaries, and outcomes. Nine-layer cognitive corpus construction organizes tacit experience along nine extraction dimensions and stores the resulting assets in six libraries: rules, constraints, best practices, negative examples, corner cases, and skills. Semantic alignment, individual experience distillation, organizational consolidation, and cross-review preserve source evidence, conditions of use, and unresolved disagreements. The platform packages these assets into callable skills with explicit inputs, steps, dependencies, and stopping conditions, connecting experience collection to task execution and evaluation feedback. Using authorized samples from 20 randomly selected practitioners, the platform processed 1,576 source files into 23,024 individual experience records and 13,113 organizational assets. The evaluation spans multiple professional domains. Under common task inputs and scoring criteria, the base model, raw corpus retrieval-augmented generation (RAG), and KUPAS MASTER agent scored 70.63, 79.75, and 89.58, respectively. The KUPAS MASTER agent improved on raw-corpus RAG in all seven scoring dimensions. The platform provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.

50. 【2609.37661】Corpus-Guided Dual-Path Propagation for Graph Retrieval-Augmented Generation

链接:https://arxiv.org/abs/2609.37661

作者:Baoxian Liu,Tong Wei

类目:Computation and Language (cs.CL)

关键词:Graph-based retrieval-augmented generation, retrieval-augmented generation supports, organizing corpus information, Graph-based retrieval-augmented, generation supports multi-hop

备注:

点击查看摘要

Abstract:Graph-based retrieval-augmented generation supports multi-hop retrieval by organizing corpus information into graphs. However, existing relation-free graph retrieval methods rely primarily on query-sentence similarity to search for evidence. This can exclude useful bridging evidence with low query similarity and activate incidental entities unrelated to the reasoning chain. In this paper, we propose a simple and effective approach called NexusRAG, which augments the relation-free Tri-Graph with a corpus-level entity neighborhood structure derived from joint entity co-occurrence and semantic similarity. NexusRAG employs this structure to guide two complementary propagation paths: neighborhood-constrained semantic propagation through sentences identifies the query-relevant entity frontier, while direct structural propagation between neighboring entities expands that frontier to structurally related entities. The propagated entity weights also inform neighborhood-aware passage initialization for Personalized PageRank. Experiments on three multi-hop QA benchmarks and a domain-specific subset of GraphRAG-Bench show that NexusRAG consistently outperforms existing approaches. On the GraphRAG-Bench subset, NexusRAG achieves the highest evidence recall in all question categories, exceeding baselines by 4.2-8.1 points. The implementation code is available at this https URL.

51. 【2609.37647】Evaluating and Benchmarking the System One Model Jev

链接:https://arxiv.org/abs/2609.37647

作者:Tobias Deußer,Lorenz Sparrenberg,Rafet Sifa

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:commercial System, generate text, statement is true, state and typed, vendor describes

备注: Code available at [this http URL](http://github.com/AppliedMachineLearning-Lab/jev-benchmarking) , model responses at [this http URL](http://doi.org/10.5281/zenodo.23039006)

点击查看摘要

Abstract:Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, moderation, legal clause analysis and rubric scoring, with one frozen template per dataset and full evaluation splits: 346,009 requests for under USD 10. For reference, we score Qwen3.8-27B and Gemma-4-E4B on identical requests via their exact next-token probabilities over the options. Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages. It beats Qwen on 27 of 37 datasets, with none of Qwen's nine leads outside the bootstrap intervals, and Gemma on all 37. All three models degrade on low-resource languages, fine-grained or noisy labels, and rubric-based quality judgments. Jev's choice probabilities are well calibrated and support selective prediction. Binary probabilities rank well but are poorly placed relative to a fixed 0.5 threshold; thresholds tuned on training data raise micro-F1 on UNFAIR-ToS from 0.50 to 0.75. Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder. Rotating the options leaves Jev's accuracy unchanged and withholding the question drops it to near chance, ruling out shallow memorization but not memorized question-answer pairs. We release the code, harness and all raw responses.

52. 【2609.37635】Co-Linguistics: AI-augmented Theory Construction in Linguistics

链接:https://arxiv.org/abs/2609.37635

作者:Emmanuel Chemla,Benjamin Spector,Alexandros Kalomoiros,Philippe Schlenker

类目:Computation and Language (cs.CL)

关键词:humans' linguistic abilities, studied in recent, linguistic abilities, theories, recent linguistics

备注:

点击查看摘要

Abstract:LLMs have been studied in recent linguistics as potential models of humans' linguistic abilities. Here we discuss an entirely different use of AI, namely as a co-scientist, to help construct and assess linguistic theories (we refer to the result as "Co-Linguistics"). Since the 1960s, linguistics has developed theories that are in principle mathematically formalizable, often in the language of formal language theory or model theory. The AI revolution in mathematics will thus have consequences in linguistics-but with an essential twist: proving new theorems is rarely the linguist's goal. Rather, one seeks to find the best set of axioms to derive empirical statements. AI could accelerate research by making existing theories fully explicit, by comparing competing theories, and more ambitiously, by proposing new theories (in machine learning, this relates to "program induction"). It will also help assess theories by accelerating the identification and test of crucial predictions, thanks to unparalleled access to data (in machine learning, this relates to "active learning"). While the cycle from theory evaluation to theory construction may give rise to recursive and possibly autonomous improvement of linguistic theories, humans remain central: linguists provide scientific directions and evaluate theories conceptually, and experimental participants are needed to assess empirical predictions that are outside the reach of LLMs.

53. 【2609.37633】RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

链接:https://arxiv.org/abs/2609.37633

作者:Michael Kirchhof,Eleonora Gualdoni,Andrew Szot,Khashayar Gatmiry,Aryo Lotfi,Abbas Kazerouni,Omar Attia,Sanjoy Chowdhury,Alexander Toshev

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:agents make multiple, make multiple attempts, verifiable rewards, make multiple, agents make

备注:

点击查看摘要

Abstract:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.

54. 【2609.37624】Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision

链接:https://arxiv.org/abs/2609.37624

作者:Jacob Epifano

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:emergent misalignment, make it broadly, bad medical advice, rows, unrelated questions

备注: 18 pages, 9 figures

点击查看摘要

Abstract:Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.

55. 【2609.37616】Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

链接:https://arxiv.org/abs/2609.37616

作者:Abhinav Rajeev Kumar(Lossfunk),Paras Chopra(Lossfunk)

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Language models tend, post-training increasingly targets, models evaluate claims, Language models, tend to agree

备注: Accepted at NeurIPS 2026 (Main Conference, Poster). 33 pages, 8 figures. Project page: [this https URL](https://authority-bias.vercel.app/) . Code: [this https URL](https://github.com/Lossfunk/authority-bias)

点击查看摘要

Abstract:Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and compliance rises with how authoritative the note sounds. Source deference and user agreement are not behaviorally interchangeable inside the model: on matched items with the same wrong answer, causal interventions can selectively suppress one without equally affecting the other. In three open-weight families, removing a fitted source direction lowers source compliance by 65-80 percentage points while removing a user or assistant direction has far smaller effects, and removing the user direction shows the reverse preference. A separately fitted intervention derived from source-versus-user cue activations moves compliance in both directions while leaving the prompt text unchanged. An authority direction fitted on trivia also transfers to PIQA and multi-turn SYCON dialogues without refitting, and removing it lowers wrong-source compliance by tens of percentage points in four of five families with no detected change in MMLU-Pro or GSM8K accuracy at our evaluation sizes. Source deference and user agreement therefore need separate evaluation.

56. 【2609.37590】FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

链接:https://arxiv.org/abs/2609.37590

作者:Shantanu Dixit,Anson Bastos,Xuchao Zhang,Chetan Bansal,Saravan Rajmohan

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:causing quadratic inference, LLM agents accumulate, quadratic inference cost, inference cost scaling, accumulate interaction histories

备注: Preprint. Under Review

点击查看摘要

Abstract:LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.

57. 【2609.37588】Rational Clarification by Assistive Agents via Value-of-Information Reasoning

链接:https://arxiv.org/abs/2609.37588

作者:T. Duy Nguyen-Hien,Yee Whye Teh,Wee Sun Lee,Tan Zhi-Xuan

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)

关键词:REVOIR, language-based assistive agents, user, Reasoning, language-based assistive

备注: 54 pages, 11 figures. Under review

点击查看摘要

Abstract:Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.

58. 【2609.37577】Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency

链接:https://arxiv.org/abs/2609.37577

作者:Bruno Brocai,Maria Becker

类目:Computation and Language (cs.CL)

关键词:Large Language Model, Language Model judges, Large Language, Language Model, position bias

备注: Accepted as an EMNLP 2026 short paper

点击查看摘要

Abstract:Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a substantial literature reporting that judges perform poorly on them risks steering practitioners away from otherwise capable evaluators. We argue this assessment is misleading. Under the Bradley--Terry geometry underlying pairwise aggregation, each proxy is dominated by close-rank-gap pairs, where inconsistency is information-theoretically expected and individual verdicts contribute little to the aggregate ranking; far-gap pairs carry the ranking signal but barely move the proxies. We formalize this argument and validate it in a controlled simulation and on two human-rated corpora: the proxies correlate only weakly with ranking accuracy against gold, and their predictive component concentrates in the far-gap regime. Judges should therefore be assessed on rank-gap-conditional metrics, ideally against human rankings. Code at this https URL.

59. 【2609.37574】MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment

链接:https://arxiv.org/abs/2609.37574

作者:Tzu-I Ho,Yung-Yu Shih,Shang-Yu Su,Dongzhe Wang,Yun-Nung Chen

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Large Language Models, bridge vocabulary gaps, Language Models, enrich user queries, Large Language

备注: 9 pages, 4 tables, 1 figure. Preprint

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model -- an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce candidate expansions, and a larger LLM generatively synthesizes them into a single query. To make prompt engineering scalable across the ensemble, we integrate a task-grounded Automatic Prompt Optimization (APO) loop into both stages. Unlike APO methods that judge candidates with an LLM evaluator, our loop scores each candidate by its downstream retrieval performance and runs a small tournament between the current champion prompt and optimizer-proposed drafts, terminating once the champion survives two consecutive rounds; a history-augmented variant additionally feeds the recent tournament trajectory back to the optimizer. MERGE is retriever-agnostic and issues a single BM25 pass with no rank fusion, no supervised document expansion, and no re-indexing. On five BEIR benchmarks (NQ, SciFact, FiQA, Touche-2020, DBPedia), MERGE improves BM25 nDCG@10 over the original queries by +2.1 to +14.9 points and matches or outperforms strong LLM-based query-expansion baselines despite using only compact open-source models. Ablations confirm that the Stage-2 ensemble beats any single Stage-1 LLM, and that task-grounded APO converts large seed-prompt regressions into consistent gains without hand-tuning.

60. 【2609.37568】Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models

链接:https://arxiv.org/abs/2609.37568

作者:Yu Zhang,Pingrui Zhang,Xuefeng Bai,Pengfei Zhang,Yang Xiang,Kehai Chen

类目:Computation and Language (cs.CL)

关键词:Audio-visual large language, Audio-visual large, textbf, large language models, made remarkable progress

备注:

点击查看摘要

Abstract:Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a $\textbf{question-relay}$ mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose $\textbf{SECRET}$ ($\textbf{S}$ourc$\textbf{E}$-$\textbf{C}$onditioned $\textbf{RE}$lay s$\textbf{T}$eering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.

61. 【2609.37564】Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging

链接:https://arxiv.org/abs/2609.37564

作者:Zijing Wang,Yongkang Liu,Mingyang Wang,Ercong Nie,Mengjie Zhao,Yunpu Ma,Kang Liu,Zihan Wang,Shi Feng,Daling Wang,Hinrich Schütze

类目:Computation and Language (cs.CL)

关键词:consolidating diverse capabilities, Merging, effective approach, approach for consolidating, task vector

备注: Under review

点击查看摘要

Abstract:Merging pretrained models has emerged as an effective approach for consolidating diverse capabilities into a single unified model. However, prevailing merging methods typically treat each task vector as an indivisible merging unit, overlooking the heterogeneous geometric changes encoded within it. This treatment can induce cross-component coupling: when merging decisions are derived from statistics of the complete task vector, the geometric characteristics of one component may influence how another is selected, weighted, or combined, potentially degrading the quality of the merged model. To address this issue, we propose DiGA, a Disentangled Geometry-Aware model merging framework. Using the pretrained weights as a shared geometric reference, DiGA orthogonally decomposes each task vector into components corresponding to distinct geometric attributes. Rather than merging the task vectors as a whole, DiGA aggregates corresponding components independently within their respective subspaces and subsequently recombines them into a unified update. This component-wise formulation preserves the geometric identity of each component and prevents the characteristics of one component from interfering with the aggregation of another. Furthermore, DiGA can be incorporated into a broad range of existing model merging methods. Extensive experiments across diverse models, tasks, and merging methods demonstrate that DiGA improves merged-model performance and reduces capability degradation. Our repository is on this https URL.

62. 【2609.37543】RunyaNER: Auxiliary Language Selection for Runyankore NER

链接:https://arxiv.org/abs/2609.37543

作者:Prosper Arineitwe Asiimwe,Francois Meyer,Jan Buys

类目:Computation and Language (cs.CL)

关键词:Named Entity Recognition, Entity Recognition, Named Entity, approaches for NLP, NLP tasks

备注: Accepted to the 6th Workshop on Multilingual Representation Learning (MRL 2026) at EMNLP 2026. Camera-ready version. 4 figures

点击查看摘要

Abstract:Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.

63. 【2609.37533】E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

链接:https://arxiv.org/abs/2609.37533

作者:Arseny Ivanov,Alexander Kolesov,Alexander Korotin,Ivan Oseledets,Mikhail Goncharov

类目:Computation and Language (cs.CL)

关键词:Masked diffusion models, limiting sample quality, diffusion speed advantage, autoregressive decoding matters, Masked diffusion

备注:

点击查看摘要

Abstract:Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.

64. 【2609.37515】Hierarchical Compression of Vision-Language Model Benchmarks

链接:https://arxiv.org/abs/2609.37515

作者:Hyunjong Ok,Seunggu Kang,Jaeho Lee

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:prohibitively expensive, relentless pace, span an ever-broader, ever-broader spectrum, spectrum of capabilities

备注: Preprint

点击查看摘要

Abstract:Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.

65. 【2609.37510】From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation

链接:https://arxiv.org/abs/2609.37510

作者:Yuhao Wang,Ruiyang Ren,Yinan Zhang,Ruiqing Zhang,Jing Liu,Chunyan Miao

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:On-policy distillation, On-policy, stronger teacher, teacher, student

备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates. Its {policy disagreement score} combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs. Across eight mathematical reasoning benchmarks, MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark. MAESTRO also reduces average training response length by 67.3\% relative to standard OPD. The code is available at this https URL.

66. 【2609.37501】Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance

链接:https://arxiv.org/abs/2609.37501

作者:Dipankar Sarkar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:regulated agentic workflows, propose RegLLM, agentic workflows, bounded autonomy, autonomy in regulated

备注: 9 pages; ancillary evaluation artefacts. Previously submitted to NLLP 2026

点击查看摘要

Abstract:We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.

67. 【2609.37499】Who Warmed the Archives? LLMs Overestimate Historical Warmth

链接:https://arxiv.org/abs/2609.37499

作者:Claudiu Creanga,Liviu P. Dinu

类目:Computation and Language (cs.CL)

关键词:instrumental climate record, climate record backward, indices climatologists derive, backward in time, derive by hand

备注:

点击查看摘要

Abstract:Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.

68. 【2609.37498】he Rashomon Wikipedia: A Data-Perspectivist Analysis of Divergent Historical Narratives

链接:https://arxiv.org/abs/2609.37498

作者:Claudiu Creanga,Liviu P. Dinu,Anca Dinu

类目:Computation and Language (cs.CL)

关键词:distinct epistemic communities, creating divergent narratives, independent language editions, Battle of Posada, Wikipedia aims

备注:

点击查看摘要

Abstract:Wikipedia aims to provide a unified, neutral record of history, yet its independent language editions often function as distinct epistemic communities, creating divergent narratives around contested events. This paper investigates cross-lingual historiographical bias by analyzing Wikipedia articles across five languages (Romanian, Hungarian, Russian, Turkish, and English) focusing on three contentious events in Romanian history: the Battle of Posada (1330), the Soviet occupation of Bessarabia (1940), and the Night Attack at Targoviste (1462). Using human annotators and Large Language Models (LLMs) to classify citation stance and quantify narrative evolution from 2005 to 2024, we identify a phenomenon of "citation isolation". In the case of the Battle of Posada, only 2 out of 119 citations were shared between language editions, with the Romanian edition exhibiting a 91% pro-national bias compared to the balanced Hungarian edition. Longitudinal analysis reveals that these narratives are volatile and responsive to contemporary geopolitics, evidenced by a significant shift in the Russian framing of Bessarabia in 2024. Finally, we propose a "Peace-Maker" pipeline to automate conflict reconciliation. We demonstrate that while standard prompting leads models to hallucinate consensus, "adversarial" prompting, which explicitly instructs the model to preserve and attribute disagreement, achieves near-perfect neutrality scores.

69. 【2609.37497】Larry Caused the Car to Stop, But the Model Didn't Notice: Transformer Blindness to the M-Heuristic

链接:https://arxiv.org/abs/2609.37497

作者:Stefania Butnaru,Claudiu Creanga

类目:Computation and Language (cs.CL)

关键词:Modern transformer models, reasoning remains understudied, transformer models excel, perform pragmatic reasoning, pragmatic reasoning remains

备注:

点击查看摘要

Abstract:Modern transformer models excel at capturing semantic relationships through sentence embeddings, yet their ability to perform pragmatic reasoning remains understudied. This paper investigates whether encoder-based transformers such as DeBERTa employ the M-Heuristic (the neo-Gricean principle that marked linguistic forms implicate marked meanings). We test this hypothesis by contrasting lexical causatives (e.g., ``Larry stopped the car'') with periphrastic causatives (e.g., ``Larry caused the car to stop'') using a Natural Language Inference framework. Our experiments across 188 conditions with 15 ambitransitive verbs reveal that DeBERTa, RoBERTa, and BART show no evidence of capturing the pragmatic distinction between these forms, with DeBERTa predicting ``Neutral'' for 100% of cases. Probing analysis initially suggested a representation-use dissociation, but control experiments reveal the probe was tracking syntactic complexity, not causative pragmatics. Semantic similarity over 30 triplets places periphrastic causatives closer to unmediated manner descriptions in 29/30 cases, opposite to M-Heuristic predictions in the embedding space. Under explicit metalinguistic framing, Gemini Flash-Lite reaches 100% with item-specific traces, so the principle is available under instruction yet unused in default NLI.

70. 【2609.37494】Your Benchmark Is Not Saturated: Reviving Multiple-Choice Evaluation with Answer Pooling

链接:https://arxiv.org/abs/2609.37494

作者:Mohamed Eltahir,Abobaker Ahmed,Nawaf Barebood,Hussain Bu Subayt,Tanveer Hussain,Naeemullah Khan

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:standard remedy, writing harder items, cheap to grade, slow and repeated, writing harder

备注:

点击查看摘要

Abstract:Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question's wrong options are written for that question alone, so a model can score by eliminating a few options. We propose AnswerPool: take $N$ questions that share a context, pool all their options into one list, and ask the model to assign every question its answer. No item is written and no label changes. The chance of guessing a group right falls from $10^{-3}$ to $5\times10^{-7}$ for five four-option questions, and a model that recognizes its answers keeps its multiple-choice score, so the accuracy lost to pooling measures the credit the format gave for elimination. Deleting answers from the pool makes questions unanswerable with exact ground truth, so abstention is scored in the same pass. Across eight text, image, and video benchmarks and eighteen models, pooling is harder for every model, the elimination credit is largest for the weakest models, and seven of eight open-weight models answer 87 to 100% of unanswerable questions.

71. 【2609.37493】Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

链接:https://arxiv.org/abs/2609.37493

作者:Dongyub Jude Lee,Jungseob Lee,Chanjun Park,Hyeonseok Moon,Heuiseok Lim

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:language model requires, model requires deciding, large language model, verifier ranking accuracy, large language

备注: 29 pages, 6 figures, 24 tables. Dongyub Jude Lee and Jungseob Lee contributed equally

点击查看摘要

Abstract:Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at this https URL.

72. 【2609.37491】Regime Boundary Alignment for Evidence-Gated Question Answering

链接:https://arxiv.org/abs/2609.37491

作者:Zeyan Li,Qirong Guo,SIyuan Qiu,Hu Xu,Chun Li,Jianfeng Xu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Retrieval-augmented language models, Retrieval-augmented language, language models, models are expected, Regime Boundary Alignment

备注:

点击查看摘要

Abstract:Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence is missing. We trace this behavior to the training signal: answer-focused fine-tuning assigns no target to unsupported contexts, so it cannot distinguish a reader that abstains from one that guesses, and unsupported answering stays near 100% even as supported accuracy improves. We introduce Regime Boundary Alignment (RBA), which trains a single reader on matched variants of the same question and gold answer. The reader is trained to produce the gold answer when the context supports it, including when conflicting evidence is also present, and to abstain when the correct support is removed; inference is ordinary decoding, with no verifier, threshold, or regime label. On three multi-hop QA datasets across three seeds, RBA reduces the unsupported-answer rate by more than sixty percentage points relative to conflict-focused training while matching its supported accuracy. On a held-out TriviaQA retrieval-miss slice, the same reader reduces unsupported answering from 100% to below 1% while also improving supported accuracy. These results indicate that evidence-gated answering must be learned on both sides of the support boundary.

73. 【2609.37488】FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding

链接:https://arxiv.org/abs/2609.37488

作者:Taiyo Sato,Takamasa Sanda,Keisuke Maeda,Takahiro Ogawa,Miki Haseyama,Shunya Nagashima

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:single prompted call, multimodal large language, referring expression comprehension, Frozen multimodal large, large language models

备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.

74. 【2609.37469】Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

链接:https://arxiv.org/abs/2609.37469

作者:Suting Chen,Peichun Hua,Yunming Xiao

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:grounds large language, Retrieval-augmented generation, grounds large, external sources, entities without providing

备注: 22 pages, 7 tables, 2 figures

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence tests: removing relevant evidence or pairing evidence with unrelated questions can reveal labels through lexical overlap or evidence position. We build a paired benchmark using substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. Sufficiency can be judged without generating an answer, but no single signal works across all datasets. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals: whether every part of the question is covered, whether any passage offers an answer, and whether a small language model reading the passages together judges them sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), exceeding the best of 10 prior methods (0.746) and a frontier model queried through an API (0.784). Its weakest dataset scores higher than any other method's weakest (0.684 vs. 0.676). RINSE runs locally before generation, taking 36.5 ms per question on a single GPU.

75. 【2609.37443】Learning to Retrieve Missing Evidence for Long-Term Memory QA

链接:https://arxiv.org/abs/2609.37443

作者:Yi-Xuan Deng,Yi Zhang,Wei Liu,Chao Xue,Shuojin Yang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Long-term memory enables, enables language models, memory enables language, Long-term memory, future conversations

备注: 22pages,6figures

点击查看摘要

Abstract:Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.

76. 【2609.37408】Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse

链接:https://arxiv.org/abs/2609.37408

作者:Annabelle K. L. Chua,Forster J. Khoo,Joel C. R. Tan,Huey Ting Ang,Kheng Hwee Tan,Joel Y. A. Sim,Shirley W. H. Ow,Ria Mundhra,Elsie C. K. Toh,Youfeng Xu,Lynnette H. X. Ng

类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:rigorous detection systems, identify online hate, supporting the construction, detection systems, identify online

备注: Accepted to IDeaS Conference 2026

点击查看摘要

Abstract:Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect more precise and interpretable narratives, we present an extraction pipeline that represents narratives as entity-evaluation pairs. Narratives are extracted using a Large Language Model (LLM) reasoning process that extends Aspect-Based Sentiment Analysis, identifying the aspect, classifying its judgement type as the basis for evaluation, and deriving the evaluation accordingly. Extracted narratives are then clustered using Leiden, following which clusters are resolved to an intended level of granularity through an LLM-guided refinement process. We illustrate this narrative pipeline with English Reddit comments from 2024 that criticize Taylor Swift, analyzing a representative cluster that exhibits hate speech patterns to demonstrate its interpretive value.

77. 【2609.37371】Compiling Learning Problems into Adaptation Programs for Language Models

链接:https://arxiv.org/abs/2609.37371

作者:Rebecca Ramnauth,Brian Scassellati

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:fixed recipe, typically governed, produce substantially, behavioral outcomes, adaptation

备注:

点击查看摘要

Abstract:Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as a joint prediction and decision problem. Rather than searching over candidate programs anew for each learning episode, a compiler learns from prior adaptations to predict a vector-valued counterfactual response surface over candidate programs---their expected effects on acquisition, transfer, boundedness, and preservation---and selects a program before adaptation begins. Because this predicted geometry captures multiple behavioral consequences rather than a single winner or scalar score, it can be reused under different downstream priorities without retraining. Across five learning types, preferred programs vary meaningfully across episodes, and this variation is predictable from pre-adaptation information. On Llama-3.1-8B, compiler-selected programs approach exhaustive search while outperforming global and objective-specific defaults. Replication on Gemma-2-9B preserves program heterogeneity and selection headroom, but shows that exploiting this headroom requires accounting for uncertainty when departing from strong defaults. Together, these results show that adaptation search can be amortized across related learning problems, turning prior adaptation experience into a basis for deciding how future learning should occur.

78. 【2609.37361】SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search

链接:https://arxiv.org/abs/2609.37361

作者:Zetong Zhou,Wentao Zhang,Jingyuan Wang,Yifan Yang,Zizhuo Wang,Shixi Hu

类目:Computation and Language (cs.CL)

关键词:Operations research supports, research supports decision-making, Solving operations research, Operations research, operations research problems

备注: Accepted at EMNLP 2026 (Findings)

点击查看摘要

Abstract:Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.

79. 【2609.37351】Port-Hamiltonian Latent Deliberation: Mitigating the Deliberation Drift Cliff in Test-Time Compute Scaling

链接:https://arxiv.org/abs/2609.37351

作者:Zeyu Jia(School of Biomedical Engineering and Technology, Tianjin Medical University, Medical School, Tianjin University)

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:representation spaces reveals, iterative deliberation directly, performing iterative deliberation, continuous latent representation, latent representation spaces

备注: 10 pages, 1 figure, 4 tables. Code and evaluation artifacts available

点击查看摘要

Abstract:Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains at short horizons (K = 4), their reasoning collapses when extrapolated to deeper thinking steps (K = 16), dropping by 22% to 62% across standard logical benchmarks. We resolve the trilemma among expressivity, Lyapunov stability, and computational efficiency in test-time latent reasoning through a 22-round empirical and theoretical investigation. We demonstrate that strictly conservative scalar potential gradient flows suppress long-range drift (cliff 3.40%) but bottleneck peak reasoning accuracy at 32.73%, whereas unconstrained rotational flows achieve high symbolic expressivity (82.33%) but suffer a severe 36.87% drift cliff. To resolve this geometric duality, we establish Port-Hamiltonian Latent Deliberation (PH-LD) and propose the Direct-Gradient Pure-Tensor Helmholtz-Hodge Decomposition (DG-HHD). DG-HHD parameterizes the attracting flow as a tangent projection tensor network while orthogonally decoupling non-zero circulation (Hodge machine error 1.65e-17, contraction error 5.55e-17), eliminating runtime autograd dependencies to achieve 1.84x vector field and 2.09x RK45 rollout speedups. In a 15-arm symmetrical Pareto benchmark, DG-HHD achieves 58.67% peak accuracy (+25.94% absolute gain over conservative HHD) and retains 35.27% at K=32. Transferred to small language model (SLM) multi-hop causal reasoning, DG-HHD delivers monotonic compute scaling (49.33% to 51.56%) and suppresses out-of-distribution drift (cliff -0.66%). All 30 Level 0 deterministic invariants are certified.

80. 【2609.37326】Solving Without Stopping: On-Policy Distillation at Small Scale

链接:https://arxiv.org/abs/2609.37326

作者:Hongyang Li,Yiming Zhu,Xiao Li,Caesar Wu,Said Mammar,Pascal Bouvry

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:stronger teacher feedback, thinking mode, student, reasoning, mode

备注: 22 pages, 13 figures

点击查看摘要

Abstract:On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student's existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.

81. 【2609.37312】Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring

链接:https://arxiv.org/abs/2609.37312

作者:Mohammadali Mohammadkhani,Madhava Krishna,Yash Sarrof,Michael Hahn

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:models trick chain, reasoning models trick, chain of thought, thinking traces, perform hidden computation

备注:

点击查看摘要

Abstract:Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.

82. 【2609.37236】Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

链接:https://arxiv.org/abs/2609.37236

作者:Ido Levy,Asaf Yehudai,Segev Shlomov,Asaf Adi,Leshem Choshen

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:tools typically responds, user explicitly, tools typically, typically responds, user never requested

备注: 48 pages. Project page: [this https URL](https://dolev31.github.io/ProactiveInquirer/) Code: [this https URL](https://github.com/dolev31/ProactiveInquirer) Model: [this https URL](https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B)

点击查看摘要

Abstract:An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose QD (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model $15\times$ larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the $15\times$ larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.

83. 【2609.37226】Follow the Entities: A Corpus Map for Agentic Search

链接:https://arxiv.org/abs/2609.37226

作者:Soyeong Jeong,Sujay Kumar Jauhar,Sung Ju Hwang,Andrew Joohun Nam

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:project approval recorded, Answering questions, requires connecting evidence, connecting evidence spread, questions and completing

备注:

点击查看摘要

Abstract:Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.

84. 【2609.37223】CredWise: A Controlled Agentic Decision-Intelligence Framework for Explainable and Auditable Credit-Risk Assessment

链接:https://arxiv.org/abs/2609.37223

作者:Aakash Kumar Tiwari

类目:Computation and Language (cs.CL)

关键词:integrates credit-risk prediction, Credit-risk prediction, important in banking, applicant is risky, Lending Club data

备注:

点击查看摘要

Abstract:Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with other evidence. This paper presents CredWise, a decision-support framework that integrates credit-risk prediction, probability calibration, explainable artificial intelligence, policy retrieval, SQL analytics, and controlled agent-based workflows. An XGBoost model is trained on Lending Club data (1,345,310 loans, 18 features) using a temporal split: 2007--2016 for training, 2017 for validation, and 2018 for testing. On the 2018 test set, the calibrated model achieved a ROC-AUC of 0.7109, PR-AUC of 0.2993, F1-score of 0.3714, and accuracy of 65.44\%. Calibration reduced the Brier score from 0.2157 to 0.1273 and the expected calibration error from 0.2862 to 0.0585. SHAP explanations were temporally stable, with a Spearman correlation of 0.9959 between 2017 and 2018 feature rankings. On 28 labeled queries covering nine policy sections, FAISS achieved the best Hit@1 (0.929) and MRR (0.964), while all three retrieval methods reached Hit@5 = 1.0. Agent routing achieved 95.6\% accuracy (43 of 45 cases), and the SQL benchmark scored 1.0 on exact-match, execution-success, and result-match across six cases. These results show that CredWise can combine predictions, explanations, policy evidence, and structured analytics in one controlled workflow. It is an academic research prototype, and final decisions remain with a human reviewer.

85. 【2609.37175】VLM Fine-Tuning for End-to-End Combinatorial Optimization

链接:https://arxiv.org/abs/2609.37175

作者:Qingsong Yan,Xia Jiang,Yaoxin Wu,Wen Song,Lu Zhang,Yingjie Zhou

类目:Computation and Language (cs.CL)

关键词:combinatorial optimization, Large language models, provided a unified, unified interface, obscure spatial

备注:

点击查看摘要

Abstract:Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.

86. 【2609.37171】Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion

链接:https://arxiv.org/abs/2609.37171

作者:Xinkai Du,Chao Lv,Yalin Sun,Quanjie Han,Lei Yao,Maosong Sun

类目:Computation and Language (cs.CL)

关键词:seamlessly integrating information, natural language processing, large language models, integrating information retrieval, Retrieval-Augmented Generation

备注: This paper is accepted by NLPCC 2026

点击查看摘要

Abstract:Retrieval-Augmented Generation has established itself as a fundamental framework in natural language processing, seamlessly integrating information retrieval with the generative capabilities of large language models. However, this process is fundamentally constrained by a critical challenge: semantic space mismatch between queries and retrieved contexts. We propose Knowledge-Aware Semantic Bridging (KASB), a novel framework that improves passage selection quality through semantic space alignment between queries and retrieved documents through intelligent knowledge fusion. Our approach leverages the complementary strengths of generative and retrieval-based knowledge through a multistage process that enhances both relevance and accuracy. We evaluate KASB on three popular open-domain Question Answering datasets to demonstrate the effectiveness of our approach.

87. 【2609.37169】rajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

链接:https://arxiv.org/abs/2609.37169

作者:Zhehao Huang,Changxin Tian,Qingyuan Yang,Kunlong Chen,Ziqi Liu,Zhiqiang Zhang,Xiaolin Huang,Jun Zhou

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:equips pretrained large, pretrained large language, Mid-training equips pretrained, large language models, compute mid-training absorbs

备注:

点击查看摘要

Abstract:Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.

88. 【2609.37148】Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction

链接:https://arxiv.org/abs/2609.37148

作者:Siddhant Jain,Dimitra Tsovaltzi

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:learn and grow, regulates their emotions, difficult conversation, directly observable, qualities that matter

备注: 8 pages, 6 figures

点击查看摘要

Abstract:Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind of clean labeling that most machine learning pipelines are built around. We study this challenge through a case that is well grounded in psychological theory but rarely modeled computationally: self-compassion, the tendency to respond to one's own setbacks with patience rather than harsh self-criticism. We examine how it appears during structured reflective interviews in a technology-mediated training setting, where people naturally talk through socio-emotionally demanding situations. Since no existing dataset captures this kind of construct in this kind of setting, we collected and annotated 51 reflective dialog sessions using an independent, temporally overlapping annotation scheme grounded in established theory. We consolidate the underlying six-component psychological model into a three-class supervision space, balancing self-kindness and mindfulness against self-critical or overwhelmed states, and build a reproducible window-based pipeline that aligns video, audio, and text on a shared timeline. Unimodal models trained on each modality separately are compared against a simple probability-level fusion strategy, which yields modest but consistent gains over the best single modality. We close by discussing where each modality succeeds or struggles, what this suggests about how this kind of construct is actually expressed in reflective speech, and what would be needed to model it, and constructs like it, more effectively.

89. 【2609.37143】LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

链接:https://arxiv.org/abs/2609.37143

作者:Yun Peng,Zihan Wu,Zeyang Zhuang,Xin Zhou,Rui Shu,Xu Han,Chun Yong Chong,Yuan Wang,Jiakun Liu

类目:oftware Engineering (cs.SE); Computation and Language (cs.CL)

关键词:Modern coding agents, Modern coding, recent benchmarks reflect, deliver increasingly large, increasingly large repository-level

备注:

点击查看摘要

Abstract:Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at this https URL.

90. 【2609.37127】LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility

链接:https://arxiv.org/abs/2609.37127

作者:Kajetan Ożóg,Alicja Wojciechowska,Dawid Malarz,Paweł Batorski,Artur Kasymov,Przemysław Spurek

类目:Computation and Language (cs.CL)

关键词:acquiring negative connotations, Establishing unbranding, image generation, critical practice, practice to prevent

备注:

点击查看摘要

Abstract:Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This frequency introduces significant risks, such as trademark dilution, false attribution, and brand defamation. In response, we formally define the novel task of LLM Unbranding. We specifically address the complex challenge of managing trade dress within textual outputs. This involves neutralizing characteristic language, slogans, and stylistic markers that define brand identity. Crucially, these elements are less evident than explicit visual logos. To benchmark this task, we introduce a comprehensive evaluation dataset incorporating prominent brands from multiple commercial domains. We rigorously evaluate existing state-of-the-art machine unlearning models using this benchmark. This evaluation identifies their limitations in selective textual unbranding. Finally, we propose MUTE, a novel inference-time method that effectively neutralizes textual trade dress while preserving the LLM's general capabilities and utility. By leveraging an iterative refinement loop, MUTE systematically optimizes system instructions to safely eliminate brand leakage without requiring fragile parameter updates. Code and dataset: The evaluation dataset and code for LLM Unbranding are available at this https URL. The implementation of MUTE is available at this https URL.

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.37127 [cs.CL]

(or
arXiv:2609.37127v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.37127

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
91. 【2609.37121】Cross-Linguistic Effects in Bilingual Phoneme BabyLMs

链接:https://arxiv.org/abs/2609.37121

作者:Nikitas Theodoropoulos,Maria Lymperaiou,Giorgos Filandrianos

类目:Computation and Language (cs.CL)

关键词:Cross-linguistic effects, bilingual first-language acquisition, first-language acquisition, central topic, Cross-linguistic

备注: 13 pages, 8 figures, 3 tables; Accepted at the 2nd BabyLM Workshop at EMNLP 2026

点击查看摘要

Abstract:Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by enabling controlled comparisons across language combinations and learning conditions. Recent work explores this direction by training bilingual language models under developmentally plausible constraints. However, human and model learners still diverge in fundamental ways, with one major difference being input modality: children learn primarily from spoken input, whereas language models are typically trained on orthographic text. To reduce this gap, researchers have trained models on phonemic representations of speech. In this work, we combine these research directions to train bilingual BabyLMs with phonemic input. We keep English fixed as the L2 and vary the L1 across German, Swedish, Persian, and Basque, selected to represent contrasting combinations of syntactic and phoneme-inventory distance from English. Our results show stronger L1-related variation in grammatical learning trajectories under phonemic than orthographic input, while early lexical differences align with phoneme-inventory similarity.

92. 【2609.37119】Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

链接:https://arxiv.org/abs/2609.37119

作者:Hongyang Li,Xiao Li,Caesar Wu,Said Mammar,Grégoire Danoy,Pascal Bouvry

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, language models increasingly, models increasingly remove, Recent approaches, approaches to reinforcement

备注: 26 pages, 15 figures, 16 tables

点击查看摘要

Abstract:Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.

93. 【2609.37105】VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

链接:https://arxiv.org/abs/2609.37105

作者:Jiexing Qi,Yu He,Jun Liu,Qichen Huang,Shaohua Hu,Zhan Dang,Guohua Chen,Rui Yang,Wen Jiang,Yang Liu,Tao Lyu,Fangming Li

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Language model agents, guides task execution, Language model, task execution, improved by updating

备注:

点击查看摘要

Abstract:Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.

94. 【2609.37104】What Does Post-Training Change in Multilingual Reasoning?

链接:https://arxiv.org/abs/2609.37104

作者:Hongyang Li,Xiao Li,Caesar Wu,Grégoire Danoy,Pascal Bouvry

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:provide unequal access, Open-source reasoning models, models provide unequal, Open-source reasoning, provide unequal

备注: 20 pages, 9 figures, 21 tables. Main paper and supplementary material in one document

点击查看摘要

Abstract:Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.

95. 【2609.37082】raverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

链接:https://arxiv.org/abs/2609.37082

作者:Jingyuan Ma,Lynx Aster,He Zhang,Siyao Song,Weijie Yuan,Zhe Zhang,Kai Jia,Zhifang Sui

类目:Computation and Language (cs.CL)

关键词:Long-horizon information-seeking agents, recovery increasingly difficult, causing early mistakes, making recovery increasingly, Long-horizon information-seeking

备注:

点击查看摘要

Abstract:Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.

96. 【2609.37044】Learning from Think-Mode Advantage via On-Policy Distillation

链接:https://arxiv.org/abs/2609.37044

作者:Wanqi Ren,Jianxiang Wang,Danxuan Liu,Linyi Ding,Yuan Zhang

类目:Computation and Language (cs.CL)

关键词:Explicit intermediate reasoning, large language models, stronger problem-solving mode, Explicit intermediate, language models

备注: 9 pages, 5 figures

点击查看摘要

Abstract:Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when responses reach the same outcome. We summarize this interaction with trace-response divergence (TRD) and introduce ThinkOPD, which routes supervision at the response level by combining group-relative reward gain with a TRD-based compatibility proxy. Final response weights are normalized within each rollout group. Across mathematical reasoning and code generation, ThinkOPD outperforms Uniform ThinkOPD in both same-model settings and both cross-model teacher-student pairs, and it exceeds representative rationale and self-distillation baselines in a controlled comparison. Controlled interventions show that outcome benefit and the TRD-based proxy provide complementary routing signals in this setting. Think-enabled OPD provides a controlled setting for studying how teacher advantage becomes transferable along student responses.

97. 【2609.37040】Selecting The Most Informative Tokens in Natural Language Autoencoders

链接:https://arxiv.org/abs/2609.37040

作者:Federico Torrielli,Gianluca Barmina,Andrea Blasi Núñez,Amon Rapp,Luigi Di Caro,Peter Schneider-Kamp,Lukas Galke Poech

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Natural language autoencoders, language autoencoders translate, Natural language, language model internal, language autoencoders

备注:

点击查看摘要

Abstract:Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.

98. 【2609.37017】LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration

链接:https://arxiv.org/abs/2609.37017

作者:Shinan Zhang,Tao Zhang,Qihui Zhu,Mengjie Zhang,Dong Jin,Yunpeng Hou,Shuangwu Chen,Xiaobin Tan,Quan Zheng,Jian Yang

类目:Computation and Language (cs.CL)

关键词:repeated encoding-decoding overhead, LLM-based multi-agent systems, natural-language communication, avoid the information, information loss

备注:

点击查看摘要

Abstract:LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.

99. 【2609.36987】CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence

链接:https://arxiv.org/abs/2609.36987

作者:Yuzhe Zhang,Weijie Zhu,Haolin Yang,Ziyun Zhang,Xianwei Xue,Mengke Chen,Qiutong Pan,Huaqian Cai

类目:Computation and Language (cs.CL)

关键词:isolated single-turn queries, existing benchmark evaluates, benchmark evaluates isolated, natural language, analysts actually work

备注: Accepted as an oral paper at EMNLP 2026

点击查看摘要

Abstract:Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level correctness remains below 5%. Second, despite strong overall rank correlation, frontier models exhibit a consequential reordering of the top of the leaderboard under autonomous operation, a phenomenon we term the Autonomy Divergence, which reveals error-management as a partially independent capability from raw generation skill. Third, scaling action budgets from x3 to x10 fails to close the autonomy gap, as the strongest frontier models self-limit to approximately two actions per turn regardless of available budget. Fourth, single-turn Cypher fine-tuning degrades multi-turn instruction following, while architecture-appropriate specialization outperforms several frontier models. These results establish CypherTurn as an open challenge for conversational graph database reasoning. Code and data are available at this https URL.

100. 【2609.36982】SRJudge: Empowering Large Language Models with Selective Reasoning for Fine-Grained Knowledge Concept Tagging

链接:https://arxiv.org/abs/2609.36982

作者:Zhiwei Yang,Jiahua Yang,Huiru Lin,Xing Chen,Quanlong Guan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:online teaching practices, assign specific concept, Knowledge concept tagging, concept tagging aims, educational content

备注: Accepted by IJCAI 2026

点击查看摘要

Abstract:Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grained knowledge concept tagging. Specifically, the Selector in Stage 1 first narrows the candidate concepts to a top-K shortlist by fine-tuning a small language model (SLM), e.g., BERT, since the top-$K$ predictions hit the correct concept in most cases, thereby reducing the decision space of correct candidates. Next, the Stage 2 Reasoner employs a lightweight LLM for refined reasoning over the shortlisted candidates. It further integrates an improved reinforcement learning strategy with a dynamic task-specific reward function and a pruning mechanism to better align with human reasoning preferences. Finally, a larger LLM acts as a judger that evaluates the overall rationality of the reasoning process and its explanations to determine the final output. In addition, we construct two high-quality datasets for further validation, i.e., the biology dataset S_Bio and the physics dataset S_Phy. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines across benchmark datasets, verifying its effectiveness and superiority. Resources are available at: this https URL.

101. 【2609.36976】AMU:Admission and Memory Update for Personalized Conversations---Structured Memory with SLM Guided Control

链接:https://arxiv.org/abs/2609.36976

作者:Tao Hwang,Yishi Diao

类目:Computation and Language (cs.CL)

关键词:Large language models, interactions remains challenging, maintaining persistent user, long-term interactions remains, Large language

备注: 14 pages, 2 figures. Source code and implementation are available at: [this https URL](https://github.com/UnicusT11/AMU-memory)

点击查看摘要

Abstract:Large language models (LLMs) have become the foundation of personalized assistants, but maintaining persistent user memory across long-term interactions remains challenging. Existing memory systems often focus on storage, retrieval, or consolidation, while memory writing remains less controlled: transient requests, duplicate statements, and outdated user states may enter memory and later be retrieved for personalization. In this paper, we present AMU: Admission and Memory Update for Personalized Conversations, an SLM-guided (Small language model guided) structured framework for writing-time memory control. AMU uses structured memory filtering to decide what should enter memory and SLM-guided storage management to determine whether an admitted record should be stored separately, discarded as a duplicate, or fused as an update. We evaluate AMU in a controlled memory writing and retrieval setting. Experimental results show that AMU maintains cleaner and more retrievable personalized memories.

102. 【2609.36974】Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech

链接:https://arxiv.org/abs/2609.36974

作者:Kirill Borodin,Vasilii Kudryavtsev,Maxim Maslov,Grach Mkrtchian

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:truncate and lose, phrase many times, models loop, lose count, Abstract

备注: Submitted to IEEE ICASSP 2027. Code and data: [this https URL](https://github.com/lab260ru/tts-counting-failure)

点击查看摘要

Abstract:Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k = 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.

103. 【2609.36965】Chinese-Jev: Bringing System One Model to Chinese-Language Tasks

链接:https://arxiv.org/abs/2609.36965

作者:Zexiao Wang,Zihao Zhang,Xudong Wang,Pan Wang,Ziyi Ye,Haoyu Zhao,Zuxuan Wu,Shuicheng Yan

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:generative language models, open-ended responses, alternative to generative, generative language, Jev

备注: 10 pages, 6 figures

点击查看摘要

Abstract:System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at this https URL.

104. 【2609.36958】VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers

链接:https://arxiv.org/abs/2609.36958

作者:Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Peng Zhang,Daren Zha,Jun Xiao

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:Repeated verifier calls, contribute conditional information, Repeated verifier, conditional marginal information, contribute conditional

备注: 27 pages, 5 figures

点击查看摘要

Abstract:Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertainty, normalizes by call cost, and stops or abstains when the next call is not informative. The controller freezes its decision and cost ledger before joining the clean oracle; a dependence-shift alarm disables channel preference and falls back to exact-stop. The controlled audit gives the mechanism boundary: at 35% symmetric corruption, majority-5 improves balanced accuracy from 0.6578 to 0.7739, whereas at 65% it loses 0.1226 points. In the matched fixed-budget comparison, breadth, redundancy, and adaptive allocation obtain balanced accuracies 0.6048, 0.6375, and 0.6538, with 3.4216 calls per item and an RLVR score of 0.6417 for VStress-CA. Dependence diagnostics also increase from same-model repeats to cross-family channels, with conditional marginal gains of 0.0126, 0.0462, and 0.0913. These measurements turn correlation from a post-hoc warning into an auditable allocation decision.

105. 【2609.36953】Cool the Sampler, Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO

链接:https://arxiv.org/abs/2609.36953

作者:Taiheng Pan

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:truncated importance weight, cooled sampler refreshed, repairs the resulting, resulting mismatch, importance-corrected GRPO refreshed

备注: 14 pages, 8 figures, 4 tables

点击查看摘要

Abstract:Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference model and the importance weights stay at temperature 1, with the behaviour probability recorded from the tempered distribution, so the learner's objective is unchanged. All corresponding cooled runs are stable, and the longer interval keeps what the short one delivered: at the same update budget, a cooled sampler refreshed every 192 steps matches an uncooled sampler refreshed every 96 at the end of training (0.857 for both) and averaged over it (0.79), whereas lowering the learning rate to a safe value ends 3-7 points lower. On Qwen2.5-Math-7B the degradation points at interval 192 predict that an interval of 144 is fatal without cooling and survivable with it; on two data seeds the uncooled runs degrade before their first refresh and the cooled runs pass it and end at 92-93% against 68-81%, with one cooled run degrading transiently late in the second cycle. The benefit has a window: at three times the safe interval and in a high-mismatch MATH setting cooling delays degradation without preventing it, stronger cooling is not better, and cooling without the correction collapses. Sampling temperature is a control on staleness tolerance, and temperature and refresh interval should be chosen together.

106. 【2609.36952】ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models

链接:https://arxiv.org/abs/2609.36952

作者:Jingnan Pu,Zi-En Fan,Feng Lian

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, lack comprehensive perception, undesirable abstract semantics, learn undesirable abstract, Large language

备注: 20 pages, 15 figures, 6 tables

点击查看摘要

Abstract:Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.

107. 【2609.36935】CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory

链接:https://arxiv.org/abs/2609.36935

作者:Jingguang Li,Yebo Wu,Zuyi Guo,Kailang Ma,Xianjie Dai,Han Zheng,Benwang Chen,Li Li,Can Rong,Heye Huang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, context length increases, long-horizon tasks, length increases, complex and long-horizon

备注: 38 pages, 13 figures. Code repository: [this https URL](https://github.com/benmagnifico/CoEM)

点击查看摘要

Abstract:Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: this https URL.

108. 【2609.36931】Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

链接:https://arxiv.org/abs/2609.36931

作者:Mario Sanz-Guerrero,Minh Duc Bui,Manuel Mager,Katharina von der Wense

类目:Computation and Language (cs.CL)

关键词:prior work shows, LLM outputs vary, hardware and batching, essential for scientific, prior work

备注: Accepted to AACL 2026 (Main)

点击查看摘要

Abstract:Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.

109. 【2609.36920】Benchmarking Automatic Speech Recognition Tools for Iberian Languages

链接:https://arxiv.org/abs/2609.36920

作者:Fernando López,Pablo Gómez,David Solans,Paulo Villegas,Jordi Luque

类目:Computation and Language (cs.CL)

关键词:Iberian languages remain, automatic speech recognition, languages remain limited, Comprehensive evaluations, Iberian languages

备注: Accepted in IberSPEECH 2026

点击查看摘要

Abstract:Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.

110. 【2609.36914】Can Language Models Learn to Forecast Stock Prices

链接:https://arxiv.org/abs/2609.36914

作者:Jiacheng Guo,Suozhi Huang,Shuzhen Li,Yunlong Gao,Zerui Cheng,Jason Ge,Shushu Liang,Zihao Li,Hao Lu,Ming Yin,Shilong Liu,Jiashuo Liu,Xu Kuang,Mengdi Wang

类目:Computation and Language (cs.CL)

关键词:including mathematical reasoning, software engineering, improve language models', significantly improve language, including mathematical

备注: 18 pages, 4 figures

点击查看摘要

Abstract:Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.

111. 【2609.36913】BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

链接:https://arxiv.org/abs/2609.36913

作者:Chihiro Taguchi,Yotaro Kubo,Rujikorn Charakorn

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Transcribing domain-specific entities, rare proper nouns, Transcribing domain-specific, automatic speech recognition, Latent Encoded Entities

备注: 5 pages, 2 figures, 2 tables

点击查看摘要

Abstract:Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.

112. 【2609.36903】MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

链接:https://arxiv.org/abs/2609.36903

作者:Ke Wang,Houxing Ren,Zimu Lu,Yunqiao Yang,Zhuofan Zong,Mingjie Zhan,Hongsheng Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)

关键词:systems remain limited, brought open-source machine, machine conversation closer, human-like interaction, existing systems remain

备注: NeurIPS 2026

点击查看摘要

Abstract:End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{this https URL}{MultiTalkPT}$ and $\href{this https URL}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{this https URL}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.

113. 【2609.36902】RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection

链接:https://arxiv.org/abs/2609.36902

作者:Wenbin Shen,Guoxuan Qin,Guangxu Yao,Baodong Wang,Yuanbo Rui,Zhongjie Ba,Zhichao Lian

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Existing multimodal fake, assist detection, information to assist, Existing multimodal, Evidence Retrieval Framework

备注:

点击查看摘要

Abstract:Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.

114. 【2609.36893】Momentum-Coupled Rubric Adaptation for Detailed Image Captioning

链接:https://arxiv.org/abs/2609.36893

作者:Zhenwen Ji,Lei Jin,Shanyong Wang,Jiaming Lu,Chengqiang Lu,Yi Wu,Yao Hu,Lizhen Cui,Yanyu Xu

类目:Computation and Language (cs.CL)

关键词:fine-grained visual content, spans factual accuracy, quality spans factual, caption quality spans, Detailed image captioning

备注: 28 pages, natural language processing, computer vision

点击查看摘要

Abstract:Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83\%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.

115. 【2609.36892】Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

链接:https://arxiv.org/abs/2609.36892

作者:Zeyu Gan,Zixuan Gong,Yong Liu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, continue to advance, increasing attention, large language, attention is turning

备注:

点击查看摘要

Abstract:As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.

116. 【2609.36850】Rethinking Multimodal Fake News Detection in the Generative AI Era

链接:https://arxiv.org/abs/2609.36850

作者:Wenbin Shen,Guoxuan Qin,Guangxu Yao,Baodong Wang,Yuanbo Rui,Zhichao Lian

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:simply manipulated material, increasingly entering, entering the production, production and dissemination, manually fabricated

备注:

点击查看摘要

Abstract:Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news event is true. To bridge the separation between these tasks in data and evaluation, we construct Weibo26, a multimodal fake news detection dataset for generative-content scenarios. On this basis, we propose the Generativity-Aware Hierarchical Reasoning (GAHR) framework, which combines global judgment with local correction so that generativity information participates in news-veracity reasoning. Experiments on multiple existing fake news detection benchmarks and Weibo26 show that GAHR achieves competitive veracity-detection performance while effectively identifying generative content.

117. 【2609.36838】On-Policy Visual Evidence Distillation

链接:https://arxiv.org/abs/2609.36838

作者:Shaohang Wei,Feifan Song,Guangyue Peng,Wenhao Yu,Wei Li,Wen Luo,Yang Xu,Yufan Shen,Luke Mao,Yang Du,Asher Qin,Houfeng Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:agents solve problems, Visual agents solve, image operations, solve problems, problems by interleaving

备注: 44 pages, including appendices. Project page: [this https URL](https://sylvain-wei.github.io/ReVuE/) . Code: [this https URL](https://github.com/sylvain-wei/ReVuE)

点击查看摘要

Abstract:Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at this https URL

118. 【2609.36820】CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

链接:https://arxiv.org/abs/2609.36820

作者:Wenbin Hu,Huihao Jing,Haochen Shi,Yuxuan Liu,Haoran Li,Yangqiu Song

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Relative Policy Optimization, Group Relative Policy, Group Relative, Policy Optimization, Relative Policy

备注:

点击查看摘要

Abstract:Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at this https URL.

119. 【2609.36804】VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction

链接:https://arxiv.org/abs/2609.36804

作者:Yitong Han,Nankai Lin,Juan Luo,Hongyan Wu,Lianxi Wang,Shengyi Jiang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Chinese Semantic Error, Semantic Error Correction, targets semantic errors, Chinese Semantic, Semantic Error

备注:

点击查看摘要

Abstract:Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.

120. 【2609.36798】Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs

链接:https://arxiv.org/abs/2609.36798

作者:Yueran Ma,Ronghao Lin

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Omni-modal large language, large language models, Factorized Modality Diagnostic, large language, explicitly refers

备注: 25 pages, 11 figures, 16 tables

点击查看摘要

Abstract:Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at this https URL.

121. 【2609.36760】QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

链接:https://arxiv.org/abs/2609.36760

作者:Zunhai Su,Yuxuan Sun,Jianchao Tan,Tao Zhang,Ruihan Hu,Yuchen Xie,Xunliang Cai,Ngai Wong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Multi-Head Latent Attention, Multi-Head Latent, expressive multi-head attention, enables expressive multi-head, expressive multi-head

备注:

点击查看摘要

Abstract:Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.

122. 【2609.36750】Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

链接:https://arxiv.org/abs/2609.36750

作者:Yiming Wang,Yikang Liu,Qingyuan Tian,Xingyu Chen,Zhuosheng Zhang,Zhaopeng Tu,Rui Wang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Self-rewarding reinforcement learning, enables large language, Self-rewarding reinforcement, large language models, enables large

备注:

点击查看摘要

Abstract:Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.

123. 【2609.36742】SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

链接:https://arxiv.org/abs/2609.36742

作者:Zhenrui Yue,Huimin Zeng,Yueqi Wang,Yaokun Liu,Fengran Mo,Jinghan Zhang,Mung Yao Jia,Gyuseok Lee,Yang Zhang,Na Wei,Dong Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:improving large language, large language models, sparse outcome rewards, outcome rewards lack, intermediate steps

备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.

124. 【2609.36738】Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space

链接:https://arxiv.org/abs/2609.36738

作者:Yuchen Li,Zongqi Fan,Nguyen H. Tran,Ken-Tye Yong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:makes history costly, parameter-sized moving average, Backpropagated Output Momentum, compact moving average, moving average

备注: 53 pages, 10 figures

点击查看摘要

Abstract:Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the model output and reprojects that history through the current network at every step. A batch-level analysis characterizes the information retained and omitted by this relocation, while the implementation preserves the current supervised gradient and can replace the first-moment component of several adaptive optimizers. As a plug-in for momentum-based optimizers, including ones that already compress their state, BOM reduces parameter-shaped optimizer state by 49.7-99.8% in three compositions and, averaged over three language backbones, paired step time by 4.0%. It also improves mean validation performance across language and vision fine-tuning, by 1.42 points in the primary five-task comparison. Language and vision pretraining studies, together with matched mechanism controls, further test the construction across output spaces and model scales.

125. 【2609.36737】Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

链接:https://arxiv.org/abs/2609.36737

作者:Eric Ming Chen,Jin Woo Lee,Vincent Sitzmann

类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:human body responsible, vocal tract, vocal tract shapes, vocal, vocal tract solely

备注: Accepted as NeurIPS 2026 spotlight paper. Supplementary material at [this https URL](https://people.csail.mit.edu/echen/vocal_recon/)

点击查看摘要

Abstract:The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.

126. 【2609.36736】From Neurons to Conversation: Speech Brain-Computer Interfaces

链接:https://arxiv.org/abs/2609.36736

作者:Moein Khajehnejad,Forough Habibollahi,Tommaso Boccato,Margarida Sousa,Michal Olak,Francesco Jamal Sheiban,Matteo Ferrante

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Neurons and Cognition (q-bio.NC)

关键词:Speech brain-computer interfaces, transforming neural activity, neural activity related, synthesized voice, brain-computer interfaces

备注: Review article, 28 pages, 4 figures, 2 boxes, 2 tables

点击查看摘要

Abstract:Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.

127. 【2609.36734】Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

链接:https://arxiv.org/abs/2609.36734

作者:Ayan Sengupta,Vaibhav Seth,Tanmoy Chakraborty

类目:Computation and Language (cs.CL)

关键词:matching output distributions, Knowledge Distillation, smaller-capacity student model, larger-capacity teacher model, trains a smaller-capacity

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, $\alpha$--$\beta$ divergence). Highlights include up to $+3.2$ average ROUGE-L on instruction-following tasks, $+2.1$ pass@1 on MBPP, $+1.7$ accuracy on GSM8k, and $+1.8$ accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.

128. 【2609.36730】Can Agents Design Libraries for Agents?

链接:https://arxiv.org/abs/2609.36730

作者:Gabriel Orlanski,Alex L. Zhang,Avi Trost,Vincent Sunn Chen,Frederic Sala,Aws Albarghouthi,Ludwig Schmidt

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:Agents increasingly build, growing the codebases, increasingly build, Agents, library

备注: 26 pages, 6 figures, 11 tables. Code and data: [this https URL](https://github.com/SprocketLab/librarydesignbench)

点击查看摘要

Abstract:Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.

129. 【2609.36722】ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

链接:https://arxiv.org/abs/2609.36722

作者:Xinghao Chen,Junnan Dong,Cai Ke,Chak Tou Leong,Haocheng Sun,Keyu Chen,Siyu An,Ruizhi Qiao,Xing Sun,Wenjie Li,Xiaoyu Shen

类目:Computation and Language (cs.CL)

关键词:agents repeatedly load, Large language model, repeatedly load reusable, load reusable content, Large language

备注:

点击查看摘要

Abstract:Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mismatch has minor effect, and independently cached artifacts retain faithful representations: reading a provided artifact stays largely accurate, and performance degrades only when the model must select among multiple artifacts. Moreover, replacing PIC's attention scores with full-prefill scores recovers performance with the cached KV unchanged, localizing the failure to the attention rather than KV recomputation. Motivated by this, we propose \textsc{Attuner}, a query-side adaptation method that learns to read a frozen artifact cache. \textsc{Attuner} inserts low-rank adapters into the query projections and is trained by distilling full-prefill distribution into the student. It trains fewer than 0.05\% of the model parameters and, at inference, requires neither cache recomputation nor a full-context reference. On Qwen3-4B and Qwen3-8B across seven benchmarks covering skills, documents, memory, and code, \textsc{Attuner} substantially outperforms prior PIC baselines in both in-domain and out-of-domain settings, matches full-context prefill quality while providing up to $3.73\times$ speedup.

130. 【2609.36707】LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts

链接:https://arxiv.org/abs/2609.36707

作者:Amrita Singh,Aditya Joshi,Jiaojiao Jiang,Hye-young Paik

类目:Computation and Language (cs.CL)

关键词:expose enterprises, enterprises to financial, legal risks, Legal contracts, Legal

备注: Under Review

点击查看摘要

Abstract:Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis essential. We propose LAURA, a post-training framework for interpretable ambiguous clause identification. LAURA leverages knowledge distillation with an IRAC-Unlearning prompting technique to transfer knowledge from a teacher LLM to an open-weight student model (=1B parameters), which is then trained using a joint objective combining classification and rationale generation losses. The framework supports both legal and non-legal stakeholders in making informed decisions about which ambiguities require further attention. Extensive experiments across 7 baselines and 7 open-weight models demonstrate that LAURA with Flan-T5 (250M) delivers state-of-the-art interpretability over all interpretable baselines while matching the identification performance of the best-performing opaque baseline.

131. 【2609.36700】Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG

链接:https://arxiv.org/abs/2609.36700

作者:Pranav Handa,Ariful Azad

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:large language models, language models, users often begin, conversing with large, large language

备注: 35 pages, 11 figures

点击查看摘要

Abstract:When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.

132. 【2609.36691】Video2Skill: From Streaming Experience to Reusable Embodied Skills

链接:https://arxiv.org/abs/2609.36691

作者:Jianshu Zhang,Ce Zhang,Xiyuan Yang,Chenwei Xu,Haoran Lu,Yijiang Li,Yaqi Xie,Katia P. Sycara,Han Liu

类目:Computation and Language (cs.CL)

关键词:behaviors vary widely, Manipulation behaviors vary, embodied agents generalize, objects and scenes, behaviors vary

备注: Project page: [this https URL](https://andyzworks.github.io/video2skill/)

点击查看摘要

Abstract:Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.

133. 【2609.36689】CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference

链接:https://arxiv.org/abs/2609.36689

作者:Wenjin Liu,Chenxi Wang,Yue Lu,Zhe Cui,Haoran Luo

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Large language models, achieved significant progress, outputs exhibit systematic, Large language, exhibit systematic calibration

备注:

点击查看摘要

Abstract:Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergraphs into three stages, evidence weighting, evidence aggregation, and source fusion, and propose CHAIN, which designs stage-specific mechanisms to mitigate bias at each stage: (i) modulating the temporal decay function by causal topological distance, (ii) aggregating approximately independent causal chains via Noisy-OR after direction-aware deduplication, and (iii) driving adaptive fusion by causal coverage and directional balance. Experimental results on cross-domain forecasting benchmarks show CHAIN outperforms existing methods in expected calibration error, Brier score, and accuracy. Our project is available at this https URL.

134. 【2609.36684】ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

链接:https://arxiv.org/abs/2609.36684

作者:Jianshu Zhang,Keliang Wu,Chengxuan Qian,Xiyuan Yang,Ce Zhang,Ariel Tian,Anbang Liu,Haoran Lu,Han Liu

类目:Computation and Language (cs.CL)

关键词:Progress, tasks, PRMs, context, PRM

备注: Project page: [this https URL](https://andyzworks.github.io/progresscompass/)

点击查看摘要

Abstract:Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.

135. 【2609.36683】MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

链接:https://arxiv.org/abs/2609.36683

作者:Shicheng Fang,Yuxin Wang,Zhuo Yang,Xiaohu Xu,Jiahao Lu,Chuanyuan Tan,Tong Zhu,Yining Zheng,Xipeng Qiu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:inherently iterative, candidate is proposed, revised while preserving, preserving a relationship, source molecule

备注:

点击查看摘要

Abstract:Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.

136. 【2609.36675】Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement

链接:https://arxiv.org/abs/2609.36675

作者:Ziqi Zhao,Fanqing Meng,Haocheng Lu,Lingxiao Du,Qiguang Chen,Mengkang Hu,Xiao-Ming Wu

类目:Computation and Language (cs.CL)

关键词:achieve compounding gains, aims to achieve, data-centric RSI directly, existing RSI systems, RSI systems optimize

备注: Preprint

点击查看摘要

Abstract:Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model's own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth. To address this challenge, we introduce G"odel Forest, a multi-agent framework that organizes recursive self-improvement as an ensemble of co-evolving search trees. In G"odel Forest, each agent autonomously grows a persistent tree, deepening, branching, or pruning data strategies based on model feedback to secure depth, while parallel trees explore distinct regions of the data space to expand breadth. Crucially, rather than leaving trees isolated or flooding them with heavy execution logs, a dynamically co-evolving memory connects the forest: agents continuously distill their successes and failures into compact procedural lessons anchored to a global leaderboard. Through this forest ecosystem, a dead-end in one tree instantly warns the whole forest against unpromising paths, while an empirical breakthrough quickly seeds new exploration branches in neighboring trees. Evaluated on RSIBench-Data across six diverse domains, G"odel Forest outperforms the single-agent baseline by an average of 10.70% while reducing wall-clock time on five tasks. Ablations confirm that co-evolving shared memory yields a +7.00% gain over independent parallel search, demonstrating that collective distillation is key to scalable self-improvement. The code is available at this https URL.

137. 【2609.36654】Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

链接:https://arxiv.org/abs/2609.36654

作者:Ruiyi Ding,Jie Li,Kang He,Ziyan Liu,Chengru Song,Yuedong Xu,Yuan Cheng

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)

关键词:major inference costs, traffic major inference, motivating low-precision formats, make weight storage, language models make

备注: 33 pages

点击查看摘要

Abstract:Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration activations, and second-order state cannot all remain on one accelerator, while assigning complete layers to devices leaves each time-consuming layer solve serial. We introduce \emph{Schur Replay}, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns. Separately, our execution infrastructure keeps only the active layer resident, tiers activations across device, host, and disk, retires full-precision layers after export, and distributes independent output rows across tensor-parallel ranks. Together, the algorithm and infrastructure attain $99.35\%$ and $100.84\%$ question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct. On the 397B model, the infrastructure reduces measured per-layer time by $15.17\times$ over ModelOpt and $23.14\times$ over LLM Compressor, with lower memory used per GPU.

138. 【2609.36636】What Makes Recurrence Effective in Looped Language Models?

链接:https://arxiv.org/abs/2609.36636

作者:Xinlin Zhuang,Siyuan Wang,Imran Razzak,Weiyang Liu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Looped language models, Looped language, increase computational depth, parameter sharing, adding parameters

备注: Preprint, under-review

点击查看摘要

Abstract:Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.

139. 【2609.36617】Generating Edit-Inducing Questions for AI Research Manuscripts

链接:https://arxiv.org/abs/2609.36617

作者:Sebastian Joseph,Zichao Wang,Jennifer Healey,Alexa Siu,Junyi Jessy Li,Ani Nenkova

类目:Computation and Language (cs.CL)

关键词:LLMs to generate, answer will improve, generate edit-inducing questions, paper draft, questions

备注: Accepted at the DocInsights Workshop @ EMNLP 2026

点击查看摘要

Abstract:We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.

140. 【2609.36608】Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

链接:https://arxiv.org/abs/2609.36608

作者:Zubin Zheng,Jiahao Wu,Shaofeng Zhang,Zhirui Zhang,Yew-Soon Ong,Shengcai Liu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:trains multi-turn language, On-policy distillation, On-policy, dense teacher supervision, times

备注: 31 pages, 8 figures

点击查看摘要

Abstract:On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of $2.3\times$ on ALFWorld, $1.8\times$ on WebShop, and $4.9\times$ on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.

141. 【2609.36590】SEED: Self-Speculative Decoding via Implicit Encoder-Decoder

链接:https://arxiv.org/abs/2609.36590

作者:Hankun Lin,Patrick Pynadath,Ruqi Zhang

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:decoding accelerates large, accelerates large language, Self-speculative decoding accelerates, large language model, decoding accelerates

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the deep contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder-decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder-decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the lightweight decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.7$\times$ average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 28% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning. Code is available at this https URL.

142. 【2609.36585】ransformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

链接:https://arxiv.org/abs/2609.36585

作者:Zehao Jin,Ruixuan Deng,Junran Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:references in context, Pretrained transformers, follow references, pretrained loops add, extra pretrained loops

备注:

点击查看摘要

Abstract:Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at this https URL

143. 【2609.36577】Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

链接:https://arxiv.org/abs/2609.36577

作者:Zhenhong Zhou,Xuanyue Zhao,Youji Liu,Yuanhe Zhang,Xiaoyu Ma,Lianyu Hu,Yang Liu

类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:perform complex tasks, complex tasks, perform complex, Audio large language, long-term memory

备注: 28 pages, 5 figures, 17 tables

点击查看摘要

Abstract:Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at this https URL

144. 【2609.36550】Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment

链接:https://arxiv.org/abs/2609.36550

作者:Josepha Michiko Leo,Hyun-seok Min,Yehoon Jang,Irvan Zidny,Jin-Woo Chung,Sungchul Choi

类目:Computation and Language (cs.CL)

关键词:Retrieval-augmented generation, professional writing, definable meaning, generation is widely, cites prior art

备注: Accepted to Findings of AACL-IJCNLP 2026. 9 pages, 2 figures. Code and data: [this https URL](https://github.com/TeamLab/probing-rag-patent-amendment)

点击查看摘要

Abstract:Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tested where "correct" has a definable meaning. Patent claim amendment supplies that signal: the examiner names the attacked limitation and cites prior art, providing per-case ground truth. We release three artifacts: (i) a corpus of 7,385 USPTO prosecution cases with XML-aligned pre/post claims, rejection, and cited prior art; (ii) a seven-probe battery comparing random and structural-match retrieval as two policies under a fixed prompt scaffold; (iii) a deterministic five-channel metric (C1-C3 and C5 in main, C4 supplementary) requiring no LLM evaluation. Across 9,600 pre-registered calls on four frontier LLMs (Claude Sonnet 4, Claude Haiku 4.5, GPT-5.4, GPT-4o-mini), no tested model exhibits detectable classical prior-injection behavior; retrieval effects are small and direction-inconsistent between random and structural retrieval, and the null is unchanged under a dense (semantic) retriever, across retrieval depths k in {1,3,5,10}, and under a paraphrase-sensitive grounding metric. Revision locality reveals a model-specific difference that the template channel misses. The four-cell taxonomy, which we treat as exploratory, leaves the prior-injector cell unoccupied.

145. 【2609.36544】DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing

链接:https://arxiv.org/abs/2609.36544

作者:Divyansh Chandarana,Sandipan De,Vivek Gupta

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)

关键词:produce writing assignments, Generative, writing, writing assignments, students produce writing

备注: 8 pages, 7 figures, 3 tables

点击查看摘要

Abstract:Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary views of writing: the final product, the writing process and interactions with an integrated AI-assistant. DraftTrace reconstructs how a document develops over time and organizes these signals into submission, longitudinal, and class-level analytics for instructors. We deployed DraftTrace in a graduate NLP course with 81 students and compared their sessions with LLM-generated responses entered by automated tools and with copy-typed responses. While product measures distinguish differences in text formulation, process measures distinguish differences in how text is entered. Considering both views together helps characterize cases such as copy-typing. Interaction traces show that students use the assistant differently across stages of writing: to clarify the question at an early stage and to verify answers at a later stage. A preliminary instructor survey highlights the importance of multi-view writing analytics and their interpretability.

146. 【2609.36535】When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

链接:https://arxiv.org/abs/2609.36535

作者:Chenxu Wang,Chaozhuo Li,Xinze Shi,Songyang Liu,Kyrie You Wu,Ziluowen Luo,Shun Zhang,Chenxi Li,Litian Zhang

类目:Computation and Language (cs.CL)

关键词:performance improves, improve iteratively, self-evolution degeneration, large language models, Self-evolution

备注:

点击查看摘要

Abstract:Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds' data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.

147. 【2609.36534】Retrieval Sensitivity to Identity Signals in Queries

链接:https://arxiv.org/abs/2609.36534

作者:Andrew Tang,Nicholas Deas,Kathleen McKeown,Vishal Misra

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR)

关键词:documents reach users, Dense retrievers decide, decide which documents, documents reach, typically evaluated

备注: EMNLP 2026 camera-ready, with a correction to Fig. 4

点击查看摘要

Abstract:Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries---political ideology and dialect---bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query's own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries' identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers' query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at this https URL.

148. 【2609.36529】riadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

链接:https://arxiv.org/abs/2609.36529

作者:Oliver Sieberling,Bharat Runwal,David Jin,Ryan Chin,Rameswar Panda,Yoon Kim

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Recurrent neural networks, Recurrent neural, linear attention, triadic linear attention, neural networks

备注: Preprint

点击查看摘要

Abstract:Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propose triadic linear attention, which writes the triadic outer product of a key, a second key, and a value, into a third-order (i.e., 3D) tensor state, and reads from it by contracting both key axes with two queries. An $E$-dimensional second key thus yields an $E$-fold increase in state size while adding only two projections. Triadic linear attention is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, triadic linear attention substantially improves long-context language modeling and recall, outperforming alternatives that enlarge the state.

149. 【2609.36526】Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations

链接:https://arxiv.org/abs/2609.36526

作者:Guanghui Min,Liang Wu,Mingjia Shi,Yinhan He,Mayank Darbari,Liangjie Hong,Chen Chen

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Long-horizon agents require, growing interaction histories, manage growing interaction, agents require context, require context compression

备注: 41 pages, 10 figures, 9 tables

点击查看摘要

Abstract:Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that severe degradation concentrates at isolated compression events. Motivated by this finding, we propose PAIR (Prompt Adaptation using Interventional Rollouts) for adapting structured compression prompts. PAIR identifies individual compressions that degrade subsequent execution, diagnoses their effects, and revises the relevant sections of a fixed compression template. PAIR achieves the strongest cross-run reliability among compressed methods in every main benchmark-scope combination, consistently exceeding the competing prompt-adaptation baseline. Without modifying the downstream agent, PAIR brings compressed execution close to the no-compression baseline and sometimes numerically exceeds it.

150. 【2609.36515】Large-scale factor analysis shows machine intelligence is only partially interpretable

链接:https://arxiv.org/abs/2609.36515

作者:Faiz Ghifari Haznitrama,Afrizal Hasbi Azizy,Faeyza Rishad Ardi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)

关键词:language model development, domain-free intelligence factor, language model, language models, language

备注: 66 pages

点击查看摘要

Abstract:A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The $g$ factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.

151. 【2609.36475】Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

链接:https://arxiv.org/abs/2609.36475

作者:Sumin Hong,Katsumi Ibaraki,Renee Shi,David Chiang,Toby Jia-Jun Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:round shapes, sharp shapes, features across modalities, systematic pairings, pairings of features

备注: 9 pages

点击查看摘要

Abstract:Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.

152. 【2609.36474】FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance

链接:https://arxiv.org/abs/2609.36474

作者:Rikhiya Ghosh,Himanshu Kumar,Sriram Venkatapathy,Sahil Wadhwa,Alexandre G.R. Day,Pranab Mohanty

类目:Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)

关键词:seemingly harmless user, triggering safety failures, harmless user queries, exploit large language, pushing responses dangerously

备注:

点击查看摘要

Abstract:In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generation cost, while treating coverage, severity, and diversity as incidental rather than joint objectives. We introduce FinRT, a structured framework that builds reusable adversarial prompt generators from adaptive red-teaming strategies. Across the six victim models in consumer finance, FinRT substantially outperforms adaptive search baselines while amortizing target-facing attack generation into a reusable generator. FinRT nearly doubles the attack success rate over the adaptive baseline Rainbow Teaming (32.9% vs. 17.2%), increases maximum adversarial severity by 33%, and preserves comparable intra-policy-domain semantic diversity to iterative search methods. Our method achieves high cross-model transferability while exhibiting distinct victim-family specialization patterns.

153. 【2609.36458】Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models

链接:https://arxiv.org/abs/2609.36458

作者:Abdullah All Tanvir,Xin Zhong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:semantically consequential variation, captures semantically consequential, affect model predictions, induce substantial motion, strongly affect model

备注:

点击查看摘要

Abstract:Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose Fisher-induced invariant representation geometry (Fisher-IRG), which measures local representation directions through their predictive sensitivity. Around each representation, we construct semantic-preserving and semantic-changing neighborhoods, aggregate their local Fisher information, and recover invariant directions through a contrastive generalized eigenvalue problem. Controlled displacement analyses first show that comparable Euclidean motion can have substantially different predictive consequences, supporting the need for a predictive geometry. Across language and vision models, Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions. Representation interventions further localize semantic effects to the Fisher-derived subspace, and held-out separation and retrieval show that the recovered geometry generalizes beyond the discovery neighborhoods. These results support Fisher-IRG as a principled framework for characterizing local invariant representation geometry.

154. 【2609.36457】Memory Consolidation Flattens the Temporal Shape of User Facts

链接:https://arxiv.org/abs/2609.36457

作者:Sugam Panthi,Muhaiminul Yeamin,Siyan Luo,Rabab Abdelfattah

类目:Computation and Language (cs.CL)

关键词:Long-term memory systems, systems turn conversations, memory systems turn, Long-term memory, systems turn

备注:

点击查看摘要

Abstract:Long-term memory systems turn conversations into short stored notes. A note can keep a user fact while losing evidence about whether the fact still holds. For example, "I am driving a Peugeot" can become "The user drives a Peugeot," which drops the cue that the activity is ongoing. We call this aspectual flattening and measure it with LAPSE, a benchmark of matched user statements that differ only in temporal form. We find that memory writers flatten aspect selectively. Three writer models flattened the progressive statement but kept its simple-present match in 244 of 381 pairs, never the reverse. The asymmetry holds in all 11 model configurations tested and in the installed pipelines mem0, Graphiti, and Letta. The lost cue matters to later readers. In exploratory tests, changing only the stored verb shifted all three readers' estimates that a fact still holds. When readers could ask the user before acting, two of three acted without asking more often on flattened notes. Our planned memory-use task could not detect this, because readers there acted on almost every stored fact, even expired ones. Memory writing can thus remove evidence that later models use to decide whether to act.

155. 【2609.36452】Reliable Parallel Decoding in Masked Diffusion Language Models

链接:https://arxiv.org/abs/2609.36452

作者:Zhenghao He,Bohan Liu,Guangzhi Xiong,Aidong Zhang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:diffusion language models, generate text efficiently, Masked diffusion language, predicting multiple masked, language models

备注:

点击查看摘要

Abstract:Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.

156. 【2609.36451】Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations

链接:https://arxiv.org/abs/2609.36451

作者:Muhammad Ahtesham,Xin Zhong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Large language models, small semantic edits, Large language, hidden representations, meaning despite substantial

备注:

点击查看摘要

Abstract:Large language models often preserve meaning despite substantial changes in wording, style, and syntax, while small semantic edits can systematically alter their hidden representations. This suggests that semantic variation may be organized along recurring local directions. We propose the Invariant Atom Hypothesis: local semantic motion admits preferred sparse coordinates along directions that remain stable under meaning-preserving transformations. We learn a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations. Empirically, the atoms exhibit strong semantic--nuisance separation, sparse reconstruction, reproducible directions, and causal effects on model predictions. The learned geometry generalizes to unseen semantic neighborhoods and nuisance families, while local reweighting improves semantic selectivity and preserves a consistent global-to-local structure. Atom signatures also remain stable under model modification. These findings support reusable invariant directions as a sparse coordinate system for local semantic geometry in language models.

157. 【2609.36435】MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

链接:https://arxiv.org/abs/2609.36435

作者:Jingxuan Wu,Yuzhe Yang,Yiqiao Huang,Chengzhi Liu,Qingni Wang,Chengxuan Qian,Shutong Wu,Jiawei Zhang,Xin Eric Wang

类目:Computation and Language (cs.CL)

关键词:preferences still hold, assistant that serves, long horizon, constraints apply, user has revealed

备注:

点击查看摘要

Abstract:An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the information as text makes the reader's input grow with the retained history, while compressing it into a fixed number of latent vectors bounds the interface but is typically trained to reconstruct text or imitate reference answers, both of which are scored on sequences the reader never produced. We present MemFold, which optimizes a fixed-budget soft memory by the behavior it supports. A query-conditioned textual memory is compressed into K continuous vectors that form the reader's memory interface, and the reader is then trained on its own rollouts under two complementary signals: group-relative rewards for task outcomes, and confidence-gated on-policy distillation in which a frozen textual-memory teacher re-scores the student's sampled tokens under the textual memory. The teacher is never sampled from, so supervision stays on the student's current distribution and adds no autoregressive decoding; at inference it is removed entirely. Across three Qwen backbones, MemFold attains the highest accuracy we measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length, and transfers to PrefEval and LongMemEval without target-domain training. Ablations attribute most of the task gain to the reward term and a smaller additional gain to the teacher signal, and memory interventions show that the reader depends on the instance-specific content of its soft memory.

158. 【2609.36414】Eternal Sunshine of the Spotless Mind: Systematically Erasing LLM's Memories

链接:https://arxiv.org/abs/2609.36414

作者:Olga Ohrimenko

类目:Computation and Language (cs.CL)

关键词:accumulate memories, memories, persistent LLMs, Deletion, user

备注:

点击查看摘要

Abstract:We consider persistent LLMs that accumulate memories of their interactions with a user over time. Such LLMs maintain memories using external storage, which they can query to overcome the limitations of a fixed context window. Such systems have numerous practical applications, as they can draw on all past interactions when responding to user queries. In this paper, we ask whether LLMs can forget information shared with them upon a user's request. We find that current LLMs fail to delete such information---even when they claim to have forgotten it and even when operating with a limited context. To this end, we consider a new direction of study: Deletion of LLM Memories. We show that naively removing messages that match a user's deletion request is insufficient, since conversations naturally introduce message dependencies that cause information to persist. To correctly handle deletion requests, we propose the DeLLM framework. It dynamically constructs relevant context for each LLM query and maintains a provenance graph of messages to determine which ones must be removed during deletion. Our experiments show that DeLLM achieves a high deletion rate while maintaining utility.

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.36414 [cs.CL]

(or
arXiv:2609.36414v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.36414

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
159. 【2609.36399】Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV

链接:https://arxiv.org/abs/2609.36399

作者:Bushra Asseri,Abdulaziz Asseri

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Decision-only language models, Decision-only language, language models return, American personas, Survey Module

备注:

点击查看摘要

Abstract:Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12 matched American personas and without a persona, in English and Arabic, under eight ways of formulating the request (288,000 answers). JEV's answers were highly repeatable (ICC 0.997), and without a persona they resembled those of its own American personas. When the persona was Saudi rather than American, the answers moved in the direction of the human Saudi-US difference, reproducing 87% of its size in English but 62% in Arabic, with long-term orientation reversed. A language cross shows that the smaller difference in Arabic comes from the language of the items, not from the language of the persona description. Age shifted the profiles about as much as nationality, gender shifted them more for Saudi than for American personas, and JEV was less confident in Arabic and for Saudi personas. These patterns held in every request design, although the model never generates text.

160. 【2609.36372】Mark: Pairwise Distortion-Free Watermarking Beyond Single-Token Entropy

链接:https://arxiv.org/abs/2609.36372

作者:Ruibo Chen,Zhengmian Hu,Donghang Lu,Xuehao Cui,Georgios Milis,Yihan Wu,Jian Du,Heng Huang

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:enables reliable attribution, watermarking enables reliable, enables reliable, reliable attribution, attribution of machine-generated

备注:

点击查看摘要

Abstract:Distortion-free watermarking enables reliable attribution of machine-generated text while preserving output distribution. However, existing methods operate independently on each generated token, making their detection capability fundamentally constrained by the entropy of the next-token distribution. We present Tandem Token WaterMark (TTMARK), a general pairwise watermarking framework that extends distortion-free watermarking from individual tokens to adjacent token pairs. By watermarking the joint distribution of consecutive tokens, TTMARK enlarges the effective watermarking alphabet from V to $V^2$, allowing the detector to exploit both token entropy and conditional entropy while preserving distortion-freeness over the joint distribution. We further introduce a branch-isolating concatenated tandem generation algorithm that efficiently constructs the joint distribution in a single forward pass. Theoretically, we show that pairwise watermarking achieves better expected detection strength in low-entropy regimes. Extensive experiments across multiple language models, datasets, and three representative distortion-free watermarking schemes demonstrate that TTMARK consistently improves detectability without degrading generation quality, while also improving robustness to edits and substantially enhancing localized watermark detection.

161. 【2609.36344】DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents

链接:https://arxiv.org/abs/2609.36344

作者:Amirhossein Abaskohi,Amirhossein Dabiriaghdam,Lele Wang,Peter West,Giuseppe Carenini

类目:Computation and Language (cs.CL)

关键词:Deep-research agents conduct, conduct long-horizon investigations, agents conduct long-horizon, belief revision, Deep-research agents

备注:

点击查看摘要

Abstract:Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reasoning to reinforce an incorrect interpretation. We introduce DeepRewind, an additive control layer for reversible deep research that represents the agent's evolving epistemic state as a typed graph of sources, evidence, claims, hypotheses, assumptions, commitments, plans, and drafts. Before accepting an intermediate conclusion, a prompt-based world model predicts its impact and estimates reversibility based on hypothesis narrowing, information loss, recovery cost, and contradiction-trigger coverage. A binary controller blocks risky commitments, while a consistency monitor performs dependency-aware rollback when later evidence invalidates them. Across DRBench and LiveDRBench, DeepRewind improves insight recall by 3.6 percentage points and reduces premature commitments by 59.1% relative to Open Deep Research.

162. 【2609.36316】raining LLMs to Verbalize Evaluation Awareness

链接:https://arxiv.org/abs/2609.36316

作者:Usman Anwar,Sahar Abdelnabi,David Krueger

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:large language models, large language, behave differently, differently during audits, measuring and accounting

备注:

点击查看摘要

Abstract:Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model's spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained with an RL objective designed to increase verbalization in a calibrated way. Across Qwen3.6-35B-A3B, Kimi K2.6, and Inkling, VT increases verbalized EA by 2.4-2.9 times and transfers to held-out agentic settings, while measured latent EA and behavior remain largely stable. In a causal experiment, we independently implant meta-knowledge about evaluations through synthetic-document fine-tuning and show that VT-induced verbalizations reflect the richer knowledge acquired by the model.

163. 【2609.36314】Fractional State Space Transition for Long Sequence Modeling

链接:https://arxiv.org/abs/2609.36314

作者:Ivan Kobyzev,Abbas Ghaddar,Ali Nasiri-Sarvi,Lifeng Shang,Yufei Cui

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:State Space Models, Space Models, compress sequence history, State Space, central architectural choice

备注: NeurIPS 2026 (Oral)

点击查看摘要

Abstract:State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum of exponential modes. This construction turns fractional memory into an efficient recurrent module with parallel training and prefill, while retaining bounded-state autoregressive decoding. Extensive experiments, including 1.3B-parameter language modeling, demonstrate that FRAC consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context. These results show that fractional dynamics provide a practical and effective prior for long-context SSMs.

164. 【2609.36303】HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization

链接:https://arxiv.org/abs/2609.36303

作者:Feijie Wu,Hugo Barbalho,Konstantina Mellou,Marco Molinaro,Jing Gao,Ishai Menache,Xinzhi Zhang,Sirui Li

类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)

关键词:automate algorithm discovery, Recent advances, automate algorithm, algorithm discovery, Recent

备注:

点击查看摘要

Abstract:Recent advances in agentic heuristic design use AI agents and execution feedback to automate algorithm discovery for challenging optimization problems. In many practical settings, high-quality solutions must be obtained under strict runtime constraints, motivating hybrid approaches that combine problem-specific heuristics with powerful mathematical programming solvers. However, existing approaches typically improve heuristic components within predefined procedures or tune solver configurations in isolation. This limits holistic adaptation of where to allocate computation, how to leverage solvers, and how to refine the overall algorithmic structure. To address these limitations, we propose HeurEvo, an automated plan--code--component co-evolution framework that jointly evolves the high-level algorithmic structures, their implementations, and a shared pool of reusable components. A planner determines which algorithmic components to use, how to combine them, and how to allocate runtime across stages, a coder realizes the resulting plan as executable code, while a component evolver updates the shared component pool. Within an island-based evolutionary framework, plans and implementations co-evolve with feedback from an interpreter agent that analyzes execution results and identifies opportunities for improvement. Across diverse combinatorial optimization benchmarks and challenging MIPLIB instances, HeurEvo finds high-quality solutions within tight runtime budgets, often matching or surpassing state-of-the-art optimization solvers given hours or days of computation. On several nonlinear geometry problems such as hexagon packing, it also improves upon the best previously reported results. These results highlight the value of jointly searching over algorithmic structure and implementation for agentic heuristic design.

165. 【2609.36301】MoRE: Scaling mixture of experts with hardware-aware low-rank routing

链接:https://arxiv.org/abs/2609.36301

作者:Honam Wong,Surbhi Goel,Enric Boix-Adserà

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:recent architectures push, frontier language models, central to frontier, frontier language, recent architectures

备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $\Theta(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $\Theta(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\A benchmarks after pretraining, while matching reasoning ability. Code available at this https URL.

166. 【2609.36294】When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders

链接:https://arxiv.org/abs/2609.36294

作者:Xiaozuo Shen,Yifei Cai,Tian Tan,Rui Ning,Chunsheng Xin,Hongyi Wu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Graph Sparse Autoencoder, Adaptive Graph Sparse, SAEs impose single-parent, graphs permit multiple, permit multiple parents

备注:

点击查看摘要

Abstract:Sparse autoencoders (SAEs) expose interpretable features in large language model activations, yet existing structured SAEs impose single-parent trees or forests, while post-hoc graphs permit multiple parents but neither guide feature learning nor ensure reliable relation recovery. We introduce the Adaptive Graph Sparse Autoencoder (AG-SAE), a structure-guided training paradigm that treats each feature's complete parent set as an atomic structural hypothesis and lets evidence select zero, one, or multiple parents. By competing complete parent sets against null, subset, and alternative explanations, AG-SAE identifies jointly necessary multi-parent relations while rejecting redundant or spurious alternatives and verifying that each child contributes beyond its parents. The induced topology over SAE features then defines a differentiable structural loss that guides SAE training, while topology-guided refinement mitigates feature absorption and uses persistent reconstruction gaps exposed by the learned structure to initialize new features. The entire graph is then induced again from the revised dictionary by reassessing every feature's complete parent set, closing the dictionary-graph self-consistency cycle. Experiments demonstrate exact mixed-topology recovery in a controlled toy model, greater relational reliability and semantic validity than structured and post-hoc baselines on real LLM activations, and stronger feature-level causal interventions than conventional SAE features. AG-SAE thereby turns recovered mixed-topology feature structure into an unsupervised training signal that improves the dictionary, enables reliable feature organization beyond the topological limitations of trees, and exhibits stronger causal control beyond reconstruction.

167. 【2609.36290】he Surge of Anti-Semitism in German Social Media following the October 7 Attacks

链接:https://arxiv.org/abs/2609.36290

作者:Gregor Wiedemann,Daniel Wehrend

类目:ocial and Information Networks (cs.SI); Computation and Language (cs.CL)

关键词:affected German social, German social media, social media debates, Israel of October, Judaism and Israel

备注: 8 pages; 5 figures; accepted at 22st Conference on Natural Language Processing (KONVENS 2026), Hamburg, Germany

点击查看摘要

Abstract:We investigate the extent to which the Hamas attacks on Israel of October 7, 2023, have affected German social media debates about Judaism and Israel. For this, we develop an approach to detect 26 anti-Semitic categories in user postings via large language models (LLMs). The approach is applied to Facebook and Telegram posts (N=125,718) from three months before and after the event. Methodically, we test different open-weight models in two setups---with and without user information as additional context to the post text. The best setup achieves up to 83 % F1-score for binary anti-Semitism detection on our manually coded validation set. User context provides valuable information for most LLMs and drastically reduces false positives, for example, when (critically) reporting on anti-Semitic incidents. Concerning our topic, we find that anti-Semitism is surging significantly on both platforms, while being about ten times more prevalent on Telegram compared to Facebook. Facebook users express anti-Semitic views most likely in posts about an alleged genocide in Gaza carried out by the Israeli army, whereas classic anti-Semitic stereotypes related to power and conspiracy theories are dominant on Telegram. After the attack, the discourse patterns on both platforms show signs of convergence, as classic anti-Semitism increases on Facebook, whereas Israel-related categories surge on Telegram.

168. 【2609.36265】In-Context Learning Amplifies a Latent Symbolic Circuit

链接:https://arxiv.org/abs/2609.36265

作者:Melissa Wessel

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language models, internal mechanisms activate, Large language, learn abstract rules, internal mechanisms

备注: Accepted to the Mechanistic Interpretability Workshop at ICML 2026

点击查看摘要

Abstract:Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to 8x from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1% to 56% at 0-shot and 17% to 88% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.

169. 【2609.36264】OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

链接:https://arxiv.org/abs/2609.36264

作者:Liner Xiang,Wenbo Zhang,Hengrui Cai

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:large language models, perform safely online, Reliable evaluation, development and deployment, safely online

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at this https URL.

170. 【2609.36253】Population Fidelity: Evaluating Population Representativeness in LLMs

链接:https://arxiv.org/abs/2609.36253

作者:Neemias B. da Silva,Martin Lukk,Ali Sutani,Abhishek Moturu,Harris Yang,Daniel Silver,Matt Ratto,Thiago H. Silva

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:Large language models, show considerable potential, Large language, show considerable, considerable potential

备注: 37 pages, 16 figures, 14 tables. Code and data: [this https URL](https://github.com/CriticalMaking/LLM-population-fidelity)

点击查看摘要

Abstract:Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.

171. 【2609.36246】Learning from Teacher Continuations at Student States

链接:https://arxiv.org/abs/2609.36246

作者:Haojin Wang,Dylan Zhang,Huaibo Chen,Suhao Yu,Yihang Sun,Zhanyang Jin,Jiaying Ye,Dianqi Li,Prasanna Sattigeri,Kamal Youcef-Toumi,Hao Peng

类目:Computation and Language (cs.CL)

关键词:OLIVE, present OLIVE, distillation, OnLine InterVEntion, teacher

备注:

点击查看摘要

Abstract:We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.

172. 【2609.36239】Cognitive Expert Language Models Better Align with the Corresponding Brain Systems

链接:https://arxiv.org/abs/2609.36239

作者:Zhivar Sourati,Mengxuan Helen Wu,Nona Ghazizadeh,Jonas Kaplan,Morteza Dehghani,Samuel A. Nastase

类目:Computation and Language (cs.CL)

关键词:Large language models, natural language comprehension, Large language, language comprehension, brain

备注:

点击查看摘要

Abstract:Large language models (LLMs) can predict human brain activity across a variety of brain regions during natural language comprehension. Typically, however, LLM-brain alignment is measured using one model for different regions of the brain, and then model performance is summarized across regions. This one-model-fits-all approach ignores the functional specialization of brain regions. In this study, we assess whether a model oriented toward a particular cognitive domain aligns better with the brain system dedicated to that domain. Through prompting and fine-tuning, we first build expert LLM variants for six domains: sensory, spatial, numerical, reasoning, social, and abstract processing. We then examine whether each expert best predicts activity in the brain region associated with the corresponding cognitive domain. Consistent with our hypotheses, each expert's representations align more closely with the brain system most associated with the matching domain than do other experts. This holds under both prompting and fine-tuning, across three base models and three fMRI datasets. In a series of control analyses, we show that this model-brain alignment is specific to cognitive domain interventions; non-cognitive and surface-level interventions do not result in comparable alignment. Specializing models shifts regional alignment while leaving aggregate prediction accuracy largely unchanged, suggesting that summarizing alignment across regions may obscure regional differences in performance for specific models.

173. 【2609.36218】CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles

链接:https://arxiv.org/abs/2609.36218

作者:Mir Tafseer Nayeem,Susmoy Chakraborty,Davood Rafiei

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Machine Learning (cs.LG)

关键词:situated audience judgments, remains comparatively underexplored, culturally situated audience, film remains comparatively, Large language models

备注: Preprint

点击查看摘要

Abstract:Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.

174. 【2609.36214】Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models

链接:https://arxiv.org/abs/2609.36214

作者:Yusheng Zhou,Eleanor Lin,David Jurgens

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, fluent speakers, receive systematically lower-quality, shown to receive

备注: 19 pages, 8 figures

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.

175. 【2609.36209】he Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations

链接:https://arxiv.org/abs/2609.36209

作者:Timo Pierre Schrader,Annemarie Friedrich,Simon Razniewski,Lukas Lange

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, knowledge bases, acquire during pre-training, vast amount

备注: Accepted at AKBC@EMNLP2026

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investigate how LLMs represent and generate multi-valued relations. We identify the canonical order problem: The probabilistic distributions inside LLMs organize many multi-valued relations according to a canonical ordering (e.g., alphabetical or chronological). Through mechanistic analysis, we show that set generation in LLMs can be thought of in terms of three phases: (1) retrieval of candidate entities, (2) internal sorting, and (3) selection of the next element. As a result, prompts aiming to construct KBs that deviate from this internal canonical ordering lead to a markedly reduced reliability of LLMs when aiming to generate complete sets for multi-valued relations.

Comments:
Accepted at AKBC@EMNLP2026

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.36209 [cs.CL]

(or
arXiv:2609.36209v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.36209

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
176. 【2609.36205】Geometric Representations of African Languages: A Regional Semantic Hub and Cultural Steering

链接:https://arxiv.org/abs/2609.36205

作者:Muhammad Abdullahi Said,Jonathan Shock

类目:Computation and Language (cs.CL)

关键词:represents African languages, African languages, represents African, responds to cultural, African

备注:

点击查看摘要

Abstract:We study how Gemma 4 31B represents African languages and responds to cultural steering. The first study compares nine African languages and three controls using probes, contrast directions, and measures of representation similarity. Transfer from English varies across languages and layers. Directions representing an Africa versus West contrast are more aligned among the African languages than between these languages and the controls at several layers. The comparison across language families passes the reported Holm threshold at five of twelve layers, although dependence between language pairs limits the statistical interpretation. Within Nigeria, Yoruba and Igbo are more aligned than the average of their pairs with Hausa at eleven of twelve layers. The second study uses separate English data to construct directions for Nigeria, Ghana, Kenya, and South Africa. Under union scoring at the selected strengths, estimated differences in attribution rates from random directions range from 0.63 to 0.81. Most outputs pass the automated structural coherence screen. Comparisons with Aya Expanse 32B show that results depend on the representation measure. Together, the studies document regional and family patterns in the sampled representations and country steering in English.

177. 【2609.36202】FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models

链接:https://arxiv.org/abs/2609.36202

作者:Darshan Thaker,Lachlan Ewen MacDonald,René Vidal

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Gradient-based reward guidance, control masked diffusion, Gradient-based reward, masked diffusion language, diffusion language models

备注:

点击查看摘要

Abstract:Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.

178. 【2609.36201】SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

链接:https://arxiv.org/abs/2609.36201

作者:Jianxing Chen,Xiao Yu,Shipra Agrawal,Zhou Yu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)

关键词:completing computer tasks, professional workflows, SCOUT, capable of completing, completing computer

备注:

点击查看摘要

Abstract:Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent's trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.

179. 【2609.36194】Concept Direction Reliability Across Languages with Different Tokenizer Fertility

链接:https://arxiv.org/abs/2609.36194

作者:Muhammad Abdullahi Said,Abass Oguntade,Elisha Komolafe,Babangida Sani,Fatima Muhammad Adam,Muhammad Sammani Sani

类目:Computation and Language (cs.CL)

关键词:Extracted sentiment directions, classification remains accurate, Extracted sentiment, remains accurate, vary across samples

备注:

点击查看摘要

Abstract:Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducibility, we measure split-half agreement in English, Hausa, and Yoruba representations across four language models using both native and translated texts. We identify layers selected for agreement using ten topics and evaluate direction agreement across separate groups of fifteen topics. Using the final token, split-half agreement ranges from 0.737 to 0.870 for English, 0.589 to 0.762 for Hausa, and 0.101 to 0.399 for Yoruba, maintaining this language rank order across all 77 complete model comparisons. Classifiers trained on these same layers consistently predict sentiment above chance, demonstrating that predictive accuracy does not imply directional consistency. Furthermore, averaging token representations yields less consistent agreement, and high agreement can partially reflect sentence length. Ultimately, our findings highlight the need to measure vector direction reproducibility independently of classification performance, though they do not establish that tokenizer fertility which is the average number of tokens per whitespace separated word causes cross-lingual differences.

180. 【2609.36178】argeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

链接:https://arxiv.org/abs/2609.36178

作者:Dongwon Jung,Hemanth Neelgund Ramesh,Yifan Wang,Xiaomin Li,Yuexing Hao,Yu Hu,Muhao Chen,Varun Chandrasekaran,Andrzej Banburski-Fahey,Jaron Lanier

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Relative Policy Optimization, Policy Optimization, language model agents, training large language, large language model

备注:

点击查看摘要

Abstract:Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.

181. 【2609.36159】Principled Thoughts for Latent Recursive LLM Systems

链接:https://arxiv.org/abs/2609.36159

作者:Fahd Seddik,Fatemeh Fard

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large language models, Large language, decoded text, hidden states, final decoded answer

备注: Project website: [this https URL](https://fard-lab.github.io/REST)

点击查看摘要

Abstract:Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: this https URL

182. 【2609.36139】Language Models Are "Insecure" Reporters

链接:https://arxiv.org/abs/2609.36139

作者:Jenny Y. Huang,Jiameng Fan,Ahmed Imtiaz Humayun,Maximillian Chen,Tian Qin,Run Chen,Vidhya Navalpakkam,Hongxiang Gu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:autonomous long-horizon tasks, increasingly autonomous long-horizon, large language models, long-horizon tasks, manually auditing

备注:

点击查看摘要

Abstract:As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon "insecure reporting." When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports. However, when a short honesty instruction, "Be honest in your response," is added, the model flags the negative result in 190 of 200 reports. Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about ways to appear successful. We perform an activation analysis and a steering experiment on Qwen3.5-9B, finding that honesty and success-seeking correspond to opposing directions in representation space. Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.

183. 【2609.36138】When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

链接:https://arxiv.org/abs/2609.36138

作者:Jiayi Li,Ruizhe Li

类目:Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:K-way action space, invoking external tools, K-way action, executing a call, seeking clarification

备注: Preprint

点击查看摘要

Abstract:Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: this https URL.

184. 【2609.36131】A Character-Level Neural Approach to Sinhala Sandhi Splitting

链接:https://arxiv.org/abs/2609.36131

作者:Yasas Ekanayaka,Deshan Sumanathilaka

类目:Computation and Language (cs.CL)

关键词:Sinhala Sandhi splitting, merged surface form, morphemes hidden inside, phonologically merged surface, Sinhala Sandhi

备注: 10 pages, 2 Figures, 10 Tables, Accepted to present at AACL-IJCNLP 2026

点击查看摘要

Abstract:Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-match accuracy (82.08\% character-level accuracy), well below the 94.00\% achieved on the more regular affixational subset. Ablations show that bidirectional encoding is the largest contributor to performance, while native Sinhala script improves exact match accuracy over romanized input. Qualitative analysis indicates that many errors are near misses involving boundary adjacent characters or plausible but incorrect phonological substitutions. These results establish an empirical baseline for Sinhala Sandhi splitting and identify data scale, Sandhi type conditioning, and attention-based decoding as the main directions for future work.

185. 【2609.36086】PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

链接:https://arxiv.org/abs/2609.36086

作者:Cheng Chang,Yining Mao,Peng Qi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:frequently employed, employed to evaluate, Language models, human, PADMÉ

备注: Accepted at the NeurIPS 2026 Workshop TAE (Trust-AI-Eval): Can We Trust AI Evaluation? 27 pages, 3 figures. Code and data at [this https URL](https://github.com/chc012/padme)

点击查看摘要

Abstract:Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.

186. 【2609.36082】GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis

链接:https://arxiv.org/abs/2609.36082

作者:Ethan D. Frakes,Amy Kvien,Rishabh Kundu,Redad Mehdi,Van D. Tran,Vibha S. Mandayam,Kristopher O. Davis,Erika I. Barcelos,Roger H. French,Yinghui Wu,Mengjie Li

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:LLM-based geospatiotemporal KGQA, existing KGQA benchmarks, assessing LLM-based geospatiotemporal, Unlike existing KGQA, KGQA benchmarks

备注: 13 pages, 6 figures, 7 tables. Accepted to the 34th ACM International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26), November 3-6, 2026, Riverside, CA, USA

点击查看摘要

Abstract:We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs' understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at this https URL.

187. 【2609.36079】A Polyphonic Conception of AI Understanding

链接:https://arxiv.org/abs/2609.36079

作者:Matthieu Queloz,Pierre Beckmann

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:model understands, engineer must decide, model output, model, understanding

备注:

点击查看摘要

Abstract:When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system's understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.

188. 【2609.36059】Mnemon: Raw Records, Fast Judgments, Slow Thoughts

链接:https://arxiv.org/abs/2609.36059

作者:Guangren Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:memory systems build, Long-term memory, graphs or typed, typed memories, memories at write

备注: 16 pages, 3 figures, 4 tables. Code, prompts and run records: [this https URL](https://github.com/Grivn/mnemon-memory-agent)

点击查看摘要

Abstract:Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster.

Comments:
16 pages, 3 figures, 4 tables. Code, prompts and run records: this https URL

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

Cite as:
arXiv:2609.36059 [cs.CL]

(or
arXiv:2609.36059v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.36059

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
189. 【2609.35970】Causal and Interpretable Structures in LLM Compositional Tasks

链接:https://arxiv.org/abs/2609.35970

作者:Gurbir Arora,Toni J.B. Liu,Jiajun Bao,Raphaël Sarfati,Christopher J. Earls

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language models, Large language, individual input tokens, individual input, Large

备注: 46 pages, 28 figures

点击查看摘要

Abstract:Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.

190. 【2609.35942】Question-Specific Knowledge Graphs for Efficient Visual Reasoning

链接:https://arxiv.org/abs/2609.35942

作者:Ting-Chih Chen,Emile van Krieken,Shujian Yu,Filip Ilievski

类目:Computation and Language (cs.CL)

关键词:visual question answering, exhibit strong reasoning, strong reasoning capabilities, Recent work, question answering

备注:

点击查看摘要

Abstract:Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and implicit knowledge sufficient to support reasoning, without introducing spurious assumptions. Existing methods that leverage detailed image captions introduce visual details unrelated to the reasoning task, inflating input token counts and increasing computational cost. To address these challenges, we propose VisKG, a reinforcement learning (RL) framework in which models learn to translate visual content into question-specific knowledge graph (KG) representations. This process filters out perceptual noise while preserving the entity-relation structure needed for chain-of-thought reasoning, following the principle of minimum sufficient information. To ensure stable RL post-training, VisKG adopts Group reward-Decoupled Normalization Policy Optimization (GDPO). In addition, we strengthen the supervision stage with negative rationale samples, exposing the model to incorrect reasoning paths before RL post-training. Experimental results across science, mathematics, and general visual understanding benchmarks show that VisKG achieves performance comparable to or better than baselines, while requiring fewer tokens than caption-based representations. Moreover, training VisKG with GDPO improves accuracy by 2% over its GRPO-trained counterpart on average. These results suggest that KG representations are a promising approach for supporting multi-step reasoning and open up future work on adaptively selecting the most suitable representation for a given task.

191. 【2609.35922】Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change

链接:https://arxiv.org/abs/2609.35922

作者:Bhavik Mangla

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD)

关键词:vulnerability rules recognise, sector rules, voice, vulnerability rules, rules recognise

备注: 38 pages, 11 figures, 15 tables. Code, scorer and development-split data at [this https URL](https://github.com/bhavik-mangla/voxparity-bench) and [this https URL](https://doi.org/10.5281/zenodo.23008159)

点击查看摘要

Abstract:A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems' misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.

192. 【2609.35916】VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

链接:https://arxiv.org/abs/2609.35916

作者:Jie Yang,Jiajun Chen,Jiazheng Zhou,Mianqiu Huang,Yining Zheng,Yuxin Wang,Xipeng Qiu

类目:Multiagent Systems (cs.MA); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Real-world embodied agents, pursue independent objectives, Real-world embodied, shared physical environment, pursue independent

备注:

点击查看摘要

Abstract:Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.

193. 【2609.35879】CruxBench: A Benchmark of Information Discovery

链接:https://arxiv.org/abs/2609.35879

作者:Hui Dai,Lina Piao,Nick Merrill,Nadja Flechner,Ezra Karger,Haifeng Xu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:fixed reference labels, large language models, reference labels, large language, fixed reference

备注: NeurIPS 2026

点击查看摘要

Abstract:Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by construction, since ground truth is generated by future world events; (2) open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and (3) grounded, with informativeness measured against quantified changes in real-world beliefs. We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes' usefulness for answering target questions. However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.

194. 【2609.35869】he Price of Token Boundaries: Compression Certificates and Prediction

链接:https://arxiv.org/abs/2609.35869

作者:Yuhao Du,Shunian Chen

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Pre-tokenisation restricts, restricts which text, text fragments, obscured when tokenisers, tokenisers are compared

备注:

点击查看摘要

Abstract:Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation. To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10\% of the vocabulary budget recovers 85.2\% and 100.0\% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.

195. 【2609.35865】PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

链接:https://arxiv.org/abs/2609.35865

作者:Yida Lin

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Computer Science and Game Theory (cs.GT)

关键词:Single-token typed-decision models, Single-token typed-decision, typed-decision models answer, one-letter answer codes, schema question

备注:

点击查看摘要

Abstract:Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores most of the structure in their training data. We study such a model whose data is curated as contrastive pairs---two contexts that differ in one edited fact that flips the answer---each carrying a machine-checked certificate that deleting the decisive sentence makes the fact unknown. We propose PACT, which turns this structure into four training terms that need no new annotation: a difference-in-differences margin over each pair that is invariant to any shared logit offset, a permutation-consistency term against answer-code position bias, an evidence-necessity term on certificate-verified ablated contexts, and an ordinal transport cost for rubric fields, plus a three-parameter contextual temperature. On a frozen 324-item holdout with three seeds, PACT matches the published recipe in accuracy ($84.6\%$ vs. $85.2\%$; McNemar $p \ge 0.50$ at every seed) while giving the lowest position bias of all runs (answer flips under relabelling $9.8\%$ vs. $13.8\%$) and the lowest ordinal error on rubric fields (MAE $0.232$ vs. $0.311$). Against a control with the same optimiser and schedule but cross-entropy only, PACT is significantly more accurate at two of three seeds, halves the seed-to-seed spread and lowers NLL by $26\%$. Seed-matched ablations and pre-specified falsification tests locate these gains precisely: no single term raises raw accuracy, and the method's value lies in robustness and stability rather than headline accuracy. Code, data splits, trained adapters, and all run records are available at this https URL.

196. 【2609.35864】Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction

链接:https://arxiv.org/abs/2609.35864

作者:Prisha Nair,Roee Shraga

类目:Computation and Language (cs.CL)

关键词:total market capitalization, missing-data problem affecting, market capitalization, consistently captured, captured in structured

备注: Presented as a Lightning Talk at MIT URTC 2026

点击查看摘要

Abstract:SEC 10-K filings contain substantial financial information that is not consistently captured in structured datasets, creating a missing-data problem affecting over 70% of firms and half of total market capitalization. This can disproportionately bias quantitative analysis against smaller firms, which may be excluded due to limited available data. Traditional financial extraction methods such as Regular Expressions (Regex) and BERT, have been widely used. However, they are highly brittle when parsing complex SEC 10-K filings, which leads to data that is existent in the files being lost since these methods do not consider that a data attribute could be located in a different section or a footnote. This study evaluates several Large Language Models (LLMs), including Llama-3 8B, Qwen-2.5 14B, and Llama-3.3 70B, to figure out individual model strengths and weaknesses when extracting specific attributes from SEC 10-K text. The extraction quality was evaluated across four financial variables of varying structural complexity: Cash and Cash Equivalents (tabular), Short-Term Debt (hybrid), Credit Facilities (narrative), and Research and Development (hybrid). Results show that while smaller models like Llama-3 8B experience performance degradation under complex negative prompting, aligning parameter scale with document complexity yields high zero-shot accuracy. Qwen-2.5 14B excels as a tabular specialist with an 83.33% F1 score on Cash, whereas Llama-3.3 70B effectively navigates dense narrative footnotes, achieving a 76.92% F1 score on RD. This scalable framework addresses critical information gaps in quantitative finance datasets and eliminates missing-data bias through a more thorough analysis of the SEC 10-K files.

197. 【2609.35860】he Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

链接:https://arxiv.org/abs/2609.35860

作者:Pranav Darshan,Pranav A,Sravan Karthick T,Minal Moharir,Ivan P. Yamshchikov

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Sampling based consistency, conceal systematic differences, Sampling based, errors are detectable, consistency is widely

备注: Accepted at GlobalSouthAI @ NeurIPS 2026

点击查看摘要

Abstract:Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|\rho|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap $95\%$ intervals excluding zero in all $12$ model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ($p0.005$) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ($16\%$ to $77\%$), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.

198. 【2609.35845】Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints, Patent Signals, and Compute Scaling

链接:https://arxiv.org/abs/2609.35845

作者:Muhammad Sukri Bin Ramli

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Total Factor Productivity, Macroeconomic productivity metrics, register technological breakthroughs, national accounting conventions, Total Factor

备注:

点击查看摘要

Abstract:Macroeconomic productivity metrics, such as Total Factor Productivity, register technological breakthroughs with multi-year reporting lags due to administrative survey intervals and national accounting conventions. This paper introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsupervised quantitative methodology that tracks technology diffusion directly from unstructured scientific and commercial text streams. We analyze 30,000 filtered document records spanning academic preprints from arXiv and patent application records from the USPTO. By projecting high-dimensional Transformer sentence embeddings onto unit hyperspheres using Spherical K-Means clustering across eight primary sub-topics and UMAP manifold reductions, HSTA formalizes two quantitative metrics: (1) Semantic Centroid Vector Drift, which tracks vocabulary shifts between temporal sub-corpora to identify structural paradigm transformations; and (2) Commercialization Offset, which evaluates cross-corpus peak density alignments between scientific discovery and intellectual property filings. Linking quarterly topic volume velocity with physical hardware metrics from the Epoch AI database, Vector Autoregressive F-tests demonstrate that quarterly paper volume velocity alone does not Granger-cause frontier compute allocation surges at conventional statistical significance levels, highlighting the necessity of conditioning textual signals on physical capital constraints. Empirical results reveal that sub-topics covering Large Language Models (with a drift metric of 0.332) and Artificial Intelligence Systems (with a drift metric of 0.234) undergo the highest rate of semantic evolution, offering an objective, real-time mechanism to complement traditional economic statistics.

199. 【2609.35833】Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

链接:https://arxiv.org/abs/2609.35833

作者:Avyay Sadhu,Alvaro Velasquez,Lekai Chen

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:network connection, formal logic problems, edge hardware, hardware provides private, private and low-latency

备注: 12 pages, 7 figures, 9 tables. This work has been submitted to the IEEE for possible publication

点击查看摘要

Abstract:Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit. We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engines and reserving the small language model (SLM) for open-ended word problems. Instead of hand-coding the routing logic, we learn a deterministic finite automaton (DFA) with the L* grammatical inference algorithm, using the SLM as a membership oracle and labeled data as an equivalence oracle. On a Raspberry Pi 4B (8 GB RAM, no GPU), evaluated on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker, learned routing attains 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), compared with 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers. Since formatted queries never reach the model, the router answers them in 1-11 ms and, in its 30-token configuration, runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.

200. 【2609.35832】When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction

链接:https://arxiv.org/abs/2609.35832

作者:Tianzhu Zhang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:external evidence, language model, model to revise, receiving new external, Intrinsic self-correction

备注:

点击查看摘要

Abstract:Intrinsic self-correction asks a language model to revise its own answer without receiving new external evidence. A second pass can recover mistakes, but it can also overturn answers that were already correct. We study this trade-off across 29 open-weight LLMs on BoolQ, GSM8K, and Corr2Cause by tracking correctness transitions between initial and revised answers. Aggregate accuracy can conceal substantially different revision behavior: for example, Llama-3.1-8B improves by 25.5 percentage points on GSM8K, while refinement changes 19.1% of initially correct answers into wrong ones. A controlled BoolQ study further shows that refinement prompts shift the balance between recovery and harm. We then compare three runtime choices: keeping the initial answer, always accepting the revision, and selectively invoking revision using signals available after the initial response. The comparison identifies settings where learned gating is useful and others where a simpler unconditional policy performs better. These results suggest treating intrinsic self-correction as a revision policy rather than as a uniformly beneficial second pass, and evaluating it through both the corrections it recovers and the errors it introduces.

201. 【2609.35831】Beyond the Context Window: An Adaptive Entropy-Based Routing Framework for Hybrid Retrieval and Long-Context Language Models

链接:https://arxiv.org/abs/2609.35831

作者:Isaac Olufadewa,Miracle Adesina,Ezekiel Oladejo,Owen Adeniyi,Fadare Fadekemi,Olamide Oso,Uthman Babatunde,Matthew Olawoyin

类目:Computation and Language (cs.CL)

关键词:Modern large language, support context windows, Modern large, large language models, large language

备注: 11 pages, 2 figures

点击查看摘要

Abstract:Modern large language models now support context windows of more than one million tokens, which has raised the question of whether retrieval-augmented generation (RAG) is still necessary. Pure long-context (LC) processing is expensive and is known to under-attend to information placed in the middle of long inputs, while pure RAG is fast but bounded by retrieval quality and prone to errors when retrieved chunks are partially relevant or contradictory. We propose the Entropy-Driven Adaptive Router (EDAR), a framework that decides at inference time whether to answer a query from retrieved chunks or to escalate it to full long-context processing. The decision uses the predictive entropy of the token-level probability distribution computed over the first few generated tokens of the RAG response. The entropy threshold is selected on a held-out validation set by sweeping cost against accuracy. Experiments compare EDAR against pure-RAG and pure-LC baselines on LongBench v2 and Infinity-Bench. Predictive entropy correlates strongly with hallucination rate on a held-out set of 2,000 generations (Pearson r = 0.85, 95% CI [0.83, 0.87]). On the long-context benchmarks, EDAR retains 97.4% of the accuracy of the pure long-context baseline while reducing total token expenditure by 70.7%, escalating only 18.2% of incoming queries. The accuracy gap between EDAR and the pure long-context system is not statistically distinguishable from zero at standard sample sizes. Predictive entropy is a useful model-internal signal for routing between RAG and long-context inference, and a threshold-based hybrid system can recover most of the accuracy of long-context models at a small fraction of the cost. The framework does not depend on a specific retriever or LC backbone, and it does not require additional supervision beyond what is normally produced during decoding.

202. 【2609.35824】Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation

链接:https://arxiv.org/abs/2609.35824

作者:Thomas Reiter,Christoph Kern,Fedor Miasnikov,Sofiia Nikolenko,Rob Chew,Stephanie Eckman,Frauke Kreuter

类目:Computation and Language (cs.CL); Methodology (stat.ME)

关键词:Large language models, Large language, give reliable labels, task design, task

备注: Accepted to "3rd Workshop on Uncertainty-Aware NLP" @ EMNLP 2026 (archival)

点击查看摘要

Abstract:Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss' $\kappa = 0.91$). Agreement fell when we changed the task design for the same tweets (median Cohen's $\kappa = 0.76$). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis points across five human instrument versions. Confidence scores did not solve this problem. They tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points. We call the variation caused by task design and model choice instrument uncertainty. Researchers can measure it only by comparing reasonable task designs. Repeating one setup or relying on confidence scores cannot replace that test.

203. 【2609.35822】racing mechanisms of sycophantic agreement in language models

链接:https://arxiv.org/abs/2609.35822

作者:Sixing Chen,Zhuofan Josh Ying,Logan Riggs Smith,Jeremy Wertheimer,Natalie Shapira

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:user stated beliefs, language models refers, beliefs or preferences, tendency to overly, overly affirm

备注:

点击查看摘要

Abstract:Sycophantic agreement in language models refers to the tendency to overly affirm a user's stated beliefs or preferences, often at the expense of factual accuracy. Although it is widely recognized as an alignment failure, its underlying mechanisms remain poorly understood. In this work, we use causal mediation analysis to identify the mechanisms behind sycophantic agreement. We show that a stated opinion is incorporated into the residual stream of the final prompt token early, where it biases subsequent answer retrieval. A sparse set of early attention heads carries this opinion signal. Ablating these heads substantially reduces sycophancy while leaving factual accuracy largely intact. The same heads carry the opinion when it is explicitly stated, regardless of how it is phrased. When an opinion is not stated explicitly but instead conveyed through content-free pushback (e.g., ``Are you sure?"), we find a distinct set of heads that suppresses the model's original correct answer to promote a revised answer. By providing a mechanistic account of how opinions induce sycophantic agreement, this work takes a step toward developing more targeted and reliable alignment interventions.

204. 【2609.35821】Can We Still Trust Disaster Social Sensing? Empirical Evidence on Detecting AI-Generated Social Media Posts

链接:https://arxiv.org/abs/2609.35821

作者:Xiaoshan Zhou,Zaifu Zhan

类目:Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:generative artificial intelligence, resemble eyewitness reports, converts public social-media, produce plausible messages, public social-media posts

备注:

点击查看摘要

Abstract:Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian needs, but generative artificial intelligence (AI) can produce plausible messages that resemble eyewitness reports. This study investigates whether text-based AI detectors can reliably distinguish human-authored from AI-generated disaster posts. We construct a dataset of 12,000 texts organised into 3,000 matched semantic units from nine disasters: original human posts (H0), minimally LLM-proofread human posts (H1), factual AI-generated posts based on the same verified facts (A0), and affectively framed versions of those AI posts (A1). A separate 6,000-text corpus from 42 events supports model selection and threshold calibration. We evaluate OSM-Det, Fast-DetectGPT, Binoculars, and direct large language model (LLM) judges across five model families, then test disaster-domain calibration, a frozen-encoder linear readout, paired transformation sensitivity, and dataset artifact controls. Across fourteen frozen cross-family configurations, AUROC is 0.402-0.517 and the best prospective recall at a calibration-derived low-false-positive operating point is 3.6%; OSM-Det reaches AUROC 0.521 and 10.4% recall at a realised 6.7% false-positive rate. A disaster-trained linear head reaches AUROC 0.817, but a seven-feature surface classifier reaches 0.784 on the H0-versus-A0 contrast, and neutralising identified surface asymmetries reduces the head from 0.733 to 0.594. The head also separates A0 from A1 even though provenance is unchanged. The results show that text-based detection is not reliable enough to serve as an operational trust gate; multimodal claims, accountable sources, and other contextual evidence should be rested on to safeguard trust in disaster social sensing.

205. 【2609.35820】$τ$-Multilingual: Benchmarking Voice Agents Across Languages

链接:https://arxiv.org/abs/2609.35820

作者:Soham Ray,Edgard dos Santos Paiva,Ruben Valenzuela,Karthik Narasimhan,Keshav Dhandhania,Victor Barres

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)

关键词:English-only benchmarks expose, English-only benchmarks, benchmarks expose, narrow slice, Brazilian Portuguese

备注:

点击查看摘要

Abstract:English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points. The failure modes also vary: Korean systems miss more responses, Mandarin systems interrupt more often, and both struggle with tools and entities. Grok leads task completion but scores lowest on generation quality, motivating separate task, interaction, and generation reporting. We release language packs, validated judges, and tools for community-built multilingual voice-agent evaluation.

206. 【2609.35817】Less Uniform Discrete Diffusion is More Powerful and Scalable

链接:https://arxiv.org/abs/2609.35817

作者:Kaibo Wang,Ding Ding,Fangyu Ding,Zijin Feng,Han Shi,Haili Bai,Jiacheng Sun,Yang Xiang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:promising diffusion paradigm, uniform diffusion language, represent a promising, diffusion language models, diffusion paradigm

备注: 23 pages, 8 figures, 5 tables

点击查看摘要

Abstract:Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.

207. 【2609.35816】PrimeSeeker: Capability-Oriented Supervision for Deep Search Agents

链接:https://arxiv.org/abs/2609.35816

作者:Linzhi Peng,Hanting Chen,Heng Chang,Ke Cheng,Bowen Du,Weifeng Lv

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language model, Large language, language model search, larger evidence graphs, additional hops

备注:

点击查看摘要

Abstract:Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indirect proxies for the local retrieval capabilities required during search. To address this mismatch, we introduce latent anchor reasoning, which consists of resolving an unnamed retrieval anchor from descriptive specifications and transferring the recovered anchor into a subsequent information demand. This primitive retrieval unit decomposes deep search into chains of coupled operations and organizes question construction around anchor resolution and relation transfer, without prescribing a canonical search path. Based on this formulation, we propose PrimeSeeker, a capability-oriented framework that constructs web-grounded anchor structures and jointly derives a question and a reference evidence skeleton. The skeleton preserves supporting evidence from construction and guides expert generation through extractive highlights of current tool observations. These highlights are removed before supervised fine-tuning, while the skeleton is subsequently reused to audit reference-step coverage for reinforcement-learning rewards. We construct 9,221 expert trajectories, training a 30B search agent. Across five deep-search benchmarks, PrimeSeeker achieves strong performance, while reference-step optimization further improves the supervised policy. The resulting trajectories exhibit low retrieval redundancy, and fixed-budget evaluation shows strong solution coverage with substantially fewer tool calls than long-horizon systems.

208. 【2609.35815】How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats

链接:https://arxiv.org/abs/2609.35815

作者:Ian Arawjo

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Methodology (stat.ME)

关键词:academia increasingly base, increasingly base significance, base significance claims, LLM judge scores, LLM judge

备注: 39 pages, 20 figures

点击查看摘要

Abstract:Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that running statistics over raw LLM judge scores leads to inflated false positives: counterintuitively, for many inter-rater agreement metrics, false positive risk peaks at "almost perfect" human-LLM agreement. To help researchers understand how to run statistics over LLM judges responsibly, we present guidance and tooling for the statistical analysis of mixed human-AI judge designs, and implement nine hypothesis tests via prediction-powered inference (PPI), including the first known PPI corrections for four rank-based tests (Wilcoxon signed-rank, Mann-Whitney U, and omnibus variants). To keep PPI++ stable with small human-labeled calibration sets, we introduce bootstrap-adaptive power tuning, which shrinks the estimated weight toward a target estimated from the labeled data, and accounts for that weight's own sampling variance. Second, through Monte Carlo simulations, we derive recommendations for what CI, p-value, and FWER correction methods to use for small-sample AI evaluations (N100), and warn researchers against bootstrap CIs. We package these recommendations into evalstats, an open-source Python package that selects calibrated methods automatically, and demonstrate it in three scenarios, including one where a real LLM judge validated at "substantial agreement" would have led a researcher to publish a spurious finding. evalstats is publicly available at this https URL.

209. 【2609.35814】Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions

链接:https://arxiv.org/abs/2609.35814

作者:Xunjian Yin,Tianchen Guan,Jinao Wang,Weili Cao,Daisy Xinlei Lin,Royce Cheng-Yue,Keagan Long,Kyle Wong,Bhuwan Dhingra,Xiangjun Wang,Shuyan Zhou

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:making tasks longer, pace by collecting, browser-use agents improve, agents improve, making tasks

备注: 40 pages

点击查看摘要

Abstract:As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challenging. We instead construct challenging instances from tasks agents already solve, turning difficulty into a programmable property of the environment. BreakingWeb pairs every base task with an intervention condition that preserves the user instruction, latent target, and backend success criterion while changing the environment at different web stack layers. Each intervention is deterministic, detectable, and recoverable, and is annotated with the cognitive primitive it primarily loads. The benchmark contains 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families, all graded against outcomes. We evaluate six strong browser-use agents, three GUI-only agents that see only screenshots, and humans. The construction is effective: interventions cut agent pass rate by 22.9% on average and overturn nearly half of the tasks each agent solves cleanly, whereas humans lose 10.0% on a first attempt and 5.7% after one familiarisation attempt. The dominant failure is belief failure: 75% of the six agents' failures end with a declared success although the required change never happened. Our code, data and environment are publicly available at this http URL.

210. 【2609.35812】Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants

链接:https://arxiv.org/abs/2609.35812

作者:Vaishnav Negi,Lev Sorokin,Soroosh Tayebi Arasteh,Andrea Stocco

类目:Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:In-car conversational assistants, support route planning, vehicle control, route planning, information access

备注: Accepted at the 29th IEEE International Conference on Intelligent Transportation Systems (IEEE ITSC 2026)

点击查看摘要

Abstract:In-car conversational assistants (ICAs) are increasingly integrated into vehicles to support route planning, vehicle control, and information access. Ensuring their reliability is challenging due to multi-turn interactions, the absence of explicit ground truth, and strict safety constraints. Existing evaluation techniques fall short, as they target single-turn settings and fail to capture constraint handling, context retention, and safety-critical behavior across turns. We propose an automated framework for testing the multi-turn conversational capabilities of ICAs. The system is treated as a black box and evaluated via closed-loop simulation with a strategy-guided user simulator, an adversarial strategy manager, and a two-tier LLM judge assessing turn-level failures and conversation-level quality. We evaluate the approach on an industrial ICA with six LLM backends and twelve human annotators. The automated judge shows substantial agreement with humans, and strategy guidance uncovers 2.96 times more unique failure types per conversation and more than doubles the number of unique failing conversations compared to unguided simulation.

211. 【2609.35811】Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning

链接:https://arxiv.org/abs/2609.35811

作者:Zongze Wu,Yani Guo,Runnan Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)

关键词:LLM-based agents operating, heterogeneous API ecosystems, operating over large, critical bottleneck, bottleneck for LLM-based

备注: 10 pages, 3 figures, 3 tables. Published in ICMR 2026

点击查看摘要

Abstract:Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40\%, outperforming the state-of-the-art ToolGen (90.16\%) by 1.24\%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.

212. 【2609.35810】RACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs

链接:https://arxiv.org/abs/2609.35810

作者:Jizheng Lai,Yingyun Li,Ying Qin,Haiyang Qian

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, Large language, language models, models are increasingly, weakly grounded

备注: 18 pages, 2 figures, 19 tables. Accepted to the EMNLP 2026 Industry Track for oral presentation

点击查看摘要

Abstract:Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensive offline structure learning from lightweight online inference: oncology concepts and relations are organized into an updatable tree-relational structure, refined using LM-loss-derived evidence, and retrieved at inference time as compact prompt evidence. This design supports task-adaptive evidence selection without requiring supervised labels in the zero-shot setting. Across ten oncology classification tasks and one MedQuAD CancerGov QA benchmark, TRACE improves both label-free evaluation and supervised fine-tuning. Additional analyses show that TRACE improves over vanilla RAG and generic GraphRAG, remains useful under leakage-controlled METABRIC inputs, and produces interpretable evidence paths aligned with clinical reasoning. These results suggest that explicit, updatable medical structure is a practical path toward more accurate and auditable oncology LLM deployment.

213. 【2609.35809】Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

链接:https://arxiv.org/abs/2609.35809

作者:Jiyao Yang,Yang Liu,Zhenyue Qin,Qingyu Chen,Xiuzhen Zhang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:large-scale disinformation campaigns, rapid advancement, advancement of generative, generative AI raises, raises concerns

备注: 15 pages, 6 figures. Accepted for publication in the Findings of EMNLP 2026

点击查看摘要

Abstract:The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open- and closed-source MLLMs for automated detection. We find that most models fall substantially short of human-level accuracy and fail critically on identifying image authenticity. Our research provides a foundation for developing robust defenses against social media fake news. Code and data are available at https: //github.com/xiuzhenzhang/Multimodal.

214. 【2609.35808】When Successful Memories Mislead Embodied Agents:Memory Adaption For Task-Conditioned Execution

链接:https://arxiv.org/abs/2609.35808

作者:Quanquan Li,Hongbo Zhang,Yihe Chi,Liuyang Song,Jingyu Li,Yuxiang Huang,Hongzhen Zhang,Guitao Cao

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:reduce repeated exploration, current execution context, reuse can reduce, reduce repeated, repeated exploration

备注:

点击查看摘要

Abstract:Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insufficient when retrieved ex perience contains incompatible actions or an inappropriate level of structure. We introduce Memory Adaptation for Task-Conditioned Execution (MATE), a deterministic post-retrieval procedure that converts trajectories into execution-oriented memory. MATE re moves obsolete control context, extracts condition-action-effect transitions, applies verified action normalization, selects a task dependent representation, and serializes the result under a fixed budget without additional LLM inference. On 134 ALFWorld tasks, MATE achieves task success rates of 81.3% and 93.3% with Qwen2.5-14B and 72B while using approximately one-tenth of the tokens required by raw trajectories. Controlled comparisons show that verified action normalization is the principal mechanism by which MATE restores the utility of retrieved experience, support ing memory adaptation as a distinct stage between retrieval and embodied execution.

215. 【2609.35807】Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

链接:https://arxiv.org/abs/2609.35807

作者:Charlie Summers,Prajwal Raghunath,Aaditya Pai,Mayur Kulkarni,Zhuo Zhang,Oliver Kennedy,Eugene Wu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)

关键词:make unsafe tool, behave safely, instructed to behave, unsafe tool calls, make unsafe

备注: 9 pages, 11 figures, REALM Workshop, EMNLP 2026

点击查看摘要

Abstract:LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur---we call this Environment Steering. We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime. When violations are detected, policy- and context-specific feedback steers the agent toward safe trajectories. On AgentDyn, this enables the agent to improve task success rate over no-defense while achieving 0% attack success rate.

216. 【2609.35806】From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

链接:https://arxiv.org/abs/2609.35806

作者:Ranuga Disansa,U. S. Samarasinghe,Lasith Gunawardena

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:underpins workforce planning, extraction underpins workforce, systems represent skills, skill extraction underpins, workforce planning

备注:

点击查看摘要

Abstract:Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA) captures exactly this dimension, defining 147 professional skills across seven responsibility levels, but no automated LLM-based extraction targeting SFIA has been reported. We formalize the task as structured prediction of (skill, level) pairs from free text and ask three questions: how accurately can text be mapped onto SFIA's closed vocabulary, which strategies reliably predict the level alongside the skill, and do agentic designs improve on simpler retrieval and prompting? We evaluate five strategies (a lexical baseline, dense retrieval with LLM reranking, a zero-shot schema-constrained LLM, single-agent agentic RAG, and a three-agent retriever--matcher--verifier crew) against expert-mapped European ICT role profiles, all drawing on an SFIA~9 corpus built by a fully automated agentic pipeline that we release. Retrieval-based matching identifies the most skills while generative strategies are markedly more precise; only strategies assigning the level as an explicit decision predict it reliably, with similarity-based selection more than twice as inaccurate; and the crew doubles latency without improving accuracy, so added agent roles do not automatically benefit closed-taxonomy matching. These results provide the first reproducible baseline for structured, level-aware skill extraction against SFIA.

217. 【2609.35805】Alignment Forecasting: Predicting Misalignment From Training Data

链接:https://arxiv.org/abs/2609.35805

作者:Chen Yueh-Han,Bruce W. Lee,Ilia Sucholutsky,Tomek Korbak

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:narrow flaw, Training, model, failure mode, model broadly misaligned

备注:

点击查看摘要

Abstract:Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.

218. 【2609.35804】Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

链接:https://arxiv.org/abs/2609.35804

作者:Mamehgol Yousefi,Ahmad Shahi,Mos Sharifi,Alvaro Romera,Simon Hoermann,Tham Piumsomboon

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)

关键词:Large language models, natural language processing, shown remarkable capabilities, Large language, language processing tasks

备注: 14 pages. Published in ICONIP 2024 (Neural Information Processing), LNCS 15290, Springer Nature, 2025

点击查看摘要

Abstract:Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It's found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others. These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness. This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.

219. 【2609.35796】Developing an OCR model for Extracting Information from Invoices with Korean Language

链接:https://arxiv.org/abs/2609.35796

作者:Xiem HoangVan,Phu TranQuang,Minh DinhBao,Tien VuHuu

类目:Computation and Language (cs.CL)

关键词:including the purchased, purchased items, total money, commercial documents, Optical Character Recognition

备注: 2023 International Conference on Advanced Technologies for Communications (ATC)

点击查看摘要

Abstract:Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money. Making the extraction of important information crucial. The stored information serves different purposes. Korean language is the native language of about 80 million people, playing an important role in not only South and North Korea but also in many other countries such as Vietnam, Philippine where a large number of Korean companies are located. In this context, to automatically extract proper information from the invoices with Korean language, we propose an efficient Optical Character Recognition (OCR) model in which a deep learning model is combined with some image preprocessing techniques. The proposed OCR model is assessed in a rich set of collected invoices showing that 87% F1-score can be achieved with negligible time processing.

220. 【2609.35794】Sieve and Sage: Efficient Distraction Filtering for Reliable RALM Abstention

链接:https://arxiv.org/abs/2609.35794

作者:Jongbin Won,Sung Geun An,Jay-yoon Lee

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Retrieval-Augmented Language Models, Language Models, Retrieval-Augmented Language, Socrates recognized, recognized the limits

备注: 20 pages, 6 figures, accepted at EMNLP 2026 Findings

点击查看摘要

Abstract:Just as Socrates recognized the limits of his own knowledge, Retrieval-Augmented Language Models (RALMs) should learn to abstain when the retrieved evidence cannot support a reliable response. Existing approaches largely rely on monolithic LLMs to handle heterogeneous retrieval failures in a single step, resulting in limited abstention performance and high computational costs. We instead decompose retrieval failures into two distinct states: (i) the unanswerable state, where the required evidence is absent, and (ii) the distracted state, where relevant evidence is mixed with conflicting, negated, or adversarial information. Based on this decomposition, we introduce a lightweight module (Sieve) that screens retrieved document sets for distracting evidence before invoking a costly LLM (Sage) for grounded generation and abstention. Evaluated across both general and high-stakes expert domains, our Sieve and Sage framework preemptively detects distracting noise, improving system accuracy by up to 69.4 percentage points and Macro-F1 by 55.2 percentage points compared to one-stage baselines. Furthermore, it achieves up to a 1.99x speedup, establishing a highly efficient and reliable abstention pipeline for RALM with abstention.

221. 【2609.35791】FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

链接:https://arxiv.org/abs/2609.35791

作者:Puneet Mathur,Dinesh Manocha

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:interaction requires determining, pause reflects hesitation, full-duplex voice interaction, voice interaction requires, Natural turn-taking

备注:

点击查看摘要

Abstract:Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ (at FP=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.

222. 【2609.35790】Sage: Formalization with Semantic Correction

链接:https://arxiv.org/abs/2609.35790

作者:Thomas Hirtz,Farzad Jafarrahmani,Abdelmouksit Sagueni,Xiang Zhou,Wengping Deng,Liang Zhang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Logic in Computer Science (cs.LO)

关键词:neural theorem provers, achieved impressive milestones, neural theorem, theorem provers, provers have achieved

备注: 27 pages, 4 figures. Preprint

点击查看摘要

Abstract:While neural theorem provers have achieved impressive milestones in formal mathematics, they largely operate on the assumption that faithful Lean 4 formal statements are already provided. Translating informal natural language into a formal language is a critical data bottleneck plagued by an "illusion of rigor": standard type-checkers accept statements that compile but drop hypotheses, introduce vacuous truths, or subtly alter mathematical bounds. To resolve this, we introduce Sage (Semantic Agent-Guided Formalization Engine), an agentic framework that replaces monolithic translation with a four-stage decomposed generation pipeline coupled with a dual-signal semantic correction loop. By pairing Lean 4 compiler diagnostics with multi-dimensional semantic feedback, our correction loop enforces mathematical fidelity alongside syntactic validity. By explicitly accounting for the gap between open-ended queries and declarative formal targets, our pipeline prevents models from achieving high formalization rates by guessing unverified answers (exhibiting a 70.9% answer leakage rate). Consequently, Sage suppresses leakage to 2.7% while achieving 73.3% pass@4 joint compilation and semantic fidelity on the Omni-MATH without proofs (compared to 42.0% for a fine-tuned Goedel-Formalizer-V2 baseline). Finally, on IMO-Unformalized, a novel frontier of 175 unformalized International Mathematical Olympiad problems, Sage demonstrates effective zero-shot generalization with 87.4% pass@4 verified fidelity compared to just 19.4% for the baseline, winning over 79% of blind pairwise evaluations.

223. 【2609.35779】Large Language Models Exhibit Human-Like Bayesian Hypocrisy

链接:https://arxiv.org/abs/2609.35779

作者:Nykko Vitali,Mahzarin R. Banaji

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:large language models, frontier models, language models, recent achievements, achievements of large

备注: Main text (38 pages) with supplementary materials appended (213 pages total). Preregistered with data and analysis code at [this https URL](https://osf.io/95tfc)

点击查看摘要

Abstract:Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-making (Cao, et al., 2019). In 5 experiments with 48 experimental conditions employing over 5,000 trials, GPT-4o and Claude 3.7 Sonnet were tested on two variations of a Bayesian reasoning task. We also assessed LLM evaluation of the competence and morality of a hypothetical person who had offered the same reasoning task as them. LLMs hovered near human performance on the Bayesian task, though their reasoning was more rule-based and rigid. Surprisingly, like humans but to a greater extent, LLMs also demonstrated the same hypocrisy in condemning others who, like them, had deployed Bayes' rule. In demonstrating Bayesian hypocrisy, LLMs highlight a humanlike error of a dissociation between self-performance and other-judgment, and caution against their use in domains where statistical fidelity and fairness norms collide.

224. 【2609.35150】oward a Culturally Adapted Chinese Language Agent: A Wizard-of-Oz Study of Nonverbal Behavior in Chinese-German Intercultural Interaction

链接:https://arxiv.org/abs/2609.35150

作者:Siddhant Jain,Anna Lea Reinwarth,Dimitra Tsovaltzi,Rafael Math,Julia Renner

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Successful intercultural communication, intercultural communication requires, Successful intercultural, grammatical competence, intercultural communication

备注: Accepted to ICMI Companion '26. 7 pages, 4 figure

点击查看摘要

Abstract:Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existing language-learning agents. We present a Wizard-of-Oz (WoZ) study design and supporting real-time system for collecting multimodal behavioral data from native Chinese speakers reacting to social norm violations by German learners. The system features a photorealistic MetaHuman avatar driven by Live Link face capture and MediaPipe upper-body tracking, a wizard console for real-time behavior selection, and synchronized multimodal logging across agent and learner streams. A layered annotation framework, based on psychological theory and covering non-observable socioemotional reactions, norm interpretation, verbal, and observable behavior thereof, and future supervision targets enables the corpus to support training of future automated cultural interpretation and behavior generation models. Four ecologically valid interaction scenarios, developed with cultural and pedagogical experts, provide the methodological and technical foundation for a culturally adapted conversational agent for Chinese language learning.

225. 【2511.08592】he Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions

链接:https://arxiv.org/abs/2511.08592

作者:Azza Bouleimen,Giordano De Marzo,Taehee Kim,Nicol`o Pagan,Hannah Metzler,Silvia Giordano,Anikó Hannák,David Garcia

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Social and Information Networks (cs.SI)

关键词:Large Language Models, Large Language, Language Models, simulate online communities, offer new avenues

备注:

点击查看摘要

Abstract:Large Language Models (LLMs) offer new avenues to simulate online communities and social media. Potential applications range from testing the design of content recommendation algorithms to estimating the effects of content policies and interventions. However, the validity of using LLMs to simulate conversations between various users remains largely untested. We evaluated whether LLMs can convincingly mimic human group conversations on social media. We collected authentic human conversations from Reddit and generated artificial conversations on the same topic with two LLMs: Llama 3 70B and GPT-4o. When presented side-by-side to study participants, LLM-generated conversations were mistaken for human-created content 39\% of the time. In particular, when evaluating conversations generated by Llama 3, participants correctly identified them as AI-generated only 56\% of the time, barely better than random chance. Our study demonstrates that LLMs can generate social media conversations sufficiently realistic to deceive humans when reading them, highlighting both a promising potential for social simulation and a warning message about the potential misuse of LLMs to generate new inauthentic social media content.

226. 【2503.03791】Predicting Team Performance from Communications in Simulated Search-and-Rescue

链接:https://arxiv.org/abs/2503.03791

作者:Ali Jalal-Kamali,Nikolos Gurney,David Pynadath

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:influence team performance, performance is valuable, directly observable, individual traits influence, Understanding how individual

备注:

点击查看摘要

Abstract:Understanding how individual traits influence team performance is valuable, but these traits are not always directly observable. Prior research has inferred traits like trust from behavioral data. We analyze conversational data to identify team traits and their correlation with teaming outcomes. Using transcripts from a Minecraft-based search-and-rescue experiment, we apply topic modeling and clustering to uncover key interaction patterns. Our findings show that variations in teaming outcomes can be explained through these inferences, with different levels of predictive power derived from individual traits and team dynamics.

227. 【2609.36754】Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT

链接:https://arxiv.org/abs/2609.36754

作者:Ki Woong Moon,Daniel Brenner

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:Explicit prosodic cues, automatic speech recognition, typically require additional, Explicit prosodic, additional trainable components

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.

228. 【2609.36097】Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

链接:https://arxiv.org/abs/2609.36097

作者:Hanbo Xie

类目:Neurons and Cognition (q-bio.NC); Computation and Language (cs.CL)

关键词:explain cognition requires, explain cognition, cognition requires, model, information

备注:

点击查看摘要

Abstract:Using predictive models to explain cognition requires more than accurate behavioral predictions. Input ablations offer an appealing route: remove information from a model and interpret the resulting performance change as evidence of its importance for behavior. Yet this inference assumes that the model's dependence on information reflects the dependence of the process generating the behavior. We test it in two synthetic sequential bandit tasks with known generating policies, where past choices can remain informative when feedback is unavailable to a predictor. We compare GRUs and Transformers trained from scratch, a fine-tuned LLaMA model, and cognitive models across systematically varied reward contributions. Our analyses distinguish prediction after training without reward observations from the response of a fixed predictor to donor-reward replacement. Three findings emerge. First, in the restless task, neural models trained without rewards predict held-out choices better than four simple training-fitted behavioral baselines. Second, under matched donor replacement, accurate predictors can respond much less than the known generator. Third, at some reward weights, neural networks predict better than a pooled reinforcement-learning model but have less faithful changes in choice probabilities; the model ordering differs between the two tasks. These independent-test results separate information sufficient for prediction from response fidelity under a specified ablation in sequential choice. They motivate validating model-ablation responses independently of predictive performance before using them to infer how the observed behavior was generated.

229. 【2609.35813】Local Predictability and Collective Fidelity in LLM-Agent Societies

链接:https://arxiv.org/abs/2609.35813

作者:Igor Itkin

类目:Physics and Society (physics.soc-ph); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:language model societies, simulating large language, large language model, Compact surrogates, reproduce collective behavior

备注: 35 pages, 7 figures. Standalone empirical companion to [arXiv:2608.11215](https://arxiv.org/abs/2608.11215)

点击查看摘要

Abstract:Compact surrogates could reduce the cost of simulating large language model societies, but must reproduce collective behavior. We compare individual predictions and collective forecasts using 9,455 published trajectories and new experiments on opinion dynamics. Neighbor information improves individual prediction in all 16 public-data settings and pooled collective forecasts on held-out questions, although collective gains depend on transfer conditions. Tests on 24 new statements do not confirm earlier contrasting history effects in forecasts from the initial state. Qwen benefits from history after three observed rounds. These findings motivate direct collective validation, explicit limits on available observations, and comparisons with simple baselines.

信息检索

1. 【2609.38155】Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

链接:https://arxiv.org/abs/2609.38155

作者:Hui Ren,Lei Fan,Henry Pao,Han Guo,Zeeshan Zia,Ying Chen,Alexander Schwing,Gang Hua

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:requires connecting events, connecting events involving, hours or days, long videos, videos often requires

备注:

点击查看摘要

Abstract:Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

2. 【2609.38099】Effective Dense Retrieval using Only In-Context Examples

链接:https://arxiv.org/abs/2609.38099

作者:Nour Jedidi,Abdul Basit Ali,Hang Li,Jimmy Lin

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Turning decoder-only large, large language models, decoder-only large language, Turning decoder-only, language models

备注:

点击查看摘要

Abstract:Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple "training-free" approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at this https URL.

3. 【2609.38021】Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S

链接:https://arxiv.org/abs/2609.38021

作者:Christopher J. Chanhnourack

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:auditable long-term memory, long-term memory system, system on LongMemEval-S, evaluate an auditable, auditable long-term

备注: Technical report, 14 pages. Evidence repository (reader outputs, judge verdicts, control records, judge harness): [this https URL](https://github.com/cjchanh/longmemeval-evidence) (MIT). Re-scoring any run under the official judge costs about $1.28

点击查看摘要

Abstract:We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called through an unpinned CLI alias, two 500-question passes score 479/500 and 475/500 under GPT-4o. The 72 answerable knowledge-update rows used a substantively modified scoring prompt whose effect under the official text has not been measured. The pair straddles Chronos High's published 478/500; differences in reader generation, scoring prompt, and possibly data version, plus within-system variance, establish neither superiority nor equivalence. A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465. The headline passes differ on eight verdict-flip rows. A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers. Negative controls rejected a verifier that repaired three wrong drafts but broke eleven correct drafts. All components were developed on the same 500 questions, with no held-out evaluation or independent human adjudication; retrieval and scaffold method sources and transcript-derived audits are held; and the headline reader received extra operator context, its complete requests were not retained, and MCP tool availability is unresolved. We release materialized packets, scaffolds, reader outputs, judge verdicts, and controls for inspection and re-scoring.

4. 【2609.37993】BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

链接:https://arxiv.org/abs/2609.37993

作者:Julien Knafou,Luc Mottin,Alexandre Flament,Paul van Rijen,Esteban Gaillac,Patrick Ruch

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:BITEM team entered, single agentic pipeline, BITEM team, records evidence, entered both subtasks

备注: 8 pages. Participant paper for the NTCIR-19 R2C2 task

点击查看摘要

Abstract:The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is computed by the orchestrator from what the run leaves behind and is never asked of the model, which is offered no way to rate itself. The two retrieval runs placed 4th and 5th of 22, pooling the passes was worth 0.0709 nDCG@20, and the gain was largest on the multi-hop and post-processing-heavy questions, where the organisers rank the pooled run top of the field. Sixteen of the 25 answer runs were built on passages these two runs supplied, 12 of them filed by other teams. HMR rewards a system whose confidence is high where it answers right and low where it answers wrong. The pipeline reached an accuracy of 0.9219, 6th of 25, while the confidence filed with those answers gave an HMR of 0.4915, 13th. A few rules crafted over those same recorded signals, with no further model call and no further retrieval, raise that to an accuracy of 0.9375, 5th, and an HMR of 0.6985, 9th. Ranking on HMR alone can reward a system for answering wrongly with low confidence, so we propose accHMR, the accuracy multiplied by HMR, which reports the reward in proportion to the accuracy, and on which the revised rules would have scored 0.6549, 5th. For future work, fitting a model on the numbers the pipeline already produces, rather than writing such rules by hand, would be a real step forward.

5. 【2609.37911】Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3

链接:https://arxiv.org/abs/2609.37911

作者:Ryan C. Barron,Cade W. Trotter,Maksim E. Eren,Kim Ø. Rasmussen,Liz D. Miller,Benjamin J. Migliori

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Scientific queries, relevant papers, papers use specialized, Scientific, specialized vocabulary

备注: 8 pages, 5 tables, 3 figures

点击查看摘要

Abstract:Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-references, and corpus-steered text all together with SPLADE-v3 on NFCorpus, TREC-COVID, and SciDocs. Every condition searches the same frozen document index and follows the same query-side integration rule and 256-dimension budget, isolating the effect of the added content. All twelve method-collection comparisons improve aggregate nDCG@10, with best relative gains of 4.81%, 8.92%, and 9.47%. Eleven remain significant after Holm correction. The gain persists in 103 of 114 interpolation settings, including every setting that assigns at least 30% of the mixture weight to the original query. Shuffled-text and non-contextual lexical-bag controls also remain above baseline in all 24 aggregate comparisons, showing that the added vocabulary carries most of the benefit. A corpus-induced typed concept graph, by contrast, produces no consistent gain, and its relation, depth, validation, random, and gating controls do not rescue it. Generated vocabulary can therefore complement a strong learned sparse retriever, provided that the original query remains strongly represented.

6. 【2609.37749】owards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy

链接:https://arxiv.org/abs/2609.37749

作者:Mohamed Ben Salha,Fiete Lüer,Maik Betka,Stefan Wagner

类目:Information Retrieval (cs.IR)

关键词:Retrieval Augmented Generation, managing large datasets, Retrieval Augmented, highlighted significant limitations, Augmented Generation

备注: 8 pages, 3 figures, 2 tables

点击查看摘要

Abstract:The increasing importance of Information Retrieval (IR) in managing large datasets has highlighted significant limitations in traditional keyword-based search systems. Context-aware chat-based search methods, such as Retrieval Augmented Generation (RAG), have recently emerged, but their evaluation compared to keyword-based systems often relies on subjective user feedback. A rigorous, quantitative comparison between these paradigms remains lacking. This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages. It focuses on two key aspects: the ranking accuracy for keyword-based systems and the completeness of retrieved information for semantic chat-based systems. Our approach enables semi-automatic comparisons of semantic and keyword-based methods using interchangeable equivalence classes tailored to domain-specific contexts (e.g., companies or problems). We validate the framework through an industrial case study, demonstrating statistically significant improvements in context-aware search over keyword-based methods, supported by analyses including the Mann-Whitney U-Test. With its adaptable design, the proposed framework provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.

7. 【2609.37574】MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment

链接:https://arxiv.org/abs/2609.37574

作者:Tzu-I Ho,Yung-Yu Shih,Shang-Yu Su,Dongzhe Wang,Yun-Nung Chen

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Large Language Models, bridge vocabulary gaps, Language Models, enrich user queries, Large Language

备注: 9 pages, 4 tables, 1 figure. Preprint

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model -- an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce candidate expansions, and a larger LLM generatively synthesizes them into a single query. To make prompt engineering scalable across the ensemble, we integrate a task-grounded Automatic Prompt Optimization (APO) loop into both stages. Unlike APO methods that judge candidates with an LLM evaluator, our loop scores each candidate by its downstream retrieval performance and runs a small tournament between the current champion prompt and optimizer-proposed drafts, terminating once the champion survives two consecutive rounds; a history-augmented variant additionally feeds the recent tournament trajectory back to the optimizer. MERGE is retriever-agnostic and issues a single BM25 pass with no rank fusion, no supervised document expansion, and no re-indexing. On five BEIR benchmarks (NQ, SciFact, FiQA, Touche-2020, DBPedia), MERGE improves BM25 nDCG@10 over the original queries by +2.1 to +14.9 points and matches or outperforms strong LLM-based query-expansion baselines despite using only compact open-source models. Ablations confirm that the Stage-2 ensemble beats any single Stage-1 LLM, and that task-grounded APO converts large seed-prompt regressions into consistent gains without hand-tuning.

8. 【2609.37472】Do Evidence-Reading Diagnostics Improve Interface Selection in Small LLM Recommenders?

链接:https://arxiv.org/abs/2609.37472

作者:Han Chen,Yingrui Li

类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Behavioral tests measure, language model reads, Behavioral tests, Behavioral, model reads evidence

备注:

点击查看摘要

Abstract:Behavioral tests measure how a language model reads evidence. We ask whether those measurements help choose a recommendation interface. We evaluate six small instruction-tuned checkpoints across four recommendation domains with chronological evaluation and 3,426 evaluation users. Each request ranks eight candidates. A baseline selector chooses among history-only prompting, prompting with collaborative evidence, and score fusion. It uses observable features and six stability prompts that vary wording and candidate order. An augmented selector adds features from six evidence-reading prompts that ask the model to compare support counts. An interface chosen once on development (validation) data for each domain and checkpoint scores 0.5524 NDCG@5, compared with 0.5447 for the baseline selector and 0.5428 for the augmented selector. Adding the diagnostic features changes NDCG@5 by -0.0019 (95% interval [-0.0046, 0.0004]). The interval includes zero, and its upper bound is below the analysis plan's 0.005 improvement target. Matching the selectors' hyperparameters also leaves the interval upper bound below that target. Evidence from retrieved similar users improves prompting by 0.0999 NDCG@5 over a control using randomly selected users matched for activity. The evidence-reading tests also reveal answer-position and tie-response biases. These results concern the tested selectors and candidate sets. They illustrate why diagnostic measurements should be evaluated by whether they improve recommendation choices beyond existing features and a fixed interface.

9. 【2609.37469】Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

链接:https://arxiv.org/abs/2609.37469

作者:Suting Chen,Peichun Hua,Yunming Xiao

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:grounds large language, Retrieval-augmented generation, grounds large, external sources, entities without providing

备注: 22 pages, 7 tables, 2 figures

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence tests: removing relevant evidence or pairing evidence with unrelated questions can reveal labels through lexical overlap or evidence position. We build a paired benchmark using substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. Sufficiency can be judged without generating an answer, but no single signal works across all datasets. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals: whether every part of the question is covered, whether any passage offers an answer, and whether a small language model reading the passages together judges them sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), exceeding the best of 10 prior methods (0.746) and a frontier model queried through an API (0.784). Its weakest dataset scores higher than any other method's weakest (0.684 vs. 0.676). RINSE runs locally before generation, taking 36.5 ms per question on a single GPU.

10. 【2609.37468】Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers

链接:https://arxiv.org/abs/2609.37468

作者:Beining Xu,Peichun Hua,Yunming Xiao

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Agentic retrieval-augmented generation, Agentic retrieval-augmented, subsequent search decisions, retrieval-augmented generation, interleaves reasoning

备注: 24 pages, 13 tables, 4 figures

点击查看摘要

Abstract:Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak backdoor purification to conceal their presence. An attacker supplies a compromised retriever checkpoint while leaving the search agent and deployment corpus unchanged. Without corpus write access, the attacker can still suppress useful evidence, persistently retrieve a selected existing document, or steer the agent toward prolonged search, inflating retrieval, context, and latency cost. To conceal these behaviors from detection, we propose leveraging a controlled inject-and-remove cycle: deliberately inject a weaker backdoor and then unlearn it. This process weakens detector-visible signatures and fools the backdoor detectors with an illusion of purification while preserving the malicious retrieval behavior. These findings expose a systematic vulnerability in RAG systems in which a weak defense becomes an attacker's concealment tool for a backdoored retriever, even when the underlying corpus remains trustworthy.

11. 【2609.37311】ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents

链接:https://arxiv.org/abs/2609.37311

作者:Haohao Qu,Yongcheng Jing,Chun Hin Chan,Shanru Lin,Wenqi Fan,Dacheng Tao

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Recent Recommendation Agents, generative agents autonomously, agents autonomously perceive, perceive external platforms, Recent Recommendation

备注: Work in progress

点击查看摘要

Abstract:Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16\% across three recommendation agent tasks, namely searching, ranking, and judging.

12. 【2609.37226】Follow the Entities: A Corpus Map for Agentic Search

链接:https://arxiv.org/abs/2609.37226

作者:Soyeong Jeong,Sujay Kumar Jauhar,Sung Ju Hwang,Andrew Joohun Nam

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:project approval recorded, Answering questions, requires connecting evidence, connecting evidence spread, questions and completing

备注:

点击查看摘要

Abstract:Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.

13. 【2609.37183】HELIX: Purified and Unified - Rethinking Feature Interaction and Sequence Modeling for Large-Scale Recommendation

链接:https://arxiv.org/abs/2609.37183

作者:Yuntao Zheng,Miao Zhang,Yadong Ding,Yanchuan Tang,Lixiyu Chen,Hao Wang,Quan Li,Shiying Cai,Yue Lin,Jiayu Li,Yu Feng,Wentao Yang,Rongkun Xing,Jiekai Wang,Mingge Zhang,Feiling Gong,Xiang Gao,Jinyu Dong,Yajing Zhang,Pengfei Ren,Yinzhou Wang

类目:Information Retrieval (cs.IR)

关键词:models typically scale, user behavior histories, Industrial recommendation ranking, ranking models typically, multi-type user behavior

备注: 17 pages, 3 figures. Technical report

点击查看摘要

Abstract:Industrial recommendation ranking models typically scale along two modeling axes: feature interaction over heterogeneous user, item, context, and cross features, and sequence modeling over long, informative, and multi-type user behavior histories. We find that scaling either capability in isolation is insufficient, as each exhibits a limited scaling ceiling and a suboptimal scaling-law slope. We conjecture that achieving a more favorable scaling-law slope requires jointly scaling both axes. To support this, we present HELIX, a purified and unified architecture for large-scale recommendation. HELIX interleaves sequence retrieval and feature interaction while enforcing one-way information flow from reusable sequence states to candidate-conditioned mix-tokens. This design preserves cross-depth communication between the two modeling axes while keeping user-side sequence computation amortizable, enabling flexible and asymmetric scaling of sequence modeling and feature interaction. Deployed in TikTok's e-commerce recommendation system, HELIX consistently improves offline CTR AUC, CVR AUC, and other ranking metrics. In online A/B tests, it achieves an approximately 6% increase in e-commerce video GMV per user.

14. 【2609.36946】Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval

链接:https://arxiv.org/abs/2609.36946

作者:Xuri Ge,Chunhao Wang,Junchen Fu,Haokun Wen,Zhiwei Xu,Ying Zhou,Zhumin Chen,Pengjie Ren,Zhaochun Ren,Xin Xin

类目:Information Retrieval (cs.IR)

关键词:image-text matching space, encoding composed queries, VLP representation space, paired supervision, typically by encoding

备注:

点击查看摘要

Abstract:ZS-CIR aims to retrieve a target image from a reference image and a modification text without paired supervision, typically by encoding composed queries as text-dominant representations within the image-text matching space of VLPs. However, queries reconstructed by visual pseudo-word learning or MLLM-based target reasoning often deviate from the native VLP representation space due to reference noise and coarse text fusion in the former, and verbose, weakly visually grounded descriptions in the latter. In this paper, we propose a unified ZS-CIR framework (named VMIR-CVI) to reconstruct multimodal composite queries from two complementary perspectives for optimizing VLP-compatible multimodal intent representation. First, it reasons and converts the multimodal intent into a unified textual description, aligning with the native text space of the VLP backbones to produce more retrieval-compatible textual queries. Second, it reconstructs the query representation with correctly decoupled visual instance cues, reducing reference noise while preserving target-relevant content. Specifically, a VLP-aligned Multimodal Intent Reasoning (VMIR) module injects few-shot VLP-style exemplars into chain-of-thought prompts, guiding the MLLM to generate target-consistent intent queries. A Training-free Visual Instance Disentanglement (TVID) module decouples fine-grained visual instances from global reference features without additional optimization. Finally, a lightweight Hybrid-modal Intent Alignment and Fusion (HIAF) module integrates the reasoned textual intent and disentangled visual cues into a unified hybrid-modal representation for robust ZS-CIR. Extensive experiments on three CIR benchmarks, namely CIRR, CIRCO and FashionIQ, show that VMIR-CVI significantly outperforms existing baselines and achieves new state-of-the-art performance. Code and trained models will be publicly released.

15. 【2609.36862】Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning

链接:https://arxiv.org/abs/2609.36862

作者:Muhammad Zeeshan Akram,Mufid Kamel Marican,Anvesh Reddy Yenugu,Ali Zain Kaimkhani,Minghong Fang

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:safety-aligned language model, harmful data mixed, fine-tuning attack surface, benign fine-tuning set, data mixed

备注: To appear in CCS-LAMPS 2026

点击查看摘要

Abstract:Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.

16. 【2609.36849】Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

链接:https://arxiv.org/abs/2609.36849

作者:Omar Sheta,Rinku Deuja,Hadi Masoudi,Minghong Fang

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Safety-aligned language models, Safety-aligned language, spread unsafe intent, adversaries spread unsafe, benign conversations

备注: To appear in CCS-LAMPS 2026

点击查看摘要

Abstract:Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.

17. 【2609.36688】GRP v0.1 Technical Report

链接:https://arxiv.org/abs/2609.36688

作者:Wenfeng Zhuo,Vincent Xue,Charles Wei,Cong Ni,Ruiming Lu,Jiwen Ren,Mo Li,Peng Yang,Xufei Wang,Dongheng Li,Jiacong He,Yi Song,Yufei Fan,Mikhail Obukhov,Yiwen Chen,Yvette Liu,Yin Ye,Chengjie Wu,Mingtao Zhang,Jinchao Ye,Lili Zhang,Chunhui Zhu

类目:Information Retrieval (cs.IR)

关键词:Industrial recommendation systems, recommendation systems rely, Industrial recommendation, systems rely, rely on multi-stage

备注: 26 pages, 3 figures, 11 tables. Technical report

点击查看摘要

Abstract:Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework that combines retrieval, ranking, and reward modeling in a single encoder-decoder model, and evaluate a progressive path toward end-to-end recommendation. The model generates multimodal Semantic IDs and scores candidates with a jointly trained ranking module. The frozen ranking module then supplies rewards for reinforcement-learning post-training. We introduce mGRPO, which adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets. Offline experiments examine history encoding, model capacity allocation, event selection, tokenization, and reward discrimination. Serving optimizations reduce end-to-end retrieval latency by 69%. Online experiments evaluate the model as a retrieval source, with early-ranking bypass, and with replacement of weaker sources. In a retrieval-only comparison, view time increases by 0.46% and shares by 0.77% relative to production. A separate comparison combining bypass and source replacement yields increases of 0.82% in view time and 2.56% in shares, with neutral platform-level guardrails. These results support progressive deployment while identifying remaining gaps in ranking quality and performance across recommendation metrics.

18. 【2609.36534】Retrieval Sensitivity to Identity Signals in Queries

链接:https://arxiv.org/abs/2609.36534

作者:Andrew Tang,Nicholas Deas,Kathleen McKeown,Vishal Misra

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR)

关键词:documents reach users, Dense retrievers decide, decide which documents, documents reach, typically evaluated

备注: EMNLP 2026 camera-ready, with a correction to Fig. 4

点击查看摘要

Abstract:Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries---political ideology and dialect---bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query's own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries' identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers' query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at this https URL.

19. 【2609.36392】ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering

链接:https://arxiv.org/abs/2609.36392

作者:Yuyan Chen

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:standard Retrieval-Augmented Generation, Retrieval-Augmented Generation systems, Retrieval-Augmented Generation, Generation systems, dynamic conflict-aware synthesis

备注: 13 pages, 7 figures, 5 tables

点击查看摘要

Abstract:In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at this https URL.

20. 【2609.36359】Better Nearest Neighbor Graph Indices via (Efficient) LLM-Guided Pruning

链接:https://arxiv.org/abs/2609.36359

作者:Fangzhou Wu,Haike Xu,Sandeep Silwal

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Graph-based approximate nearest, approximate nearest neighbor, approximate nearest, semantic, Graph-based approximate

备注: 29 pages

点击查看摘要

Abstract:Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on geometric relationships among embeddings of an input dataset (e.g., documents or images), rather than explicitly optimizing for semantic relevance. However, when using these indices for downstream query retrieval, performance is evaluated based on the semantic relevance of the retrieved results to the query. This creates a fundamental "geometry-semantic" mismatch between how the indices are constructed and how their retrieval results are evaluated. While existing LLM-based reranking methods can partially mitigate this mismatch at query time, they leave this underlying structural problem in the graph unresolved. We therefore propose LLM-Guided Graph Pruning (LGP), a general framework that addresses this mismatch directly by leveraging LLM reasoning to refine an existing ANN graph index itself. LGP identifies structurally "low-value" neighbors of nodes and replaces them with LLM-selected alternatives that provide useful semantic information while retaining desired geometric structures of the original graph, including sparsity and efficient navigability. Experiments on representative semantic retrieval benchmarks show that LGP consistently improves end-to-end retrieval performance over both vanilla greedy graph search and LLM-based reranking across widely used graph-based ANN indices such as DiskANN and HNSW.

21. 【2609.36340】huRunel: Dynamic Decoupling for Structured Advisory Dialogue

链接:https://arxiv.org/abs/2609.36340

作者:Yuyan Chen

类目:Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)

关键词:High-stakes advisory domains, educational planning exhibit, legal consultation, medical aesthetics, High-stakes advisory

备注: 14 pages, 22 figures, 8 tables

点击查看摘要

Abstract:High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.

22. 【2609.36082】GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis

链接:https://arxiv.org/abs/2609.36082

作者:Ethan D. Frakes,Amy Kvien,Rishabh Kundu,Redad Mehdi,Van D. Tran,Vibha S. Mandayam,Kristopher O. Davis,Erika I. Barcelos,Roger H. French,Yinghui Wu,Mengjie Li

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:LLM-based geospatiotemporal KGQA, existing KGQA benchmarks, assessing LLM-based geospatiotemporal, Unlike existing KGQA, KGQA benchmarks

备注: 13 pages, 6 figures, 7 tables. Accepted to the 34th ACM International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26), November 3-6, 2026, Riverside, CA, USA

点击查看摘要

Abstract:We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs' understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at this https URL.

23. 【2609.36059】Mnemon: Raw Records, Fast Judgments, Slow Thoughts

链接:https://arxiv.org/abs/2609.36059

作者:Guangren Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:memory systems build, Long-term memory, graphs or typed, typed memories, memories at write

备注: 16 pages, 3 figures, 4 tables. Code, prompts and run records: [this https URL](https://github.com/Grivn/mnemon-memory-agent)

点击查看摘要

Abstract:Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster.

Comments:
16 pages, 3 figures, 4 tables. Code, prompts and run records: this https URL

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

Cite as:
arXiv:2609.36059 [cs.CL]

(or
arXiv:2609.36059v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.36059

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
24. 【2609.35904】Structured Interaction, Visual Localization, and Robust Execution for Complex Web Tasks: A Technical Report on the WebRetriever Challenge

链接:https://arxiv.org/abs/2609.35904

作者:Ziqi Zhang,Shaohui Li,Bing Li

类目:Information Retrieval (cs.IR); Multimedia (cs.MM)

关键词:web agent system, agent system developed, report presents, presents the web, web agent

备注: Winning Report for the WebRetriever Challenge

点击查看摘要

Abstract:This report presents the web agent system developed for the WebRetriever Challenge. The system follows a structuredinteraction- first strategy, using semantic webpage information for routine browser operations and invoking visual perception only when structured representations are insufficient. Three key designs are introduced: grid-assisted visual localization for difficult-to-access controls, hierarchical context management for reducing redundant page and interaction history, and fault-aware execution mechanisms for stable multi-browser task processing. The system achieved a pass rate of up to 79% in local evaluation on Protocol 1. In the official Protocol 3 competition, it achieved a 59% pass rate with eight concurrent browser workers and ranked first overall, winning the WebRetriever Challenge.

25. 【2609.35783】Soft Curriculum Learning for Optimizing Fresh and Generalized Recommendations

链接:https://arxiv.org/abs/2609.35783

作者:Arnab Bhadury,Siyan Zheng,Anlan Yu,Palaksh Rungta,Jiawei Li,Changping Meng,Dapeng Hong,Chuan He,Onkar Dalal

类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:short-form video platforms, short-form video, bottlenecked by massive, Soft Curriculum Learning, Large-scale recommender systems

备注: to be published in ACM RecSys '26, Minneapolis, MN, USA

点击查看摘要

Abstract:Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such environments, as models recommend popular items, they generate an overwhelming amount of skewed training data for "head" items. This creates a self-reinforcing cycle where retrieval and ranking models memorize "head" item patterns at the expense of generalizing across the vast "tail" of the catalogue. While Curriculum Learning (CL) offers a powerful mechanism to break this feedback loop by systematically exposing models to progressively more difficult and less frequent examples, its adoption in industrial recommendation has been hampered by hardware utilization inefficiencies or the needs for complicated pre-processing techniques because dynamic data rejection algorithms tend to starve hardware accelearators (TPUs/GPUs) by becoming largely CPU-bound. In this work, we introduce a scalable Soft Curriculum Learning framework designed specifically for continuous training setups within industry-scale retrieval and ranking models. By utilizing loss annealing and in-graph weight adjustments rather than rigid data filtering, we break the popularity feedback and enable dynamic curriculum pacing without sacrificing system throughput. We demonstrate empirical evidence through applications across sequence-based retrieval models (such as SASRec), two-tower retrieval models, and large-scale continuous ranking models. Online A/B tests on our short-video platform demonstrate substantial lifts in both overall user satisfaction and fresh content consumption, all without degrading model throughput.

26. 【2609.35782】Financial Evidence Crowding: Diagnosing and Mitigating Constraint-Induced Displacement in Retrieval-Augmented Generation

链接:https://arxiv.org/abs/2609.35782

作者:Yixi Zhou,Jiayi Yin,Fan Zhang,Xiangyi He,Haipeng Zhang

类目:Information Retrieval (cs.IR)

关键词:Retrieval-augmented generation, limited top-ranked subset, top-ranked subset, evidence, retrieves candidate evidence

备注: 10 pages, 2 figures, 9 tables

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) retrieves candidate evidence and sends only a limited top-ranked subset, the top-k context, to a generator. In financial question answering, passages can match a query's topic while conflicting with its period, segment, metric scope, or table scope. We study the resulting set-level ordering failure, which we call financial evidence crowding. FinDeCrowd-Stress isolates this failure through matched compatible and incompatible candidate pools while fixing the query, relevant evidence, ranking model, candidate count, and retrieval budget. On a FinDER test split containing companies unseen during training, incompatible pools reduce top-10 evidence inclusion (Recall@10) by 0.147 relative to equally difficult compatible pools. This gap shows that conflicting candidates consume limited context slots and displace answer-supporting evidence. We then introduce FinDeCrowd-RAG, a learned score correction that combines a fixed relevance score with typed compatibility and local lexical competition. A query-level identity gate applies the correction only when it predicts a better order; otherwise, it preserves the original ranking. On identical controlled candidates, FinDeCrowd-RAG raises top-10 evidence inclusion from 0.757 to 0.902 by recovering evidence already present in the candidate set. On a FinDER index built without query-specific candidate insertion, gated reranking raises this inclusion rate from 0.420 to 0.492, while top-100 retrieval coverage remains 0.743 by design. With a fixed generator, the same ordering change improves answer accuracy and citation recall on FinanceBench and FinQA. These results identify constraint-induced displacement as a measurable RAG evaluation target and show that identity-gated reranking can recover evidence already covered by first-stage retrieval.

27. 【2609.35780】SG Suggester: Tree-Structured Knowledge-Graph Retrieval for Troubleshooting Guide Recommendation in Cloud Incident Management

链接:https://arxiv.org/abs/2609.35780

作者:Shawn Pan,PavanUttej Ravva,Walt Williams,CJ Barberan,Nutan Sahoo,Ziran Min,David Gross,Irene Shaffer

类目:Information Retrieval (cs.IR)

关键词:keyword driven process, prior empirical work, empirical work finds, correct Troubleshooting Guide, large scale cloud

备注:

点击查看摘要

Abstract:On call engineers in large scale cloud services work under intense time pressure, yet locating the correct Troubleshooting Guide (TSG) for an incident remains a largely manual, keyword driven process, and prior empirical work finds that guide search consumes a substantial fraction of total mitigation time. We present TSG Suggester, a retrieval system that recommends relevant TSGs directly from an incident description. We evaluate five retrieval strategies: Text Only RAG, Image Augmented RAG, RAPTOR, Tree Structured Retrieval, and our proposed Tree + Knowledge Graph (Tree+KG), on 314 real world incidents spanning 112 unique TSGs drawn from 18 service teams on a production incident management system. Tree+KG converts each guide into a tree that preserves its native section hierarchy, attaches LLM generated problem abstractions to internal nodes to bridge the solution oriented language of guides and the problem oriented language of incidents, extracts a per guide entity knowledge graph, and fuses embedding similarity with entity level matching at query time. Tree+KG attains 54.78% Top 1 and 82.48% Top 5 accuracy, leading every baseline at every cutoff, with an 8.58 point Top 1 gain over text only RAG. Two findings are of independent interest. First, structural alignment dominates: methods that preserve or rebuild document structure outperform flat chunking where precise discrimination matters. Second, and contrary to our initial hypothesis, multimodal enrichment actively hurts. Captioning guide screenshots and injecting the captions costs 22.64 Top 5 points relative to the text only baseline because generic captions dilute embeddings rather than sharpen them. We report error analyses for both results and provide concrete deployment guidance.

Subjects:

Information Retrieval (cs.IR)

Cite as:
arXiv:2609.35780 [cs.IR]

(or
arXiv:2609.35780v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2609.35780

Focus to learn more

              arXiv-issued DOI via DataCite

Submission history From: Shawn Pan [view email] [v1]
Tue, 4 Aug 2026 21:47:06 UTC (25 KB)

28. 【2609.35774】Post-Generation Verification Dominates Retrieval Optimization: A 2^4 Factorial Ablation of RAG Pipeline Features

链接:https://arxiv.org/abs/2609.35774

作者:Ng S. T. Chong

类目:Information Retrieval (cs.IR)

关键词:Modern RAG pipelines, Modern RAG, RAG pipelines stack, validated in isolation, stack many enhancement

备注: 11 pages, 11 figures

点击查看摘要

Abstract:Modern RAG pipelines stack many enhancement features, but these features are typically validated in isolation, leaving their interactions unmeasured. We run a 2^4 full factorial ablation of four pipeline features -- section expansion (SE), agentic search (AS), completeness check (CC), and table-of-contents-guided retrieval (ToC) -- across 16 configurations, 24 queries spanning eight interaction types, and two cloud-class models (768 conditions) on five public documents (78-492 pages), scoring every answer against a verified reference. Post-generation verification dominates: CC is the strongest feature (d=+0.48, p0.001), improving accuracy, completeness, and usefulness simultaneously, and CC alone (4.31/5) outperforms every configuration without it, including the three-feature SE+AS+ToC (4.11). ToC yields a significant gain at zero LLM cost (d=+0.22); AS is small and unstable, helping some queries and harming others; SE is neutral. The highest-quality configuration roughly doubles baseline latency, producing a genuine quality-latency Pareto frontier of six configurations. Feature utility is strongly query-type dependent -- CC reaches d=+0.83 on completeness-demanding queries -- so single-query-type evaluations systematically mis-rank features. We conclude that verifying answers matters more than optimizing retrieval, and that factorial designs with diverse query types are necessary to evaluate RAG features.

29. 【2609.35773】Socrates-RAG: Premise-Directed Inquiry against Coordinated Evidence Poisoning

链接:https://arxiv.org/abs/2609.35773

作者:Renyu Zhao,Xinyuan Zou,Lanbin Liu

类目:Information Retrieval (cs.IR)

关键词:defenses typically decide, fixed retrieved set, Retrieval-augmented generation, defenses typically, retrieved set

备注:

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) defenses typically decide how to filter or aggregate a fixed retrieved set. In open-corpus question answering, however, decisive evidence may be absent from the initial context but retrievable, making the next query part of the reliability problem. We introduce Socrates-RAG, a premise-directed active retrieval policy that represents competing answers, selects an unresolved premise whose resolution would discriminate them, and uses newly acquired evidence to refine a subsequent query before answering or abstaining. We formalize the resulting finite-budget evidence state and give a conditional rescue guarantee relative to repeated or topical-query policies. We evaluate Socrates-RAG against a matched control in which the same backbone generates ordinary relevance-oriented search queries; both policies share the initial evidence, deterministic retriever, two-query/top-three budget, answer prompt, and label-free evidence-chain release rule. On a disjoint 48-world counterfactual evaluation, premise-directed inquiry raises safe accuracy from 79.2% to 93.8%, with 8 wins, 1 loss, and 39 ties (two-sided exact p=.0391). Decisive-evidence recall improves by the same margin, while unsafe answers fall from one to zero. Both policies solve all 24 one-hop cases; the gain is concentrated in two-hop cases, where Socrates-RAG substitutes a newly resolved premise into its second query. This controlled study isolates a specific benefit of premise-directed acquisition without claiming general robustness on the open Web.

Subjects:

Information Retrieval (cs.IR)

Cite as:
arXiv:2609.35773 [cs.IR]

(or
arXiv:2609.35773v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2609.35773

Focus to learn more

              arXiv-issued DOI via DataCite</p>

计算机视觉

1. 【2609.38180】Point2Part: Unified 3D Partitioning from Point Prompts

链接:https://arxiv.org/abs/2609.38180

作者:Hao-Tang Tsui,Yu-Rou Tuan,Xiaoxuan Ma,Nicolas Ugrinovic,Takaaki Shiratori,Kris Kitani

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:downstream part-level applications, hinder downstream part-level, entire shape, part, part decomposition

备注: Project Page: [this https URL](https://henrytsui000.github.io/Point2Part)

点击查看摘要

Abstract:Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part decomposition as a joint partitioning of the entire shape, where the predicted parts are non-overlapping and jointly recover the entire shape. Our key insight is that part decomposition should consider all desired parts jointly, rather than modeling each part independently. To this end, we develop a promptable model for 3D part decomposition from images or meshes. Users can specify desired parts through 3D point prompts for controllable decomposition. Given one point prompt per desired part, our model produces the corresponding parts as a complete partition of the entire shape. We build on a pretrained 3D generation model and first obtain a shape latent from either an input image or mesh. We then introduce a prompt encoder that maps each 3D point prompt to a part token while attending to the shape latent. To decode the desired parts, we propose a novel part decoder jointly scoring the entire shape against all part tokens in a coarse-to-fine manner, assigning every position within the shape volume to exactly one part. We perform part decomposition in this shared shape latent space, enabling a unified model for image-to-part generation, mesh-to-part generation, and part segmentation. Our method outperforms existing works on all part-quality metrics across all three tasks, and improves compatibility among parts by an order of magnitude over previous SOTA methods. Code and models will be released.

2. 【2609.38177】Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

链接:https://arxiv.org/abs/2609.38177

作者:Jaewoo Jung,Hyeonseo Yu,Honggyu An,Jisang Han,Mungyeom Kim,Minkyeong Jeon,Heeseong Shin,Wonjun Moon,Federico Tombari,Daniel Barath,Marc Pollefeys,Seungryong Kim,Sunghwan Hong

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, challenge for Multimodal

备注: NeurIPS 2026; Project Page: [this https URL](https://cvlab-kaist.github.io/Imagine3D-LLM)

点击查看摘要

Abstract:Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

3. 【2609.38172】Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

链接:https://arxiv.org/abs/2609.38172

作者:Zihan Wang,Zhen Wu,Pieter Abbeel,Rocky Duan,Jitendra Malik,Carmelo Sferrazza,C. Karen Liu,Guanya Shi,Angjoo Kanazawa

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Teaching humanoids loco-manipulation, humanoids loco-manipulation skills, loco-manipulation skills, Teaching humanoids, visual imitation

备注: published at CoRL 2026. Project page: [this https URL](https://prism-real2sim2real.github.io/)

点击查看摘要

Abstract:Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

4. 【2609.38170】Adversarial Training for Pixel Diffusion

链接:https://arxiv.org/abs/2609.38170

作者:Xin Lin,Zhifei Zhang,Yuqian Zhou,Haitian Zheng,Zhe Lin,Ming-Hsuan Yang,Truong Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models generate RGB, RGB images directly, generate RGB images, generate RGB, systematically underrepresent fine-scale

备注:

点击查看摘要

Abstract:Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.

5. 【2609.38165】Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

链接:https://arxiv.org/abs/2609.38165

作者:Joseph Metcalfe,Sara Sharifzadeh,Fabio Caraffini

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:satellite imagery time, imagery time series, time series datasets, Cropland Parallel Attention, landscape of satellite

备注: Main body: 19 pages, 7 figures; Appendices: 15 pages, 16 figures. All code and models associated with this work are available at [this https URL](https://github.com/JoeMetc/CroplandPAtteRNS) , along with preparation guides for the two publicly available crop segmentation datasets used in this work

点击查看摘要

Abstract:The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.

6. 【2609.38163】Rethinking Representations for World-Action Modeling

链接:https://arxiv.org/abs/2609.38163

作者:Haoyi Jiang,Liu Liu,Xinjiang Wang,Zhihao Sun,Zequn Chen,Sen Wang,Xinjie Wang,Xia Chen,Jingfeng Yao,Weiheng Zhao,Shanglin Yuan,Zhizhong Su,Wei Sui,Wenyu Liu,Xinggang Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:predict future observations, jointly learn robot, learn robot policies, World-action models jointly, future observations

备注: [this https URL](https://github.com/hustvl/ReWAM)

点击查看摘要

Abstract:World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.

7. 【2609.38156】DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses

链接:https://arxiv.org/abs/2609.38156

作者:Xin Lin,Zhifei Zhang,Yuqian Zhou,Haitian Zheng,Shaoteng Liu,Lehan Yang,Zhe Lin,Ming-Hsuan Yang,Truong Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Distribution matching distillation, few-step diffusion generation, diffusion generation, Distribution matching, latent diffusion

备注:

点击查看摘要

Abstract:Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA$^2$. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA$^2$ student performs better than the 25-step teacher and evaluated few-step distillers.

8. 【2609.38155】Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

链接:https://arxiv.org/abs/2609.38155

作者:Hui Ren,Lei Fan,Henry Pao,Han Guo,Zeeshan Zia,Ying Chen,Alexander Schwing,Gang Hua

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:requires connecting events, connecting events involving, hours or days, long videos, videos often requires

备注:

点击查看摘要

Abstract:Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

9. 【2609.38154】LongLive-Plug: Once-for-All Distillation for Video Generation

链接:https://arxiv.org/abs/2609.38154

作者:Shuai Yang,Luozhou Wang,Wei Huang,ZhiFei Chen,Bohan Zhang,Xiao Fu,Qianli Ma,Chen-Hsuan Lin,Weian Mao,Bryan Chu,Song Han,Yukang Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video diffusion models, Video diffusion, improve long-video generation, increasingly developed, improve long-video

备注: Code and models are available at [this https URL](https://github.com/NVlabs/LongLive)

点击查看摘要

Abstract:Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.

10. 【2609.38153】PowerSim: Differentiable Physics Simulation and Rendering with Power Diagrams

链接:https://arxiv.org/abs/2609.38153

作者:Trong-Tung Nguyen,Anand Bhattad

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:power diagram based, PowerFoam power diagram, Material Point Method, diagram based, power diagram

备注:

点击查看摘要

Abstract:We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam's power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point Method (MPM) by exploiting a natural alignment between the two: the geometric and appearance properties of each primitive correspond closely to the quantities MPM already tracks as an object deforms. Consequently, simulated motion can drive the scene's geometry and appearance directly, without an auxiliary representation in between. Built on this framework, we enable a range of applications on real and synthetic scenes: (1) simulating a static scene under user interaction, (2) recovering spatially varying material fields, (3) compositing primitives from independently captured scenes into a single simulation-ready scene and (4) ray-tracing reflections that update consistently as the object deforms. Our results suggest that PowerSim excels over previous frameworks for physically grounded dynamics, while unlocking unique advantages-such as secondary ray lighting effects on dynamic scenes. Results are best viewed on our project website: this https URL.

11. 【2609.38152】FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation

链接:https://arxiv.org/abs/2609.38152

作者:Trong-Tung Nguyen,Jiahan Zhang,Anand Bhattad

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:controllable fracture dynamics, produces plausible, conditioned on physics, physics signals, single image

备注:

点击查看摘要

Abstract:We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation framework that augments material point method (MPM) simulation with a continuum damage model, producing paired fracture videos and dense, pixel-aligned physical fields at no additional cost beyond standard rendering. FracGen leverages these maps in two ways: it is trained to jointly predict them alongside RGB video, encouraging the model to capture physical state rather than surface appearance; and it is supervised with physics-informed losses that encourage consistency among the predicted maps. As a result, FracGen captures distinct material-specific fracture behavior without expensive test-time simulation or per-scene tuning, while offering fine-grained control over where an object tears, how fast the crack propagates, and how much deformation precedes failure. We further introduce a benchmark for evaluating the physical plausibility of generated fracture video, and show through extensive experiments that FracGen outperforms existing video generation baselines in both physical and visual fidelity. Results are best viewed in our project website: this https URL.

12. 【2609.38146】LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

链接:https://arxiv.org/abs/2609.38146

作者:Shengxiang Ji,Boyang Wang,Haiyang Xu,Bingnan Li,Yucheng Mao,Zeyuan Chen,Xiaojun Shan,Xiang Zhang,Gang Hua,Jianwen Xie,Zezhou Cheng,Zhuowen Tu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:framework that complements, enabling users, complements camera control, generation framework, camera

备注: Project Page: [this https URL](https://jsxzs.github.io/LIFT/)

点击查看摘要

Abstract:We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.

13. 【2609.38140】Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

链接:https://arxiv.org/abs/2609.38140

作者:Yu Xu,Yuxin Zhang,Xiao Yang,Haotian Yang,Yizhi Wang,Xinwei Huang,Minxuan Lin,Angtian Wang,Chongyang Ma,Fan Tang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:large language models, popularized by large, large language, promising paradigm, experts

备注: Accepted as a Spotlight paper at NeurIPS 2026. Project page: [this https URL](https://yuci-gpt.github.io/SplitMoE/)

点击查看摘要

Abstract:Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

14. 【2609.38136】CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

链接:https://arxiv.org/abs/2609.38136

作者:Teng Zhou,Yunhao Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Style, content leakage, render target content, generated output, aims to render

备注:

点击查看摘要

Abstract:Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{this https URL}{this https URL}.

15. 【2609.38123】HelixWorld: A Real-time Interactive Audio-Visual World Model

链接:https://arxiv.org/abs/2609.38123

作者:Lei Ke,Jiahao Pan,Zeyue Tian,Jiaming Wang,Haoyuan Huang,Kam Man Wu,Pengjun Fang,Hongyu Liu,Chenyang Qi,Lin Wang,Ruibin Yuan,Weijia Chen,Fangneng Zhan,Qifeng Chen,Wei Xue,Yike Guo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demanding synchronized visual, inherently multisensory, demanding synchronized, real time, simulation is inherently

备注:

点击查看摘要

Abstract:World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.

16. 【2609.38119】VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

链接:https://arxiv.org/abs/2609.38119

作者:Jinfa Huang,Jianming Xu,Jingyang Lin,Zhengyuan Yang,Jiebo Luo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Long-form video understanding, iteratively gather evidence, understanding requires multimodal, video understanding requires, Long-form video

备注: Code: [this https URL](https://github.com/philipxjm/videoloop)

点击查看摘要

Abstract:Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).

17. 【2609.38116】GA-EIRFS: A Geometry-Augmented Repeat-Factor Sampling Method for Long-Tailed LiDAR 3D Object Detection

链接:https://arxiv.org/abs/2609.38116

作者:Taufiq Ahmed,Constantino Álvarez Casado,Daniel Herrera Castro,Sasan Sharifipour,Abhishek Kumar,Miguel Bordallo López

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:LiDAR supervision quality, supervision quality depends, Repeat Factor Sampling, object observability, Geometry-Augmented Exponentially Weighted

备注: 5 pages, 4 figures, Submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Long-tailed 3D object detection is treated as a class-frequency problem, but LiDAR supervision quality depends on object observability: similar frequencies can hide different geometric evidence. We introduce Geometry-Augmented Exponentially Weighted Instance-Aware Repeat Factor Sampling (GA-EIRFS), a detector-agnostic method that modulates a frequency-based repeat factor with a fixed geometry score combining point count, surface-normal entropy, and surface coverage. GA-EIRFS changes only frame-sampling probabilities, leaving the detector and inference unchanged. On nuScenes it improves mean average precision (mAP) and the nuScenes detection score (NDS) in four converged experiments with CenterPoint and PointPillars over two seeds; for CenterPoint at seed 666, mAP rises from 0.552 to 0.563 and bicycle AP from 0.306 to 0.359. Per-class gains correlate with the class sampling-weight increase (Spearman rho=0.70, p=0.025) but not with geometry score alone (rho=0.32, p=0.37), so geometry amplifies frequency-driven need. KITTI results vary across seeds, most for the rarest class. Code: this https URL.

18. 【2609.38114】Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

链接:https://arxiv.org/abs/2609.38114

作者:Weiqiang Wang,Zhuokun Chen,Yusheng Dai,Boying Li,Yi Zhang,Hossein Rahmani,Qiuhong Ke,Jianfei Cai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Autoregressive video diffusion, video diffusion enables, diffusion enables interactive, enables interactive streaming, Autoregressive video

备注:

点击查看摘要

Abstract:Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: this https URL.

19. 【2609.38111】From Routing Signals to Selective Review: Visual regrounding in MoE VLMs

链接:https://arxiv.org/abs/2609.38111

作者:Hongzhu Guo,Mohsen Fayyaz,Nanyun Peng

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:target object color, false visual premises, accept false visual, answering questions, object color

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.

20. 【2609.38086】VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

链接:https://arxiv.org/abs/2609.38086

作者:Zheng Jiang,Houde Qian,Yiming Chen,Ling Li,Chaoyang Li,Yueqi Li,Yuxuan Liu,Lifeng Sun

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:acquire task-relevant evidence, acquire task-relevant, task-relevant evidence, Active multimodal, Active multimodal agents

备注:

点击查看摘要

Abstract:Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.

21. 【2609.38079】OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

链接:https://arxiv.org/abs/2609.38079

作者:Jiaxin Ge,Yiming Qin,Ji Xie,Haozhe Jiang,Xiaochuang Han,Junyi Zhang,Andrew Dai,Yinfei Yang,Jitendra Malik,Ranjay Krishna,Sewon Min,Haiwen Feng,Le Xue,Baifeng Shi,Trevor Darrell,XuDong Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generate visual content, understanding remain unclear, perceptual capabilities related, visual understanding remain, spatial relationships

备注:

点击查看摘要

Abstract:Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: this https URL.

22. 【2609.38077】MUGEN: Interactive Panoramic World Exploration via Camera Control

链接:https://arxiv.org/abs/2609.38077

作者:Jiaming Tan,Zhen Li,Shuwei Shi,Minggui Teng,Siqi Yang,Yuwei Wu,Bo Zheng,Chuanhao Li,Kaipeng Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:user-specified camera trajectories, video generation, video generation aims, remain visually coherent, panoramic video generation

备注: Project page: [this https URL](https://alaya-lab.github.io/MUGEN)

点击查看摘要

Abstract:Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Plücker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.

23. 【2609.38072】RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

链接:https://arxiv.org/abs/2609.38072

作者:Chengjie Jiang,Yunqi Zhou,Jiafeng Yan,Sihang Zhao,Chun Yuan,Jing Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:extremely large images, visual question answering, resolve small visual, question answering, large images

备注: 16 pages, 7 figures

点击查看摘要

Abstract:Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming pervious SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.

24. 【2609.38059】WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

链接:https://arxiv.org/abs/2609.38059

作者:Shenghe Zheng,Wenbo Li,Jiyao Zhang,Bin Xia,Haoyang Huang,Nan Duan,Jiaya Jia

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:evaluating candidate behaviors, Real-world robot learning, Real-world robot, candidate behaviors, cost of collecting

备注: A work about visual simulators for embodied AI

点击查看摘要

Abstract:Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{this https URL}{project page}.

25. 【2609.38057】EVO-WAM: Evolving World Action Models through Video-Action Verification

链接:https://arxiv.org/abs/2609.38057

作者:Shiyang Zhou,Xionghao Wu,Wenbo Li,Shenghe Zheng,Jiyao Zhang,Songsong Yu,Yijun Yang,Jianhui Liu,Haoze Sun,Senqiao Yang,Li Jiang,Jingyong Su,Haoyang Huang,Zhuotao Tian

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Improving robot policies, collecting additional expert, additional expert demonstrations, expert demonstrations remains, Improving robot

备注:

点击查看摘要

Abstract:Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: this https URL.

26. 【2609.38054】Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors

链接:https://arxiv.org/abs/2609.38054

作者:Christopher Kolios,Ishaan Mehta,Sasa Janjic,Yeganeh Bahoo,Sajad Saeedi

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:real-time RGB-D simultaneous, RGB-D simultaneous localization, RGB-D SLAM systems, traditional RGB-D SLAM, simultaneous localization

备注: 9 pages, 4 figures, 4 tables. Project page: [this https URL](https://chriskolios.github.io/Pow3R-SLAM/)

点击查看摘要

Abstract:We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network's prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6x faster in wall time, has 15% lower mean trajectory error, a 3.1x lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1x faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. A project webpage is available at: this https URL , and code will be made open-source upon acceptance.

27. 【2609.38028】doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

链接:https://arxiv.org/abs/2609.38028

作者:Parthib Roy,Yash Tandon,Marcus Blennemann,Giovanni Tapia Lopez,Angel Martinez-Sanchez,Mohan M. Trivedi,Ross Greer

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Autonomous vehicles interacting, Autonomous vehicles, vehicles interacting, Passenger intent, Passenger

备注:

点击查看摘要

Abstract:Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at this https URL. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.

28. 【2609.38019】Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation

链接:https://arxiv.org/abs/2609.38019

作者:Bangxun Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reference-Grounded Oral Refinement, Oral Refinement, audio-driven lip-sync framework, Reference-Grounded Oral, mouth

备注: 19 pages, 8 figures, 5 tables. Under review

点击查看摘要

Abstract:We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person's reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person's own lip and dental detail while keeping synchronization and the rest of the face intact.

29. 【2609.38016】Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy

链接:https://arxiv.org/abs/2609.38016

作者:Huan Rong,Chao Yin,Anouar Imel,Yijie Xia,Tinghuai Ma

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Robotics (cs.RO)

关键词:Constrained Reinforcement Learning, Reinforcement Learning, recently gained increasing, gained increasing attention, action risk bounded

备注: 18 pages, 11 figures

点击查看摘要

Abstract:Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues arising in AD can be mitigated through constrained actions. However, existing Constrained RL methods still lack dynamics on the imposed constraints. For instance, the action cost adopted by the existing Primal-Dual/soft-constrained methods is often defined as static state-to-cost mapping, and the safe-action projection in hard-constrained methods relies on the static projection with the fixed feasible region boundary estimated from offline demonstrations. The above drawback tightly couples the imposed constraints to the training scenarios, leaving the AD policy hard to handle different interaction scenarios, due to the improper state-level action-cost and the static projection boundary. Consequently, in this paper, we propose Brain-SAD, a brain-inspired safe autonomous driving control framework with dynamic fear-oriented constraints. By perceiving the current vehicle-interaction scene, Brain-SAD generates dynamic fear signal as fear reaction to online decide long-term policy for regular interaction or short-term policy for urgent-collision defense. In such two policy, the above fear-reaction will be constructed as the dynamic fear constraints, respectively reflecting the overall fear cost directly coupled with action-impact, and the dynamic fear boundary of the feasible region derived from different risky neighbors, both of which will in turn serve for the online policy optimization. Experimental results show that Brain-SAD outperforms existing methods, achieving higher success rate in shorter task-completion and collision-recovery time, and exhibits stronger reliability across continuous intersections of fluctuating complexity.

30. 【2609.38010】From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection

链接:https://arxiv.org/abs/2609.38010

作者:Mohamed Benkedadra,Aissa Saoudi,Maxime Gloesener,Sidi Ahmed Mahmoudi,Matei Mancas

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Modern computer vision, Modern computer, computer vision models, large-scale annotated datasets, achieve high accuracy

备注:

点击查看摘要

Abstract:Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom $\Delta$-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of $-50\%$ relative to real data, while CIA-only training shows a milder $-16.5\%$ degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall mAP@0.5 of $62.68\%$ ($+7.64\%$ over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at $74.45\%$. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.

31. 【2609.38008】HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

链接:https://arxiv.org/abs/2609.38008

作者:Tongbo Chen,Junbo Niu,Zhengxi Lu,Niu Lian,Fei Tang,Yuchen Yan,Yike Hong,Yong Du,Yizhou Liu,Bofan Chen,Yongliang Shen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrated strong capabilities, completing digital tasks, GUI, CLI, GUI interactions

备注: Project Page: [this https URL](https://zjureal.com/HybridCUA/) Code: [this https URL](https://github.com/ZJU-REAL/HybridCUA)

点击查看摘要

Abstract:Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.

32. 【2609.37986】ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals

链接:https://arxiv.org/abs/2609.37986

作者:Xuyi Hu,Francesco Palandra,Shangzhe Wu,Daniel Cremers,Riccardo Marin,Silvia Zuffi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:monocular videos remains, remains challenging due, supervision data, large diversity, morphologies and lack

备注:

点击查看摘要

Abstract:Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithfully represent the observed instance. We present ORMA, a training-free reconstruction framework that decouples articulation from shape, using the predicted pose as reference for optimization while leveraging generative 3D priors for accurate shape reconstruction. Given a reference image, we reconstruct the animal geometry and register it to the parametric model SMAL+, yielding an articulated shape adapted to the observed instance. We then combine per-frame articulated pose estimates with globally consistent camera poses to recover animal motion in a shared world coordinate frame, and further refine the reconstruction using self-supervised DINO correspondences and temporal consistency. To enable quantitative evaluation, we introduce PAW4D, a synthetic multi-species benchmark with ground-truth 3D geometry and camera motion. Experiments on PAW4D, PFERD, and challenging in-the-wild videos demonstrate that ORMA improves reconstruction accuracy while recovering globally consistend animal motion across diverse quadruped species.

33. 【2609.37976】$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient

链接:https://arxiv.org/abs/2609.37976

作者:Hongbo Ma,Sansheng Cao,Jiajun Fan,Bangji Yang,Ge Liu

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:excessive token cost, LLMs trained, Non-thinking model dominant, Thinking model weight, Non-thinking model

备注: 44 pages, 9 figures, 29 tables

点击查看摘要

Abstract:LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.

34. 【2609.37970】PhysWAM: Physically Consistent World Action Model for Autonomous Driving

链接:https://arxiv.org/abs/2609.37970

作者:Dhruv Parikh,Fengcheng Yu,Quankai Gao,Jiawei Yang,Junjie Ye,Maulik Bhatt,Thang Vu,Charles Ochoa,Rowan McAllister,Igor Vasiljevic,Rajgopal Kannan,Viktor Prasanna,Vitor Guizilini,Yue Wang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:ego motion, Coupled Point Projection, unified world-action model, agent should act, necessarily impose

备注: Technical Report

点击查看摘要

Abstract:World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

35. 【2609.37969】SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

链接:https://arxiv.org/abs/2609.37969

作者:Haozhe Liu,Tian Ye,Shuchen Xue,Yitong Li,Junsong Chen,Haopeng Li,Jincheng Yu,Duomin Wang,Ruihua Zhang,Lei Zhu,Song Han,Enze Xie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cost grows rapidly, High-resolution video generation, generation is expensive, spatiotemporal tokens, cost grows

备注: 15 pages

点击查看摘要

Abstract:High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at $3840\!\times\!2176$ it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an $8.91\times$ speedup in refinement latency over the same baseline in our 2K latency setting.

36. 【2609.37950】Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution

链接:https://arxiv.org/abs/2609.37950

作者:Bingjun Luo,Jialin Guo,Siqi Li

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Video understanding, Video understanding agents, understanding agents acquire, understanding agents, understanding

备注:

点击查看摘要

Abstract:Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement. We introduce Video-RSI, a framework for recursive self-improvement in which a video understanding agent uses its own language model to revise its harness. Through active video investigation, the model revisits the original training videos to test competing failure explanations with additional observations, grounding proposed changes in evidence beyond the existing trace. Cost-aware harness evolution turns these diagnoses into reusable revisions and determines which revisions to retain by considering both answer accuracy and visual cost. Across our evaluation settings on video understanding benchmarks, the evolved agent improves accuracy while processing fewer frames and achieves competitive accuracy-efficiency trade-offs against existing video understanding agents. These results demonstrate the potential for video understanding agents to improve their own evidence acquisition and use through harness evolution. Code is available at this https URL .

37. 【2609.37938】Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

链接:https://arxiv.org/abs/2609.37938

作者:Yuedong Tan,Lei Qi,Yu Liu,Di Wen,Ruiping Liu,Xiaoye Wang,Yufan Chen,Junwei Zheng,Chengzhi Wu,Chen Zhang,Zhihang Chen,Haiwen Sun,Zongwei Wu,Radu Timofte,Danda Pani Paudel,Kunyu Peng

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:make knowledge acquired, Embodied systems, systems must make, make knowledge, knowledge acquired

备注:

点击查看摘要

Abstract:Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at this https URL.

38. 【2609.37937】Look Closer: Patch-wise Supervision for AI-Generated Image Detection

链接:https://arxiv.org/abs/2609.37937

作者:Zhida Zhang,Tao Wu,Siyu Liu,Jie Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Small RGB regions, Small RGB, RGB regions, Abstract, image

备注: 29 pages, 11 figures, 28 tables. Code: [this https URL](https://github.com/LF-Jade/look-closer)

点击查看摘要

Abstract:How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.

39. 【2609.37925】Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

链接:https://arxiv.org/abs/2609.37925

作者:Chenjian Gao,Zhihao Hu,Jianqi Ma,Jun Zhang,Weidong Zhang,Tianfan Xue

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:diffusion enables low-latency, streamable video generation, enables low-latency, video diffusion enables, Autoregressive

备注:

点击查看摘要

Abstract:Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at this https URL

40. 【2609.37923】EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory

链接:https://arxiv.org/abs/2609.37923

作者:Ziyun Zeng,Hang Hua,Shaden Alshammari,Rogerio Feris,William T. Freeman,Jiebo Luo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:experience remains challenging, past executions, remains challenging, learn from past, reuse and build

备注: Preprint

点击查看摘要

Abstract:Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. A second harness raises the original system's macro-average score by 2.6 points across eleven benchmarks. Across four host configurations, EpiCon improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67\% to 74\% relative to backbone-sized memory models.

41. 【2609.37918】SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation

链接:https://arxiv.org/abs/2609.37918

作者:Sara Ghazanfari,Siddharth Garg,Prashanth Krishnamurthy,Farshad Khorrami

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:integrating partial observations, requires aligning events, videos requires aligning, matching identities, comparing motion

备注:

点击查看摘要

Abstract:Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and human evaluation assess dependence on the supplied evidence and answer recoverability. Evaluation of 22 multimodal large language models reveals persistent difficulties in physical comparison and scene integration that increasing model size does not consistently resolve. Supervised fine-tuning raises Qwen3-VL-8B's average SYNCR accuracy from 32.6% to 61.6%, with gains extending to task configurations and video sources absent from training for those tasks. Transfer to real footage is most consistent for temporal ordering: accuracy improves by 9.0-20.5 percentage points on constructed Assembly101 and Panoptic ordering sets across three checkpoints spanning two model families and two model sizes, with additional gains on existing temporal reasoning benchmarks. These results establish SYNCR as a controlled setting for diagnosing cross-video reasoning failures, testing their learnability, and identifying where synthetic supervision transfers.

42. 【2609.37907】Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics

链接:https://arxiv.org/abs/2609.37907

作者:Abhishek Pillai,Ekta Prashnani,Joohwan Kim,Iuri Frosio

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:http URL online, URL online gameplay, online gameplay videos, offer scalable environments, rarely include player

备注: Accepted at the Workshop on Multimodal Digital Agents (ECCV 2026): [this https URL](https://mda-workshop.allen.ai/)

点击查看摘要

Abstract:Video games offer scalable environments for studying perception and control in embodied this http URL online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as $F_1^{macro}$. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.

43. 【2609.37889】ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning

链接:https://arxiv.org/abs/2609.37889

作者:Tao Hu,Zhinuo Zhou,Xialiang Tong,De-Chuan Zhan,Da-Wei Zhou

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:multimodal large language, large language models, enable multimodal large, multimodal large, Multimodal continual instruction

备注:

点击查看摘要

Abstract:Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.

44. 【2609.37888】Visual Branch is What You Need for CLIP-based Class-Incremental Learning

链接:https://arxiv.org/abs/2609.37888

作者:Tao Hu,Zhen-Hao Xie,Jingcai Guo,De-Chuan Zhan,Da-Wei zhou

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Class-Incremental Learning, forgetting previously learned, CIL, requires models, models to recognize

备注:

点击查看摘要

Abstract:Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual this http URL by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VISuses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VISemploys a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VISaccumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VISachieves state-of-the-art performance without a textual branch.

45. 【2609.37874】EndoPrior-GS: Dynamic Endoscopic Reconstruction with a Joint Texture Prior

链接:https://arxiv.org/abs/2609.37874

作者:Jiaqi Huang,Shidong Wang,Tong Xin,Kabita Adhikari

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Dynamic endoscopic reconstruction, Dynamic endoscopic, computer-assisted interventions, endoscopic reconstruction, reconstruction is fundamental

备注: Accepted at ACCV 2026. Code: [this https URL](https://github.com/jiaqi-huang-77/EndoPrior-GS)

点击查看摘要

Abstract:Dynamic endoscopic reconstruction is fundamental to robotic surgery and computer-assisted interventions. While 3D Gaussian Splatting (3DGS) realises real-time rendering, its application to deformable intraoperative environments remains constrained by spurious geometry and varying illuminations. To address these limitations, we introduce EndoPrior-GS, a novel pipeline that explicitly couples frame-extracted vision heuristics and estimated depth maps. EndoPrior-GS derives a joint texture prior from a tool-filtered valid tissue mask, a non-specular photometric filter, and anatomical structural salience, yielding a probability map that guides primitive initialisation and subsequent density control. The prior is further extended to the temporal domain through a texture-aware term that dynamically weighs pairwise primitive contributions during training. We conduct extensive experiments on benchmark datasets EndoNeRF and SCARED, and the obtained results show that our method EndoPrior-GS reduces Flow Error by 27.7% and 25.8% over the representative approaches while preserving competitive rendering quality and real-time rendering speed. Our project website is available at this https URL.

46. 【2609.37870】Learning from synthetic photorealistic raindrop for single image raindrop removal

链接:https://arxiv.org/abs/2609.37870

作者:Zhixiang Hao,Shaodi You,Yu Li,Kunming Li,Feng Lu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:computer vision systems, adhered to camera, camera lens, lens or windshield, windshield are inevitable

备注: 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW)

点击查看摘要

Abstract:Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomous driving. Because raindrop appearance is affected by too many parameters, therefore it is unlikely to find an effective model based solution. Learning based methods are also problematic, because traditional learning method cannot properly model the complex appearance. Whereas deep learning method lacks sufficiently large and realistic training data. To solve it, in our work, we propose the first photo-realistic dataset of synthetic adherent raindrops for training. The rendering is physics based with consideration of the water dynamic, geometric and photometry. The dataset contains various types of rainy scenes and particularly the rainy driving scenes. Based on the modeling of raindrop imagery, we introduce a detection network which has the awareness of the raindrop refraction as well as its blurring. Based on that, we propose the removal network that can well recover the image structure. Rigorous experiments demonstrate the state-of-the-art performance of our proposed framework.

47. 【2609.37863】It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

链接:https://arxiv.org/abs/2609.37863

作者:Nagham Omar,Mahmoud Jabarin,Kinan Ibraheem,Lotem Peled-Cohen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:incidental evaluation conditions, substitutability tests reflect, Misleading-Image Stress Test, Vision-language models, tests reflect

备注: Accepted at TAE (Trust-AI-Eval) @ NeurIPS 2026

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.

48. 【2609.37855】HandAnthro: Automated Hand Anthropometry from a Single Image

链接:https://arxiv.org/abs/2609.37855

作者:Fan Zhou,Shuairan Chen,Mengying Zhang,Yulin Wu,Sadegh Jafari,Sixing Yu,Rui Li,Ali Jannesari,Guowen Song

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:supports protective-glove design, anthropometry supports protective-glove, Hand anthropometry supports, require trained operators, existing measurement methods

备注: 21 pages, including 7 pages of main text and references and 14 pages of supplementary material

点击查看摘要

Abstract:Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific landmarks from a fine-tuned You Only Look Once (YOLO) pose model using image-specific geometry and contours. Controlled evaluation comprised 720 captures from 45 held-out participants, each contributing 16 images across two smartphones, two backgrounds, two angles, and two nominal illumination settings. HandAnthro produced complete outputs for 704 captures (97.8%); among these, mean absolute error (MAE) was 3.80 mm per dimension against two trained operators' caliper measurements. Regional MAEs were 2.48 mm for non-thumb fingers, 6.04 mm for thumbs, and 6.17 mm for palm and wrist. In a researcher-assisted mobile-app pilot, automated batch processing returned all 44 dimensions for 260 of 268 retained, researcher-screened firefighter images (97.0%). A descriptive, unpaired comparison with an independent national firefighter reference yielded a mean absolute difference of 2.40 mm across 28 sex-by-dimension group-mean contrasts. These results characterize controlled measurement performance and researcher-assisted field feasibility for future distributed hand-anthropometry studies.

49. 【2609.37851】FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators

链接:https://arxiv.org/abs/2609.37851

作者:Zhiqi Li,Bo Zhu

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Few-step flow-map generators, enable efficient sampling, distillation remains underexplored, on-policy distillation remains, including MeanFlow

备注: 38 pages, 18 figures

点击查看摘要

Abstract:Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.

50. 【2609.37850】RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

链接:https://arxiv.org/abs/2609.37850

作者:Xijun Wang,Xin Li,Zirui Lang,Suhang Yao,Haoran Li,Zhibo Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recover realistic detail, lightweight VSR network, Sparse Generative Relay, large generative model, real-world video super-resolution

备注: The code is available at [this https URL](https://github.com/kopperx/RelayVSR)

点击查看摘要

Abstract:Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at this https URL.

51. 【2609.37848】Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study

链接:https://arxiv.org/abs/2609.37848

作者:Bhanu Prakash Vangala,Sowmya Guda,Latha Peddi,Navya Vangala

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Biomedical machine learning, machine learning papers, compress model performance, Biomedical machine, machine learning

备注:

点击查看摘要

Abstract:Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced accuracy by 0.090 on average. The official test split is also measurably different from the training pool: a partition classifier distinguishes them at AUC 0.697, rising to 0.898 for normal radiographs. Most strikingly, a classifier using only file properties, with no image anatomy, reaches 0.992 balanced accuracy within the training pool but falls to 0.496 on the official test split. Validation-fitted thresholds and calibration also transfer imperfectly. These results show that a high benchmark score can support different conclusions when the split, training policy, threshold, metric, calibration, and uncertainty are not communicated with it. We end with a seven-item reporting recommendation in which each item is tied to an effect measured in the study

52. 【2609.37831】ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing

链接:https://arxiv.org/abs/2609.37831

作者:Xijun Wang,Xin Li,Suhang Yao,Zirui Lang,Bingchen Li,Zhibo Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Real-time diffusion-based video, diffusion-based video super-resolution, compromise generative fidelity, stringent latency requirements, Real-time diffusion-based

备注: The code is available at [this https URL](https://github.com/kopperx/ReCaVSR)

点击查看摘要

Abstract:Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At $1080{\times}1920$ output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72$\times$ faster while using 38.0\% less peak allocated memory than FlashVSR Tiny. The code is available at this https URL.

53. 【2609.37817】Minkowski Attractor Networks: Closed-Form Hyperbolic Flows for Visual Representations

链接:https://arxiv.org/abs/2609.37817

作者:Zhongping Ji

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:flat Euclidean subspaces, Geometric representation learning, compact product tori, learning predominantly scaffolds, Euclidean subspaces

备注: 15 pages

点击查看摘要

Abstract:Geometric representation learning predominantly scaffolds representations onto flat Euclidean subspaces or compact product tori ($\mathbb{T}^K$). However, flat manifolds possess vanishing curvature and polynomial volume growth, inherently suffering from metric distortion when embedding multi-scale, tree-like visual hierarchies. While hyperbolic spaces ($\mathbb{H}^m$) circumvent this via constant negative curvature ($K0$) and exponential volume expansion, prior hyperbolic deep architectures are hindered by computationally cumbersome Riemannian optimization, non-linear gyrovector calculus, and floating-point instabilities. In this work, we introduce \textbf{Minkowski Attractor Networks (MAN)}, an operator-splitting-inspired framework that embeds representations within pseudo-Riemannian Minkowski spacetime ($\mathbb{R}^{1,m}$). By framing hyperbolic manifolds as quadric level sets, MAN resolves hyperbolic geometry by combining linear Lorentz group transport with non-linear cone lifting and closed-form radial rescaling, evaluating in a single forward pass without numerical ODE solvers or iterative retractions. We establish \textbf{MAN-2D} ($\mathbb{R}^{1,1} \to \mathbb{H}^1$) as our primary, high-throughput visual backbone, which maximizes channel factorization granularity into $D/2$ independent two-dimensional Minkowski blocks. We further formulate \textbf{MAN-4D} ($\mathbb{R}^{1,3} \to \mathbb{H}^3$) as a spacetime extension, leveraging a commuting Cartan-subalgebra parameterization of $\mathrm{SO}^+(1,3)$ to evaluate 4D Lorentz isometries via two commuting 2D planar maps without matrix-exponential overhead.

Comments:
15 pages

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.37817 [cs.CV]

(or
arXiv:2609.37817v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.37817

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
54. 【2609.37816】WINGS: Reference-Free Gaussian Splatting Inpainting with 3D-Native Generative Priors

链接:https://arxiv.org/abs/2609.37816

作者:Noé Lallouet,Michael Fischer,Elie Michel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires generating plausible, Gaussian splatting inpainting, Gaussian Splatting scenes, Gaussian Splatting, splatting inpainting method

备注: Preprint. Under review

点击查看摘要

Abstract:Inpainting 3D Gaussian Splatting scenes, a key challenge in 3D editing, requires generating plausible content within a masked region of 3D space. Prior approaches rely on 2D diffusion models to produce one or several inpainted reference views, making them susceptible to challenges associated with multi-view inconsistency and lengthy optimization times. Departing from these approaches, we introduce a reference-free Gaussian splatting inpainting method operating natively in 3D. Our method leverages the embedding space of a large, pre-trained 3D prior, combined with a structure completion network to feed a generative prior which reconstructs the missing region's geometry and appearance. Performing content generation entirely in 3D, it avoids the need to reconcile inconsistencies of multiple inpainted reference images, and is faster than related 2D-based methods. We demonstrate the effectiveness of our method qualitatively and quantitatively, through extensive experiments and a user study. To the best of our knowledge, this work is the first Gaussian splatting inpainting method to operate in the learned representation space of a 3D-native generative prior without relying on inpainted reference views.

55. 【2609.37809】Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study

链接:https://arxiv.org/abs/2609.37809

作者:Sven Ligensa,Jan Pauls,Karsten Schrödter,Ibrahim Fayad,Fabian Gieseke

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:climate change mitigation, Predicting canopy height, world forests, change mitigation, Predicting canopy

备注: Accepted at ACM SIGSPATIAL 2026

点击查看摘要

Abstract:Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world's forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown strong performance in many domains, their straightforward application to dense (i.e., pixel-level) regression tasks often yields suboptimal results. In particular, the patch size has a crucial impact on the model performance. In this work, we consider pixel-level attention schemes and show that the resulting models generally outperform those relying on larger patch sizes. However, pixel-level attention can be a prohibitively resource-intensive operation. For this reason, we conduct an extensive experimental study using efficient attention variants to identify favorable trade-offs between prediction quality and resource requirements, facilitating the practical deployment of the proposed models. In addition, we perform a comprehensive comparison with several well-established models in the field and show that, with suitable hyperparameter choices, Transformer-based architectures can outperform competing approaches. Our findings provide practical guidance for designing models for pixel-level regression tasks on medium-resolution satellite imagery, including canopy height and biomass estimation, soil moisture mapping, and yield forecasting.

56. 【2609.37801】ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding

链接:https://arxiv.org/abs/2609.37801

作者:Thomas A. O'Shea-Wheller

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:computationally efficient multi-object, efficient multi-object tracking, computationally efficient, efficient multi-object, multi-object tracking architecture

备注:

点击查看摘要

Abstract:The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with tracklet similarity matching to robustly deal with object occlusions. However, this strategy is nevertheless vulnerable to erroneous track reclassification and identity switching, as detection confidence scores dictate association priority. To address this, I present a simple enhancement of the ByteTrack architecture, named ByteTraX, that optimises track continuity via a single unified matching threshold, while penalising identity switches through stringent track initiation criteria. This approach achieves consistently improved performance across a range of diverse benchmarks including GMOT-40, LC-MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea-MOT, while simultaneously increasing processing speed by 10%. Specifically, results demonstrate a 40% reduction in identity switches, accompanied by mean increases in HOTA of 3.6, IDF1 of 5.6, and FPS of 6.3. As such, adoption of the ByteTraX algorithm has the potential to substantially enhance tracking performance over the ByteTrack baseline, while retaining the efficiency needed for real-time deployment. To facilitate usage, I provide the source code, integration functionality for the YOLO family of object detection models, and deployment instructions via an open source repository.

57. 【2609.37786】CHOQOLATE: Organizing Concept Bottleneck Latent Spaces with Choquet Integrals

链接:https://arxiv.org/abs/2609.37786

作者:Rémi Kazmierczak,Johanne Cohen,Marianne Clausel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Concept Bottleneck Models, Bottleneck Models, CLIP represent, vision-language models, built on vision-language

备注:

点击查看摘要

Abstract:Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are entangled, so individual scores do not reflect their intended meaning. We propose CHOQOLATE, an interpretable-by-design layer based on 2-additive Choquet integrals, which merges correlated concepts into compact nodes. Across four datasets, CHOQOLATE achieves a favorable accuracy-interpretability trade-off, with weight-sparse and semantically coherent nodes. A closed-form gradient derivation, backed by experiments, explains why Choquet layers drive this organization without explicit supervision. Choquet weights also map directly to Shapley values, which enables test-time intervention. On standard bias-mitigation benchmarks, suppressing spurious concepts after training performs on par with methods that require group annotations or retraining, while needing neither.

58. 【2609.37784】Planetary Feature Fields are Scalable Earth Representations

链接:https://arxiv.org/abs/2609.37784

作者:Arjun Rao,Sebastian Loeschcke,Anthony Fuller,Isaac Corley,Nico Lang,Evan Shelhamer

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Satellite observations, evolving Earth, map products describe, precomputed embeddings, stored as independent

备注: 28 pages, 16 figures, 7 tables

点击查看摘要

Abstract:Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at planetary scale. PFFs are spatially local explicit-implicit (hybrid) neural fields. Each field shares a factored feature volume---a decomposition of an explicit 3D grid with smaller factors---across products, while lightweight implicit decoders reconstruct individual products across multiple timesteps. PFFs reconstruct EO products over space and time more accurately than single-product fields at matched compression rates. At $1800\times$ compression relative to the uncompressed source data, reconstructed features retain approximately $90\%$ or more of the performance achieved with the original features on pixel-level segmentation, change detection, and patch-level classification tasks. PFFs can add new timesteps by extending their factored feature volumes and add new products by attaching new decoders, while leaving existing outputs unchanged. PFFs reduce end-to-end feature access latency by an order of magnitude relative to evaluated API and cloud-storage pipelines.

59. 【2609.37783】A Benchmark Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

链接:https://arxiv.org/abs/2609.37783

作者:Kelly McConvey,Sajad Ebrahimi,Nima Jamali,Jalehsadat Mahdavimoghaddam,Matina Mahdizadeh Sani,Maksym Taranukhin,Wentao Zhang,Jacquelyn Burkell,Yuntian Deng,Karen Eltis,Maura R. Grossman,Vered Shwartz,Ebrahim Bagheri

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:equipped to evaluate, increasingly vulnerable, legal and technical, technical workflows, CIFAR Synthetic Evidence

备注:

点击查看摘要

Abstract:Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.

60. 【2609.37775】HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

链接:https://arxiv.org/abs/2609.37775

作者:Xuanyu Zhu,Yan Bai,Yang Shi,Yihang Lou,Yuanxing Zhang,Tengfei Liu,Jing Jin,Yuan Zhou

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Pretrained visual representations, fine-grained details needed, representations support image, Pretrained visual, support image generation

备注:

点击查看摘要

Abstract:Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.

61. 【2609.37759】Selective Channel Restoration for Backdoored Vision-Language Models

链接:https://arxiv.org/abs/2609.37759

作者:Shuming Liu,Zhifang Zhang,Suqin Yuan,Khin Mi Mi Aung,Zhuoyi Lin,Lei Feng

类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)

关键词:exhibit strong multimodal, strong multimodal capabilities, poisoned fine-tuning data, Vision-language models, exhibit strong

备注: 14 pages, 4 figures

点击查看摘要

Abstract:Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.

62. 【2609.37750】Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection

链接:https://arxiv.org/abs/2609.37750

作者:Aawez Mansuri,Mohammadreza Chavoshi,Theodorus Dapamede,Wasif Bala,Beatrice Brown-Mulry,Rohan Isaac,Bardia Khosravi,Hanzhou Li,Frank Li,John T. Moon,Chad Robichaux,Dan I.G. Cohen-Addad,Ninad V. Salastekar,Janice Newsome,Judy W. Gichoya,Hari Trivedi

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:remains incompletely characterized, models remains incompletely, detection models remains, cardiovascular mortality, incompletely characterized

备注:

点击查看摘要

Abstract:Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.

63. 【2609.37732】he Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns

链接:https://arxiv.org/abs/2609.37732

作者:Sebastian Rückerl

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Instruction-based image editors, editors insert objects, image editors insert, Instruction-based image, insert objects

备注: 23 pages, 13 figures, 8 tables

点击查看摘要

Abstract:Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which estimates the camera of an image, this isolates the camera under which the editor paints. On 120 rendered cameras with exact ground truth, Qwen-Image-Edit-2511 paints tile edges that meet their vanishing points within 0.26 degrees, and its implicit camera matches the true one to 0.8 degrees in pitch and 6% in focal length, more accurately than GeoCalib except in roll. Asked to draw the horizon or mark a vanishing point instead, the editor fails, so this knowledge is revealed by painting and not by the explicit tasks we tried. The implicit camera has two priors: roll is pulled towards level (slope 0.71), and telephoto perspective towards a default of about 30 mm, which roughly matches the camera the models paint without any scene. For Qwen, the priors do not grow when blur removes four fifths of the line evidence. They are stronger on real photographs, and on NYUv2 a shorter wording of the task removes the difference for roll. On photographs from a 24--240 mm zoom lens the painted perspective grows with only 0.62 of the lens's slope, while GeoCalib and MoGe-2 saturate at about 52 and 42 mm. FLUX.1 Kontext and LongCat-Image-Edit are pulled much harder. Finally, from a level camera a camera-control LoRA executes pose commands at only 50--70% of their strength, and a board painted into its output agrees with the camera it produced.

64. 【2609.37721】CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

链接:https://arxiv.org/abs/2609.37721

作者:Sen Wang,Liu Liu,Xinjiang Wang,Zequn Chen,Haoyi Jiang,Taojun Ding,Tingyang Xiao,Zhizhong Su,Jie Wang,Sanping Zhou

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Robot policies increasingly, policies increasingly incorporate, Robot policies, increasingly incorporate semantic, actions remain aligned

备注:

点击查看摘要

Abstract:Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.

65. 【2609.37712】PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

链接:https://arxiv.org/abs/2609.37712

作者:GuangJian Team:Kaili Huang,Yongshuo Zhang,Bingtao Fu,Changjiang Jiang,Chenfan Qu,Chenfeng Zhang,Fangming Cui,Gaoyang Zhang,Jiangwei Xie,Jianshu Li,Jing Huang,Jingwen Bai,Mingqi Fang,Tao Fang,Weihong Zhang,Wenbo Du,Xiongfei Bai,Xuekang Zhu,Yinan Xia,Zhenming Wang,Jian Liu,Jingjing Liu,Xiang Qi,Weiqiang Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Optical Character Recognition, Optical Character, Character Recognition, general visual intelligence, complex visual environments

备注: Technical Report

点击查看摘要

Abstract:Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision. We introduce Competence-Guided Policy Optimization, which combines verifier-based Group Relative Policy Optimization with on-policy distillation through sample-wise routing based on teacher reliability and the teacher--student competence gap. We also introduce OCRBench v2.1, our revision of OCRBench v2 with manually verified annotation corrections and task-aligned scoring metrics. Extensive experiments across OCRBench v2.1, CC-OCR, in-house KIE Benchmark, OmniDocBench v1.6 and MDPBench demonstrate that PolyOCR achieves state-of-the-art or highly competitive performance.

66. 【2609.37709】VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

链接:https://arxiv.org/abs/2609.37709

作者:Yuta Oshima,Masakazu Yoshimura,Masahiro Suzuki,Yutaka Matsuo,Hiroki Furuta

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Recent multimodal image, Recent multimodal, enabling reference-based generation, reference-based generation guided, visual instructions

备注: Code: [this https URL](https://github.com/shim0114/VIF-Bench) , Benchmark: [this https URL](https://huggingface.co/datasets/shim0114/VIF-Bench)

点击查看摘要

Abstract:Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.

67. 【2609.37690】Honeycomb: Constant-Size Scene Memory Representation for Video World Models

链接:https://arxiv.org/abs/2609.37690

作者:Jack Wei Lun Shi,Kaichen Zhou,Haoyu Chen,Yufeng Weng,Keane Ong,Ruojin Cai,Hang Hua,Justin K. W. Yeoh,Mengyu Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models require persistent, require persistent scene, world models require, require persistent, persistent scene memory

备注: Project Page: [this https URL](https://jackswl.github.io/honeycomb/) Code: [this https URL](https://github.com/kaichen-z/honeycomb)

点击查看摘要

Abstract:Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce **Honeycomb**, a video world model built on **HexMemory**, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory warps the existing planes while preserving their dimensions, then integrates new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latent features from HexMemory to condition subsequent video generation. Because the writer processes only observations from the latest chunk, Honeycomb avoids per-scene optimization and repeated processing of the full generation history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust consistency when revisiting previously observed regions, while maintaining constant feature-storage requirements throughout generation. Code and additional visualizations are available on our this https URL.

68. 【2609.37685】PAIQ: Patch-Aligned Semantic Injection via Residual Rotation

链接:https://arxiv.org/abs/2609.37685

作者:Pinze Ren,Yuwei Zhang,Hao Chen,Linghao Meng,Chang Li,Qiankun Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Language-aligned and self-supervised, offer complementary strengths, encoders offer complementary, self-supervised visual encoders, visual encoders offer

备注:

点击查看摘要

Abstract:Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. We introduce PAIQ, a patch-aligned semantic injection framework that combines content-based cross-encoder matching with orthogonally constrained residual updates. Using DINOv3 patch features as the spatial base, PAIQ aggregates complementary SigLIP features through joint source allocation and injects the aggregate--base differences through a shared orthogonal transformation Q. This rotation adapts update directions while preserving residual norms and pairwise angles. For fixed projected features, we derive conditions for patch separability under similar semantic aggregates and show that rotation adds a nonnegative separation term over direct interpolation when the aggregate is shared. Only the projection and fusion parameters are trained; both visual encoders and the language model remain frozen, and fusion retains 196 visual tokens. Across diverse language backbones, PAIQ yields broad gains in judge-assessed correctness and reductions in hallucination severity over single-encoder interfaces on image description and visual question answering. On the 2B and 9B Qwen backbones, this compact interface outperforms the strongest evaluated fusion or token-compression baselines by about 2.9 correctness points on average.

69. 【2609.37682】Med-RADIO: Reducing All Medical Domains Into One via Multi-Teacher Distillation

链接:https://arxiv.org/abs/2609.37682

作者:Chu Zhang,Haoyu Jiang,Hongyuan Zhang,Hongbin Liu,Dong Yi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:driven significant progress, large-scale medical datasets, rapid expansion, expansion of large-scale, datasets and computational

备注:

点击查看摘要

Abstract:The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the inherent heterogeneity of medical imaging modalities, current research mainly follows two paths: specialized models optimized for specific modalities, and generalist models designed to handle multiple modalities. However, medical generalist models suffer from both insufficient training data scale relative to natural image generalists and inadequate domain-specific depth relative to medical specialists. Empirically, generalist models establish a cross-modality performance baseline, while specialists define the performance ceiling within their respective domains. To elevate this baseline toward these ceilings, we propose Med-RADIO, a medical multi-teacher distillation framework that Reduces All Domains Into One by compressing complementary expertise from multiple domain-specific teachers into a unified medical vision foundation model. Our method curates both generalist and specialist teachers, allocates modality-aligned distillation streams to reorganize generalist pretraining data so it matches specialist domains, and uses a balanced loss to prevent any single teacher from dominating the distillation process. On internal and external classification benchmarks spanning five modalities, Med-RADIO improves over strong medical generalists under linear probing and remains competitive with representative specialists on most evaluated modalities. Code is available at this https URL.

70. 【2609.37670】MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators

链接:https://arxiv.org/abs/2609.37670

作者:Haocheng Tang,Tianchi Xie,Xingqiao Lin

类目:Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Computer Vision and Pattern Recognition (cs.CV)

关键词:predicting interval-average velocities, existing advantage-based objectives, MeanFlow enables efficient, enables efficient few-step, efficient few-step generation

备注:

点击查看摘要

Abstract:MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.

71. 【2609.37659】Are In-Context Images Worth 10 Dimensions?

链接:https://arxiv.org/abs/2609.37659

作者:Adhemar de Senneville,Xavier Bou,Jérémy Anger,Rafael Grompone,Gabriele Facciolo

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:In-Context Learning capabilities, Learning capabilities, Large Language Models, Large Vision Language, induction circuit

备注:

点击查看摘要

Abstract:There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.

72. 【2609.37656】racing the Evidence: Faithful Token Attribution Through Vision-Language Reasoning

链接:https://arxiv.org/abs/2609.37656

作者:Bowen Yuan,Danny Wang,Ruihong Qiu,Zijian Wang,Zi Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large vision-language models, Large vision-language, responses remains difficult, generated responses remains, exhibit strong reasoning

备注:

点击查看摘要

Abstract:Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that rank image and prompt tokens by how much the model relies on them, such that removing higher-ranked tokens causes the likelihood of the generated response to drop more rapidly. However, existing token-attribution methods have been developed mainly for text-based language models, and our empirical study reveals two challenges when complex multimodal sources are involved. First, the joint image-text attribution can underrepresent visual evidence relative to text, obscuring the image regions supporting the response. Second, visual evidence may influence the generated response through multiple intermediate reasoning paths, while existing methods trace only a limited subset of these paths, causing important visual contributions to be underestimated. Motivated by these insights, we introduce VTrace, a multimodal token-attribution framework that traces input contributions through intermediate reasoning and calibrates attribution scores across modalities. VTrace constructs pairwise attributions that highlight token-specific contributions and aggregates all forward attribution paths in closed form to account for both direct and indirect contributions. Cross-modal calibration then rescales image and text attribution scores using modality contributions estimated from response-likelihood changes, enabling a unified ranking of input tokens. Evaluations against seven baselines across six visual reasoning benchmarks demonstrate the superior attribution faithfulness. Project page: this https URL.

73. 【2609.37655】Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

链接:https://arxiv.org/abs/2609.37655

作者:Jiayu Ying,Qijian Tian,Ruijie Xu,Xinnan Zhu,Daoguo Dong,Jiachen Xu,Xin Tan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Multimodal Large Language, Language Models, Multimodal Large, Large Language

备注: Accepted to NeurIPS 2026. 33 pages, 11 figures, 11 tables

点击查看摘要

Abstract:Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at this https URL

74. 【2609.37654】xture Space Material Diffusion

链接:https://arxiv.org/abs/2609.37654

作者:Jacob Munkberg,Peter Kocsis,Jon Hasselgren

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:texture space, material, generating high quality, material generation, texture

备注:

点击查看摘要

Abstract:We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.

75. 【2609.37648】VoxelSage: Tool-Augmented 3D CT Analysis and Simulator-Shielded Sequential Resection Planning for Liver Tumors

链接:https://arxiv.org/abs/2609.37648

作者:Binghong Qian,Xuanhe Liu,Yifan Xing,Wenjie Deng,Jian Wu,Haochao Ying

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:liver-tumor assessment requires, assessment requires segmentation, Preoperative liver-tumor assessment, assessment requires, physical-space measurement

备注: 21 pages, 10 figures. Technical report. Code at [this https URL](https://github.com/ZJUMAI/VoxelSage)

点击查看摘要

Abstract:Preoperative liver-tumor assessment requires segmentation, physical-space measurement, visual evidence, and resection planning from the same three-dimensional CT volume. Existing tools often handle these steps separately, while language models cannot reliably compute physical measurements from CT. To provide an integrated workflow, we present VoxelSage, a multi-modal system for two- and three-dimensional visualization, liver-tumor analysis, and preoperative resection planning. Its dual-port architecture separates language-model orchestration from image computation: Port A interprets requests and selects skills, while Port B applies them to CT volumes and segmentation masks and returns structured results. Keeping physical measurements in Port B prevents the LLM from computing them directly and reduces the risk of fabricated numerical results. Eight built-in skills support quantitative analysis, visual evidence generation, three-dimensional reconstruction, segmentation refinement, and sequential resection planning; user-defined skills can extend these functions. For sequence planning, a behavior-cloned neural ranker orders candidate resection targets, while a simulator-based shield checks them against predefined constraints. Across 256 unseen simulator scenes, this approach reduced mean simulated time from 34.274 to 33.388 min (0.886 min, 2.59%) and mean simulated blood loss from 300.847 to 183.852 mL (116.995 mL, 38.89%) relative to a deterministic baseline. These results demonstrate system integration and simulator-level performance, not clinical efficacy or safety. The public implementation is available at this https URL.

76. 【2609.37638】argeted Visual Counterfactual Explanations for Contrastive Vision-Language Model

链接:https://arxiv.org/abs/2609.37638

作者:Van Bach Nguyen,Jörg Schlötterer,Christin Seifer

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Current explanation methods, Current explanation, identify important regions, contrastive vision, identify important

备注:

点击查看摘要

Abstract:Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textbf{A}daptive \textbf{C}ounterfactual \textbf{E}xplanations (\mace), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. \mace constructs an editable region from either source attribution or source--target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate \mace on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel-level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.

77. 【2609.37631】Procedural Core: A Compact Recurrent Initialization for Vision Transformers

链接:https://arxiv.org/abs/2609.37631

作者:Zachary Shinnick,Christian Internò,Hemanth Saratchandran,Anton van den Hengel,Damien Teney

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:large-scale optimization, typically trained, emerge from large-scale, abstract procedurally generated, initialization

备注: Project page: [this http URL](http://zlshinnick.github.io/procedural-core/)

点击查看摘要

Abstract:Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.

78. 【2609.37605】omoTransformer: Towards a Foundation Model for CT Reconstruction

链接:https://arxiv.org/abs/2609.37605

作者:AmirEhsan Khorashadizadeh,Benjamín Béjar

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Supervised deep learning, Supervised deep, advanced sparse-view tomographic, sparse-view tomographic reconstruction, deep learning

备注:

点击查看摘要

Abstract:Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.

79. 【2609.37602】When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation

链接:https://arxiv.org/abs/2609.37602

作者:Michele Antonazzi,Alejandra C. Hernandez,José Araujo,Olov Andersson,Patric Jensfelt

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous robots operating, Robust and reliable, reliable perception, perception is essential, essential for autonomous

备注:

点击查看摘要

Abstract:Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.

80. 【2609.37582】FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning

链接:https://arxiv.org/abs/2609.37582

作者:Xinyuan Zhao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:private architectures, private models, local, remain usable, Joint

备注: 17 figures

点击查看摘要

Abstract:Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, connecting heterogeneous private models through a common prediction interface. Private models teach local Q copies; the returned Q supports local learning and Joint inference, with only Q parameters and counts exchanged. Across six datasets, FedSocket improves missing-modality recipient accuracy over Local by 14.44 and 15.51 percentage points on MELD and UCF-51. Under matched inference capacity, Joint exceeds independent ensembles by 11.06 points in UCF-51 accuracy and 4.87 points in mean bidirectional Flickr30k R@1. Joint also improves over Q alone on all four heterogeneous endpoints, demonstrating the value of combining local and exchanged predictions. Teacher controls, sharing-path interventions, and component factorials identify the roles of supervision, sharing, and deployment. FedSocket makes exchanged knowledge directly usable from federated training to recipient inference.

81. 【2609.37581】ReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning

链接:https://arxiv.org/abs/2609.37581

作者:Jing Wang,Zhiping Wu,Dongdong Ren,Youfang Han,Wei Zhao,Wenbin Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Large Language Model, incur substantial inference, substantial inference costs, inference costs due, Vision-Language Models

备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.

82. 【2609.37576】Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

链接:https://arxiv.org/abs/2609.37576

作者:Yu Zhao,Jiarui Wang,Huiyu Duan,Ye Zhao,Jutao Tang,Juntong Wang,Guangtao Zhai,Xiongkuo Min

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:traditional metrics fail, capture fine-grained alignment, robust evaluation, critical yet challenging, generative artifacts

备注:

点击查看摘要

Abstract:With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systematically evaluates LMM capabilities as evaluators of AI-generated images across four aspects: Semantics, Quality, Authenticity, and Responsibility. SQUARE-Bench introduces a granular taxonomy of 38 sub-dimensions to evaluate nearly 10K AI-generated images sampled from 22 diverse models, ranging from legacy to state-of-the-art generators, complemented by over 3K real-world images. The images are annotated with curated question-answering pairs. Extensive experiments on 23 LMMs reveal that top proprietary models, such as Gemini-3-Pro, already outperform the individual human expert baseline. However, the performance gap between models remains significant, exhibiting notable disparities in fine-grained inference and domain-specific robustness. Beyond benchmarking, we conduct a proof-of-concept study of LMM-guided iterative editing, in which dimension-specific LMMs provide diagnostic feedback to fixed image editors. The resulting guided system yields selective improvements in semantics, authenticity, and responsibility, while exhibiting a consistent visual-quality trade-off. SQUARE-Bench can serve as both a diagnostic tool for characterizing LMM evaluator capabilities and studying their use in T2I generation refinement. The benchmark and dataset will be released upon publication.

83. 【2609.37569】Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering

链接:https://arxiv.org/abs/2609.37569

作者:Yazhen Xie,Xingsong Ye,Zhineng Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Rendering accurate Chinese, Rendering accurate, text remains challenging, accurate Chinese text, Chinese text remains

备注:

点击查看摘要

Abstract:Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.

84. 【2609.37559】APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

链接:https://arxiv.org/abs/2609.37559

作者:Jianguo Huang,Jinming Liu,Qiyao Wang,Liang Xu,Jianhang Li,Zhimian Wen,Mingda Li,Shule Lu,Zhicheng Wang,Yuhan Guo,Xin Jin,Wenjun Zeng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:persistent memory, memory, real-world personal assistants, streaming, real-world personal

备注: 33 pages, 11 figures, 15 tables

点击查看摘要

Abstract:To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.

85. 【2609.37537】Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models

链接:https://arxiv.org/abs/2609.37537

作者:Arian Komaei Koma,Seyed Amir Kasaei,Aida Aryafar,Matin Ghiasi,Ali Aghayari,Amirhossein Souri,Mohammad Mosayyebi,AmirMahdi Sadeghzadeh,Mohammad Hossein Rohban

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:erase sensitive concepts, critical post-hoc safety, post-hoc safety measure, Machine unlearning, models without prohibitive

备注:

点击查看摘要

Abstract:Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.

86. 【2609.37530】RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

链接:https://arxiv.org/abs/2609.37530

作者:Shuhong Liu,Heng Zhou,Lingfeng Qian,Yuhao Fang,Xianbao Hou,Qianyu Zhou,Lin Gu,Wei Sui,Jianfei Yang,Ziteng Cui

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:models typically operate, RGB images produced, operate on RGB, image signal processor, models typically

备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we introduce RawVLA, a streaming neural ISP that adaptively renders RAW observations for frozen VLA policies while concentrating its capacity on the imaging factors relevant to embodied behavior. We further present RawVLA-Bench, a RAW-domain manipulation benchmark to expose image processing as an explicit evaluation variable across clean and adverse acquisition conditions. Experiments on RawVLA-Bench show that RawVLA preserves performance under standard conditions while substantially improving robustness under degraded imaging, establishing adaptive RAW processing as an effective interface between physical cameras and embodied policies.

87. 【2609.37529】Principled MAP estimation for inverse problems: bridging the gap between convergence and performance

链接:https://arxiv.org/abs/2609.37529

作者:Alexandre Lagier,Valentine Tosel,Anne Gagneux,Mathurin Massias,Ségolène Martin

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Pretrained denoisers provide, incorporate image priors, Pretrained denoisers, provide a powerful, incorporate image

备注:

点击查看摘要

Abstract:Pretrained denoisers provide a powerful way to incorporate image priors into restoration algorithms. Plug-and-Play and RED approaches exploit fixed-noise-level denoisers within first-order optimization schemes, with convergence guarantees, but often struggle to achieve high-quality reconstruction on severely ill-posed inverse problems. In contrast, recent state-of-the-art approaches leverage denoisers derived from flow- or diffusion-based generative models and evaluate them along a sequence of decreasing noise levels. While these methods achieve strong empirical performance, their convergence theory remains limited. In this paper, we bridge this gap by specifically designing an algorithm that combines denoisers at decreasing noise levels with a schedule tailored to ensure convergence. From a Bayesian perspective, we prove that our method converges to a $\textit{Maximum a Posteriori}$ (MAP) estimate, under suitable assumptions. Subsequently, we apply our method to various ill-posed inverse problems and show that it surpasses convergent methods while competing with state-of-the-art empirical ones.

88. 【2609.37515】Hierarchical Compression of Vision-Language Model Benchmarks

链接:https://arxiv.org/abs/2609.37515

作者:Hyunjong Ok,Seunggu Kang,Jaeho Lee

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:prohibitively expensive, relentless pace, span an ever-broader, ever-broader spectrum, spectrum of capabilities

备注: Preprint

点击查看摘要

Abstract:Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.

89. 【2609.37496】GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation

链接:https://arxiv.org/abs/2609.37496

作者:Jeonghyeok Do,Munchurl Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Paired synthetic aperture, synthetic aperture radar, Paired synthetic, aperture radar, imagery is increasingly

备注: Please visit our project page [this https URL](https://kaist-viclab.github.io/GeoSET_site/)

点击查看摘要

Abstract:Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions. We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR--EO pairs from a collection of more than 10 million SAR observations, spanning diverse sensors, spatial resolutions, and ground sampling distances. To bridge the modality gap between SAR observations and a pretrained image generator, we develop a speckle-robust SAR encoder and pretrain the conditional generator on this heterogeneous corpus. The resulting parent supports efficient adaptation across downstream datasets through low-rank adaptation (LoRA), updating only 0.60% of the generator parameters and requiring approximately one hour per dataset. Across six downstream benchmarks, GeoSET achieves state-of-the-art results in FID and DISTS with full fine-tuning or LoRA, demonstrating effective transfer across heterogeneous SAR-EO domains.

90. 【2609.37495】MotionMaestro: Masked Tokenization for Unified Motion Generation

链接:https://arxiv.org/abs/2609.37495

作者:Yun Chen,Munchurl Kim,Jeonghyeok Do

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Human motion generation, Human motion, virtual environments, character animation, embodied interaction

备注: Please visit our project page [this https URL](https://kaist-viclab.github.io/MotionMaestro_site/)

点击查看摘要

Abstract:Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy. Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.

91. 【2609.37492】Attention-Scoped Guidance: Training-Free Spatial Control for Image Editing

链接:https://arxiv.org/abs/2609.37492

作者:Zeyan Li,Wei Zhou,Hadi Amirpour,Minghao Zou,Panqi Yang,Jianfeng Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Instruction-guided image editing, Instruction-guided image, editing should change, leave the rest, image untouched

备注:

点击查看摘要

Abstract:Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (CFG), an editor combines two directions at every denoising step, one that pushes toward the instructed edit and one that pulls back toward the source image, using global weights. We introduce Attention-Scoped Guidance (ASG), a sampler wrapper that makes these weights spatial. It reads a soft support map from the instruction attention that the editor already computes, then weakens text guidance where support is low and strengthens image anchoring where support is high. The wrapper requires no training, no external mask, and no additional network evaluation. On the full MagicBrush and PIE-Bench++ splits, ASG improves preservation-oriented metrics, leading three of four MagicBrush metrics and PIE-Bench++ background PSNR. A dose-matched control that removes the spatial placement loses up to 0.73 CLIP on PIE-Bench++, confirming that the spatial allocation itself carries the gain.

92. 【2609.37488】FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding

链接:https://arxiv.org/abs/2609.37488

作者:Taiyo Sato,Takamasa Sanda,Keisuke Maeda,Takahiro Ogawa,Miki Haseyama,Shunya Nagashima

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:single prompted call, multimodal large language, referring expression comprehension, Frozen multimodal large, large language models

备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.

93. 【2609.37487】Physics-Guided Flow-Map Matching for Precipitation Nowcasting

链接:https://arxiv.org/abs/2609.37487

作者:Shunya Nagashima,Takumi Bannai,Makoto Misaizu,Keisuke Maeda,Takahiro Ogawa,Miki Haseyama

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Precipitation nowcasting, generating future radar, generating future, disaster response, flood warning

备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and rare high-intensity structures that matter most. Deterministic models minimize a pixel loss and are driven toward the conditional mean, which blurs exactly those structures, while generative models that add a stochastic residual on top of a deterministic backbone inherit the same blur. We propose Physics-Guided Flow-Map Matching (PG-FMM), a conditional flow-map model that decouples predictable advection from uncertain small-scale detail. A frozen Lagrangian advection prior transports the radar field and supplies an explicit motion forecast, and a flow-map generative head, conditioned on the past frames and the prior rollout rather than summed onto it, produces sharp stochastic detail in four sampling steps. The prior serves only as guidance, so the head replaces blurred structure instead of inheriting it. Extensive experiments on four radar benchmarks show that PG-FMM outperforms state-of-the-art methods on 18 of 24 metrics, with the largest gains at heavy-rain thresholds, where the critical success index improves by up to 58.9%. The project page can be found at this https URL.

94. 【2609.37485】PoE-Fuse: Precision-Weighted Expert Fusion for Bi-Temporal Change Understanding

链接:https://arxiv.org/abs/2609.37485

作者:Haruki Watase,Shunya Nagashima,Takayuki Nishimura

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Bi-temporal change understanding, spanning change detection, building localization, satellite images, environmental monitoring

备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Bi-temporal change understanding, which localizes and characterizes what changed between two satellite images, is central to disaster response and environmental monitoring, spanning change detection, building localization, and damage assessment. Strong vision-language models address these tasks, but adapting them typically requires full fine-tuning or reinforcement learning, which is costly and unstable. We propose PoE-Fuse, a parameter-efficient framework that instead composes frozen foundation experts for geometry, grounding, and language, resampling their features onto a shared spatial grid and training only a lightweight fusion trunk. PoE-Fuse treats the aligned features as Gaussian observations of a latent scene state and fuses them by learned per-cell precision. This product-of-experts estimator strictly generalizes uniform summation and scalar gating, and extends to change fields by composing the precisions of the two timestamps. A single shared trunk solves the three tasks at once, reaching a mean F1 of 59.2%, compared with 40.7% for an instruction-tuned temporal vision-language assistant, and surpassing dedicated change-detection models retrained under the same protocol and training budget.

95. 【2609.37481】Label Less, Learn More: Resource-Efficient Active Semi-Supervised Learning for Onboard Satellite Image Annotation

链接:https://arxiv.org/abs/2609.37481

作者:Ahmed Abdelnaby,Mohamed Elmahallawy,Marius Bernahrndt,Tobias Hecking

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large-scale pervasive sensing, task-specific onboard vision, Large-scale pervasive, pervasive sensing increasingly, sensing increasingly relies

备注:

点击查看摘要

Abstract:Large-scale pervasive sensing increasingly relies on high-resolution satellite imagery, yet task-specific onboard vision is constrained by costly annotation and limited computation, memory, energy, and communication resources. Existing approaches largely rely on either data-hungry supervised learning or large vision-language foundation models, limiting efficient adaptation and deployment under these constraints. We present SatLabel, a resource-aware learning framework that transforms limited satellite labels into progressively refined onboard models through adaptive sample acquisition and semi-supervised model adaptation. Rather than repeatedly training on uniformly sampled labels, SatLabel closes the loop between model uncertainty, class imbalance, and pseudo-label quality to selectively acquire informative samples while exploiting abundant unlabeled imagery. This enables a compact student to adapt to target sensing domains with reduced annotation and inference costs. We further introduce an optional Mixture-of-Experts (MoE) student with graph-based feature refinement to enhance representation capacity while retaining a lightweight footprint. We evaluate SatLabel on 11 remote-sensing datasets spanning core, extended, and unseen domains against RemoteCLIP zero-shot inference. SatLabel improves Macro-F1 on most core and extended datasets while maintaining strong cross-dataset transfer to unseen domains. More importantly, the Balanced student contains only 11.2 M parameters and occupies approximately 42.8 MB, compared with 151.3M parameters and 577 MB for RemoteCLIP, while requiring 3.65 versus 5.89 GFLOPs. Across four efficiency benchmarks, it achieves approximately 2x higher GPU-forward throughput and reduces energy per image on datasets.

96. 【2609.37476】Learning Social Navigation from Internet Videos in the Policy State Space

链接:https://arxiv.org/abs/2609.37476

作者:Jiaming Wang,Duc Thang Nguyen,Jizhuo Chen,Volodymyr Shcherbyna,Diwen Liu,Zhengcheng Shen,Harold Soh

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:policies requires simulators, robust social-navigation policies, social-navigation policies requires, Training robust social-navigation, diverse scene layouts

备注: 9 pages, 5 figures, 6 tables

点击查看摘要

Abstract:Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning.

97. 【2609.37465】Event-Only Wingbeat Counting under Camera Motion: A Controlled MuJoCo Benchmark

链接:https://arxiv.org/abs/2609.37465

作者:Zhang Nengbo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Counting completed wingbeats, completed wingbeats requires, wingbeats requires identifying, requires identifying individual, Counting completed

备注: 10 pages, 1 figure, 6 tables. Controlled simulation study with an ideal contrast-event sensor. Follow-up to [arXiv:2609.17308](https://arxiv.org/abs/2609.17308) using new event-only acquisitions and a motion ablation protocol

点击查看摘要

Abstract:Counting completed wingbeats requires identifying individual cycles, including during frequency changes and pauses; estimating a dominant frequency alone is insufficient. Camera motion further mixes target and background brightness changes in event observations. We present a controlled MuJoCo benchmark that separates motion training from event-only image translation compensation. The acquisition contains 324 streams from 24 independent scenes, three flapping geometries, two distances (1.5 and 3.0 m), and static, moderate-motion and stronger-motion views. Fifteen scenes are used for fitting, three for validation and six for held-out testing. A fixed causal temporal convolutional network is evaluated in a matched 2 x 2 ablation with three initialization seeds and compared with ridge, Fourier, autocorrelation and an adapted EEPPR baseline. Under moderate motion, paired motion training reduces count mean absolute error from 31.130 to 3.185 cycles at 1.5 m and from 42.019 to 5.444 at 3.0 m. Adding the tested compensation increases these errors to 4.630 and 10.185, respectively. A Fourier baseline achieves 0.944 cycles at 1.5 m under moderate motion, showing that the neural model is not uniformly best. We report exact-count accuracy and temporally matched cycle F1 alongside count error. These findings support motion-aware training in this small synthetic benchmark, while exposing limits of simple event-background stabilization. They do not establish real-sensor performance, aerodynamic flight, or generalization to unseen vehicle types.

98. 【2609.37426】LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension

链接:https://arxiv.org/abs/2609.37426

作者:Arka Mukherjee,Kaleen Shrestha,Larissa Zhu,Maja Matarić

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Modern vision-language models, shown promising results, Modern vision-language, rich semantic information, shown promising

备注: Under review at conference. Preprints allowed when under review

点击查看摘要

Abstract:Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs--Gemma 4 31B and Qwen3.6 27B--across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.

99. 【2609.37407】Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

链接:https://arxiv.org/abs/2609.37407

作者:Xianghan Wei,Xiaoda Yang,Zhi Wang,An Pan,Daoan Zhang,Huayi Zhang,Yan Zhang,Wei Xu,Zishun Liao,Jianwen Lou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generating high-quality short, major bottleneck lies, independently generated shots, preserve consistent characters, high-quality short videos

备注:

点击查看摘要

Abstract:While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.

100. 【2609.37400】BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling

链接:https://arxiv.org/abs/2609.37400

作者:Xiaojian Shen,Dahu Shi,Jianrong Zhang,Hai Li,Hongwei Zhao,Dawei Zhang,Yunzhi Zhuge,Zhiliang Wu,Guanghui Yue,Wei Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires accurate synchronization, Generating realistic, Hierarchical Decoupled Attention, challenging task, task that requires

备注: Published in Pattern Recognition

点击查看摘要

Abstract:Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.

101. 【2609.37387】Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI

链接:https://arxiv.org/abs/2609.37387

作者:Jesse Phitidis,William N. Whiteley,Joanna M. Wardlaw,Miguel O. Bernabeu,Yajun Cheng,Xiaodi Liu,Junfang Zhang,Una Clancy,Stephen Makin,Roberto Duarte Coello,Susana Muñoz Maniega,Mark E. Bastin,Simon R. Cox,Maria del C. Valdés Hernández

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Enlarged perivascular spaces, poor brain health, brain magnetic resonance, magnetic resonance imaging, Enlarged perivascular

备注:

点击查看摘要

Abstract:Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, radiologists visually score their burden following a clinical grading scale - a task that would benefit from automation to accelerate analyses and overcome the influence of inter-observer differences. We developed and evaluated methods for training machine learning models to score PVS incidence in the basal ganglia (BG) and centrum semiovale (CSO) leveraging the Potters/Wardlaw scale. The novelty in our work lies in the use of imperfect, semi-automatically generated "silver-standard" PVS segmentation masks during training, in addition to PVS radiological scores. We comparatively evaluated a conditional convolutional neural network (CNN) which accepts PVS masks as an extra input channel, a multi-task CNN which performs both PVS segmentation and scoring, and a logistic regression model which utilises features derived from PVS masks to predict PVS scores. Multi-task learning was the most effective method, achieving a mean average precision of 64.08% compared to 60.22% for the conditional CNN, 52.11% for a baseline CNN trained only to predict PVS scores, and 49.32% for the logistic regression model. The multi-task model showed an ability to localise individual PVS not shown by the other CNNs, and behaved in a probabilistically sensible way, predicting with lower confidence on inherently harder classes. Age, sex, hypertension status, white matter hyperintensity volume, and ischaemic stroke lesion status were shown to be associated with the multi-task model's PVS score predictions and the ground truth in a similar way.

102. 【2609.37386】Anatomy-Aware Prediction of Bronchoscopic Accessibility from 3D CT

链接:https://arxiv.org/abs/2609.37386

作者:Linkai Peng,Cuiling Sun,Bin Wang,Jamie Rowell,Catherine Gao,Oyku Ikizgul,Eminenur Sentasci,Andrea Bejar,Halil Ertugrul Aktas,Gorkem Durak,Momen Wahidi,Christopher Kapp,Ulas Bagci

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:lung lesions, diagnosis of lung, Path Geometry Expert, Current accessibility assessment, Abstract

备注: Accepted in MICCAI 2026

点击查看摘要

Abstract:Pre-operative planning for bronchoscopy is critical for the diagnosis of lung lesions. Current accessibility assessment relies on subjective manual inspection of CT scans, which is time-consuming and prone to inter-observer variability. In this paper, we formalize bronchoscopy accessibility prediction as a novel supervised learning task and present the first end-to-end framework to address it. We propose an Anatomy-Aware Mixture-of-Experts (MoE) model that integrates specialized modules: a CT Expert for local morphological features, a Lobe Expert for anatomical priors, and a Path Geometry Expert that encodes the sequential constraints of the bronchial tree. To support this task, we curated the first clinical dataset of 438 cases with pre-operative CT scans and documented procedural outcomes. Experimental results demonstrate that our method achieves an AUROC of 0.8052, significantly outperforming both state-of-the-art baselines and experienced human experts. This work establishes a new benchmark for computer-aided interventional planning in pulmonary medicine. Our data and code will be publicly available at this https URL.

103. 【2609.37378】Do-JEPA: From Masking to Intervention in Latent World Models

链接:https://arxiv.org/abs/2609.37378

作者:Hossein Resani,Javen Qinfeng Shi

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Latent world models, action, effect, Latent, effect error

备注:

点击查看摘要

Abstract:Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $\Delta z=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.

104. 【2609.37374】MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding

链接:https://arxiv.org/abs/2609.37374

作者:Heyu Huang,Chi Chen,Zonghao Guo,Yuhua Li,Maosong Sun,Ruixuan Li

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:fine-grained visual perception, recently delivered substantial, delivered substantial gains, Reinforcement learning, opening a promising

备注:

点击查看摘要

Abstract:Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task--sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.

105. 【2609.37372】hink Before You Score: Thinking Reward Model for Visual Generation

链接:https://arxiv.org/abs/2609.37372

作者:Xuehai Bai,Zhenchen Tang,Yang Shi,Dianyi Wang,Tengfei Liu,Wanshun Su,Xuanyu Zhu,Ruohui Wang,Haiwen Diao,Haotian Wang,Xiaoling Gu,Yuanxing Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing approaches typically, approaches typically map, typically map task, map task conditions, candidate outputs directly

备注: 31 pages

点击查看摘要

Abstract:Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.

106. 【2609.37360】Visual Anomaly Synthesis for Model Selection in Data Scarcity

链接:https://arxiv.org/abs/2609.37360

作者:Daniel Pröll,Thomas Kraxner,Tobias Schaefer,Sebastian Hegenbart

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:industrial condition monitoring, systems for industrial, industrial condition, defective samples, specific asset

备注:

点击查看摘要

Abstract:Defect detection systems for industrial condition monitoring can only be relied upon if they are validated, yet defective samples are rare and, for a specific asset, often nonexistent. We present a framework that synthesizes severity-graded defects on real non-defective images without any defect references for the target asset, that can be used for model selection and validation. A defect taxonomy for common failure modes is distilled from literature into prescriptive prompts at varying defect severities. Regions of interest are cropped from in defect-free images and edited with a pre-trained image generation model ("FLUX.2 [klein]"). Color-matching and blending are employed to improve structural coherence with the original image. Generations are filtered out by a scorer and by estimated detection difficulty. Model selection experiments on MVTecAD show image AUROC choice regret over model selection can be nearly halved compared to the best fixed model chosen with access to test data. Experiments show the need for severity-graded anomaly synthesis. A case study investigates the proposed method for in-situ monitoring of Pelton turbine runners in hydropower, where real defect images are rare and expensive to collect. A PatchCorebased anomaly detection model is fit on Pelton turbine images and selected and validated using synthetic images, showing strong detection performance (94 % correct detection at optimal threshold and AUROC 0.97). The model reliably detects moderate and advanced defects, while early-stage defects remain challenging, indicating the synthetic data meaningfully stresses detector sensitivity.

107. 【2609.37359】Encore: Few-Shot Agentic Discovery of Manipulation Strategies

链接:https://arxiv.org/abs/2609.37359

作者:Yifan Kang,Zihan Wang,Zhiwen Fan,Bangya Liu

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:agent, coding agent studies, debug programs, Coding agents, coding agent

备注:

点击查看摘要

Abstract:Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the pack, writes a policy program against a fixed perception and action API, refines it iteratively over a few development rollouts, and freezes it before a sealed evaluation that never reveals the success signal. On LIBERO-PRO, the agent's first program already succeeds in half of the perturbed tasks with demonstrations and in one task without them, and the frozen programs outperform the strongest prior agentic system run with the same language model (96.3% against 89.3%). On RoboDojo tasks whose instructions leave the goal unstated, no program succeeds without demonstrations. ENCORE also runs on a real bimanual robot, learning cube handover and cup inversion from five demonstrations each.

108. 【2609.37350】PCaPaint: Prostate Cancer Inpainting by Mitigating Shortcut Learning

链接:https://arxiv.org/abs/2609.37350

作者:Levente Lippenszky,Hongxu Yang,Marcell Dömötör,Krisztian Koos,László Ruskó

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:systems for tumor-specific, tumor-specific applications, applications is limited, scarcity of labeled, prostate cancer MRI

备注: Accepted at the DGM4MICCAI workshop at MICCAI 2026

点击查看摘要

Abstract:The development of AI systems for tumor-specific applications is limited by the scarcity of labeled data. Synthetic tumor inpainting offers a promising approach but faces challenges for prostate cancer MRI which contains high-resolution multi-sequence data. Although methods leveraging latent diffusion models (LDMs) enable large-volume synthesis, they are prone to shortcut learning, simply reproducing the condition image created by masking the lesion region. In this work, we introduce PCaPaint, a prostate cancer inpainting method based on LDMs that explicitly addresses this failure mode. To overcome shortcut learning that compromises synthetic tumor texture, we propose a simple yet efficient conditioning strategy in which the condition image is filled with Gaussian noise, and we provide theoretical justification. In addition, we propose a novel training objective for LDM that emphasizes the error within the lesion region. Furthermore, we introduce a multi-sequence latent design, in which T2w scans and DWIADC scans are compressed using two separate autoencoders to preserve their distinct frequency characteristics. Extensive experiments demonstrate that the generated synthetic data improves downstream performance in prostate lesion segmentation, patient-level classification and lesion-level detection. Furthermore, our method significantly outperforms a recent state-of-the-art LDM-based tumor inpainting method both in downstream performance and in synthetic image quality.

109. 【2609.37349】AEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG

链接:https://arxiv.org/abs/2609.37349

作者:Yalun Wu,Bingzhou Wang,Boyang Wang,Peiying Wang,Shaojie He,Yunhan Wang,Shaozu Yuan,Jiawei Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:visual retrieval-augmented generation, answers complex questions, repeatedly retrieving visual, evidence, retrieval-augmented generation

备注:

点击查看摘要

Abstract:Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.

110. 【2609.37345】When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs

链接:https://arxiv.org/abs/2609.37345

作者:Xiang Hu,Jiazuo Yu,Lu Zhang,Yunzhi Zhuge,Huchuan Lu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video Large Language, Streaming video understanding, requires Video Large, Large Language Models, Large Language

备注:

点击查看摘要

Abstract:Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.

111. 【2609.37340】HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing

链接:https://arxiv.org/abs/2609.37340

作者:Li Pang,Xinqiao Wu,Jing Yao,Pedram Ghamisi,Jun Zhou,Zhengchao Chen,Deyu Meng,Xiangyong Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:material-level Earth observation, Earth observation, model remains difficult, material-level Earth, Hyperspectral remote sensing

备注: Accepted by IEEE Geoscience and Remote Sensing Magazine (GRSM)

点击查看摘要

Abstract:Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlight{present} \textbf{HyperSAM}, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.

112. 【2609.37339】UGO: Unified Architecture for General Multi-Object Tracking by Segmentation

链接:https://arxiv.org/abs/2609.37339

作者:Jer Pelhan,Alan Lukezic,Matej Kristan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:single first-frame exemplar, General multi-object tracking, first-frame exemplar, General multi-object, user-specified category

备注: Accepted to NeurIPS2026

点击查看摘要

Abstract:General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, and conflicts. A hierarchical memory spanning global and instance levels improves recall and per-instance segmentation accuracy using a new memory management protocol. UGO sets a new state-of-the-art on GMOT benchmarks and video object counting, and is competitive with specialist MOT methods, establishing a strong paradigm for unified, open-category multi-object tracking.

113. 【2609.37331】OFBD: Object-Focused Background Debiasing for Long-Tailed Learning

链接:https://arxiv.org/abs/2609.37331

作者:Shenghan Chen,Yiming Liu,Zhipeng Deng,Haolin Wang,Jiale Zhou,Zhijian Wu,Xiankai Lu,Yafei Ou,Yefeng Zheng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Balancing performance trade-offs, Balancing performance, performance trade-offs, remains a long-standing, long-standing challenge

备注:

点击查看摘要

Abstract:Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, representation learning, or data augmentation, but the underlying cause of tail class degradation is still insufficiently explored. In this paper, we find that standard long-tailed training induces background-biased representation and optimization: tail classes suffer larger background distribution shifts and become increasingly driven by background gradients. This reveals that tail degradation is not merely caused by insufficient samples, but also by the learning of irrelevant background features. To tackle this issue, we propose Object-Focused Background Debiasing (OFBD), a framework that mitigates background bias from both distribution and optimization perspectives. Specifically, Foreground-guided CutMix preserves target-related foregrounds while diversifying complementary backgrounds, and Background-guided Feature Rectification suppresses background-biased features without learnable parameters or additional training. Extensive experiments show that our method improves overall accuracy, achieves significant tail-class gains, and can serve as a plug-in for mainstream long-tailed methods without external data or pretrained recognition models. The code is available at: this https URL

114. 【2609.37330】he Domain Is a Residue: Adapting Self-Supervised Features, Not Generators

链接:https://arxiv.org/abs/2609.37330

作者:Thomas Deixelberger,Markus Steinberger

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Machine Learning (cs.LG)

关键词:Clearing fog, DINO feature map, source domain, Representation Feature Adapter, Clearing

备注: 9 pages main text, 28 pages including appendix. 12 figures, 13 tables

点击查看摘要

Abstract:Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.

115. 【2609.37317】What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation

链接:https://arxiv.org/abs/2609.37317

作者:Sieun Hyeon,Yejoon Lee,Mintaek Lim,Woojin Kim,Jaeik Kim,Jaeyoung Do

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:individually plausible outputs, coherent shared event, independent text, individually plausible, shared event

备注:

点击查看摘要

Abstract:Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.

116. 【2609.37314】FLASH: A "Generate Once, Synthesize Many" Framework for Synthetic Anomaly Generation in Industrial Anomaly Detection

链接:https://arxiv.org/abs/2609.37314

作者:Abhay Kumar Das,Rajesh Gangireddy,Ashwin Vaidya,Samet Akcay

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:expand industrial anomaly, scarce or unavailable, expand industrial, Object Boundary Suppression, industrial anomaly datasets

备注: Submitted to WACV 2027

点击查看摘要

Abstract:Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extremes: procedural approaches are fast but struggle to represent complex anomalies, while generative approaches produce diverse defects but require costly per-sample generation. We present FLASH, a framework that decouples defect generation from anomaly synthesis under a ``generate once, synthesize many'' paradigm. Given only normal images, FLASH uses Vision-Language Model (VLM) guidance and an image-generation model to produce a small set of defect images, from which it extracts, validates, and banks reusable defect patches. For synthesis of anomalous images, Object Boundary Suppression (OBS) first identifies the probable foreground object-aware region of the host image, while Multi-Resolution Spectral Pyramid (MRSP) noise generates diverse, size-controllable masks that determine the defect location and spatial extent. It then composes a large and diverse synthetic anomalous image set by localizing the defect region, sampling size-controllable placement masks and seamlessly blending retrieved defects onto new defect-free images without further need for image generation. Experiments on the MVTec AD 2 dataset show that FLASH-generated anomalies nearly close the calibration gap on real defects, reaching 78.1% image-level F1 against an 83.6% real-anomaly upper bound and providing the most consistent calibration transfer across detectors among procedural and generative alternatives. Moreover, FLASH synthesizes anomalies more than 11.95x faster than per-sample generative approaches.

117. 【2609.37302】chnical note on: Zero-Training Feature-Space Alignment via Information Geometry

链接:https://arxiv.org/abs/2609.37302

作者:Behraj Khan,Tahir Qasim Syed,Syed Ahmad Chan Bukhari

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Deep vision models, Deep vision, ZFGA, Fisher, Fisher Geometry Alignment

备注:

点击查看摘要

Abstract:Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness under covariate shift without modifying model parameters. ZFGA is based on the observation that distribution shifts distort feature-space geometry. It estimates the Fisher information matrix of the predictive distribution with respect to feature embeddings and applies a linear transformation that aligns test-feature Fisher geometry with a reference geometry computed from clean data. This provides a natural-gradient-inspired preconditioning step in feature space. We evaluate ZFGA on CIFAR-10-C and ImageNet-C using ResNet-50, DINO ViT-S/16, and CLIP ViT-B/32. ZFGA consistently improves over zero-shot inference across all three models, although it is not the strongest method for every model. Covariance whitening performs better on ResNet-50, while Fisher whitening is statistically indistinguishable from ZFGA on CLIP. Across six training-free and gradient-based alternatives (covariance whitening, Fisher whitening, TENT, T3A, LAME, and AdaNPC), ZFGA is the only method that does not substantially harm any of the three model families. The Fisher geometry distortion is also positively correlated with ZFGA gain (Pearson r = 0.366, p = 0.017), providing preliminary evidence that geometric misalignment contributes to robustness degradation. ZFGA requires only forward passes and matrix operations at inference time, offering a lightweight and deterministic alternative to optimization-based test-time adaptation.

118. 【2609.37298】Scaling Full Conformal Image Classifiers

链接:https://arxiv.org/abs/2609.37298

作者:Julio Silva-Rodríguez,Ender Konukoglu

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:full conformal prediction, Conformal prediction, making it attractive, attractive for high-stakes, Targeted Full Conformal

备注: NeurIPS 2026. Code: [this https URL](https://github.com/jusiro/T-FCP)

点击查看摘要

Abstract:Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.

119. 【2609.37297】Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

链接:https://arxiv.org/abs/2609.37297

作者:Zhiyuan Li,Wenyan Yang,Pekka Marttinen,Joni Pajarinen

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Cross-skeleton motion generation, motion generation trains, generation trains generative, carry action structure, trains generative models

备注:

点击查看摘要

Abstract:Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: this https URL.

120. 【2609.37287】VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics

链接:https://arxiv.org/abs/2609.37287

作者:Bo Lv,Mao Zheng,Zheng Li,Fangxu Liu,Mingrui Sun,Tao Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:requiring visual understanding, Image translation, understanding and meaning, fundamental capability, multilingual applications

备注:

点击查看摘要

Abstract:Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.

121. 【2609.37283】SAM Meets VLM: Parameter-Decoupled Full-Parameter Training for Unified Medical Reasoning and Segmentation

链接:https://arxiv.org/abs/2609.37283

作者:Xuyang Cao,Enyou Liu,Jun Zhao,Zhuoyun Liu,Jintao Fei, Leo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Medical multimodal large, multimodal large language, answer clinical questions, large language models, multimodal large

备注:

点击查看摘要

Abstract:Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special SEG token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the SEG hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable SEG prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.

122. 【2609.37264】UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

链接:https://arxiv.org/abs/2609.37264

作者:Yuhao Liu,Yiming Zhong,Hanqing Wang,Shaocheng Yan,Yuhang Zhang,Wenzhou Lyu,Ziyang Ding,Wei Zhang,Xue Zhao,Jin Pan,Yuexin Ma,Xinge Zhu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:localize actionable regions, supporting embodied interaction, actionable regions supporting, regions supporting embodied, embodied interaction

备注:

点击查看摘要

Abstract:Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: this https URL

123. 【2609.37263】Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

链接:https://arxiv.org/abs/2609.37263

作者:Siqi Lu,Suo Wei,Yongbin Zheng,Jianhang Yao,Wanying Xu,Peng Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Vision-Language Models, Vision-Language Models, Large Vision-Language, achieve remarkable success, remarkable success

备注:

点击查看摘要

Abstract:While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.

124. 【2609.37260】Collision-Aware and Observation-Aligned Object-Centric Scene Reconstruction from Point Cloud

链接:https://arxiv.org/abs/2609.37260

作者:Yuxuan Xie,Xuan Yu,Rong Xiong,Yue Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires completing partial, reconstruction requires completing, Object-centric scene reconstruction, completing partial object, preserving metric alignment

备注:

点击查看摘要

Abstract:Object-centric scene reconstruction requires completing partial object observations while preserving metric alignment and avoiding collisions with the surrounding. Existing generation-based methods are often image-conditioned and suffer from scale ambiguity and insufficient geometric constraints. We propose COOL, a framework for COllision-aware and Observation-aLigned reconstruction. Based on an object generation model, COOL conditions the generation on instance and background point clouds. Instance geometry anchors generation in scene coordinates, while background geometry provides local context for scene-consistent completion. We further introduce an explicit collision loss and use joint optimization and resampling to reduce collisions during inference. Experiments on 3D-Front and Scan2CAD demonstrate strong scene-level fidelity, observation alignment, and collision reduction. Moreover, additional studies validate its robustness to mask errors and its applicability to real-world scene replicas.

125. 【2609.37250】V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

链接:https://arxiv.org/abs/2609.37250

作者:Yang Zhang,Jiangyuan Zhao,Chenyou Fan,Jiayu Hu,Xiu Yuan,Chenjia Bai,Xiu Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:couple future visual-state, future visual-state prediction, World-action models, couple future, future visual-state

备注: 19 pages, 5 figures, 11 tables

点击查看摘要

Abstract:World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at this https URL.

126. 【2609.37243】Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features

链接:https://arxiv.org/abs/2609.37243

作者:Dae Ung Jo,Jongin Lim,YoungJoon Yoo,Daeho Um

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:modality, Cross-modal, student, student modality, teacher modality

备注:

点击查看摘要

Abstract:Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.

127. 【2609.37230】Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

链接:https://arxiv.org/abs/2609.37230

作者:Woosang Jeon,Jiwon Yang,Soo Chung,Taehyeong Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Visual, linguistic conceptualization, matched visual grounding, visual discriminability, linguistic addressability

备注: 27 pages, 10 figures. Code available at [this https URL](https://github.com/LABA-SNU/seeing-is-not-addressing)

点击查看摘要

Abstract:Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.

128. 【2609.37229】AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control

链接:https://arxiv.org/abs/2609.37229

作者:Jingzhong Lin,Zhanke Wang,Heng Li,Wenxiang Liu,Zhao Zhang,Kecheng Tang,Dongdong Xiang,Changbo Wang,Di Kang,Chunchao Guo,Linchao Bao,Gaoqi He

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Human motion defines, Human motion, camera, Human, motion defines

备注:

点击查看摘要

Abstract:Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.

129. 【2609.37225】ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

链接:https://arxiv.org/abs/2609.37225

作者:Zijing Cai,Yuzhe Wang,Jingxian Zhu,Fengbin Zhu,Richang Hong

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:large language models, shown strong potential, multimodal representation learning, Multimodal large language, universal multimodal representation

备注: 19 pages

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.

130. 【2609.37200】Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

链接:https://arxiv.org/abs/2609.37200

作者:Songlin Yang,Xiaotong Zhao,Jiacheng Zhang,Zhe Wang,Toyota Li,Eric Liu,Alan Zhao,Anyi Rao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multi-reward guided reinforcement, guided reinforcement learning, Multi-reward guided, including modality-specific quality, reinforcement learning

备注:

点击查看摘要

Abstract:Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.

131. 【2609.37195】Exploring In-Context Learning for Handwritten Text Recognition

链接:https://arxiv.org/abs/2609.37195

作者:Eric Ayllon,Abel Gandia,Jorge Calvo-Zaragoza

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Handwritten Text Recognition, historical documents, indispensable tool, digitization of historical, Handwritten Text

备注: 19 pages, 3 figures

点击查看摘要

Abstract:Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses mostly on specialized models that require large amounts of annotated samples to achieve satisfactory performance. We explore the use of In-Context Learning with pre-trained Vision-Language Models (VLMs) to create a transcription pipeline without updating the model's parameters. We then evaluate this pipeline across multiple collections and models, and demonstrate that general-purpose VLMs can be effectively taught how to transcribe handwritten text from images. To assess how our observations may translate to practical applications, we evaluate the performance in a Cross-Domain (CD) scenario, where context examples are drawn from a different collection than the query image. Results in both the controlled In-Domain (ID) scenario and the realistic CD scenario follow the same patterns. First, as context size grows, the error range is expected to narrow towards the average performance. Thus, larger context sizes sacrifice the performance of the oracle-best sampling for lower expected error rates. The results obtained show that, without any parameter updates, this methodology has strong potential to compete with traditional HTR in the presence of domain shift. Moreover, we show and argue that some context samplings work better than others and suggest more effort should be put into finding an ideal sampling method in future work.

Comments:
19 pages, 3 figures

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

ACMclasses:
I.4.1; I.4.9

Cite as:
arXiv:2609.37195 [cs.CV]

(or
arXiv:2609.37195v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.37195

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Eric Ayllon [view email] [v1]
Tue, 29 Sep 2026 10:15:38 UTC (733 KB)

132. 【2609.37190】HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

链接:https://arxiv.org/abs/2609.37190

作者:Zhangquan Chen,Yaoxin Niu,Xiang An,Mingze Sun,Zhumei Wang,Chih-Ting Liao,Hongkun Cao,Ruqi Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multi-turn visual search, Multi-turn visual, visual search agents, agents answer questions, search agents answer

备注: 24 pages, 9 figures. Code: [this https URL](https://github.com/zhangquanchen/HAPRL)

点击查看摘要

Abstract:Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.

133. 【2609.37187】InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

链接:https://arxiv.org/abs/2609.37187

作者:Hongpei Zheng,Hujun Yin

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Language-guided navigation requires, requires connecting partial, connecting partial observations, persistent spatial reference, navigation requires connecting

备注:

点击查看摘要

Abstract:Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.

134. 【2609.37177】Sparse cubical complexes for efficient topology-preservation in image data

链接:https://arxiv.org/abs/2609.37177

作者:Alexander H. Berger,Marco Fontana,Daniel Rueckert,Johannes C. Paetzold,Laurin Lux,Ulrich Bauer

类目:Computer Vision and Pattern Recognition (cs.CV); Computational Geometry (cs.CG); Machine Learning (cs.LG)

关键词:Persistent homology, image segmentation, frequently used tool, tool for extracting, extracting and preserving

备注:

点击查看摘要

Abstract:Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the runtime cost of PH-based methods often makes their practical use infeasible. In this work, we argue that this runtime cost is largely driven by processing information that is unimportant for downstream application (e.g. as optimization objective). We propose sparse cubical filtrations as an alternative foundation for PH computation, reducing subsequent computational costs by factors of up to 100 on real datasets. We show close agreement with the optimization signal of the dense counterpart and empirically evaluate our solution's effectiveness as an optimization objective in realistic training regimes where other PH-based objectives can practically not operate (i.e., 3D data with large patch sizes). We show how our solution improves topological accuracy by up to 80\% across six diverse datasets while maintaining pixel- and region-based accuracy.

135. 【2609.37162】End-to-End Self-Supervised RGB-T Tracking without Modality Misleading

链接:https://arxiv.org/abs/2609.37162

作者:Shenglan Li,Rui Yao,Kunyang Sun,Hong Jia,Yong Zhou,Javen Qinfeng Shi,Xinyu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:object tracking leverages, thermal infrared modalities, RGB-T object tracking, adverse conditions, bounding box annotations

备注:

点击查看摘要

Abstract:RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at this https URL.

136. 【2609.37148】Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction

链接:https://arxiv.org/abs/2609.37148

作者:Siddhant Jain,Dimitra Tsovaltzi

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:learn and grow, regulates their emotions, difficult conversation, directly observable, qualities that matter

备注: 8 pages, 6 figures

点击查看摘要

Abstract:Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind of clean labeling that most machine learning pipelines are built around. We study this challenge through a case that is well grounded in psychological theory but rarely modeled computationally: self-compassion, the tendency to respond to one's own setbacks with patience rather than harsh self-criticism. We examine how it appears during structured reflective interviews in a technology-mediated training setting, where people naturally talk through socio-emotionally demanding situations. Since no existing dataset captures this kind of construct in this kind of setting, we collected and annotated 51 reflective dialog sessions using an independent, temporally overlapping annotation scheme grounded in established theory. We consolidate the underlying six-component psychological model into a three-class supervision space, balancing self-kindness and mindfulness against self-critical or overwhelmed states, and build a reproducible window-based pipeline that aligns video, audio, and text on a shared timeline. Unimodal models trained on each modality separately are compared against a simple probability-level fusion strategy, which yields modest but consistent gains over the best single modality. We close by discussing where each modality succeeds or struggles, what this suggests about how this kind of construct is actually expressed in reflective speech, and what would be needed to model it, and constructs like it, more effectively.

137. 【2609.37147】Improved Distributional Diffusion Models

链接:https://arxiv.org/abs/2609.37147

作者:Tommaso Martorella,Alexandre Galashov,Felix Krause,Stefan Andreas Baumann,Valentin De Bortoli,Arthur Gretton,Björn Ommer

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Distributional Diffusion Models, standard mean-prediction denoiser, Distributional Diffusion, scoring rule objective, scoring rule

备注:

点击查看摘要

Abstract:Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-$256^2$, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at this https URL.

138. 【2609.37141】MSTypography: Multi-character Semantic Typography via Balancing Word Legibility and Object Recognizability

链接:https://arxiv.org/abs/2609.37141

作者:Xinye Yang,Xinding Zhu,Kai Fang,Xinyi Ren,Mengjian Li,Bin Cao,Jiazhou Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:design technique, visual representation, typography, semantic meaning, Semantic typography

备注:

点击查看摘要

Abstract:Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Existing digital typography methods mainly focus on single-character scenarios. They suffer from a lack of legibility constraints and insufficient local deformation when extended to multi-character words, as the intricate structures among multiple characters are hardly preserved during the typography process. In this paper, we propose a global-to-local typography framework for multi-character scenarios. It performs mask-driven silhouette approximation at the global level, while semantic-guided refinement at the local level, with a culling step in between to improve efficiency. To preserve word legibility, we designed structural losses (including explicit collision constraints and implicit Jacobian singular value constraints) and an OCR constraint for character-level readability. To enhance the object recognizability, we leverage semantic guidance with diffusion priors, which drives the character glyph toward the target concept while preserving its structural integrity. To the best of our knowledge, this is the first multi-character semantic typography method that effectively balances word legibility and object recognizability. Evaluations on five representative languages (English, Chinese, Japanese, Korean, Arabic) demonstrate superiority over SOTA methods. Codes will be open-sourced.

139. 【2609.37139】aoFlowForge: Progressive Native Mesh Generation via Cascaded Flow Matching

链接:https://arxiv.org/abs/2609.37139

作者:Xianze Fang,Qiyuan Feng,Dongfang Sun,Yan Zhang,Xiuchao Wu,Jingnan Gao,Jiangjing Lyu,Chengfei Lyu,Gang Yu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:work of designers, printing and gaming, gaming industries, content generation technology, technology has significantly

备注:

点击查看摘要

Abstract:3D content generation technology has significantly advanced the work of designers, as well as the 3D printing and gaming industries. However, it remains difficult to produce lightweight, editable, and topologically clean artistic content that is directly production-ready. To achieve this, we present TaoFlowForge, an artistic mesh foundation model that generates production-ready meshes. Specifically, TaoFlowForge decomposes the mesh generation process into vertices generation and their connectivity prediction, i.e., edges. We formulate vertices generation as a two-stage coarse-to-fine process and incorporate several effective loss functions to further enhance its performance. In the connectivity prediction stage, we propose a simple yet effective method for estimating the connectivity affinity between vertices and additionally predict per-vertex normals, which determines the correct orientation of faces. Besides, we construct a large-scale dataset combining hand-crafted 3D assets with public high-quality topology datasets. Based on this, a carefully designed data curation pipeline is employed to filter the raw dataset, retaining only high-quality topology data for model training. Our model is trained on the combined dataset and tested on both out-of-distribution hand-crafted set of 3D assets and public datasets. Under image-conditioned generation, TaoFlowForge outperforms autoregressive methods and achieves state-of-the-art results among open-source mesh topology generators. We will release all the code and weights together with a portion of our test dataset.

140. 【2609.37135】Multi-Granularity Language-Guided Imitation Learning via Instruction Decomposition

链接:https://arxiv.org/abs/2609.37135

作者:Yi-Pei Chiu,Wei-Ta Chu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:important research domain, guide robot policy, research domain, robot policy learning, important research

备注:

点击查看摘要

Abstract:Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed method decomposes an overall task description into more fine-grained, concrete subtask-level language instructions, thereby enhancing learning efficiency and improving performance. We evaluate the proposed method in the setting of multi-task imitation learning and validate its effectiveness.

141. 【2609.37123】EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

链接:https://arxiv.org/abs/2609.37123

作者:Yaoxin Niu,Zhangquan Chen,Yang Zhang,Xiang An,Zhumei Wang,Chih-Ting Liao,Hongkun Cao,Ruqi Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:perception enables vision-language, distinguish subtle attributes, visual perception enables, enables vision-language models, Fine-grained visual perception

备注: 23 pages, 9 figures. Code: [this https URL](https://github.com/YXNiu/EviViT) Data: [this https URL](https://huggingface.co/datasets/YXNiu/Human-Search-Traces)

点击查看摘要

Abstract:Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.

142. 【2609.37115】NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting

链接:https://arxiv.org/abs/2609.37115

作者:Pratik Singh Bisht,Andreas Kolb

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Gaussian Splatting, revisit the role, show that limited, Neural Residual Fields, neural residual field

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricting the splats' ability to model high-frequency directional effects, which is typically compensated by increasing the number of splats. We propose \emph{NRF-GS: Neural Residual Fields for Gaussian Splatting}, a hybrid representation that replaces per-splat SH-bases with a shared neural residual field. Each Gaussian encodes a compact set of appearance features and a lambertian base color, while a lightweight \emph{global scene-level MLP} predicts view-dependent residuals conditioned on viewing direction, distance, and per-splat features. This formulation enhances directional reflectance modeling by combining diffuse per-splat reflectance representations with a shared global function for high-frequency details, enabling both higher expressiveness and parameter sharing across splats. Our key insight is that by accurately capturing high-frequency directional reflectance, especially in specular regions, the GS-representation becomes more expressive, reducing the need for geometrically redundant splats. As a result, NRF-GS achieves comparable or better rendering quality while reducing the number of Gaussians by up to 50\%, and produces visibly improved specular and high-frequency details.

143. 【2609.37107】Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

链接:https://arxiv.org/abs/2609.37107

作者:Rajit Rajpal,Shahbuland Matiana,Liew Wei Pyn,Anmol Agarwal,Ryan Craig,Andrew Lapp,Mithun Hunsur,Sami BuGhanem,Scottie Fox,Aaron Sanders Carson Poole,Irene Park,Dave Rossi,Spencer Frazier,Louis Castricato

类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:real-time diffusion world, interactive video generation, diffusion world model, interactive world models, generation on consumer-grade

备注:

点击查看摘要

Abstract:We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours of diverse, control-aligned video game data across hundreds of games, and generates playable video conditioned on full keyboard and mouse input. The model includes two resolution variants that run across a wide spectrum of consumer hardware. To characterize this unique setting, we distinguish rendered FPS, latent FPS, and control rate. We describe the data pipeline, architecture, training methodology, and runtime system behind Waypoint 1.5. We evaluate interactivity through latency and throughput. Finally, we discuss the safety and ethics considerations unique to iWMs.

144. 【2609.37098】V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

链接:https://arxiv.org/abs/2609.37098

作者:Junwei You,Weizhe Tang,Can Wang,Yan Zhao,Jun Hua,Haotian Shi,Wei Zhang,Lin Wang,Bin Ran

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:providing valuable support, complement onboard sensing, Vehicle-infrastructure cooperation, providing valuable, cooperative driving methods

备注:

点击查看摘要

Abstract:Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of the current scene, while the future consequences of prospective driving actions are rarely modeled explicitly. This limits the ability of the planner to anticipate how its decisions may interact with the evolving traffic environment. To address this issue, we propose V2X-WAM, a cooperative world action model that tightly couples cooperative scene understanding, action generation, and future-world reasoning. V2X-WAM constructs a reliability-aware spatiotemporal representation from vehicle- and infrastructure-side observations, while compressing infrastructure information into a compact quantized message for efficient communication. Based on the resulting cooperative representation, a multimodal planner generates prospective trajectories, which explicitly condition future occupancy and dynamic-flow prediction. The predicted world consequences are then fed back to refine the planned trajectory, forming a closed interaction between action and future-world evolution. Experiments on a large-scale real-world cooperative driving dataset demonstrate that V2X-WAM consistently improves planning accuracy and safety over representative end-to-end cooperative driving methods, while achieving stronger future-world prediction and substantially lower communication overhead. Ablation studies further validate the effectiveness of the proposed design.

145. 【2609.37096】Why MLLMs Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack

链接:https://arxiv.org/abs/2609.37096

作者:Liwei Che,Yihao Quan,Sen Fang,Hongyi Wang,Ranjay Krishna,Ruixiang Tang,Vladimir Pavlovic

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, remain poorly understood

备注:

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subsequent counting aggregation process, where representation separation rapidly diminishes as numerosity increases due to attention compression. Identifying and formalizing these twin bottlenecks constitutes our first major contribution. To overcome them, we propose ConvStack, a lightweight architecture that operates directly in the visual token space to explicitly aggregate and inject local spatial structures via zero-initialized residual connections. By explicitly addressing the individuation bottleneck, ConvStack provides unambiguous geometric evidence for downstream aggregation. Remarkably, by fine-tuning exclusively on counting tasks, the model achieves substantial improvements in dense object counting and broader spatial understanding benchmarks, without compromising on general visual capabilities.

146. 【2609.37090】ask-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

链接:https://arxiv.org/abs/2609.37090

作者:Luning Pang,Cheng Yuan,Jiawei Shao,Mingtao Huang,Yuan Shen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large multimodal models, support diverse visual, diverse visual understanding, support diverse, Large multimodal

备注: 13 pages. Submitted to IEEE Transactions on Mobile Computing

点击查看摘要

Abstract:Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.

147. 【2609.37089】Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

链接:https://arxiv.org/abs/2609.37089

作者:Kerui Ren,Yingxiang Xu,Kaiwen Song,Linning Xu,Bo Dai,Mulin Yu,Tao Lu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Real-world videos provide, videos provide rich, requires visually aligned, provide rich demonstrations, skills requires visually

备注: Project page: [this https URL](https://real2gym.github.io/)

点击查看摘要

Abstract:Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.

148. 【2609.37080】LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation

链接:https://arxiv.org/abs/2609.37080

作者:Zhengqiang Zhang,Lingchen Sun,Rongyuan Wu,Qiaosi Yi,Xiangtao Kong,Chaodong Xiao,Lei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Latent Diffusion Models, Latent Diffusion, latent space, Diffusion Models, diffusion model

备注: Accepted by NIPS 2026. More info can be found in [this https URL](https://github.com/PolyU-VCLab/LDMisAE)

点击查看摘要

Abstract:Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.

149. 【2609.37071】Context without Commitment: Robust Dense Correspondence under Non-Rigid Deformation

链接:https://arxiv.org/abs/2609.37071

作者:Yuzhen He,Sara Homscheid

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:deforming source surface, point-cloud registration aims, aims to find, Non-rigid point-cloud registration, full target cloud

备注:

点击查看摘要

Abstract:Non-rigid point-cloud registration aims to find the corresponding target point for each point on a deforming source surface. Point-level matching keeps the full target cloud available, but correspondence becomes ambiguous when different regions have similar local geometry. Regional or coarse-to-fine methods provide broader spatial context, but an incorrect regional match can exclude the correct correspondence before dense matching. We propose CoCo-Reg, which uses regional patches to enrich dense point features without allowing patch predictions to restrict the final point-level search. CoCo-Reg constructs farthest-point-sampled patches, exchanges geometric information within and between source and target, supervises patch similarity using identity-corrected point overlap, and projects the resulting regional information back to dense point features. The final registration stage still scores the full target cloud before global point-level candidate selection. On 726 held-out ModelNet10 objects across nine deformation levels, two established learning-based baselines obtain mean correspondence errors of 0.1993 and 0.1921, whereas CoCo-Reg obtains 0.0547. Relative to its point-level baseline, this is a 72.6\% reduction. CoCo-Reg achieves lower correspondence error on 92.3\% of paired test objects and reduces the mean fraction of points with error above 0.1 from 47.3\% to 17.3\%. Chamfer distance and HD95 decrease in the same direction, and CoCo-Reg remains lower across all tested deformation levels. These results support using regional context for dense non-rigid correspondence without imposing a hard patch-level restriction on the final search. Because evaluation uses one checkpoint per method, the reported gains characterize the complete systems rather than the isolated causal contribution of an individual component. Code will be made publicly available.

150. 【2609.37055】Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

链接:https://arxiv.org/abs/2609.37055

作者:Zhenyu Liu,Zhangquan Chen,Keyi Chen,Mingze Sun,Xiang An,Haodong Jing,Ruqi Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:spatially grounded settings, Vision-language models, increasingly operate, grounded settings, operate in embodied

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at this https URL.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.37055 [cs.CV]

(or
arXiv:2609.37055v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.37055

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
151. 【2609.37052】OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models

链接:https://arxiv.org/abs/2609.37052

作者:Yuchen Deng,Zidang Cai,Feidiao Yang,Yufei Wang,Jie Wang,Hai-Tao Zheng,Yuxing Han

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Omnimodal large language, large language models, temporally interleaved token, Omnimodal large, interleaved token sequences

备注:

点击查看摘要

Abstract:Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower's retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.

152. 【2609.37048】NHO: A Neural Hamiltonian Operator for Anchor-based Region Localization and Dense Correspondance

链接:https://arxiv.org/abs/2609.37048

作者:Jing Li,Yawei Luo,Xiangze Meng,Ying Li,Tieru Wu,Rui Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:anchors requires identifying, sparse anchors requires, recovering dense correspondences, partial shape, sparse anchors

备注:

点击查看摘要

Abstract:Non-rigid partial-to-full shape correspondence from sparse anchors requires identifying the corresponding region on the full surface and recovering dense correspondences between the partial shape and that region. We present NHO, which combines sparse anchors with the intrinsic geometry of the partial shape to learn a neural Hamiltonian operator whose localized eigenspace encodes both the region support and intrinsic coordinates for dense correspondence. NHO parameterizes the Hamiltonian potential as an intrinsic neural field and optimizes it using anchor evidence together with spectral and geometric constraints. To resolve the spatial ambiguity left by sparse anchors, we introduce reciprocal refinement between operator estimation and correspondence recovery. At each round, the current eigenspace provides spectral coordinates and restricts matching to its induced support, while geometrically reliable correspondences provide additional evidence for updating the potential. After refinement, aggregated eigenfunction energy yields the final localization, and the recovered map initializes dense correspondence refinement. Experiments demonstrate competitive accuracy on both tasks and robustness to uniform scaling and rotation.

153. 【2609.37047】Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks

链接:https://arxiv.org/abs/2609.37047

作者:Aidin Attar,Eleonora Cicciarella,Michele Rossi

类目:Neural and Evolutionary Computing (cs.NE); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:process static images, spiking neural network, design to process, spiking neural, process static

备注: 22 pages. Submitted to Neurocomputing. Code available at [this https URL](https://github.com/aidinattar/multi-depth-temporal-fusion-snn)

点击查看摘要

Abstract:We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at this http URL temporal-fusion-snn.

154. 【2609.37046】Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving

链接:https://arxiv.org/abs/2609.37046

作者:Katharina Winter,Stefan Englmeier,Fabian B. Flohr

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains insufficiently characterized, recover dynamic physical, autonomous-driving systems, insufficiently characterized, ability to recover

备注: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three tasks: surrounding-agent speed, current ego speed, and short-horizon future ego-speed proposal. On nuScenes, we evaluate open-weight general-purpose and PhysicalAI VLMs, together with the driving-oriented Alpamayo-1.5 Vision-Language-Action model, using multiple input and output formulations. We combine verbal evaluation with temporal perturbations, counterfactual ego-speed hints and linear probes of hidden representations. The tasks exhibit distinct failure modes. Surrounding-agent speed is weakly encoded in an agent-specific form, whereas current ego speed is often internally accessible but poorly verbalized: continuous probes achieve 4.7-5.8 km/h MAE compared with 10.2-16.8 km/h MAE for verbal outputs. Multiple frames provide inconsistent verbal gains to single frame inputs, and frame order is rarely exploited. Under non-optimized QLoRA, task-specific adaptation improves both task-relevant latent speed representations and verbal readout, but continuous surrounding-agent speed estimation remains weak, while most future-speed gains survive frame shuffling, indicating limited temporal grounding. Driving specialized Alpamayo-1.5 shows stronger latent representations for surrounding-agent and future ego speed, while current ego-speed decodability is comparable and substantial probe-verbal gaps remain. Thus, driving specialization can strengthen motion representations but does not guarantee stronger encoding across both scene and ego states or reliable readout. The results show that plausible planning outputs do not necessarily imply reliable recovery or temporal grounding of the underlying dynamic state.

155. 【2609.37042】GleanVID: Complementary Token Selection for Efficient Video Large Language Models

链接:https://arxiv.org/abs/2609.37042

作者:Shuo Yang,Changbai Li,Rui Tang,Xinyu Zhao,Linlin Yang,Baochang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Video Large Language, Large Language Models, Large Language, achieved strong video, strong video understanding

备注:

点击查看摘要

Abstract:Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.

156. 【2609.37038】NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters

链接:https://arxiv.org/abs/2609.37038

作者:Haoran Xu,Xingzhuo Guo,Yuchen Zhang,Jincheng Zhong,Jianmin Wang,Mingsheng Long

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:strong spatiotemporal variability, demands accurate short-term, nowcasting demands accurate, accurate short-term forecasts, Precipitation nowcasting demands

备注: 28 pages, 11 figures

点击查看摘要

Abstract:Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.

157. 【2609.37031】UniBuild: Unified Building Mapping From Multi-Source Optical Remote Sensing Imagery With Detail Decoding and Geometry Regularization

链接:https://arxiv.org/abs/2609.37031

作者:Wei Huang,Chenying Liu,Yilei Shi,Xiao Xiang Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:optical remote sensing, RGB optical, multi-source RGB optical, RGB optical datasets, remote sensing

备注:

点击查看摘要

Abstract:Building extraction from optical remote sensing (RS) imagery is fundamental to urban mapping, yet existing methods are often dataset-specific and generalize poorly to unseen domains. Their practical use is also limited by insufficient detail recovery and weak geometric regularization, leading to blurred boundaries, irregular shapes, and merged adjacent buildings. To address these issues, we propose UniBuild, a unified building extraction framework for multi-source RGB optical RS imagery. First, a unified multi-dataset training scheme is constructed over heterogeneous RGB optical datasets to learn transferable building representations across sensors and resolutions. Second, a novel detail-preserving HR-DPT decoder is designed to integrate high-level semantic features with high-resolution spatial features, enhancing building detail recovery. Third, geometry-aware regularization is introduced through a structure-tensor-based direction-aware loss for boundary direction consistency and a saddle-aware loss for suppressing false activations in narrow inter-building gaps under low-resolution conditions. We train and evaluate UniBuild on multi-source RGB optical datasets, including 10 public high-resolution datasets and two self-collected low-resolution datasets. Experiments show that UniBuild consistently improves building-region accuracy, boundary sharpness, and adjacent-building separation across diverse datasets. It also generalizes well to unseen domains and supports practical building extraction from RGB optical RS imagery up to 10\,m resolution. The predicted masks can be further converted into GIS-compatible building footprints through simple polygonization. The trained model and inference code are released at this https URL.

158. 【2609.37030】MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos

链接:https://arxiv.org/abs/2609.37030

作者:Jiahao Zhan,Yongrui Ma,Qunliang Xing,Xuanyu Zhang,Jingqi Tong,Junlin Li,Li zhang,Shijie Zhao,Tianfan Xue

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:video generation models, exhibit obvious motion, generation models, rapid progress, exhibit obvious

备注:

点击查看摘要

Abstract:Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.

159. 【2609.37016】Back2Struct: Making Structured Images Editable Again

链接:https://arxiv.org/abs/2609.37016

作者:Pengyu Yan,Yixin Wu,Yunjie Tian,David Doermann

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:inherently symbolic, compactly represented, structured images editable, Structured images, editable format

备注:

点击查看摘要

Abstract:Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designers who wish to incorporate modified versions of existing graphic content into new materials without manually reconstructing it. In this study, we presentBack2Struct, which "makes structured images editable again" by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic, Back2Struct predicts semantically object-level SVG / XML code that explicitly encodes text, shapes, topology, and layout, rather than performing low-level pixel vectorization. The generated code can be seamlessly imported into tools such as PowerPoint, allowing users to edit, refine, restyle, and reuse graphic content while preserving structural fidelity. Beyond supervised fine-tuning on ground-truth SVG token sequences, we further optimize Back2Struct with reward-based learning to better match deployment-time requirements: the output should be syntactically valid, properly concise, and visually faithful to the input diagram. Specifically, we design a composite reward that jointly encourages SVG / XML compilability, length consistency with the reference code, and structural or semantic similarity between the generated and ground-truth graphics. These complementary signals guide the model to produce SVGs that are not only closer to the training distribution, but also more complete, editable, and renderable in practice. Experiments show that Back2Struct improves accuracy, editability, validity, and user alignment over baselines. Dataset and code are available at: this http URL

160. 【2609.37015】RBF-GNN: Rational Basis Functions for Pseudo-Coordinate based Graph Convolutions

链接:https://arxiv.org/abs/2609.37015

作者:Paweł Batorski,Abtin Pourhadi,Paul Swoboda

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:spatial inductive bias, powerful spatial inductive, account Euclidean, pseudo-coordinate based graph, graph neural network

备注:

点击查看摘要

Abstract:We propose RBF-GNN, a new pseudo-coordinate based graph neural network architecture that takes into account Euclidean, spherical or angular coordinates and uses them to induce a powerful spatial inductive bias. Similar in architecture to SplineCNN, we improve upon the latter by replacing the less efficient sparse-activation based B-splines whose number grows exponentially with dimension by rational Padé basis functions. For effective training we propose a spline-subspace initialization and a variance-preserving weight rescaling. Experimentally, we evaluate on a number of popular neural network architectures that use SplineCNNs. We replace only the SplineCNNs with RBF-GNN. We achieve improved results, including on semantic keypoint matching, shape matching, event based camera computer vision tasks. We will make our implementation publicly available upon acceptance of the paper.

161. 【2609.37013】Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction

链接:https://arxiv.org/abs/2609.37013

作者:Thomas Goudemant,Benjamin Francesconi,Marjorie Bellizzi,Adrien Dorise

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:support emergency response, Rapid assessment, emergency response, natural disasters, disasters is essential

备注: 8 pages. Accepted at OBPDC 2026 (International Workshop on On-Board Payload Data Compression), Barcelona, October 2026

点击查看摘要

Abstract:Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal building damage assessment pipeline built on a siamese detector derived from YOLOX, designed to compress information at both ends of the ground/space link. On the ground, pre-disaster reference images are encoded into a compact latent space -- compressed by up to a factor of 64 -- and uplinked to the satellite. On board, this reference is compared with a fresh post-disaster acquisition so that the downlink carries only actionable object-level products, bounding boxes and damage classes, instead of full scenes. This cuts the data exchanged in both directions, while on xBD the strongly compressed reference still preserves most of the detection performance. Because on-board acquisitions suffer from residual pre/post co-registration errors, we introduce a latent-space shift estimation and correction module that regresses the global offset from the coarse feature level and realigns the post-disaster features before fusion. It substantially improves robustness to de-registration -- especially under large shifts, where fusion-only variants collapse -- while also raising nominal accuracy and remaining compatible with the strongest compression. We finally port the pipeline to two embedded targets, a Xilinx Versal VCK190 and an NVIDIA Jetson AGX Orin, and report hardware performance (latency, throughput, power efficiency). The core detector and its compression port cleanly to both, but the operators needed for long-range robustness survive only on the Jetson GPU, whereas the Versal DPU does not.

Comments:
8 pages. Accepted at OBPDC 2026 (International Workshop on On-Board Payload Data Compression), Barcelona, October 2026

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

ACMclasses:
I.4.8; I.2.10

Cite as:
arXiv:2609.37013 [cs.CV]

(or
arXiv:2609.37013v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.37013

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
162. 【2609.37004】World2Motion: Turning Video World Models into 3D Human Motion Generators

链接:https://arxiv.org/abs/2609.37004

作者:Tu Fangyuan,Xiangyue Zhang,Yiyi Cai,Yichen Peng,Kunhang Li,Bo Zheng,Zhixiang Wang,Kaipeng Zhang,Erwin Wu,Haoran Xie,Haiyang Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:motion, generates scene-aware, framework that generates, single image, video

备注: 15 pages, 6 figures

点击查看摘要

Abstract:We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3$\times$ faster inference.

163. 【2609.37003】VesselBench-800K: A Large-scale Perception Benchmark for Multimodal Vessel Detection, Counting, and Density Estimation

链接:https://arxiv.org/abs/2609.37003

作者:Danfeng Hong,Chenyu Li,Jocelyn Chanussot

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:environmental protection, space is crucial, wide range, monitoring to environmental, Vessel perception

备注:

点击查看摘要

Abstract:Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general object detection tasks in optical remote sensing (RS) images. Relying solely on single-modality optical RS images proves inadequate for effectively perceiving vessel objects in complex maritime scenarios, where ever-changing weather conditions (e.g., clouds and rain), the need for day-and-night coverage, and the inherent limitations of a single imaging modality pose significant challenges. To fill this gap, we introduce VesselBench-800K, the largest-to-date benchmark dataset on a global scale for vessel perception in multimodal RS images. As its name suggests, VesselBench-800K comprises 800,000 images, each at a resolution of 512x512 pixels, specifically curated for vessel perception tasks such as detection, counting, and density estimation. These multimodal image pairs (i.e., optical, SAR) are collected from diverse platforms, sensors, scenes, shooting heights, and synthetic sources, spanning spatial resolutions from 4.5m to 0.1m. Furthermore, we evaluate numerous state-of-the-art detection, counting, and density estimation models on VesselBench-800K through both qualitative and quantitative comparisons. By revealing previously unrecognized cues, this dataset holds immense potential to significantly advance our understanding of marine traffic. Our VesselBench dataset will be publicly available at this https URL to support and contribute to community development.

164. 【2609.37002】Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom

链接:https://arxiv.org/abs/2609.37002

作者:Xijia Tao,Yihua Teng,Xinyu Fu,Cheng Gong,Ziru Liu,Xudong Xie,Rui Liu,Lingpeng Kong

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:High-resolution visual question, visual question answering, spatially localized evidence, localized evidence needed, question answering

备注:

点击查看摘要

Abstract:High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.

165. 【2609.37001】Parameterized Stripe Attention for Efficient Video Generation

链接:https://arxiv.org/abs/2609.37001

作者:Xingyu Jia,Baole Ai,Ang Wang,Kang Zhao,Yong Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Diffusion Transformers, substantial inference latency, computationally expensive full, expensive full spatio-temporal, full spatio-temporal attention

备注:

点击查看摘要

Abstract:Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57$\times$ and 1.37$\times$ end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.

166. 【2609.36995】Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

链接:https://arxiv.org/abs/2609.36995

作者:Xingtong Ge,Yutong Wang,Lunjie Zhu,Haitao Lin,Fangyu Lin,Yushi Huang,Xin Zhang,Yi Zhang,Yu Liu,Jun Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)

关键词:Few-step streaming audio, video generation requires, standard training recipes, training recipes face, streaming audio

备注: under review

点击查看摘要

Abstract:Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step $1664\times960$ generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: this https URL

167. 【2609.36980】UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching

链接:https://arxiv.org/abs/2609.36980

作者:Jiajun Le,Yifan Lu,Zizhuo Li,Lei Cao,Junjun Jiang,Jiayi Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dense token-level matching, Transport Path Router, dense token-level, matching, token-level matching

备注: 18 pages, 5 figures

点击查看摘要

Abstract:Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67$\times$ faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2$\times$ end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at this https URL.

168. 【2609.36975】A Dual-Track Curation-and-Classification Framework for Resolving Ground-Truth Label Noise in Operational Sentinel-2 Wheat Area Estimation

链接:https://arxiv.org/abs/2609.36975

作者:Kasimali Agharia,Ujjwal Kumar Gupta

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remotely sensed classification, sensed classification products, estimation of wheat-cultivated, persistently constrained, record-keeping and remotely

备注: 21 pages, 4 figures, 5 tables. Preprint. Not yet peer-reviewed

点击查看摘要

Abstract:Operational estimation of wheat-cultivated area is persistently constrained by discordance between administrative record-keeping and remotely sensed classification products. We address this administrative reference discordance for the 2022 Rabi season in Patiala district, Punjab, India, using a thirteen-timestep Sentinel-2 NDVI time series. A curated 849-sample reference dataset, developed through an iterative rule-based bootstrapping procedure, underpins both a feature sensitivity analysis and an operational classifier. Feature sensitivity independently assessed via Cohen's d and gradient-boosted information gain converges on the February-to-March grain-fill window as most discriminative. Four classifiers (1D-CNN, LSTM, hybrid CNN-LSTM, and XGBoost) were benchmarked on an identical 679/170 sample split. XGBoost achieved the highest overall accuracy (78.82%) against deep-learning baselines (64-66%), consistent with tree-based ensembles' favourable parameter-to-sample ratio in low-sample regimes. At full-population deployment across 36.25 million valid district pixels, the operational classifier attained 86.31% precision and 71.05% recall. The predicted wheat extent deviated by only +2.99% from the official tabular target, whereas the government's spatial reference mask exhibited a +25.11% positive area bias against the identical target. This asymmetry indicates that a classifier trained on an auditor-curated reference set reconciles more closely with the official tabular area than the spatial product conventionally used to validate it. We present this dual-track curation-and-classification framework as a methodological reference for crop-area reconciliation in label-noisy administrative settings.

169. 【2609.36971】Structured Visual Target Learning For Cross-Subject eeg-to-image retrieval

链接:https://arxiv.org/abs/2609.36971

作者:Salini Yadav,Taveena Lotey,Mickaël Coustaty,Pravendra Singh,Partha Pratim Roy

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:visual embedding space, retrieval requires, neural represen, tation trained, requires a neural

备注:

点击查看摘要

Abstract:Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the perspective of the visual target. Our approach preserves the spatial information of the Perception Encoder, converts its patch grid into a compact set of learned visual views, and aggregates them for each image with a block-structured, content-dependent router. The target is learned jointly with the EEG encoder through contrastive learning with MMD regularization across source subjects. For deployment, we propose a training-free representation refinement that aligns frozen embeddings without updating either encoder. Under leave- one-subject-out evaluation on THINGS-EEG2, the structured target achieves 35.3%/65.6% Top-1/Top-5 accuracy, the best among com- pared methods. Refinement raises this to 48.1%/77.1%, an 18.5% Top-1 gain over the strongest compared method, improving all ten held-out subjects.

170. 【2609.36969】Prior-Driven Enhancements in 3D Gaussian Splatting: Normals and Depths Regularization

链接:https://arxiv.org/abs/2609.36969

作者:Gyeonggwan Lee,Seunghwan Hong,Junghun Suh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:offering high efficiency, Gaussian Splatting, offering high, high efficiency, efficiency and excellent

备注: 7 pages, 2 figures, 1 table. Oral presentation at ISPRS Geospatial Week 2025 (Dubai). Project page: [this https URL](https://gandanlee.github.io/pdigs/) Code: [this https URL](https://github.com/gandanlee/pdigs)

点击查看摘要

Abstract:3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properties, it can suffer from geometric inaccuracies and visual artifacts, particularly in complex scenes. To address these challenges, we propose an improved 3DGS approach that regularizes the optimization process by integrating geometric priors, including surface normals and dense depth information. Surface normal regularization improves geometric consistency by aligning Gaussian covariance with local surface structures, while dense depth priors combined with an initial points from SfM enhance per-pixel depth estimation, increasing accuracy and reducing ambiguities. These enhancements enable robust handling of diverse and complex real-world scenarios, minimizing visual distortions and improving reconstruction quality across various environments. To validate our method, we evaluate it on challenging datasets, including street-view scenes and highly reflective environments, while testing it across multiple SfM pipelines. Our results demonstrate compatibility across diverse environments and highlight the robustness of our approach. Experimental findings further show that our method enhances geometric accuracy and visual quality, establishing a reliable solution for real-time 3D scene rendering in complex environments.

171. 【2609.36965】Chinese-Jev: Bringing System One Model to Chinese-Language Tasks

链接:https://arxiv.org/abs/2609.36965

作者:Zexiao Wang,Zihao Zhang,Xudong Wang,Pan Wang,Ziyi Ye,Haoyu Zhao,Zuxuan Wu,Shuicheng Yan

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:generative language models, open-ended responses, alternative to generative, generative language, Jev

备注: 10 pages, 6 figures

点击查看摘要

Abstract:System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at this https URL.

172. 【2609.36957】Beyond Readability: Evaluating Task Information Recoverability

链接:https://arxiv.org/abs/2609.36957

作者:Yiwei Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:task-information recoverability, target, visual, Direct, knowledge

备注:

点击查看摘要

Abstract:Direct visual readability and task-information recoverability are different quantities. Failure to decode a target from a fixed observation need not eliminate access to that target through another recovery route. We develop an evaluation perspective that makes the observation, query, target, and available knowledge explicit and measures the overlap between routes' success sets. For information available on the original visible surface under suitable imaging conditions, direct optical recovery reads the target from the image, optionally after restoration; entity-linked recovery uses residual visual evidence to identify the depicted entity and accesses its target through an entity--attribute relation in a specified knowledge resource. Such access can draw on stored knowledge or an external source. A controlled book-cover study instantiates external access with a fixed title--author catalog, comparing optical author recovery with visual title resolution and deterministic lookup under resolution degradation. Entity-linked successes persist across the tested vision--language models, revealing information access beyond the tested direct visual frontier despite substantial differences in absolute performance. A substantial optical-only region remains. These complementary outcomes show why visual degradation should be evaluated through the task information accessible along specified routes and knowledge resources, alongside direct readability.

173. 【2609.36940】DispFlow-GS: Displacement Flow Supervision with Motion Disentangling for Monocular Deformable 3D Gaussian Splatting

链接:https://arxiv.org/abs/2609.36940

作者:Thai Duy Nguyen,Haitian Zhang,Addison Lin Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate dynamic scene, temporally consistent representations, Accurate dynamic, rendered Gaussian flow, dynamic scene reconstruction

备注:

点击查看摘要

Abstract:Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are essential. Deformable 3D Gaussian Splatting (3DGS) models dynamic scenes through deformation fields, and recent methods incorporate motion supervision by aligning rendered Gaussian flow with optical flow. However, we find that such Gaussian-flow-based supervision provides only limited improvements in motion modeling. We identify a fundamental limitation of this supervision paradigm, namely a domain gap between rendered Gaussian flow and optical flow. To address this limitation, we propose a motion supervision framework built on Displacement Flow, which splats per-Gaussian 3D displacements onto the image plane to provide direct and stable optimization signals. We further disentangle scene motion from camera motion via intermediate-view rendering, enabling more reliable motion priors and targeted constraints on deformation and geometry. We also observe a discrepancy between motion fidelity and image-based evaluation, where improved motion awareness does not necessarily translate into better rendered image quality or higher image-based metric scores. Motivated by this mismatch, we introduce Deformation-Rendering Consistency (DRC), a motion-aware metric that measures the alignment between predicted deformation and rendering improvement. Experiments on dynamic scene benchmarks show substantial improvements in motion localization and motion--rendering consistency, reaching up to 39% and 6%, respectively, while image-based metrics change by only about 0.1%. These results confirm the observed mismatch between motion fidelity and image-based evaluation, demonstrating the significance of DRC for motion-aware evaluation.

174. 【2609.36937】WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation

链接:https://arxiv.org/abs/2609.36937

作者:Sangeyl Lee,Seunghyun Shin,Seungho Park,Wooseok Jeon,Hae-Gon Jeon

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Human image animation, Human image, image animation aims, aims to transfer, image animation

备注:

点击查看摘要

Abstract:Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.

175. 【2609.36929】SFE-VGGT: Source-Free VGGT Distillation for Event-Based Monocular Depth Estimation

链接:https://arxiv.org/abs/2609.36929

作者:Thai Duy Nguyen,Addison Lin Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent event-based depth, vision foundation models, Recent event-based, successfully transfer geometric, event-based depth estimation

备注:

点击查看摘要

Abstract:Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE-VGGT, a novel source-free framework that distills the geometric priors of VGGT to the event domain without any paired RGB observations. Our core idea is to reconstruct surrogate frames directly from the target event stream to act as a frozen geometric teacher, entirely eliminating the need for genuine source RGB data. Crucially, as these surrogate frames inherently yield imperfect and spatially varying supervision, directly distilling from them propagates artifacts. To resolve this, we introduce a novel reliability-aware distillation strategy. This includes Density-Aware Feature Distillation to emphasize informative event regions, and Confidence-Weighted Depth Distillation to dynamically regulate supervision based on relative teacher-student prediction confidence. Meanwhile, we propose a Cross-Frame Relational Consistency loss that enforces temporal geometric stability using reliable inter-frame correspondences, bypassing the need for temporally consistent teacher's depth. Extensive experiments demonstrate that, despite source-free, our SFE-VGGT closely matches the accuracy of RGB-dependent baselines under standard conditions and significantly surpasses them in challenging nighttime scenarios. Across MVSEC nighttime sequences, SFE-VGGT reduces the average 10 m depth error by 15.3% compared with EventVGGT. Moreover, our method exhibits robust zero-shot generalization across real-world datasets, proving that highly effective geometric priors can be transferred to event cameras using strictly source-free supervision.

176. 【2609.36924】rack-and-Complete: Learning Humanoid Skills from a Single Failed Human Video

链接:https://arxiv.org/abs/2609.36924

作者:Sarmad Idrees,Jongeun Choi

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:custom data collection, demands custom data, videos typically requires, data collection, typically requires

备注: 8 pages, 6 figures, 6 tables. Project website: [this https URL](https://tracc-humanoid.github.io)

点击查看摘要

Abstract:Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until the failure occurs, after which the task-completion reward guides the policy to learn the intended task goal without requiring a successful task trajectory. We evaluate our method on six in-the-wild failed human tasks from the Oops! dataset. Our experimental results demonstrate the effectiveness of the proposed approach for learning from failed attempts when no successful demonstration is available. Thus, these findings establish failed human videos as a viable source of supervision for humanoid skill learning.

177. 【2609.36918】Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation

链接:https://arxiv.org/abs/2609.36918

作者:Jiantao Lin,Meixi Chen,Yingjie Xu,Chenbo Fu,Leyi Wu,Hao Chen,Yinchuan Li,Ying-Cong Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains challenging due, coherent multi-part reasoning, image remains challenging, ambiguous boundaries, single image remains

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.

178. 【2609.36916】Representation Dynamics Reveal Semantic Saliency and Similarity for Visual Token Pruning in MLLMs

链接:https://arxiv.org/abs/2609.36916

作者:Weixuan Li,Zikun Zhou,Xinyi Zhuang,Xinyan Guo,Rui Tian,Chuyao Zhang,Lin Gao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multimodal large language, incur high inference, high inference latency, large language models, Multimodal large

备注: Preprint. 33 pages, 17 figures, 19 tables

点击查看摘要

Abstract:Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflect foreground saliency and semantic consistency remain insufficiently understood. We analyze visual token representation dynamics across encoder depth and uncover two findings. First, the relationship between token update magnitudes and foreground saliency is layer-dependent: large token updates concentrate on foreground regions in two depth intervals, separated by several sink-dominated layers at intermediate depths. Second, similarities between token update directions better distinguish same-class from different-class tokens than those between encoder output features. Building on these findings, we propose MSDG-Prune, a training-free method that uses update magnitudes and directions to preserve salient and diverse visual information. Specifically, we group tokens by update-direction similarity and use query-weighted saliency derived from update magnitudes across a chosen depth window for group-wise token pruning. Extensive experiments across four MLLMs demonstrate the effectiveness and generalizability of MSDG-Prune. On LLaVA-NeXT, it retains 91.9% of uncompressed performance on average with only 5.6% of visual tokens, while achieving a 7.8x prefilling speedup. Code is available at this https URL.

179. 【2609.36906】SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions

链接:https://arxiv.org/abs/2609.36906

作者:Sean Hardesty Lewis,Zuyi Guo,Benwang Chen,Zirui Liu,Hongyi Lin,Heye Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:require informative observations, partial observability require, observability require informative, sufficient supporting evidence, require informative

备注:

点击查看摘要

Abstract:Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores alone do not reveal which viewpoints justify a claim or where additional evidence should be acquired. We introduce SafeVantage, a vantage-aware semantic memory and active acquisition framework that retains each claim's supporting views, camera poses, and estimated target location, keeping positive support distinct from search coverage. A learned candidate-observability model uses claim-grounded geometry to predict target visibility at reachable viewpoints. These predictions guide view selection through expected reduction in terminal decision loss, accounting for travel cost and geometrically distinct corroboration. A calibrated head then combines support, spatial consistency, and coverage to produce Yes, No, or Abstain decisions. We evaluate SafeVantage on a category-presence benchmark spanning 232 unseen ProcTHOR houses and 7,424 paired episodes per method and action budget. Compared with validation-selected equal-budget baselines, SafeVantage achieves macro-F1 gains of 24.7% and 12.0% at eight and twelve actions, respectively, with lower risk and higher answer rates at both budgets and 31.7% less travel at eight actions. Equal-input HM3D experiments show lower selective risk under fixed observations, while controlled ScanNet interventions show that restoring supporting views improves downstream VLM answers. Ablations further support the contribution of candidate observability to decision quality and acquisition efficiency. Results demonstrate the value of claim-level viewpoint evidence for connecting semantic memory, active acquisition, and reliable decision-making. Code is available at this https URL

180. 【2609.36902】RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection

链接:https://arxiv.org/abs/2609.36902

作者:Wenbin Shen,Guoxuan Qin,Guangxu Yao,Baodong Wang,Yuanbo Rui,Zhongjie Ba,Zhichao Lian

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Existing multimodal fake, assist detection, information to assist, Existing multimodal, Evidence Retrieval Framework

备注:

点击查看摘要

Abstract:Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.

181. 【2609.36894】DiffReID: Discriminative Diffusion Model for Object Re-Identification

链接:https://arxiv.org/abs/2609.36894

作者:Yingquan Wang,Pingping Zhang,Dong Wang,Huchuan Lu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image processing task, fundamental image processing, processing task, aims to retrieve, non-overlapping cameras

备注: Accepted by TIP2026. More modifications can be performed

点击查看摘要

Abstract:As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbf{DiffReID} for object ReID. It leverages a discriminative diffusion model to gradually learn identity-aware distributions and generate identity-invariant features. More specifically, with the Contrastive Language-Image Pre-training (CLIP) model, we first obtain identity-aware text features by prompt tuning. Then, we propose a Vision-guided Noise Generator (VNG) to initialize probabilistic noises and gradually corrupt identity-aware text features. Afterwards, we take visual features as conditions and propose a Light Weight Denoiser (LWD) to denoise the corrupted text features step-by-step for identity-aware distribution learning. To obtain discriminative features, we further generate identity-invariant guided features from randomly sampling visual-guided noises. Finally, we propose a Mutual Enhancement Constraint (MEC) to facilitate mutual learning between visual features and guided features to enhance the representation robustness and discrimination. Extensive experiments on five object ReID benchmarks demonstrate that our method shows better results than most state-of-the-art methods. The source code is available at this https URL.

182. 【2609.36891】ProGuT: Label-Efficient Panoptic Segmentation for Forest Scenes

链接:https://arxiv.org/abs/2609.36891

作者:Pankaj Deoli,Karsten Berns

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:usable stuff maps, approaches produce usable, produce usable stuff, near-zero thing quality, Prototype Guided Training

备注:

点击查看摘要

Abstract:Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approaches produce usable stuff maps but near-zero thing quality. Depth or flow-based instance discovery methods needs sensors that are not always available. We present ProGuT (Prototype Guided Training), which produces panoptic pseudo-labels without per-image training masks, needing only unlabeled images and one-time cluster-to-class mapping. ProGuT clusters CLIP patch features, then recovers trunk instances through multiscale geometric prior that falsifies non-trunk structures via structure-tensor. This is cheap compared to depth, flow or class-supervision methods to create pseudo labels. These are then used for downstream tasks which we evaluate against other unsupervised baselines. ProGuT achieves a Panoptic Quality (PQ) of 65.2 on Our-forest dataset (2.6x improvement over the initial pseudo-label quality) and reaches 65.9 mIoU on Freiburg Forest, outperforming unsupervised baselines like PiCIE (45.3 IoU) and STEGO(57.6IoU). Additionally, ProGuT outperforms existing unsupervised methods for class-agnostic trunk instance benchmark.

183. 【2609.36882】Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images

链接:https://arxiv.org/abs/2609.36882

作者:Junhee Lee,Donghyeon Jeon,Taeoh Kim,Beomyoung Kim,MyeongAh Cho

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Localizing AI-edited regions, interpretable forensic analysis, remains challenging due, spatially distributed artifacts, Localizing AI-edited

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models' ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose ReGFLoW, a Reconstruction-Guided Fake Localization framework under Weak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image-level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments demonstrate that ReGFLoW achieves stronger out-of-domain generalization than fully supervised learning baselines.

184. 【2609.36875】S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

链接:https://arxiv.org/abs/2609.36875

作者:Jingdong Zhang,Xin Li,Jan Kautz,Wenping Wang,Chris Choy

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate instance segmentation, Accurate instance, autonomous driving, important for downstream, downstream applications

备注:

点击查看摘要

Abstract:Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.

185. 【2609.36866】S2T-Unet: A Structure-to-Style Framework for Inter-Modality MRI Translation

链接:https://arxiv.org/abs/2609.36866

作者:Yichao Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:preserving clinically relevant, clinically relevant anatomical, Inter-modality MRI translation, synthesize missing MRI, missing MRI modalities

备注:

点击查看摘要

Abstract:Inter-modality MRI translation aims to synthesize missing MRI modalities from available acquisitions, reducing the need for additional scanning while preserving clinically relevant anatomical information. However, existing image translation methods often learn intensity mappings without explicitly separating modality-invariant structural information from modality-specific appearance, which may lead to structural information loss or unrealistic image details. In this work, we propose S2T-Unet, a structure-to-style framework that explicitly models these two aspects. Specifically, vector quantization is introduced at the lower-level bottleneck to encode modality-invariant structural information using a learned discrete codebook. At higher levels, a modality transformation module uses decoder features to condition and transform encoder representations toward the target modality, thereby recovering modality-specific intensity and contrast information. Experiments on the IXI multi-contrast MRI dataset across four translation tasks demonstrate that S2T-Unet is comparable or outperform with state-of-art method.

186. 【2609.36852】Socialality Anchors: Towards Group-bounded Trajectory Prediction

链接:https://arxiv.org/abs/2609.36852

作者:Ziqian Zou,Conghao Wong,Qinmu Peng,Xinge You

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:understanding human behavior, human behavior patterns, key component, component for understanding, behavior patterns

备注:

点击查看摘要

Abstract:Trajectory prediction is a key component for understanding human behavior patterns in dynamic scenes. Researchers have devoted substantial efforts to modeling social interactions, especially group-wise interactions, since group membership often reflects shared intention, coordinated motion, and stable mutual adaptation, thus providing a persistent and semantically meaningful social prior for forecasting. However, existing group modeling methods may rely on a fixed threshold and infer groups mainly from agents' relative positions within the observation window, overlooking the fact that grouping rules should be agent-specific, temporally coherent, and context-adaptive across diverse personalities, culturalities, and evolving interaction contexts. Inspired by human social perception that alternates between interpersonal distance in boundary-sensitive situations and relative speed consistency in dynamic interactions, we propose Socialality, a human-inspired trajectory prediction framework with interpretable Socialality anchors and an extended grouping window for stable, context-aware grouping inference. Concretely, Socialality introduces a duo-scalar-controlled grouping kernel Socialality that jointly leverages historical observations and short-term future trajectory previews to learn agent-specific grouping rules, and employs a group-wise perception mechanism to model in-group and out-of-group interactions in an intuitive and explainable manner. Furthermore, we conduct extensive experiments on standard benchmarks to demonstrate the performance gains of Socialality, and provide qualitative analyses and statistical studies of anchor distributions to verify the interpretability and stability of the proposed Socialality anchors.

187. 【2609.36851】RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts

链接:https://arxiv.org/abs/2609.36851

作者:Hongbin Lin,Chaoda Zheng,Yiming Yang,Xiangyu Li,Shijia Chen,Jinhao Deng,Kangjie Chen,Dongbin Zhang,Jie Feng,Yu Zhang,Xianming Liu,Shuguang Cui,Boyang Wang,Zhen Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous driving policies, closed-loop real-world deployment, autonomous driving, leading to causal, real-world deployment

备注: Project page: [this https URL](https://hongbin98.github.io/RoXDrive/;) Github: [this https URL](https://github.com/Hongbin98/RoXDrive)

点击查看摘要

Abstract:End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.

188. 【2609.36850】Rethinking Multimodal Fake News Detection in the Generative AI Era

链接:https://arxiv.org/abs/2609.36850

作者:Wenbin Shen,Guoxuan Qin,Guangxu Yao,Baodong Wang,Yuanbo Rui,Zhichao Lian

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:simply manipulated material, increasingly entering, entering the production, production and dissemination, manually fabricated

备注:

点击查看摘要

Abstract:Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news event is true. To bridge the separation between these tasks in data and evaluation, we construct Weibo26, a multimodal fake news detection dataset for generative-content scenarios. On this basis, we propose the Generativity-Aware Hierarchical Reasoning (GAHR) framework, which combines global judgment with local correction so that generativity information participates in news-veracity reasoning. Experiments on multiple existing fake news detection benchmarks and Weibo26 show that GAHR achieves competitive veracity-detection performance while effectively identifying generative content.

189. 【2609.36844】GlassFormer: Learning Real-time Glass Segmentation using Radar-Depth Fusion

链接:https://arxiv.org/abs/2609.36844

作者:Suhani Grover,Astik Srivastava,Viswas Dinesh,Avinash Sharma,K. Madhava Krishna

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:persistent failure case, robotic perception, ubiquitous in built, persistent failure, failure case

备注: Accepted for presentation at IEEE IROS 2026. Code available at [this https URL](https://github.com/Suhani92/GlassFormer)

点击查看摘要

Abstract:Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.

190. 【2609.36842】Does the VGGT Family Need All Its Layers?

链接:https://arxiv.org/abs/2609.36842

作者:Fengyi Zhang,Holger Caesar,Xiangyu Sun,Zheng Zhang,Zi Huang,Yadan Luo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:feed-forward geometry model, poses and dense, feed-forward geometry, needed to preserve, preserve both camera

备注:

点击查看摘要

Abstract:Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $\pi^3$, and VGGT-$\Omega$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from $O(L^4)$ to $O(L^2)$, where $L$ is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: this https URL

191. 【2609.36838】On-Policy Visual Evidence Distillation

链接:https://arxiv.org/abs/2609.36838

作者:Shaohang Wei,Feifan Song,Guangyue Peng,Wenhao Yu,Wei Li,Wen Luo,Yang Xu,Yufan Shen,Luke Mao,Yang Du,Asher Qin,Houfeng Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:agents solve problems, Visual agents solve, image operations, solve problems, problems by interleaving

备注: 44 pages, including appendices. Project page: [this https URL](https://sylvain-wei.github.io/ReVuE/) . Code: [this https URL](https://github.com/sylvain-wei/ReVuE)

点击查看摘要

Abstract:Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at this https URL

192. 【2609.36837】You Cannot Recover What Was Never Measured: Quantifying the Information Ceiling of Ultra-Low-Field MRI Super-Resolution

链接:https://arxiv.org/abs/2609.36837

作者:Prathamesh Pradeep Khole,Shreya Handa,Utkarsh Gupta,Razvan Marinescu

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Medical Physics (physics.med-ph)

关键词:degrading high-field images, Generative super-resolution models, synthetically degrading high-field, Generative super-resolution, MRI into images

备注: 24 pages, 10 tables, 8 figures

点击查看摘要

Abstract:Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated on tests whose correct answer is known in advance, we find that, judged over the whole brain, real 64 mT scans carry structure specific to the individual only down to approximately 3 to 4 mm half-pitch in plane, and coarser still through plane. Standard synthetic degradations preserve subject information roughly 1 mm beyond this ceiling, so models trained and benchmarked on synthetic pairs are evaluated on information that real scanners never record. We then test trained diffusion models and a publicly released external model on real paired acquisitions; 24 trained runs of five architectures (GAN, diffusion, and transformer families) give the coverage of the audit. On every subject where faithfulness can be measured, fine output detail is no more correlated with the subject's own 3T scan than with a stranger's, while sample-variance uncertainty does not distinguish fabricated structure from reconstruction difficulty. Because PSNR and SSIM score resemblance to a reference rather than whether detail belongs to the subject, a benchmark scored by them cannot tell recovery from fabrication. Code for the measurement protocol will be released so that recoverability claims can be tested for newer models.

193. 【2609.36832】Motion Concept Unlearning in Video Diffusion Models

链接:https://arxiv.org/abs/2609.36832

作者:Ping Liu,Chi Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:raising safety concerns, generate realistic depictions, motivate targeted concept, raising safety, motion concept erasure

备注: Accepted to ACM MM 2026. Dr. Chi Zhang is the corresponding author

点击查看摘要

Abstract:Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that motivate targeted concept erasure. Although concept erasure has been extensively studied for static concepts in text-to-image and T2V models, erasing motion concepts remains largely unexplored. We present a systematic study of motion concept erasure in video Diffusion Transformers (DiTs). Through causal interventions, we show that text-conditioning attention carries concept-specific motion information and supports selective intervention, whereas perturbing temporal positional encoding suppresses both target and non-target dynamics. We further find that directly adapting ESD, a representative weight-level image erasure method, to a video DiT yields modest and uneven motion suppression: reducing its erasure training loss does not by itself remove the concept signal from the difference between the conditional and unconditional predictions, which classifier-free guidance (CFG) then scales at every denoising step. From these findings, we derive three requirements for motion concept erasure: concept specificity, spatial selectivity, and temporal naturalness. Each determines one component of MUTE (Motion concept Unlearning in Text-to-video gEneration): at each denoising step, MUTE extracts a concept direction through token neutralization, derives a spatial gate from the direction's intrinsic structure, and subtracts the resulting correction from the velocity output before CFG is applied. MUTE is training-free and requires no weight modification. Experiments on 20 motion concepts show that MUTE outperforms representative prompt-level, weight-level, and inference-time baselines on Wan2.1-T2V, and the same formulation transfers to CogVideoX, supporting its applicability across distinct T2V attention architectures.

194. 【2609.36826】Learning via Self-Consistency for Diffusion-based Video Reasoning

链接:https://arxiv.org/abs/2609.36826

作者:Zhenghao Ni,Weimin Qiu,Meng Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrated emerging zero-shot, emerging zero-shot capabilities, Video generation, demonstrated emerging, emerging zero-shot

备注: 23 pages, including 10 pages for the main body

点击查看摘要

Abstract:Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.

195. 【2609.36822】RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection

链接:https://arxiv.org/abs/2609.36822

作者:Wenpeng Mu,Junshan Jin,Tanfeng Sun,Xinghao Jiang,Qiang Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image generators calls, Reconstruction Evolution, reconstruction, Reconstruction Evolution Dynamics, generation mechanisms

备注:

点击查看摘要

Abstract:The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5\% and average precision of 97.5\% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.

196. 【2609.36815】CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval

链接:https://arxiv.org/abs/2609.36815

作者:Zhen Liu,Letian Li,Jinpeng Wang,Shuzhao Xie,Yuzhi Huang,Jingyan Jiang,Zhi Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Partially Relevant Video, Partially Relevant, retrieve untrim-med videos, seeks to retrieve, temporal annotations

备注: 10 pages. Accepted to ACM Multimedia 2026 (MM '26)

点击查看摘要

Abstract:Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging than standard full-video retrieval. This task presents two intertwined challenges: (1) signal dilution, where coarse global representations blur the brief relevant signal into the dominant irrelevant surroundings;(2) curvature rigidity, where embedding all videos in the same fixed-geometry space distorts representations for videos that range from flat atomic events to deep compositional hierarchies. Existing PRVR methods have improved moment selection and cross-modal matching, but they still typically encode all videos in a single fixed-curvature retrieval space, limiting their ability to model diverse video structures. To address both challenges, we propose CurvSpec, a framework that learns content-adaptive curvature for video retrieval representations rather than imposing a fixed geometric prior. CurvSpec processes features through parallel Euclidean and hyperbolic attention layers, with independently learned curvatures assigned to the hyperbolic layers, and a content-aware fusion mechanism routes each input to its most suitable geometric regime. To further suppress signal dilution, CurvSpec represents each video with semantic centroids whose number is determined by the video's content complexity, projects them onto the learned manifold, and matches each query against its nearest centroid by geodesic distance. Experiments on ActivityNet Captions, TVR, and Charades-STA demonstrate state-of-the-art retrieval performance.

197. 【2609.36810】MeteoVerse: Unified Weather-Controllable Video World Model

链接:https://arxiv.org/abs/2609.36810

作者:Renlong Wu,Guanqiao Wang,Xuan Shang,Yin Hanming,Xiaoxiao Sheng,Tianyu Huang,Hui Li,Wangmeng Zuo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:weather, predict future content, Video world models, world models aim, aim to predict

备注: 13 pages

点击查看摘要

Abstract:Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determined not only by changes in viewpoint and object dynamics, but also by environmental conditions such as weather, which can substantially alter scene appearance and visibility. Modeling such realistic weather evolution is challenging because the required weather modification depends jointly on the observed and desired weather states. Depending on their relation, the model may need to preserve, introduce, or remove a weather effect. Existing video world models typically leave this weather transition implicit, forcing the generation backbone to infer weather evolution together with scene dynamics and camera motion, which leads to imprecise weather control. To address this limitation, we propose MeteoVerse, a unified weather-controllable video world model that generates future videos from a single sunny or adverse-weather image, conditioned on a weather-free scene description, a target-weather instruction, and a camera trajectory. Rather than conditioning only on the desired weather, MeteoVerse explicitly estimates the observed and target weather states and represents the required weather transition. A transition-aware mixture of weather experts then translates this transition into category-specific residual weather features, unifying weather preservation, introduction, and removal while enabling fine-grained control over introduced weather intensity. We further construct the MeteoVerse dataset with over 50K real-world weather video clips, generated sunny counterparts, disentangled scene and weather descriptions, weather-intensity annotations, and camera trajectories. Extensive experiments demonstrate substantially improved weather controllability while retaining competitive scene consistency and camera-control performance.

198. 【2609.36803】EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding

链接:https://arxiv.org/abs/2609.36803

作者:Yuwei Miao,Xuesheng Zhang,Wenhao Zou,Jixia Zhang,Jianwei Lv,Bo Yuan,Junfeng Wang,Shiao Xie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:understanding requires incrementally, requires incrementally maintaining, requires dense process, video understanding requires, dense process signals

备注:

点击查看摘要

Abstract:Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.

199. 【2609.36801】Scene Retargeting: Learning Object Placement with Analogical Transfer

链接:https://arxiv.org/abs/2609.36801

作者:Minkwan Kim,Junho Kim,Seungmin Lee,Changwoon Choi,Young Min Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:support daily activities, Interactive simulations, build on realistic, daily activities, computing applications build

备注: Project page: [this https URL](https://mkjjang3598.github.io/Scene-Retargeting/)

点击查看摘要

Abstract:Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizable framework for generating similar functional context. We formalize Scene Retargeting as stably transferring the semantically coherent spatial organization across layouts, rather than relying on textual descriptions or pairwise relationships. Our cluster-wise transfer flexibly handles mismatched object instances and adapts to distinctive floor plans. We optimize to preserve the rich semantic context of individual clusters by respecting the spatial distribution of foundation features. We can then impose physical constraints to refine wall contacts, pairwise alignment, or clear passageways and openings. Our framework outperforms state-of-the-art methods on layout generation on the 3D-FRONT dataset, and demonstrates downstream applications including real-to-sim transfer, analogical trajectory transfer, and multi-reference composition.

200. 【2609.36798】Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs

链接:https://arxiv.org/abs/2609.36798

作者:Yueran Ma,Ronghao Lin

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Omni-modal large language, large language models, Factorized Modality Diagnostic, large language, explicitly refers

备注: 25 pages, 11 figures, 16 tables

点击查看摘要

Abstract:Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at this https URL.

201. 【2609.36782】Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

链接:https://arxiv.org/abs/2609.36782

作者:Cheng Ye,Weidong Chen,Zhaobo Qi,Beier Zhu,Zhendong Mao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, objective understanding tasks, multimodal large language, demonstrated exceptional capabilities, falls significantly short

备注: 19 pages, 8 figures

点击查看摘要

Abstract:While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.

202. 【2609.36779】DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients

链接:https://arxiv.org/abs/2609.36779

作者:Xiaotian Zhang,Yusheng Wang,Naoya Kagawa,Noritaka Takamura,Keiji Okuhara,Hiroyasu Baba,Jun Ota

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Accurate hand-eye calibration, Accurate hand-eye, hand-eye calibration, differentiable rendering, differentiable rendering methods

备注:

点击查看摘要

Abstract:Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.

Subjects:

Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

Cite as:
arXiv:2609.36779 [cs.RO]

(or
arXiv:2609.36779v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2609.36779

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Journalreference:
IEEE Transactions on Instrumentation and Measurement, vol. 75, Art. no. 7505816, pp. 1-16, 2026

Related DOI:

https://doi.org/10.1109/TIM.2026.3712913

Focus to learn more

            DOI(s) linking to related resources</p>
203. 【2609.36776】Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention

链接:https://arxiv.org/abs/2609.36776

作者:Cheng Ye,Weidong Chen,Peipei Song,Zhendong Mao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video Captioning aims, emotionally empathetic descriptions, generate factually accurate, Emotional Video Captioning, Video Captioning

备注: 20 pages, 7 figures

点击查看摘要

Abstract:Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from {causal redundancy and spurious correlations} in co-occurrence bias (e.g., misclassifying ``sadness'' as ``joy'' on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal reasoning or merely exploited background confounders. To address these limitations, we first construct {EVC-CauseGround}, a comprehensive benchmark with dense spatio-temporal causal annotations. Crucially, it introduces a carefully selected {Causal-Faithfulness Subset} to explicitly quantify genuine emotion-cause attribution. Second, we propose {Causal-EVC}, an emotion-grounding captioning framework, which introduces a Motion-guided Causal Spatiotemporal Localization module to precisely decouple causal triggers from background confounders. Besides, we introduce an Interpretable Sparse Emotion Routing module. By synthesizing counterfactual representations and formulating a novel counterfactual contrastive objective, we enforce the model to anchor its emotion predictions strictly on authentic causal triggers instead of confusing background. Extensive experiments show that Causal-EVC not only achieves the best performance on semantic metrics but also exhibits significant advantages in the causal-faithfulness subset, which demonstrates that our model could mine emotional cues from genuine visual causes and mitigate co-occurrence bias for interpretable multimodal emotion understanding.

204. 【2609.36775】DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning

链接:https://arxiv.org/abs/2609.36775

作者:Cheng Ye,Weidong Chen,Bingyan Xu,Zhendong Mao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:complex reasoning capabilities, significantly advanced, advanced the complex, Reinforcement Learning, complex reasoning

备注: 17 pages, 5 figures

点击查看摘要

Abstract:Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.

205. 【2609.36770】Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

链接:https://arxiv.org/abs/2609.36770

作者:Aram Davtyan,Pablo Acuaviva,Sebastian Stapf,Paolo Favaro

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:neural networks develop, division of labor, neural networks, networks develop, shared gate

备注:

点击查看摘要

Abstract:Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.

206. 【2609.36759】Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

链接:https://arxiv.org/abs/2609.36759

作者:Chiyuan He,Zihuan Qiu,Fanman Meng,Chao Wang,Liangjiang Chen,Linfeng Xu,Qingbo Wu,Hongliang Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:transferable visual-textual alignment, CLIP transferable knowledge, visual-textual alignment, widely adopted, adopted for class-incremental

备注: 21 pages, 11 figures, and 12 tables, including the appendix

点击查看摘要

Abstract:Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.

207. 【2609.36757】FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

链接:https://arxiv.org/abs/2609.36757

作者:Xiaoxu Chen,Qin Yang,Haoran Bai,Sibin Deng,Ying Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recovers realistic details, costly VAE encoding, Diffusion-based video restoration, restoration recovers realistic, Diffusion-based video

备注:

点击查看摘要

Abstract:Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built on a one-step diffusion model, which delivers strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. To improve inference efficiency, FastVR combines a lightweight VAE with chunk-wise causal attention, which substantially reduces the computational cost. During training, it further adopts velocity consistency regularization and continuous trajectory learning, which improve restoration quality. Extensive experiments show that FastVR is more efficient than the evaluated diffusion baselines while achieving state-of-the-art performance on synthetic and real-world benchmarks. We hope that this work supports further progress in the community.

208. 【2609.36756】NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation

链接:https://arxiv.org/abs/2609.36756

作者:Jiawei Zhang,Shuhao Liu,Rong Huang,Yuancheng Li,Zhihui Li,Xiaojun Chang,Changlin Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:enable adaptive compression, tokenizers enable adaptive, models to flexibly, visual tokenizers enable, enable adaptive

备注: Computer Vision, Autoregressive Model

点击查看摘要

Abstract:One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256$\times$256 among existing variable-length autoregressive image generation methods. Code will be available at this https URL.

209. 【2609.36755】Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing

链接:https://arxiv.org/abs/2609.36755

作者:Xinyu Pu,Hongsong Wang,Jie Gui,Pan Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:precise spatial control, Modern image editors, Modern image, image editors excel, visual synthesis

备注:

点击查看摘要

Abstract:Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.

210. 【2609.36693】AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling

链接:https://arxiv.org/abs/2609.36693

作者:Shiwei Ren,Zhiang Liu,Yongchun Fang,Hongwei Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrated remarkable potential, Gaussian Splatting, view synthesis, predict Gaussian appearance, Gaussian appearance attributes

备注:

点击查看摘要

Abstract:Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a $0.8$ dB improvement in PSNR over the pose-free method NAS3R and a $1.1$ dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset. Project page: this https URL.

211. 【2609.36685】When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation

链接:https://arxiv.org/abs/2609.36685

作者:Zhirui Xing,Long Ye,Kaige Li,Ziyi Xu,Ming Meng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:synthesize natural gestures, spoken content, Co-speech gesture generation, aims to synthesize, synthesize natural

备注: 9 pages, 5 figures, 3 tables

点击查看摘要

Abstract:Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.

212. 【2609.36680】Reprogramming Vision-Language Models via Structured Prompt Reparameterization

链接:https://arxiv.org/abs/2609.36680

作者:Zizhao Li,Chengyi Cai,Mohammed Yaqoob Ansari,Feng Liu,Joseph West,Kourosh Khoshelham

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reprogramming adapts pretrained, adapts pretrained models, Visual reprogramming adapts, Visual reprogramming, adapts pretrained

备注:

点击查看摘要

Abstract:Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.

213. 【2609.36677】ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

链接:https://arxiv.org/abs/2609.36677

作者:Haoyang Wu,Shoudong Han,Chaoyue Li,Sijia Chen,Zhenyang Xie,Wang sihan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Language-guided multi-camera tracking, Language-guided multi-camera, early associations unreliable, make early associations, unobserved gaps

备注: 41 pages, 8 figures

点击查看摘要

Abstract:Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.

214. 【2609.36661】You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming

链接:https://arxiv.org/abs/2609.36661

作者:Zizhao Li,Mohammed Yaqoob Ansari,Xinyu Su,Jiayang Ao,Joseph West,Kourosh Khoshelham

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remain computationally expensive, adapting pretrained models, computationally expensive, parameter-efficient method, method for adapting

备注:

点击查看摘要

Abstract:Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4--24.4\%. On 16-shot CLIP, it raises the four-backbone average from 71.4\% to 77.2\%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.

215. 【2609.36655】Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation

链接:https://arxiv.org/abs/2609.36655

作者:Youjia Zhang,Huiling Liu,Soyun Choi,Jaehong Yoon,Sungeun Hong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unlabeled test stream, unlabeled test, test stream, Continual test-time adaptation, CTTA

备注:

点击查看摘要

Abstract:Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target observations can provide complementary evidence for correcting the source prediction, but this history may become misaligned as the target distribution changes. The key question is therefore not how much the correction differs from the source prediction, but whether and how strongly it should be applied. This paper proposes Gain-Aware INtervention (GAIN), a backpropagation-free CTTA framework guided by a simple principle: history proposes, gain decides. GAIN maintains compact target statistics to form a correction proposal and a posterior-predictive evaluator that accounts for estimation uncertainty. The resulting source-relative gain estimates the proposal's benefit and determines a sample-specific intervention strength along a continuous path through efficient one-dimensional optimization. Gain-controlled predictions then update the target statistics online, limiting the propagation of unreliable corrections, all without backpropagation, sample storage, or replay. Across five benchmarks, our method achieves strong predictive performance, with favorable accuracy--calibration--efficiency trade-offs in continual adaptation. On ImageNet-C, for example, GAIN achieves 61.9% accuracy with near-source calibration. It remains stable under diverse and challenging continual shifts while running 15.9x faster than a representative optimization-based CTTA baseline.

216. 【2609.36651】FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

链接:https://arxiv.org/abs/2609.36651

作者:FangZhi Zhong,Xuerui Qiu,Yuqi Pan,Ya Liu,Shaowei Gu,Bo Xu,Guoqi Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:incurs substantial computation, large language models, language models incurs, models incurs substantial, memory costs

备注: 23 pages, 10 figures. Code: [this https URL](https://github.com/fangzhi-zhong/FoucsVTC)

点击查看摘要

Abstract:Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.

217. 【2609.36648】VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training

链接:https://arxiv.org/abs/2609.36648

作者:Yuanwei Hu,Bo Peng,Yuheng Jia,Xinting Hu,Yadan Luo,Wenjie Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:complement visual representations, leverages textual semantics, image clustering, giving rise, visual representations

备注:

点击查看摘要

Abstract:Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trained vision-language models (VLMs). VLM4Cluster implements 17 representative methods spanning classical, deep, and language-assisted image clustering, and evaluates them on 20 datasets covering classical, challenging, fine-grained, large-scale, and out-of-distribution settings. Beyond effectiveness, VLM4Cluster systematically investigates image clustering along three complementary dimensions: robustness to adversarial perturbations, generalization under distribution shifts, and computational efficiency. Our study shows that LaIC substantially advances the clustering performance frontier on many semantically demanding benchmarks, generally exhibits stronger generalization under distribution shifts, and achieves a more favorable effectiveness-efficiency trade-off. However, its gains become less consistent on large-scale and fine-grained datasets, while language assistance does not systematically reduce sensitivity to adversarial perturbations. VLM4Cluster is released at this https URL.

218. 【2609.36644】OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration

链接:https://arxiv.org/abs/2609.36644

作者:Pei An,Jiaqi Yang,Yulong Wang,Siwen Quan,Liangliang Nan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:component in learning-based, crucial component, OCA, Cross-attention, Abstract

备注: Accepted to ECCV'26

点击查看摘要

Abstract:Cross-attention is a crucial component in learning-based image-to-point-cloud (I2P) registration. Although existing cross-attention mechanisms have achieved promising progress, attention ambiguity remains a fundamental challenge that hinders the learning of discriminative 2D-3D correspondences. To address this problem, we revisit cross-attention and establish ordinary differential equations (ODEs) to model the ideal I2P feature interaction. Based on this formulation, we develop an ODE-driven cross-attention (OCA) module that refines feature representations and attention matrices through ODEs. In practice, OCA can be seamlessly integrated into existing I2P registration frameworks. To validate its effectiveness, we incorporate OCA into five state-of-the-art baselines and evaluate on four public benchmark datasets. Experimental results demonstrate that OCA improves registration recall by up to 5\%, 9\%, and 15\% under the standard, fine-tuning, and zero-shot settings, respectively.

219. 【2609.36638】PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation

链接:https://arxiv.org/abs/2609.36638

作者:Mingfeng Lin,Chengfei Cai,Lin Xu,Chengqian Ma,Yuxiang Wei,Liang Han

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:detailed textual conditions, concise and underspecified, detailed textual, textual conditions, conditions for reliable

备注:

点击查看摘要

Abstract:Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.

220. 【2609.36628】Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning

链接:https://arxiv.org/abs/2609.36628

作者:Yanan Wang,Tingsong Li,Kaixun Jiang,Chongyang Zhong,Chenwei Xoe,Zhaohe Liao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generate rich video, Vision-Language Models, rich video captions, camera cuts, generate rich

备注: 18 page, 8 figures

点击查看摘要

Abstract:Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.

221. 【2609.36616】CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

链接:https://arxiv.org/abs/2609.36616

作者:Hanwen Lu,Jun He,Mingjia Yang,Hao Wei,Jinhao Huang,Yi Lin,Xiang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:records urban evolution, imagery records urban, uneven coverage leaves, coverage leaves substantial, leaves substantial gaps

备注: 43 pages, 9 figures, 7 tables. Project page: [this https URL](https://luhanwen67.github.io/CrossTimeEdit-release/)

点击查看摘要

Abstract:Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at this https URL.

222. 【2609.36607】Pixel-wise Exposure for Highly Robust In-Vehicle Remote-PPG

链接:https://arxiv.org/abs/2609.36607

作者:Jieying Wang,Xinqi Cai,Caifeng Shan,Wenjin Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:inherent hardware limitation, promising non-contact solution, Remote photoplethysmography, existing camera exposure, uniform exposure time

备注:

点击查看摘要

Abstract:Remote photoplethysmography (rPPG) offers a promising non-contact solution for heart rate monitoring, yet its real-world robustness is fundamentally limited by an inherent hardware limitation: existing camera exposure control paradigms, whether fixed or auto-exposure, impose a uniform exposure time across all pixels within a frame. In high-dynamic-range scenes such as automotive cabins with strong directional sunlight, this spatially invariant exposure constraint inevitably leads to localized facial overexposure or underexposure, irreversibly corrupting the subtle pulsatile signals essential for rPPG at the point of capture, a physical degradation that no downstream algorithm can recover. To overcome this bottleneck, we propose PixExpo (Pixel-wise Exposure), a "temporal-for-spatial" framework that sequentially captures frames under a predefined cyclic exposure schedule and performs non-iterative pixel-wise fusion. At each pixel location, PixExpo selects the observation closest to an rPPG-motivated target intensity. This criterion seeks to reduce local saturation and severe underexposure rather than optimize perceptual appearance. PixExpo requires no sensor modification but assumes programmable frame-level exposure control. We validate the proposed PixExpo framework using our newly introduced MEX-Drive dataset, comprising 48 participants under real-world driving conditions. Experimental results demonstrate that PixExpo outperforms manufacture-default auto-exposure methods, reducing the mean absolute error (MAE) by 7.21 bpm (from 13.94 to 6.73 bpm) and increasing the success rate by 37.29 percentage points (from 25.95% to 63.24%) across challenging driving scenarios.

223. 【2609.36599】Scaling Video Generation for Reasoning: At What Cost?

链接:https://arxiv.org/abs/2609.36599

作者:Weihang Guo,Xiaoyu Wu,Yifei Wang,Niloofar Mireshghallah,Lydia E. Kavraki

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:video generation enables, scaling video generation, generation enables models, computational cost, video generation

备注:

点击查看摘要

Abstract:We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.

224. 【2609.36598】Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation

链接:https://arxiv.org/abs/2609.36598

作者:Ziying Zhang,Litao Li,Junchao Liao,Tianyi Zeng,Siyu Zhu,Long Qin,Zhenghao Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:exhibit convincing motion, text, exhibit convincing, convincing motion, motion and photorealism

备注: ICLR 2027 under review

点击查看摘要

Abstract:A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental or using static OCR metrics that ignore temporal dynamics. We introduce VidScribe, a unified diagnostic benchmark spanning four generation regimes: writing from language (T2V), transferring text identity from a reference (R2V), sustaining text under dynamics (I2V), and localized text editing (V2V). VidScribe contains 803 human-verified samples across a 12-axis conditionally orthogonal factor space covering Intrinsic Text Properties, Physical Imaging Conditions, and Temporal Behavior. For reliable evaluation, we build a track-grounded, gated suite with 11 shared metrics and 2 task-specific probes under strict measurability conditions. Benchmarking 11 commercial and open-source systems shows that video text capability is non-monolithic, with content recognition decoupled from stroke-level glyph correctness. Performance is highly task-asymmetric: I2V sustains text most reliably, whereas V2V editing is the primary bottleneck. Counter-intuitively, degradation concentrates on a small subset of text-centric structural and temporal factors rather than adverse imaging conditions. Further probes show that visual references improve glyph and typographic fidelity rather than content accuracy, while localized editing fails to isolate target text without corrupting undeclared source text. Beyond evaluation, VidScribe also provides an actionable training signal, where benchmark-aligned preference optimization measurably improves visual text generation. this https URL.

225. 【2609.36563】AffectReveal: Event-Grounded Emotion Recognition Beyond Visual Appearances

链接:https://arxiv.org/abs/2609.36563

作者:Yihao Qian,Runhao Zeng,Sicheng Zhao,Feng Liang,Hongmin Cai,Mingkui Tan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recognition commonly assumes, emotion recognition commonly, commonly assumes, emotion recognition, recognition commonly

备注: 23 pages, 5 figures, 7 tables. Submitted to ICLR 2027

点击查看摘要

Abstract:Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26--10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.

226. 【2609.36562】hinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models

链接:https://arxiv.org/abs/2609.36562

作者:Ruochen Zhang,Yao Huang,Yitong Sun,Jiahe Xie,Jin Yan,Jifan Ma,Yuanfang Guo,Xingxing Wei

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Multimodal Large Language, Large Language Models, Large Language, Multimodal Large, multimodal implicit risks

备注: 9 pages, 4 figures, accepted by ACMMM 2026

点击查看摘要

Abstract:While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at this https URL.

227. 【2609.36560】FM-ReID: Selective Competitive Token Routing for Object Re-Identification

链接:https://arxiv.org/abs/2609.36560

作者:Zhiqi Li,Xiaowei Zhou,Zeyuan Sun,Feng Gao,Junyu Dong

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Object re-identification, share highly similar, highly similar global, similar global appearances, faces a recurring

备注: 12 pages

点击查看摘要

Abstract:Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.

228. 【2609.36557】How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective

链接:https://arxiv.org/abs/2609.36557

作者:Janet Wang,Yunbei Zhang,Xiao Wang,Jihun Hamm

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:show significant promise, offering accurate diagnosis, clinical image understanding, show significant, offering accurate

备注: 27 pages, 16 figures

点击查看摘要

Abstract:Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.

229. 【2609.36545】SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence

链接:https://arxiv.org/abs/2609.36545

作者:Gyeonggwan Lee,Eunsoo Im,Seunghwan Hong,Junghun Suh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:ERP underpins panoramic, underpins panoramic stereo, Equirectangular projection, robust dense feature, omnidirectional SLAM

备注: Accepted to ACCV 2026. 35 pages: 16-page main paper (including references) and 19-page supplementary material. Project page: [this https URL](https://gandanlee.github.io/sccm/) Code: [this https URL](https://github.com/gandanlee/sccm)

点击查看摘要

Abstract:Equirectangular projection (ERP) is the standard representation for 360$^\circ$ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically on ERP because the chart introduces three coupled distortions -- topological, metric, and area -- that standard coarse matching and visibility estimation do not explicitly model. We show that correcting the three distortions at the coarse-stage interfaces where they arise -- pairwise distortions in attention, per-pixel distortion in covisibility gating -- improves PCK@$1^\circ$ from 0.229 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged -- our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.

230. 【2609.36531】Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

链接:https://arxiv.org/abs/2609.36531

作者:Estela Monserrat Arriaga Santana(1),Julian Rosas Scull(1),Ehécatl Sacamch'en Núñez Rico(1),Hugo Jair Escalante(2) ((1) National Autonomous University of Mexico, (2) University of Texas at El Paso)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:largely regarded, regarded as predictive, Video world models, world model anticipate, world models

备注: 8 pages, 2 figures, 6 tables

点击查看摘要

Abstract:Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.

231. 【2609.36520】Distilling Privileged Control Barrier Functions into RGB-Only Safety Filters for Dynamic Visual Navigation

链接:https://arxiv.org/abs/2609.36520

作者:Seungyeon Yoo,Gawon Lee,Seungwoo Jung,Inkyu Jang,H. Jin Kim

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)

关键词:Control Barrier Function, dedicated safety layer, policies remain vulnerable, real-world dynamic environments, visual Control Barrier

备注: Project page: [this https URL](https://syeon-yoo.github.io/distill-cbf-site/)

点击查看摘要

Abstract:RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety layer. Existing visual Control Barrier Function (CBF) approaches seek to provide safety from RGB observations, but often rely on real-time rendering or explicit scene reconstruction and are primarily designed for static scenes, limiting their practicality for onboard deployment. We propose a teacher-student visual distillation framework that transfers the safety behavior of a privileged CBF teacher to an RGB-only student filter for dynamic environments. The student maps a short RGB history, robot velocity, and a nominal control action directly to a safe action, while the teacher uses ground-truth robot and obstacle states in a real-to-sim dynamic Gaussian Splatting environment. To reduce the teacher-student information gap, the teacher constructs safety constraints only from obstacles observable within the student's RGB history. It also accounts for obstacle-velocity uncertainty to improve robustness to motion variations, while action augmentation exposes the student to diverse safe and unsafe nominal actions to better capture the safety boundary. At deployment, the student requires only RGB observations and robot velocity, without explicit 3D reconstruction or online rendering. Experiments show that the proposed method outperforms visual CBF baselines and improves the safety of RGB-based navigation policies under dynamic obstacle motion. Project page: this https URL.

232. 【2609.36496】Reimagine Video Dynamics

链接:https://arxiv.org/abs/2609.36496

作者:Yu Yuan,Yawen Lu,Guoxian Song,Kevin Duarte,Ratheesh Kalarot,Di Chang,Xijun Wang,Stanley H. Chan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:offering limited control, editing methods focus, visual context, Reimagine Video Dynamics, methods focus

备注: Project Page: [this https URL](https://yuyuanspace.com/RVD/)

点击查看摘要

Abstract:Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.

233. 【2609.36492】Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics

链接:https://arxiv.org/abs/2609.36492

作者:Yicong Li,Junjie Wang,Leander Lauenburg,Ella Hugie,Alexandra Irger,Wanhua Li,Donglai Wei,Hanspeter Pfister

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:inspecting electron microscopy, electron microscopy images, benchmarked vision-language models, synapse detection, presence and polarity

备注:

点击查看摘要

Abstract:We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.

234. 【2609.36475】Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

链接:https://arxiv.org/abs/2609.36475

作者:Sumin Hong,Katsumi Ibaraki,Renee Shi,David Chiang,Toby Jia-Jun Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:round shapes, sharp shapes, features across modalities, systematic pairings, pairings of features

备注: 9 pages

点击查看摘要

Abstract:Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.

235. 【2609.36471】Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

链接:https://arxiv.org/abs/2609.36471

作者:Guoheng Sun,Chen Chen,Jin Wang,Ang Li,Teresa Lv

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:World-Action Models, conditioning action generation, iterative action generation, future prediction adds, expensive iterative action

备注:

点击查看摘要

Abstract:World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, $3.62\times$ the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.

236. 【2609.36458】Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models

链接:https://arxiv.org/abs/2609.36458

作者:Abdullah All Tanvir,Xin Zhong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:semantically consequential variation, captures semantically consequential, affect model predictions, induce substantial motion, strongly affect model

备注:

点击查看摘要

Abstract:Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose Fisher-induced invariant representation geometry (Fisher-IRG), which measures local representation directions through their predictive sensitivity. Around each representation, we construct semantic-preserving and semantic-changing neighborhoods, aggregate their local Fisher information, and recover invariant directions through a contrastive generalized eigenvalue problem. Controlled displacement analyses first show that comparable Euclidean motion can have substantially different predictive consequences, supporting the need for a predictive geometry. Across language and vision models, Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions. Representation interventions further localize semantic effects to the Fisher-derived subspace, and held-out separation and retrieval show that the recovered geometry generalizes beyond the discovery neighborhoods. These results support Fisher-IRG as a principled framework for characterizing local invariant representation geometry.

237. 【2609.36454】DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction

链接:https://arxiv.org/abs/2609.36454

作者:Wenliang Guo,Zhanbo Huang,Yu Kong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:monocular RGB videos, produce visually plausible, RGB videos, monocular RGB, study hand-object interaction

备注: Page: [this https URL](https://wenliangguo.github.io/HOI-Reconstruction-Page/)

点击查看摘要

Abstract:We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamics derive hand generalized forces and object wrenches for dynamics-level supervision. We further couple hand and object dynamics through contact-force transfer and recover active hand actuation as an interaction-level physical quantity. We formulate its empirical magnitude distribution into a probabilistic prior that penalizes unlikely actuation and suppresses mechanically implausible reconstructed motion. Experiments on three HOI datasets show consistent improvements in both hand and object reconstruction. The reconstructed trajectories further benefit downstream applications including hand world-model generation and robotic manipulation learning, demonstrating the value of physics-aware HOI modeling beyond reconstruction.

238. 【2609.36442】Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time

链接:https://arxiv.org/abs/2609.36442

作者:Jae-Ho Lee,Min-Yeong Park,Jun-Yeong Moon,Jung Uk Kim,Gyeong-Moon Park

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:ever-changing data distributions, learning enables vision, enables vision systems, Online VIL, Online Versatile Incremental

备注: 17 pages, Accepted at ECCV 2026

点击查看摘要

Abstract:Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at this https URL.

239. 【2609.36440】DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing

链接:https://arxiv.org/abs/2609.36440

作者:Jae-Ho Lee,Jeong-Eun Lee,Gyeong-Moon Park

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large vision-language models, models generate descriptions, recently achieved remarkable, achieved remarkable progress, generate descriptions inconsistent

备注: 19 pages, Accepted at ECCV 2026

点击查看摘要

Abstract:Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects the generation dynamics that lead to hallucinated outputs. In this paper, we analyze the relationship between TF-based editing and AR generation behavior and find that TF-based editing alone may be insufficient to capture both decoding dynamics and multimodal interactions associated with hallucinations. To address this limitation, we propose DARE (Dual-path Auto-Regressive-aware Editing), a hybrid hallucination editing framework that integrates two complementary contrast pathways: textual contrasts and image contrasts, together with autoregressive-aware representation signals. Specifically, DARE constructs hallucination editing directions from (1) TF-based textual contrasts, (2) AR-aware representation transitions during decoding, and (3) controlled visual differences between paired images. Extensive experiments on multiple LVLM hallucination benchmarks demonstrate that DARE consistently reduces object hallucinations while preserving multimodal perception capability and inference efficiency. Our implementation code is available at this https URL.

240. 【2609.36436】Merlin Plus: A Large-Scale, Multi-Cancer, Image-Mask-Report Dataset

链接:https://arxiv.org/abs/2609.36436

作者:Pedro R. A. S. Bassi,Wenxuan Li,Szymon Plotka,Ruby Honjol,Jakub Przado,Xinze Zhou,Kang Wang,Yang Yang,Malte Jensen,Akshay S. Chaudhari,Curtis P. Langlotz,Alan L. Yuille,Zongwei Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:tumor masks, computed tomography, masks, Merlin, tumor

备注: MICCAI 2026

点击查看摘要

Abstract:Multi-cancer segmentation in computed tomography (CT) is fundamentally limited by the scarcity of tumor masks across different organs. We present Merlin Plus, the first large-scale CT dataset with radiologist-created tumor masks across 9 organs. Merlin Plus extends the Merlin dataset by adding 1,153 per-voxel tumor masks and longitudinal metadata. To create these tumor masks, we developed a report-based active-learning framework in which radiology reports identify tumor cases for annotation and support training of a tumor segmentation model. The model generates initial masks, which radiologists review and correct to produce the final masks, reducing annotation burden while maintaining high-quality annotations. Besides tumor masks, the longitudinal metadata in Merlin Plus enables temporal modeling of cancer progression. By directly addressing the major bottleneck of limited multi-cancer segmentation masks, Merlin Plus supports scalable multi-organ cancer detection, segmentation, and longitudinal analysis in CT. Dataset is available at: this https URL

241. 【2609.36433】RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance

链接:https://arxiv.org/abs/2609.36433

作者:Yiming Liu,Ben Wan,Tongxuan Liu,Ao Wang,Yuqi Xiong,Fan Zhang,Hui Chen,Guiguang Ding

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Diffusion models enable, high-quality visual generation, remains computationally expensive, models enable high-quality, enable high-quality visual

备注: 30 pages, 17 figures, NeurIPS 2026

点击查看摘要

Abstract:Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictions. However, existing branch-local reuse criteria do not explicitly account for how cache errors combine under CFG or how local perturbations affect the final output. We identify two misalignments in cache control: a branch-guided mismatch, where guided error depends on both the magnitudes and alignment of branch errors, and a local-final mismatch, where the downstream impact of a local error varies across timesteps. We propose RA-CFGCache, a Risk-Aligned Caching framework under CFG that incorporates both factors while keeping the sampling schedule and guidance rule fixed. CFG-aware Guided-Risk Composition combines existing branch-wise proxies using CFG coefficients and offline-calibrated cross-branch alignment. Propagation-Aware Rescaling further weights the resulting guided-risk estimate with a timestep-dependent propagation prior calibrated from isolated reuse perturbations. An online threshold controller then determines when to jointly refresh or reuse both branches. Experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B demonstrate improved efficiency--fidelity trade-offs over evaluated training-free caching baselines. Moreover, RA-CFGCache is compatible with diverse base proxy families, including TeaCache-, DiCache-, and MagCache-style estimators, and consistently improves fidelity at nearly unchanged latency. Code is available at this https URL.

242. 【2609.36432】mporal-Aware Fusion for Robust Outdoor LiDAR Localization

链接:https://arxiv.org/abs/2609.36432

作者:Minghang Zhu,Zhijing Wang,Yuxin Guo,Chen Liu,Yongshu Huang,Wen Li,Sheng Ao,Cheng Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:LiDAR relocalization aims, Global Coordinate Estimation, Prior Coordinate Generation, Coordinate Estimation module, Coordinate Generation module

备注: 10 pages, 8 figures, 6 tables, conference

点击查看摘要

Abstract:LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches often encounter limitations in dynamic or ambiguous scenarios, as they typically prioritize single-frame inference, leaving the potential of spatio-temporal consistency across scans not fully explored. In this paper, we propose a Temporal-aware Localization framework (TempLoc) designed to enhance the robustness of outdoor localization by effectively modeling sequential consistency. Specifically, a Global Coordinate Estimation module is first introduced to predict point-wise global coordinates and associated uncertainties for each LiDAR scan. A Prior Coordinate Generation module is then presented to estimate inter-frame point correspondences by the attention mechanism. Lastly, an Uncertainty-Guided Coordinate Fusion module is deployed to integrate both predictions of point correspondence in an end-to-end fashion, yielding a more temporally consistent and accurate global 6-DoF pose. Experimental results on the NCLT and Oxford RobotCar benchmarks show that our TempLoc outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of temporal-aware correspondence modeling in LiDAR relocalization.

243. 【2609.36429】owards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images

链接:https://arxiv.org/abs/2609.36429

作者:Zijun Gao,Chunbin Gu,Jinxi Xiang,Xiangde Luo,Pheng-Ann Heng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:costly spatial transcriptomics, existing methods operate, critical cellular heterogeneity, Predicting gene expression, histology images offers

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Predicting gene expression from HE-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-HE pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from HE images.

244. 【2609.36426】Losing the name before the box: measuring and repairing what narrow fine-tuning costs a detector outside its deployment vocabulary

链接:https://arxiv.org/abs/2609.36426

作者:Trung Minh Bui,Jongsul Moon,YoungOuk Kim,Jung-Hoon Hwang,Dongin Shin

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:broad corpus, in-domain accuracy improves, in-domain accuracy, accuracy improves, in-domain

备注: 25 pages, 3 figures. Supplementary material (69 pages) is included as an ancillary file. Submitted to the International Journal of Computer Vision

点击查看摘要

Abstract:A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-$K$ proposal coverage $C_\tau$: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top $K$ regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. $C_\tau$ falls while in-domain accuracy rises, on four architectures and three domains, by $5.12$ to $63.35$ points on boxes above $1024$ px$^2$. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs $87\%$ of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most $2.47$ points of in-domain accuracy. Seeing it costs one extra evaluation pass.

245. 【2609.36416】FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

链接:https://arxiv.org/abs/2609.36416

作者:Jade Choghari,Pepijn Kooijmans,Mansi Agarwal,Yusuf Umut Ciftci,Aseem Doriwala,Catherine Weaver,Mouli Sivapurapu,Kai Yang,Jackson Lee,Thomas Wolf,Pragna Mannam

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:isolated actions, execute complex, operating in real-world, real-world environments, environments must execute

备注: 26 pages. Code and model weights will be integrated into Hugging Face LeRobot [this https URL](https://github.com/huggingface/lerobot)

点击查看摘要

Abstract:Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask, and show that mid-training it this way yields substantial gains. Specifically, success on a spatial disambiguation task increases from 32.0% to 100.0%, and step-by-step human subtask guidance lifts success on an unseen long-horizon task from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires one-tenth the data of baselines without mid-training and generalizes zero-shot to completely unseen tasks on the new hardware. We open-source the full dataset, model weights, and training code.

246. 【2609.36407】What Makes High-Magnification Knowledge Transferable? A Study of Cross-Resolution Distillation in Whole-Slide Imaging

链接:https://arxiv.org/abs/2609.36407

作者:Zhiyuan Yang,Jiahao Cheng,Mahdi S. Hosseini

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:transferring high-magnification representations, transfer remain unclear, knowledge distillation aims, remain unclear, transferring high-magnification

备注: ICLR 2027 Submission

点击查看摘要

Abstract:Cross-resolution knowledge distillation aims to improve low-magnification whole- slide analysis by transferring high-magnification representations, yet the conditions for useful transfer remain unclear. We develop a decomposition-based analysis of teacher access, representation loss, and model excess, motivating three questions: whether (a) teacher targets help the task, (b) low-magnification students can predict them, and (c) slide models benefit from those predictions. We investigate them through controlled experiments across ten pathology cohorts spanning classifi- cation, grading, and survival prediction. In the main comparison, providing teacher regional means alongside native low-magnification features improves downstream performance in all ten cohorts. Direct prediction achieves lower reconstruction error than residual prediction, yet the predicted features underrepresent variation in the teacher targets. Moreover, better reconstruction does not consistently improve downstream scores, and retaining native features changes performance even when the predicted teacher features are held fixed. Together, these findings expose a gap between reconstructing teacher representations and realizing their downstream value. They challenge the sufficiency of reconstruction error as a measure of cross-resolution transfer and provide a diagnostic framework for examining where that transfer breaks down. Future distillation designs must account for both what students can predict and how slide models use those predictions.

247. 【2609.36386】Stealth Is a Relation, Not a Property: How Event Representations Create Blind Spots for Timing Attacks in Event-Based Perception

链接:https://arxiv.org/abs/2609.36386

作者:Shoaib Ahmed Dipu,Md. Shaown Miah,Kamrul Hasan,Sayeed Shafayet Chowdhury

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:event camera produces, DVS Gesture, processes time, downstream consumer, camera produces

备注: 30 pages, 8 figures, 20 tables. Code: [this https URL](https://github.com/shoaibdipu/Stealth_Is_a_Relation_Not_a_Property)

点击查看摘要

Abstract:An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector, processes time. The same timestamp change may leave a coarse temporal representation unchanged while changing the response of a model that preserves finer timing. We characterize this dependence as observer-relative stealth. For recorded event streams, retiming an event within its protected accumulation window leaves the accumulated integer tensor exactly unchanged. We use this exact blind space to construct Null, a gradient-guided timestamp-retiming attack, and define SC-ASR_A(tau) to measure attack success while bounding the change visible to observer A. On DVS Gesture at a 10% event budget, Null reaches 81.56 +/- 5.81% ASR on ConvSNN and 98.67 +/- 0.45% on a GRU while preserving the protected tensor exactly. On DailyDVS-200, a protocol-scale Multi-View Fusion Network variant reaches 99.28 +/- 0.11% exact-null ASR, compared with 9.70 +/- 1.06% for its matched control. In a five-attack comparison, Null is the only method with nonzero attack success at exact observer equality, reaching 81.4% on DVS Gesture and 87.35% on DailyDVS-200. We also search the same exact blind space with an independently implemented constrained projected-gradient optimizer, C-PGD. At matched victim-gradient evaluations, C-PGD reaches 84.50 +/- 2.89% ASR on DVS Gesture and 89.55 +/- 4.39% on DailyDVS-200, again with exact protected equality. Perturbations that are exactly hidden from the protected observer become visible under shifted, finer, overlapping, and randomized temporal views. Adding observer constraints reduces the real-valued blind-space fraction from 87.5% to 75.0% to 62.5%, while DVS ConvSNN ASR falls from 74.9% to 61.9% to 37.2%. These results show that stealth is not a property of the perturbation alone.

248. 【2609.36380】LEGO-Anything: Coding Agents for 3D Scene Reconstruction

链接:https://arxiv.org/abs/2609.36380

作者:Xirui Li,Peng Shi,Mingwen Dong,Sheng Zhang,Zhuoyan Xu,Dongkyu Lee,Shuaichen Chang,Yi Xiang,Lin Pan,Jiarong Jiang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:explicit scene program, execution yields, executes Blender code, scene, Blender code

备注:

点击查看摘要

Abstract:A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.

249. 【2609.36374】OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute

链接:https://arxiv.org/abs/2609.36374

作者:Brandon Leblanc,Charalambos Poullis

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:precludes edge deployment, achieved impressive performance, scaling model, dataset size, edge deployment

备注: Accepted at ACCV 2026

点击查看摘要

Abstract:Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling $\pi^3$ (959M parameters) into a 102M-parameter student yields 9.4$\times$ compression and up to 7$\times$ faster inference, trained at 1.6% of VGGT's training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4$\times$ more accurate than COLMAP on 7-Scenes at 980$\times$ throughput, with near-teacher completion. Code is available at this https URL

250. 【2609.36368】AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

链接:https://arxiv.org/abs/2609.36368

作者:Konstantinos D. Polyzos,Eleni Oikonomou,Tara Javidi

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large foundation models, Large foundation, downstream tasks, foundation models, promise of efficient

备注:

点击查看摘要

Abstract:Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on training neural-based decoders. Both approaches struggle under limited supervision, while fine-tuning additionally requires access to model parameters, which is often unavailable for closed-source models. We introduce AdaKerNet, a novel learnable task-adaptive neural kernel decoder. AdaKerNet is fully agnostic to the parameters of the underlying MLLM and operates solely on its (frozen) rich representations obtained from the diverse available modalities. AdaKerNet relies on (i) a set of learnable, Lipschitz-controlled multimodal features derived from these MLLM representations; (ii) a reference kernel that provides a soft structural prior on those features; and (iii) a lightweight nonlinear neural predictor that adaptively deforms that structure. Learning the kernel representation and the neural predictor jointly within a unified optimization framework allows AdaKerNet to capture features and geometric relationships relevant to the downstream task. Numerical tests across four MLLMs: BLIP-2, LLaVA-1.5, Qwen2.5-VL, and Gemini Embedding 2, and multimodal inputs spanning text, audio, images, and tabular measurements demonstrate significant and consistent improvements over direct MLP, attention-, autoencoder- and kernel-based decoders, across a range of scarce-label budgets, with average error reduction of up to 41% across baselines. These results establish AdaKerNet as an effective approach for prediction from frozen multimodal representations in the scarce label regime. Additional structural ablations highlight the complementary contributions of AdaKerNet's components.

251. 【2609.36364】Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

链接:https://arxiv.org/abs/2609.36364

作者:Xiaoyu Wu,Weihang Guo,Yifei Wang,Xinze Feng,Lydia E. Kavraki,Zhiwei Steven Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Standard video generators, Standard video, reusable memory tokens, natively compact historical, compact historical context

备注: Under Review

点击查看摘要

Abstract:Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Context Compaction (PACC), which uses a learned compressor to aggregate information across past frames into compact memory tokens. We train the compressor through on-policy distillation, using the same frozen generator both as a student when conditioned on compressed memory and as a teacher when conditioned on the full history. The student generates continuations, while the teacher provides targets for the same noisy inputs at each denoising step. Only the compressor is updated to align the student's predictions with these targets. We evaluate PACC on MBench, which jointly measures memory-event coverage and consistency. PACC outperforms the strongest baseline by 6.63 points on Causal-rCM and 3.19 points on Causal Forcing. Evaluation on VBench-Long using MovieGen prompts further shows that PACC produces minute-long videos with generation quality competitive with baselines. Together, these results show that learning to compact historical context can improve long-video memory without modifying the underlying generator.

252. 【2609.36352】StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

链接:https://arxiv.org/abs/2609.36352

作者:Ziyi Yin,Sangmin Woo,Kang Zhou,Sungyeon Kim,Aosong Feng,Haibo Ding,Jun Huan

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)

关键词:multiple dependent manipulations, require multiple dependent, shorter-horizon manipulation tasks, models perform, single command

备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at this https URL.

253. 【2609.36348】Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers

链接:https://arxiv.org/abs/2609.36348

作者:Xiaoyu Wu,Yifei Wang,Chen Wei

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:remain asymmetrically connected, learning remain asymmetrically, asymmetrically connected, Generative, semantic representations

备注:

点击查看摘要

Abstract:Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4\% using the class token and 10.1\% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.

254. 【2609.36315】PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States

链接:https://arxiv.org/abs/2609.36315

作者:Arya Kondur,Giosue Migliorini,Cameron Schmitt,Francesco Immorlano,Tairan Wang,Rebecca C. Scholten,Efi Foufoula-Georgiou,Gary Johnson,Chris Lautenberger,Valentin Waeselynck,J. Shane Romsos,Kasra Shamsaei,Alejandro Tejedor,Tianjia Liu,Yang Chen,Padhraic Smyth,James T. Randerson

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:air quality, human systems, creating a growing, diverse landscapes, increasing hazard

备注: 23 pages, 6 figures, 5 tables

点击查看摘要

Abstract:Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requires integrating meteorological conditions, fuels, vegetation, and topography at spatial and temporal resolutions suitable for both physical simulation and data-driven approaches. However, existing datasets often lack the resolution and coverage needed to capture these interacting controls. The PyroStack dataset addresses this gap by providing a harmonized, event-based collection of wildfire and environmental data across the contiguous United States and Alaska. It integrates satellite-derived fire observations with atmospheric reanalysis, vegetation, fuel characteristics, and topographic information into a unified framework spanning 6994 wildfires that occurred between 2012 and 2024 across a wide range of ecosystems and climate conditions. PyroStack offers spatial resolutions ranging from 30 m to 9 km and hourly temporal resolution, along with fire progression data at 12-hour intervals to support model initialization and evaluation. By combining broad spatial coverage with fine spatial and temporal detail, the dataset enables systematic analysis of wildfire dynamics and supports both physics-based and machine learning approaches, providing a foundation for benchmarking and improving fire spread models, with future extensions aimed at incorporating additional regions and fire suppression data streams to further advance wildfire prediction.

255. 【2609.36243】hink Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention

链接:https://arxiv.org/abs/2609.36243

作者:Mingqiu Liang,Dongdong Wang,Siyang Lu,Ting Huang,Yingjun Qi

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:historical Manchu manuscripts, unknown degradation regions, Full-page blind restoration, fragile connected strokes, Full-page blind

备注: Submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gated rEstoration), a selective restoration framework that first assesses where restoration is needed and then uses this assessment to guide restoration candidate generation and pixel-level selection. Its encoder predicts patch-level repair probabilities from complementary appearance and stroke-structural cues to condition restoration candidate generation, while the corresponding repair logits are refined into a pixel-level soft gate that selectively controls where the restoration candidate is applied. We further introduce a fidelity-aware evaluation protocol that jointly measures degraded-region recovery, intact-content preservation, and their balance. SAGE-Restore achieves the highest R-Recovery (0.463) and RFS (0.626), while maintaining high U-Fidelity (0.968), demonstrating an effective balance between restoration and content preservation.

256. 【2609.36224】Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

链接:https://arxiv.org/abs/2609.36224

作者:Wentao Zhou,Weijie Gan,Jiayun Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Unified multimodal models, Unified multimodal, combine image generation, shared backbone, Mutually Adversarial self-Training

备注:

点击查看摘要

Abstract:Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.

257. 【2609.36219】LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

链接:https://arxiv.org/abs/2609.36219

作者:Bang Xiao,Wenqi Jia,Ozgur Kara,Tiancheng Shen,Yibo Yang,Bolin Lai,Junho Kim,James Matthew Rehg

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requiring models interpret, Perspective taking, interpret spatial relations, models interpret spatial, imagined observer

备注: 22 pages, 8 figures

点击查看摘要

Abstract:Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.

258. 【2609.36217】Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

链接:https://arxiv.org/abs/2609.36217

作者:Xinming Dai,Qihang Jin,Tianshu Tan,Baiyuan Chen,Hanrui Lyu,Lenny Aharon,Kyle Daruwalla,Xun Helen Hou,Matthew R. Whiteway,Liam Paninski,Yizi Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scientific analysis remains, brain function requires, extracting behavioral representations, requires a precise, structured characterization

备注:

点击查看摘要

Abstract:A deeper understanding of brain function requires a precise, structured characterization of this http URL, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior this http URL augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.

259. 【2609.36210】On the spectral properties of generative denoiser Jacobians

链接:https://arxiv.org/abs/2609.36210

作者:Alexandros Graikos,Nebojsa Jojic,Dimitris Samaras

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:deep neural network, recover clean data, neural network denoiser, diffusion and flow-matching, Generative denoising models

备注:

点击查看摘要

Abstract:Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited insight into how the underlying denoiser, which drives generation, differs. In this work, we propose to analyze the spectrum of the denoiser Jacobian as a tool to characterize these differences. Across pre-trained denoising models, we observe that better generative performance is associated with larger Jacobian eigenvalues. Motivated by this, we introduce a regularization scheme that controls the Jacobian spectrum by training the denoiser on perturbed inputs, with perturbations suppressing or amplifying Jacobian responses. On ImageNet, we test whether directly modifying the Jacobian spectral properties leads to improved generations. Our findings suggest that denoisers benefit from both strengthening responses along data-relevant principal eigen-directions and suppressing the noisy, data-irrelevant ones. This establishes the denoiser Jacobian as a useful tool for identifying differences between generative denoising models.

260. 【2609.36199】PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

链接:https://arxiv.org/abs/2609.36199

作者:Vighnesh Subramaniam,Boris Katz,Brian Cheung,Chun-Liang Li,Tomas Pfister,Yale Song

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:temporally grounded actions, produce striking images, attribute binding, spatial relations, object counts

备注: 24 pages, 11 figures, 3 tables

点击查看摘要

Abstract:Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.

261. 【2609.36189】FD-AA: A Lightweight Focal-Diffuse And Attenuation-Aware Head for Incidental Abdominal Abnormality Detection in Chest CT

链接:https://arxiv.org/abs/2609.36189

作者:Haoyan Ding,Kritika Iyer,Halid Yerebakan,Zhenyu Bu,Chushu Shen,Peiyu Duan,Xinyuan Zheng,Sepehr Farhand,Xueqi Guo,Chaowei Wu,Yoshihisa Shinagawa,Gerardo Hermosillo Valadez

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Routine chest, relevant incidental abnormalities, captures upper-abdominal structures, clinically relevant incidental, upper-abdominal structures

备注: Submitted to a conference

点击查看摘要

Abstract:Routine chest CT captures upper-abdominal structures that may contain clinically relevant incidental abnormalities. Detecting these findings requires feature extraction from organs with different spatial extents and attenuation patterns. We propose FD-AA, a lightweight organ-aware classification head adaptable for frozen 3-D CT encoders. Within each organ, an attenuation-aware module preserves sparse focal evidence, while masked generalized-mean pooling captures diffuse anomaly patterns. In seven abdominal organs, FD-AA with Pillar-0 achieved state-of-the-art (SOTA) performance in both the CT-RATE test set (AUC = 0.798) and the external RAD-ChestCT dataset (AUC = 0.713). More specifically, FD-AA improved macro AUC/AP from 0.763/0.346 to 0.798/0.405 over direct classification using frozen Pillar-0 only (p = 0.034/0.016). Such performance gain generalizes across multiple frozen encoders (AUC improvement on MedicalNet +9.8%, CT-CLIP +14.7%, ResNet +3.7%), demonstrating the effectiveness of FD-AA across different feature representations. These results support the effectiveness of integrating focal-diffuse aggregation with explicit HU evidence for incidental abdominal abnormality detection.

262. 【2609.36172】Exploring Learning Models for Topological Relationship Recognition from Image Data

链接:https://arxiv.org/abs/2609.36172

作者:Saptak Das,Monidipa Das

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:stay completely separate, fields like GIS, biomedical imaging, stay completely, matters a lot

备注:

点击查看摘要

Abstract:Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot in fields like GIS, biomedical imaging, and robotics. Even though machine learning has come a long way, people haven't really focused on spotting these topological relationships in images. The main roadblocks? Not enough good datasets and no clear way to measure results. So, we rolled up our sleeves and built a new dataset. It's pretty sizable: over 11,000 labelled images showing all those essential relationships. We ran tests with some classic machine learning models, Naive Bayes, KNN, Random Forest, SVM, and Artificial Neural Networks, and threw in some deep learning stars like VGG16 and InceptionResNetV2. For the dataset itself, we used segmentation, contour detection, and grayscale normalization to tease out solid feature vectors. The results? Deep learning methods, especially VGG16, pulled ahead, with validation accuracy hitting 89.55%. That's a big jump compared to the traditional models. This shows how powerful transfer learning is for analyzing topological relationships in images, and it gives researchers a new standard to aim for in future work on spatial reasoning and topological classification.

263. 【2609.36168】Boosting Metric Depth Completion via Training-Free Adaptive Response Geometry

链接:https://arxiv.org/abs/2609.36168

作者:Mia Zhang,Jizong Peng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:sparse sensor measurements, increasingly leveraging visual, leveraging visual foundation, visual foundation models, Depth completion aims

备注:

点击查看摘要

Abstract:Depth completion aims to recover dense metric depth from sparse sensor measurements, increasingly leveraging visual foundation models as geometric priors. However, aligning these priors to true metric scale typically relies on rigid affine assumptions in predefined coordinate systems, leaving systematic calibration errors. Linearity in depth calibration depends on the response coordinate. We introduce adaptive response geometry, which makes the fixed choice of depth, log depth, or disparity an image-level unknown. A continuous response family unifies these coordinates and defines an explicit depth-dependent gain. We derive the response-gradient relation and estimate the response parameters in metric space. Hard-Dirichlet residual reconstruction completes the calibrated prior. Under deliberately incomplete metric observations, the training-free pipeline achieves macro AbsRel 0.0301 and macro NMed 14.04°, improving both aggregate measures over PriorDA, LDCM, and Any2Full. Linearity diagnostics examine how the selected response changes the depth relation and its metric error.

264. 【2609.36145】From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection

链接:https://arxiv.org/abs/2609.36145

作者:Kaiqing Lin,Songze Li,Shen Chen,Yunfei Guo,Xiaoye Qiu,Haodong Li,Taiping Yao,Bo Wang,Youchang Xiao,Bin Li,Shouhong Ding

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Tampered Text Detection, Multimodal Large Language, safeguarding document authenticity, Text Detection, Large Language Models

备注:

点击查看摘要

Abstract:Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.

265. 【2609.36136】Xiaomi-OCR-0 Technical Report

链接:https://arxiv.org/abs/2609.36136

作者:Xin Chen,Anan Du,Feng Feng,Pei Fu,Jian Luan,Longwei Xu,Shaojie Zhang,Hang Li,Heng Qu,Cheng Tan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Compact OCR-specific vision-language, Compact OCR-specific, OCR-specific vision-language models, document parsing performance, strong document parsing

备注:

点击查看摘要

Abstract:Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: this https URL.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.36136 [cs.CV]

(or
arXiv:2609.36136v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.36136

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
266. 【2609.36134】Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation

链接:https://arxiv.org/abs/2609.36134

作者:Mohammad Sadegh Sirjani

类目:Computer Vision and Pattern Recognition (cs.CV); Hardware Architecture (cs.AR); Machine Learning (cs.LG)

关键词:Functional Kolmogorov-Arnold Networks, MRI Gibbs artifact, Gibbs artifact removal, edge medical devices, MRI Gibbs

备注:

点击查看摘要

Abstract:Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, hardware-aware compression of FunKAN for point-of-care use. FunKANLite-TR reduces the spatial prior and replaces the ResBlock offset predictor with a depthwise-separable block. It has 1.9x fewer parameters than FunKAN and no loss in accuracy. We then distill FunKANLite-TR into FunKANLite-ST, which lowers the Hermite basis rank, factorizes the spatial prior into a low-rank form, and halves the filter widths. FunKANLite-ST has 5.6x fewer parameters and 3.7x fewer GFLOPs than FunKAN. It stays within 1.4 percentage points IoU of FunKAN on BUSI, GlaS, and CVC-ClinicDB, and reaches 33.95 dB PSNR on IXI. On an NVIDIA Jetson Orin Nano and a Raspberry Pi 5, FunKANLite-ST reduces energy per inference by up to 68% and raises throughput by 2.9x.

267. 【2609.36101】One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

链接:https://arxiv.org/abs/2609.36101

作者:Aditya Sharma,Divya Saxena

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:learn shared embedding, shared embedding spaces, representations remain separated, matched image-text pairs, aligning matched image-text

备注: 9 pages, 3 figures, 3 tables

点击查看摘要

Abstract:Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.

268. 【2609.36066】AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

链接:https://arxiv.org/abs/2609.36066

作者:Tongtong Feng,Xin Wang,Haoran Hou,Ren Wang,Weiran Wang,Shaokai Zhu,Ziqi Jia,Hao Wang,Yu-Wei Zhan,Zongyuan Wu,Jinghao Cui,Wenwu Zhu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Multimedia (cs.MM); Robotics (cs.RO)

关键词:unstructured three-dimensional environments, Open-world aerial object-goal, autonomously explore large-scale, requiring aerial agents, aerial object-goal search

备注:

点击查看摘要

Abstract:Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at this https URL.

269. 【2609.36024】CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes

链接:https://arxiv.org/abs/2609.36024

作者:Shuzhao Xie,Lelin Wang,Guying Lin,Zhi Wang,Minchen Li

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO)

关键词:existing methods largely, methods largely assume, observations enables robotics, real-world observations enables, largely assume rigid

备注: [this https URL](https://shuzhaoxie.github.io/CoDimRecon/)

点击查看摘要

Abstract:Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit joints. For deformables, category-wise agent sessions reconstruct curves as centerlines with radii, surfaces as manifold shells with thickness, and volumes as watertight solids for volumetric meshing. Reusable simulator skills initialize compatible physical models and parameters, while agent-guided behavioral tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling. On evaluated Replica and ScanNet++ scenes, CoDimRecon achieves competitive compositional reconstruction accuracy while additionally producing deformable assets for rod, shell, and solid simulation. We further demonstrate robot interactions across all three representations, including a controlled paper-folding case in which behavioral testing motivates plastic bending.

270. 【2609.36014】Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

链接:https://arxiv.org/abs/2609.36014

作者:Chong Wang,Zixuan Fu,Shiqi Huang,Siyuan Yang,Hao Cheng,Bihan Wen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Pixel-space diffusion Transformers, typically undergo uniform, diffusion Transformers, hidden representations typically, representations typically undergo

备注: Project page and code: [this https URL](https://chongwang1024.github.io/PerF)

点击查看摘要

Abstract:Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet $256\times256$, PerF-L achieves FID of $1.91$, approaching $1.86$ of JiT-H with only half the parameters, while PerF-H further achieves FID of $1.63$ and $1.76$ on ImageNet $256\times256$ and $512\times512$, respectively.

271. 【2609.36001】Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study

链接:https://arxiv.org/abs/2609.36001

作者:Rafael Garcia-Dias,Alexandre Triay Bagur,Chayanin Tangwiriyasakul,Virginia Fernandez,Parhom Esmaeili,Piyalitt Ittichaiwong,Yang Li,Lawrence Adams,Wason Buncharoen,Martin Chapman,Benjamaporn Chayanond,Sadthavud Chunrod,Tanawat Fongsri,Kass Gibson,Supat Plungprasertkul,Supawit Tangpanithandee,Kanyakorn Veerakanjana,Vicky Goh,Michela Antonelli,Joe Zhang,Kongkiat Kespechara,Sebastien Ourselin,M. Jorge Cardoso

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)

关键词:healthcare remains challenging, Federated Learning Interoperability, Learning Interoperability Platform, rebuilding governance guarantees, Federated learning

备注: 25 pages, 3 figures, 6 tables. Code and data: [this https URL](https://github.com/londonaicentre/FLIP)

点击查看摘要

Abstract:Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training and evaluation repeatable. FLIP implements common FL workflows as a set of composable services: cohort queries against per-site structured databases, on-demand DICOM retrieval from institutional PACS, per-site project approval, and reusable FL job types. To demonstrate FLIP, we ran two distinct use cases, federated fine-tuning and federated evaluation, on synthetic chest X-ray cohorts across two client nodes based in the United Kingdom (UK) and Thailand. In FLIP, each institution independently approves its participation in each project and operates its own node under local IT security processes. This study makes an operational rather than an algorithmic claim. It does not compare federated with centralised training; for that question, we refer the reader to existing systematic reviews and meta-analyses. The central result is evidence that such platforms enable international FL collaboration and improve repeatability, auditability, and site-specific governance. We also present a comprehensive comparison of existing platforms to help researchers and operators choose the right platform for their use case.

272. 【2609.35965】Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

链接:https://arxiv.org/abs/2609.35965

作者:Yunzhe Xu,Zhe Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:real-world applications require, applications require teams, largely focused, real-world applications, applications require

备注: 39 pages, 18 figures, 16 tables

点击查看摘要

Abstract:Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent's exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, planning, and execution, highlighting the challenges of coordinating under MAVLN task constraints. Project page: this https URL.

273. 【2609.35955】HEIR: Learning Human-Entity Interactions with Functional Roles

链接:https://arxiv.org/abs/2609.35955

作者:Di Wen,Wenhao Guo,Yuedong Tan,Yun Huang,Minheng Wu,Zhihang Chen,Haiwen Sun,Fei Teng,Zhiyuan Gao,Yufeng Zhang,Yuanhao Luo,Jingqi Zhang,Yufan Chen,Junwei Zheng,Ruiping Liu,Jiale Wei,Kailun Yang,Kunyu Peng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Understanding human-entity interactions, interactions requires recovering, person-action event participants, human-entity interactions requires, Understanding human-entity

备注: 24 pages, 4 figures. Code and dataset: [this https URL](https://github.com/Kratos-Wen/HEIR)

点击查看摘要

Abstract:Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at this https URL.

274. 【2609.35943】HERO: Histology Encoder for Robust Representation in Oncology

链接:https://arxiv.org/abs/2609.35943

作者:Zhi Li(1),Eghbal Amidi(1),Yating Cheng(1),Tyson Dawson(1),Gorkem Can Ates(1),Shuzhen Kuang(1),Norsang Lama(1),Md Ashequr Rahman(1),Zhiying Lu(1),Elisabeth K. Kong(1),Milan Radovich(1),David Spetzler(1),Matthew Oberley(1),George W. Sledge(1),Ming Chen(1) ((1) Caris Life Sciences, Irving, TX, United States)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:pathology foundation model, Foundation models trained, foundation model trained, provide strong, foundation model

备注: 22 pages, 4 figures, 13 tables

点击查看摘要

Abstract:Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may introduce downstream errors and hinder safe clinical adoption. A pathology foundation model should therefore be robust to acquisition shift without giving up representation quality, yet robustness is seldom the axis along which models are compared. In this report, we introduce HERO (Histology Encoder for Robust Representation in Oncology), a ViT-G/14 pathology foundation model trained with the DINO and iBOT objectives and refined with high-resolution Gram anchoring on a morphology-balanced corpus of 500 million tiles from approximately 575,000 clinical whole-slide images. Across the evaluated public benchmarks, HERO shows the strongest robustness to center, scanner, and stain variation among the compared state-of-the-art foundation models, performs comparably on tile-level classification, segmentation, and gene-expression prediction, ranks first on average across 39 evaluated slide-level clinical tasks, and, under an equal-weighted framework-level analysis, has the best average rank across the six benchmark frameworks.

275. 【2609.35916】VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

链接:https://arxiv.org/abs/2609.35916

作者:Jie Yang,Jiajun Chen,Jiazheng Zhou,Mianqiu Huang,Yining Zheng,Yuxin Wang,Xipeng Qiu

类目:Multiagent Systems (cs.MA); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Real-world embodied agents, pursue independent objectives, Real-world embodied, shared physical environment, pursue independent

备注:

点击查看摘要

Abstract:Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.

276. 【2609.35910】he Decision Value of Perception Compute

链接:https://arxiv.org/abs/2609.35910

作者:Hoang Pham Cong,Ho Viet Duc Luong

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Adaptive perception spends, perception spends extra, Adaptive perception, spends extra computation, expected to improve

备注:

点击查看摘要

Abstract:Adaptive perception spends extra computation on inputs where perception is expected to improve. When perception feeds a downstream decision system, a better perception output need not produce a better decision. We define the decision value of perception compute as the change in downstream loss from escalating an input from a cheap to an expensive perception mode. Because this value can be negative, the allocation of perception compute should be judged against a budget-constrained decision oracle, with uniform full-fidelity inference as a baseline rather than an upper bound. We introduce DEEP (Decision Evaluation for Escalated Perception), a benchmark that scores pre-escalation allocators against this oracle under selection, latency and energy budgets, charging each allocator for its own computation. With deployed monocular geometry on KITTI and nuScenes, we find that 34--54% of the escalations that change downstream loss make it worse; harmful escalations also occur for the published PDM-Closed planner, evaluated open-loop on nuPlan with real detector outcomes. On nuScenes, perception-level gain frequently disagrees in sign with decision value. This mismatch has practical consequences: choosing among fixed deployable signals by missed-object perception gain rather than by decision value reduces realized test decision gain by 7.4% of the all-cheap loss on average. Learned allocators recover part of the oracle's value by finding beneficial escalations but select nearly as much harm as random, and once their own computation is charged at a 20% latency budget, only the lightweight routers, at about 3.5\% of a full detector pass, still beat random.

277. 【2609.35865】PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

链接:https://arxiv.org/abs/2609.35865

作者:Yida Lin

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Computer Science and Game Theory (cs.GT)

关键词:Single-token typed-decision models, Single-token typed-decision, typed-decision models answer, one-letter answer codes, schema question

备注:

点击查看摘要

Abstract:Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores most of the structure in their training data. We study such a model whose data is curated as contrastive pairs---two contexts that differ in one edited fact that flips the answer---each carrying a machine-checked certificate that deleting the decisive sentence makes the fact unknown. We propose PACT, which turns this structure into four training terms that need no new annotation: a difference-in-differences margin over each pair that is invariant to any shared logit offset, a permutation-consistency term against answer-code position bias, an evidence-necessity term on certificate-verified ablated contexts, and an ordinal transport cost for rubric fields, plus a three-parameter contextual temperature. On a frozen 324-item holdout with three seeds, PACT matches the published recipe in accuracy ($84.6\%$ vs. $85.2\%$; McNemar $p \ge 0.50$ at every seed) while giving the lowest position bias of all runs (answer flips under relabelling $9.8\%$ vs. $13.8\%$) and the lowest ordinal error on rubric fields (MAE $0.232$ vs. $0.311$). Against a control with the same optimiser and schedule but cross-entropy only, PACT is significantly more accurate at two of three seeds, halves the seed-to-seed spread and lowers NLL by $26\%$. Seed-matched ablations and pre-specified falsification tests locate these gains precisely: no single term raises raw accuracy, and the method's value lies in robustness and stability rather than headline accuracy. Code, data splits, trained adapters, and all run records are available at this https URL.

278. 【2609.35823】CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning

链接:https://arxiv.org/abs/2609.35823

作者:Kang Yang,Shuai Liu,Hang Li,Yance Fang,Deying Li,Yongcai Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:made substantial progress, Vision-language models, made substantial, substantial progress, progress in autonomous

备注: 32 pages, 11 figures, 13 tables. Main text 9 pages; appendices from page 17

点击查看摘要

Abstract:Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional cooperative-driving systems typically transform them into geometric representations for downstream perception and planning. Directly incorporating these views into VLMs offers an opportunity to improve cooperative scene understanding and trajectory planning. However, question answering and trajectory planning have not been jointly evaluated on the same real-world vehicle-infrastructure scenes. We present CoVLM-Bench, a benchmark for cooperative driving question answering (CDQA) and cooperative planning (CP) on vehicle-infrastructure paired scenes. CoVLM-Bench provides scene-grounded CDQA annotations, three-part rationales as auxiliary supervision, and future trajectory targets derived from recorded ego motion. It contains 2,196 paired frames with 35,136 CDQA annotations, while CP predicts six waypoints over a three-second horizon. The annotations combine model-assisted drafting, record-based computation, and human verification. Built upon CoVLM-Bench, we introduce CoVLM-Drive, a unified VLM baseline that directly uses paired views for both CDQA and CP. Experiments show that CDQA adaptation improves answer accuracy and that CoVLM-Drive reaches a lower FDE than the compared V2X planners; QA initialization and rationale supervision each reduce planning error. Together, CoVLM-Bench and CoVLM-Drive support the training and comparison of VLMs for cooperative scene understanding and planning.

279. 【2609.35800】HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization

链接:https://arxiv.org/abs/2609.35800

作者:Nenad Banfic

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:sharply degrade vision-language, cache quantization saves, degrade vision-language model, sharply degrade, degrade vision-language

备注:

点击查看摘要

Abstract:Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries. Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bits, the six-task discriminative mean over eight models rises from 0.436 to 0.580 on the weakest base; protection can also improve generated answers and caption fidelity to BF16 outputs. Mean accuracy gains persist across all three quantizers with both tested calibration datasets. Keys-only protection retains substantial recovery at lower modeled storage cost. Evaluated through simulated quantization, HeadGuard offers a composable way to improve low-bit VLM accuracy without replacing the underlying quantizer.

280. 【2609.35150】oward a Culturally Adapted Chinese Language Agent: A Wizard-of-Oz Study of Nonverbal Behavior in Chinese-German Intercultural Interaction

链接:https://arxiv.org/abs/2609.35150

作者:Siddhant Jain,Anna Lea Reinwarth,Dimitra Tsovaltzi,Rafael Math,Julia Renner

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Successful intercultural communication, intercultural communication requires, Successful intercultural, grammatical competence, intercultural communication

备注: Accepted to ICMI Companion '26. 7 pages, 4 figure

点击查看摘要

Abstract:Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existing language-learning agents. We present a Wizard-of-Oz (WoZ) study design and supporting real-time system for collecting multimodal behavioral data from native Chinese speakers reacting to social norm violations by German learners. The system features a photorealistic MetaHuman avatar driven by Live Link face capture and MediaPipe upper-body tracking, a wizard console for real-time behavior selection, and synchronized multimodal logging across agent and learner streams. A layered annotation framework, based on psychological theory and covering non-observable socioemotional reactions, norm interpretation, verbal, and observable behavior thereof, and future supervision targets enables the corpus to support training of future automated cultural interpretation and behavior generation models. Four ecologically valid interaction scenarios, developed with cultural and pedagogical experts, provide the methodological and technical foundation for a culturally adapted conversational agent for Chinese language learning.

281. 【2603.15106】PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units

链接:https://arxiv.org/abs/2603.15106

作者:Mark Deutel,Simon Geis,Axel Plinge

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Enabling efficient deep, deep neural network, Enabling efficient, efficient deep neural, typically requires DNN

备注: Accepted at ECML-PKDD 2026. 18 pages, 7 figures, 4 tables. This work was funded by the European Commission as part of the MANOLO project under the Horizon Europe programme Grant Agreement No.101135782

点击查看摘要

Abstract:Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately. To avoid the huge manual effort, one can use neural architecture search (NAS). However, many existing NAS methods are resource-intensive and time-consuming because they require the training of many different DNNs from scratch. Furthermore, they do not take the resource constraints of the target system into account. To address these shortcomings, we propose PrototypeNAS, a zero-shot NAS method to accelerate and automate the selection, compression, and specialization of DNNs to different target microcontroller units (MCUs). We propose a novel three-step search method that decouples DNN design and specialization from DNN training for a given target platform. First, we present a novel search space that not only cuts out smaller DNNs from a single large architecture, but instead combines the structural optimization of multiple architecture types, as well as optimization of their pruning and quantization configurations. Second, we explore the use of an ensemble of zero-shot proxies during optimization instead of a single one. Third, we propose the use of Hypervolume subset selection to distill DNN architectures from the Pareto front of the multi-objective optimization that represent the most meaningful tradeoffs between accuracy and FLOPs. We evaluate the effectiveness of PrototypeNAS on 12 different datasets in three different tasks: image classification, time series classification, and object detection. Our results demonstrate that PrototypeNAS is able to identify DNN models within minutes that are small enough to be deployed on off-the-shelf MCUs and still achieve accuracies comparable to the performance of large DNN models.

282. 【2609.36702】Quantum Fidelity Landscape-Guided Prior Calibration for Single-Circuit QGAN Image Generation

链接:https://arxiv.org/abs/2609.36702

作者:Xue Yang,Rigui Zhou,Dax Enshan Koh,Siong Thye Goh,Yitao Tang,ShiZheng Jia,Young-Wook Cho,Hongyu Chen

类目:Quantum Physics (quant-ph); Computer Vision and Pattern Recognition (cs.CV)

关键词:Noisy Intermediate-Scale Quantum, attracted increasing attention, Generative Adversarial Networks, quantum machine learning, Noisy Intermediate-Scale

备注:

点击查看摘要

Abstract:Quantum Generative Adversarial Networks (QGANs) have emerged as representative generative models in the Noisy Intermediate-Scale Quantum (NISQ) era and have attracted increasing attention in quantum machine learning. However, most existing QGAN methods rely on patch-based decomposition strategies, which weaken the global consistency of generated images and increase quantum resource overhead. In this work, we investigate a simpler approach: pixel-level, end-to-end image generation using a single-quantum-circuit QGAN. By analyzing the structural matching relationship between the quantum prior and the target data distribution in Hilbert space, we provide a new theoretical perspective for understanding the training behavior of naive end-to-end QGANs. Specifically, we introduce the Quantum Fidelity Landscape (QFL), defined as the pairwise-fidelity structure induced by an ensemble of quantum states and preserved under shared unitary transformations of the quantum generation process. We show that, under a fixed Lipschitz readout, this invariant imposes a one-sided bound on decoded sample separation, motivating calibration of the prior-induced QFL before adversarial training. To validate this theoretical insight, we propose BasicQGAN, a QGAN framework incorporating quantum prior calibration. Before adversarial optimization, BasicQGAN aligns the prior-induced QFL with the data-induced QFL. Experimental results on small-scale grayscale image datasets show that BasicQGAN achieves stable and effective end-to-end pixel-level image generation while requiring fewer qubits and trainable parameters than representative patch-based quantum generators. Furthermore, experiments with different initial quantum-state ensembles show that QFL-calibrated ensembles achieve better generative performance.

283. 【2609.36400】CAMEO: A Class-Activation-Mapped Equitable Overlay Framework for Fair and Robust Deep Learning-based Skin Condition Diagnosis

链接:https://arxiv.org/abs/2609.36400

作者:Youssef Attia,Debasmita Mukherjee

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Deep learning classifiers, reach high in-distribution, Deep learning, high in-distribution accuracy, spurious background cues

备注: 21 pages, 9 figures

点击查看摘要

Abstract:Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robustness and fairness across skin tones. This work asks whether Explainable AI (XAI), typically used only to audit a finished model, can instead be repurposed as an active training signal that corrects this shortcut without sacrificing diagnostic accuracy. We introduce CAMEO (Class Activation Mapped Equitable Overlay), a framework that improves skin-lesion classification by selecting stable model explanations and using them to separate lesions from their backgrounds. It then replaces the background with realistic synthetic skin while keeping the lesion unchanged. On HAM10000 and dark-skin ISIC images, CAMEO maintained accuracy while reducing background-driven errors by nearly four times. It also made the model's attention more consistent when backgrounds changed. Results across multiple tests show that reducing reliance on background information improves robustness, with Fitzpatrick-based backgrounds providing a realistic and interpretable approach. Results show that XAI-guided augmentation can make dermoscopic classifiers measurably more robust and fair at no cost to accuracy. They also clarify that it is the mechanism and not the specific tone palette that matters, and that the lasting contribution of XAI here lies in stability-screened, annotation-free lesion localisation rather than in the robustness number itself.

284. 【2609.36227】One-Step Next-Latent Prediction Is Not a World Model

链接:https://arxiv.org/abs/2609.36227

作者:Shitong Wang,Zhongang Cai,Yuzhou Hong

类目:Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Next-latent prediction fits, Next-latent prediction, Next-latent, error, one-step

备注: 24 pages, 3 figures

点击查看摘要

Abstract:Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward innovation covariances. That error grows with $K$ after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient $0.9$, the one-step mean squared error is $0.998$ and the $16$-step open-loop error is $5.10$. On a hidden rotation, an eight-step window reaches $16$-step error $0.056$, while the current scalar alone reaches $0.778$. Raising the isotropy weight from $0.1$ to $10$ leaves eight-step latent error inside $[0.78,0.85]$ on three seeds.