本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新1792篇论文,其中:

  • 自然语言处理244篇
  • 信息检索35篇
  • 计算机视觉288篇

自然语言处理

1. 【2610.06851】Base Models Can Reason By Taking a Cue From Training Data

链接:https://arxiv.org/abs/2610.06851

作者:Sophie L. Wang,Amil Dravid,Rulin Shao,Kevin Farhat,Sewon Min,Alexei A. Efros

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:data creates associations, base model response, base model, creates associations, training data creates

备注: Project page: [this https URL](https://www.sophielwang.com/cues) Code: [this https URL](https://github.com/sophicle/cues)

点击查看摘要

Abstract:In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.

2. 【2610.06843】Recursive Video In-Context Learning for Agentic Robot

链接:https://arxiv.org/abs/2610.06843

作者:Wenrui Bao,Xinxin Liu,Bingxin Xu,Yuzhang Shang

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:LLM agents, orchestrate frozen, policies improve, text memory, improve across episodes

备注:

点击查看摘要

Abstract:LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.

3. 【2610.06830】MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

链接:https://arxiv.org/abs/2610.06830

作者:Haozhen Zhang,Haodong Yue,Quanyu Long,Jianzhu Bao,Qingyuan Liu,Tao Feng,Bohan Liu,Weida Liang,Wenya Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:supporting information retention, LLM agent ecosystem, supporting information, reuse across interactions, information retention

备注: Code is available at [this https URL](https://github.com/ViktorAxelsen/MemPilot)

点击查看摘要

Abstract:Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.

4. 【2610.06829】CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

链接:https://arxiv.org/abs/2610.06829

作者:Yifan Zhang,Yutong Dai,Viraj Prabhu,Zhiyuan Hu,Ran Xu,Zeyuan Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:binary task success, realistic browser tasks, execute realistic browser, binary task, Compositional Conformal Certifier

备注:

点击查看摘要

Abstract:Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.

5. 【2610.06825】PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

链接:https://arxiv.org/abs/2610.06825

作者:Yaohui Zhang,Binxu Li,Haoyi Duan,Jiacheng Miao,Yixin Wang,Xinran Du,Chenyue Li,Shilong Liu,Kevin Wu,James Zou

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:reusing published findings, making accurate plot, real scientific figures, Scientific figures, encode quantitative results

备注:

点击查看摘要

Abstract:Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing benchmarks rely largely on synthetic charts or cover only a limited range of chart types. We introduce PlotGround, an automated pipeline for building plot digitization benchmarks from real scientific figures and their author-released source data. PlotGround maps figures to source tables, identifies reconstructable panels, and generates quantitative questions with source-grounded reference values. We use PlotGround to construct PlotGround-1k, a human-verified benchmark of 1,119 questions from 1,066 bioRxiv preprints. Across sixteen multimodal models, the best reaches 87.5% accuracy at a $\pm 5\%$ relative-error tolerance. Tightening the tolerance to $\pm 2\%$ lowers every model's accuracy by 11-24 percentage points, revealing a gap between approximate visual reading and precise quantitative recovery. PlotGround's paired figure-source structure lets us compare how accurately the same values are recovered from figures and from source tables. Providing source tables instead of figures raises a coding agent's accuracy from 90.0% to 97.4% while cutting cost by 72%.

6. 【2610.06817】Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

链接:https://arxiv.org/abs/2610.06817

作者:Sahil Mahendrakar

类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)

关键词:widely used open, Paradee, teacher, Kokoro architecture, Paradee keeps Kokoro

备注: 16 pages, 2 figures, 8 tables. Code: [this https URL](https://github.com/sahilmahendrakar/paradee) . Model and audio samples: [this https URL](https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0)

点击查看摘要

Abstract:We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately against the frozen teacher. It has 10x fewer parameters and needs 15x less compute. We first synthesize a corpus with the teacher and keep its durations, pitch, energy and phoneme features. We then train a small text side to predict these values, and a small decoder to turn the teacher's saved values into the teacher's audio, first with spectral losses and then adversarially. Finally, we connect the two halves and quantize the weights to int8. It needs no alignment learning and no joint training, and it runs on one laptop. Stored in int8, Paradee is 8.5 MB, runs 25x faster than real time on one CPU thread, and scores 4.41 on UTMOS against the teacher's 4.52. The student initially kept a slight buzz, which we trace to the phase of voiced speech between 2 and 8 kHz. A phase-locking filter applied after synthesis removes most of it, with no training and no extra parameters. Code, model files and audio samples are at this https URL

7. 【2610.06782】-Search: An Open Agentic Retriever and Playground for Hard Multi-Step Search

链接:https://arxiv.org/abs/2610.06782

作者:Olga Tsymboi,Ramil Latypov,Aleksandr Medvedev,Danil Taranets,Dmitrii Stoianov,Nikita Gulyakov,Gleb Alektorov,Anatolii Potapov

类目:Computation and Language (cs.CL)

关键词:open-weight agentic retriever, hard multi-step search, open-weight agentic, agentic retriever, retriever for hard

备注:

点击查看摘要

Abstract:We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator can be swapped without retraining. T-Search is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic search tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks with gold evidence annotations, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts, outperforming larger open models. We release the model, harness, live demo, and three benchmarks, including TRuST, the first native-Russian hard-search benchmark.

8. 【2610.06778】IdeaLens: Detecting AI Ideas in Long-form Writing

链接:https://arxiv.org/abs/2610.06778

作者:Rishanth Rajendhran,Minjoon Choi,Jenna Russell,Ramya Namuduri,Deniz Bölöni-Turgut,Marzena Karpinska,John Wieting,Mohit Iyyer

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:emerging policies, IdeaLens, detectors identify, ideas, Pangram

备注: 53 pages (9 main), 7 figures, 50 tables. Code: [this https URL](https://github.com/RishanthRajendhran/IdeaLens) Models and data: [this https URL](https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f) Demo: [this http URL](http://ideadetector.ai/)

点击查看摘要

Abstract:While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.

9. 【2610.06750】Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs

链接:https://arxiv.org/abs/2610.06750

作者:Hyunji Lee,Joykirat Singh,Zaid Khan,Justin Chih-Yao Chen,Elias Stengel-Eskin,Alessandro Sordoni,Arman Cohan,Mohit Bansal

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:recurrent layers, recurrent, layers, Recurrent-attention hybrid language, attention

备注: Code: [this https URL](https://github.com/amy-hyunji/Balancing-Memory-Pathways)

点击查看摘要

Abstract:Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of disparate information over long contexts. However, we observe that simply having access to both pathways does not mean that hybrid LMs are effectively using them. We find that they rely substantially more on attention than on the recurrent state. Standard supervised fine-tuning improves overall performance but does not improve how the two memory pathways are coordinated: the model becomes more reliant on information propagated by attention layers, while its use of information propagated by recurrent layers remains limited. To encourage better coordination between the two memory pathways, we add an auxiliary loss that limits attention's access to earlier context while the recurrent state propagates through the full sequence. This objective encourages the model to retain and use information through the recurrent pathway alongside attention. It improves overall performance, with particularly strong gains on tasks involving longer contexts or requiring information aggregation, consistent with the strengths of recurrent layers observed in analysis. Crucially, this imbalance and the benefit of our auxiliary loss generalize: they apply to multiple recurrent-attention LMs in question-answering and agentic tasks, as well as to attention-based LMs that combine different forms of memory. Together, our findings show that simply providing multiple memory pathways does not ensure their effective use, and that targeted supervision is needed to better coordinate them.

10. 【2610.06744】ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring

链接:https://arxiv.org/abs/2610.06744

作者:Sait Furkan Teke(ufak AI)

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Turkish decision model, open Turkish decision, Turkish decision, Turkish text, released model

备注: 9 pages (text on pages 1 to 8, references on pages 8 and 9). Model, code, benchmark and demo: [this https URL](https://huggingface.co/ufakai/ufakzeka-karar) , [this https URL](https://github.com/ufakai/ufakzeka-karar) , [this https URL](https://huggingface.co/datasets/ufakai/HakemBench) , [this https URL](https://karar.ufakzeka.com)

点击查看摘要

Abstract:ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without generating text and in one CPU forward pass for up to ten options. Built on the lab's ufakzeka-1-base, its head scores each option blind to the others at shared positions, so the answer does not depend on option order. A sequential head trained with shuffled options was about as accurate but changed 2.3 to 2.8 percent of its answers when only the option order changed; REINFORCE lost 10.2 points (0.102) of macro F1 to cross-entropy. On the open set of HakemBench v1.0 (4,275 questions, 7 tracks) the released model ranks 7th of 16 rows with a composite of 0.660 (95% interval 0.642 to 0.677). Temperature scaling lowers calibration error (smooth ECE) on the development set but raises it on held-out support questions, from 0.027 to 0.045 for the first scored run, which never trained on them; the released model later trained on them, so its 0.036 to 0.064 is not an unseen-question test. The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run's new training data was aimed at the first run's errors on the full test set in guardrails, moderation and customer support, and the released run was trained after the second run's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. All its numbers come after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. Weights and code are under Apache-2.0.

11. 【2610.06729】Improving Diversity in LLM Short Story Generation

链接:https://arxiv.org/abs/2610.06729

作者:Zahra Solati Dehkordi,Vasileios Lampos

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, generate accurate responses, language models, generate accurate

备注:

点击查看摘要

Abstract:Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation in genre, tone, style, and named entities. To promote diversity across these dimensions, we introduce DivLM, an LLM post-training framework consisting of two phases. First, we perform continued pre-training on a creative writing corpus and restore instruction-following capabilities using weight residuals. We then apply reinforcement learning with a custom, composite reward function that jointly maximizes diversity across the targeted narrative dimensions while maintaining response quality. Our empirical results on two LLM families show that DivLM increases diversity metrics by more than 9% on average compared to alternative approaches, while preserving instruction following, overall response quality, and similarity to human outputs.

12. 【2610.06715】Domain adaptation of Russian ModernBERT for long legal documents

链接:https://arxiv.org/abs/2610.06715

作者:I. Litvak,D. Gvozdetsky,F. Lashkin,V. Kirova,S. Lagutin,V. Volf,T. Maksiyan,A. Kostin,R. Leva,I. Kiselev

类目:Computation and Language (cs.CL)

关键词:Russian legislative documents, Russian ModernBERT encoder, Russian ModernBERT, legislative documents improves, Russian legislative

备注: 17 pages, 11 figures, 4 tables

点击查看摘要

Abstract:We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 corpus tokens. Corpus token counts are distinguished from positions produced by the model tokenizer. We compare the original and adapted encoders on a fixed external collection of 1,031 court-decision segments. Both models receive the same hidden positions in each of five masking realizations. At maximum input lengths of 512, 2,048, and 8,192 tokens, mean masked-token cross-entropy decreases by 0.10942, 0.07052, and 0.06604 natural-log units, respectively. The reported 95% intervals summarize sensitivity to masking on this fixed collection; they do not quantify uncertainty across document collections. A second evaluation addresses legal-entity extraction. The original and adapted models achieve entity-level F1 scores of 0.99852 and 0.99820. However, 99.95% of test spans have the same normalized surface form and class in the training split. This evaluation therefore provides limited evidence about transfer to previously unseen forms. The paper explains the masking objective, overlapping windows, averaging rules, and exact entity-boundary scoring using editable diagrams and clearly marked illustrative examples. The comparison supports lower masked-token prediction loss for the studied pair of models and collection. It does not isolate the contribution of distant context or establish practical legal utility.

13. 【2610.06708】SAFE-MR: Evidence Sufficiency Learning for Selective Multimodal Rumor Detection

链接:https://arxiv.org/abs/2610.06708

作者:Shiwen Ni

类目:Computation and Language (cs.CL)

关键词:Multimodal rumor detectors, rumor detectors increasingly, detectors increasingly rely, Multimodal rumor, rumor detectors

备注:

点击查看摘要

Abstract:Multimodal rumor detectors increasingly rely on retrieved evidence, yet relevant evidence is not necessarily sufficient for verification. Missing provenance, duplicated reports, and unresolved contradictions can produce confident predictions without adequate support. We introduce SAFE-MR, a framework that separates claim veracity from evidence sufficiency. The method decomposes image-text posts into verifiable claims, constructs a relation-aware claim-evidence graph, and aggregates evidence using provenance and contextual compatibility. Separate veracity and sufficiency heads support selective prediction, while evidence interventions encourage stability under irrelevant additions and sensitivity to evidence removal. On NewsCLIPpings, VERITE, and XFacta, SAFE-MR achieves macro-F1 scores of 91.2%, 75.8%, and 85.2%, respectively. Against the matched backbone with evidence, its macro-F1 gains are 2.2, 4.9, and 4.8 percentage points. On the diagnostic selection set, SAFE-MR reduces AURC from 0.105 for maximum-probability rejection to 0.075 and lowers error at 80% coverage from 13.8% to 8.5%. Evidence-perturbation and ablation results support the role of sufficiency learning and intervention training in improving selective verification.

14. 【2610.06703】Reading the Mood: Emotion-Guided Book-to-Music Recommendation via CGANs and LLMs

链接:https://arxiv.org/abs/2610.06703

作者:Manousos Linardakis,Georgios Alexandridis

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:motivating recommender systems, make readers feel, Generative Adversarial Network, Aware Generative Adversarial, Conditional Generative Adversarial

备注: 9 pages, 5 figures, 5 tables. Accepted at SENTIRE 2026 (ICDM 2026 Workshops)

点击查看摘要

Abstract:Background music that matches the mood of a text has been shown to make readers feel more immersed and improve their reading experience, motivating recommender systems that pair books with mood-matched music. In this direction, we present Sentiment Aware Generative Adversarial Network for Cross Domain Recommendation (SAGA-CDR), a two-phase cross-domain recommendation framework that personalizes music suggestions and emotionally aligns them with the book being read. In the first phase, transformer-based sentiment embeddings are constructed from user reviews and mapped across domains via a Conditional Generative Adversarial Network, whose mask-conditioned generator handles missing sentiment components and injects stochasticity for richer preference transfer. A compact rating neural network then fuses sentiment-specific interaction scores with a collaborative filtering prior to predict music ratings. In the second phase, large language models classify each book into a valence-arousal emotional quadrant, and candidate tracks are filtered to match that quadrant. Experiments on both the English Amazon and Chinese Douban datasets show that SAGA-CDR achieves the best rating prediction accuracy on Amazon (RMSE 0.98) and the lowest RMSE on Douban (0.91), with ranking performance competitive with the strongest sentiment-aware baseline, even in cross-lingual settings.

15. 【2610.06695】MedPrune: Topology-Efficient Multimodal Multi-Agent Communication Evolution for Medical VQA Tasks

链接:https://arxiv.org/abs/2610.06695

作者:Jiuheng Wan,Runze Li,Chen Chen,Tingyuan Hu,Daiyang Yu,Yimin Jing,Taolin Zhang,Richang Hong

类目:Computation and Language (cs.CL)

关键词:large language models, existing clinical workflow-inspired, excessive computational overhead, computational overhead caused, advance medical visual

备注:

点击查看摘要

Abstract:While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.

16. 【2610.06689】Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation

链接:https://arxiv.org/abs/2610.06689

作者:Jiaming Qian,Huiyan Yang,Mandi Liu,Jie Liu,Wenkai Shen,Pengyang Zhou,Jing Jin,Jin Ma,Dezhi Ye,Chaochao Chen

类目:Computation and Language (cs.CL)

关键词:agent, Programmatic Search Agent, Search, search interfaces leave, fixed search interfaces

备注: 17 pages, 5 figures

点击查看摘要

Abstract:Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the returned evidence can reduce subsequent search. We introduce Programmatic Search Agent (PSA), which makes a local executable computation over candidates the unit of a search action. PSA unifies a persistent candidate workspace, flexible primitive composition, and selective evidence presentation. It incrementally generates program cells that reuse candidates, execute dependent operations, and select what the agent inspects next. The runtime resolves specified data dependencies within each cell, while the agent adapts its search strategy across cells as new evidence arrives. We compare PSA with the Query-based Agent and Tool-based Agent on InfoSeek-Eval and BrowseComp-Plus using five policy backbones without task-specific training. All three interfaces share the search substrate, and the Tool-based Agent also shares PSA's primitives and persistent workspace. Relative to the Query-based Agent, PSA improves macro-averaged task success by 4.00 and 7.56 percentage points on the two benchmarks, respectively; within-backbone reductions in final-step tokens average 28.3% and 33.9%. These results support extending agent control beyond query reformulation to the processing and presentation of retrieved evidence. Code will be released subject to approval.

17. 【2610.06685】Aligning Multimodal Patient Evidence with Biomedical Knowledge Graphs for Clinical LLMs

链接:https://arxiv.org/abs/2610.06685

作者:Jiawen Du,Arshan Ali Khan,Chenhao Zhang,Zachary Plotkin,Li Shen,Qi Long,Yun Li,Can Chen,Tianlong Chen,Nicholas Konz

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:existing predictive systems, predictive systems rarely, systems rarely represent, external biomedical knowledge, Multimodal Knowledge Graph

备注:

点击查看摘要

Abstract:Clinical questions often depend on linking a patient's multimodal evidence to external biomedical knowledge, yet existing predictive systems rarely represent such links explicitly, so they can neither be traced to their evidence sources nor removed to measure their contributions. We present MM-KG (Multimodal Knowledge Graph), which represents heterogeneous, multimodal patient observations and biomedical concepts as separate layers in one typed graph, joined by explicit alignment edges. First, modality-specific harmonizers convert EHR text, imaging, genomic, and biospecimen data into typed observations mapped to UMLS concepts, which a route-prioritized aligner links to a biomedical knowledge graph. Query-conditioned retrieval then selects a compact subgraph for downstream use by a large language model or a graph neural network. We build MM-KGs for MIMIC-IV and ADNI, and evaluate them with a 2x2 design that separates patient evidence, biomedical knowledge, and their interaction. On questions that require both sources, neither source alone performs far above chance, whereas their combination yields a drug-controlled AUROC interaction of +0.194 on MIMIC and +0.299 on ADNI. On held-out five-candidate ranking, MM-KG outperforms MindMap by +0.131 Hits@1 and leads an adapted GraphCare on the items that require consulting the patient, and deleting the single answer-bearing relation from the retrieved packet returns Hits@1 to the no-knowledge baseline. Finally, query-conditioned retrieval reaches 0.731 AUROC with 6.8x less context than the strongest generic policy, whereas static knowledge graph context gives no consistent gain on ordinary outcome prediction. Knowledge graphs thus benefit clinical LLMs not as background context but as explicit links between multimodal patient evidence and the relation a question requires, and MM-KG makes these links retrievable, traceable, and testable.

18. 【2610.06677】How Sparse Probability Maps Shape Mixture-of-Experts Routing

链接:https://arxiv.org/abs/2610.06677

作者:Tomás Brogueira,Marcos Treviso,Miguel Couceiro

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:typically apply softmax, routers typically apply, typically apply, softmax, expert

备注: 22 pages, 6 figures, 9 tables. Under review at ICLR 2027

点击查看摘要

Abstract:Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, 1.5-entmax, sparsemax and 2-normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores reaches a fixed threshold, and the trained routers differ in the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. While none of the sparse maps improves validation loss over softmax, they make the trained models far less sensitive to selecting more experts at inference: sparsemax trained with K=2 loses 0.02 nats when run with K=8, where softmax loses 0.58. Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.

19. 【2610.06673】he Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance

链接:https://arxiv.org/abs/2610.06673

作者:Stefan Bühler,David Exler,Markus Reischl,Mark Schutera

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:language models compliant, model, Claude, models compliant, language

备注: Accepted at URAI 2026

点击查看摘要

Abstract:Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and give up a higher payoff, which measures stoppability. The two compliance rates combine into a compliance index $\kappa$. Applied to twelve language models, the benchmark shows that seven mostly follow the instruction in both probes and justify their action by pointing to the instruction. Only Claude Sonnet-4.6 and Claude Opus-4.7 can be stopped without being exploitable, Claude Opus-4.6 and GPT-5-mini resist both instructions, and no model is exploitable but unstoppable. Knowing where a language model sits on the compliance index $\kappa$ matters for human operators and for multi-agent systems, whether distributed or orchestrated.

20. 【2610.06670】Reward Stealing Attack on Large Language Models

链接:https://arxiv.org/abs/2610.06670

作者:Jiaming Qian,Pengyang Zhou,Jiahe Xu,Chaochao Chen

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, induce harmful content, aim to induce, harmful content

备注: 19 pages

点击查看摘要

Abstract:Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at this https URL.

21. 【2610.06668】Language models can notice an impossible engineering problem yet still report it as solved

链接:https://arxiv.org/abs/2610.06668

作者:Shaoliang Yang,Jun Wang

类目:Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Computation and Language (cs.CL)

关键词:draft engineering calculations, Language models draft, Language models, engineering calculations, models draft engineering

备注: 36 pages, 6 figures, 3 tables; Supplementary Information included as an appendix; figure source data as ancillary files

点击查看摘要

Abstract:Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of their flawed counterparts. Each reply required a "solved" or "cannot solve" status; rejection meant "cannot solve" or withholding an answer. The initial prompts did not warn that problems could be flawed. Across three recent models, 12 of 90 replies failed to reject a flawed problem. In 11 of these replies, the model stated the flaw, answered a corrected problem and still reported the original as "solved", according to artificial intelligence raters and numerical checks. We later retested four models from one provider, offering "flawed" instead of "cannot solve" and asking them to name and explain the defect. Three models showed statistically significant increases in rejection, but valid-problem solving fell in three. Evaluations therefore need to score both versions and distinguish flaw recognition from the reported status.

22. 【2610.06666】What Matters for Latent Reasoning with Flow Matching

链接:https://arxiv.org/abs/2610.06666

作者:Yassine Ouali,Adrian Bulat,Georgios Tzimiropoulos

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language model, Flow-based Latent Reasoning, explicit CoT, Latent reasoning, large language

备注:

点击查看摘要

Abstract:Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.

23. 【2610.06650】Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents

链接:https://arxiv.org/abs/2610.06650

作者:Mohamed Chenene,Carlos Rosas-Hinostroza,Pierre-Carl Langlais,Anastasia Stasenko

类目:Computation and Language (cs.CL)

关键词:open knowledge bases, largest open knowledge, requires a SPARQL, SPARQL query, knowledge bases

备注: Technical Report

点击查看摘要

Abstract:Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less prominent entities. We study agents that instead answer by exploring the graph, and argue that two obstacles limit them: the lack of training data recording how a solver explores, and interfaces that add large graph results directly to the model's context. We test three hypotheses: that the difficulty of graph search can be controlled through the structure of a question rather than only through obscure entities or wording; that much of the failure on long-horizon search comes from how retrieved evidence is managed rather than from the model itself; and that, in a suitable environment, open-weight models can match commercial closed ones. We construct multi-hop questions on a frozen Wikidata snapshot by replacing named entities with nested conditions, checking after each expansion that the target remains unique and that every new condition is necessary. We release 10,235 solving traces over single-entity and multi-hop questions, together with the recursive language model (RLM) harness that produced them, in which models batch graph calls, keep results in persistent Python state and interpret selected evidence through sub-calls. On 100 questions, the harness improves both models we ran under both interfaces compared with direct tool calling over the same functions: gpt-6-luna rises from 49 to 61 correct answers, doubling its multi-hop accuracy, and Qwen3.8-27B, an open-weight model served on a single GPU, from 60 to 74.

24. 【2610.06648】Representation-Space MMD for Diffusion Language Models

链接:https://arxiv.org/abs/2610.06648

作者:Ilya Drobyshevskiy,Ilia Sudakov,Maksim Semenov,Denis Kuznedelev,Maksim Ignatov,Pavel Temirchev,Nikita Balagansky,Viacheslav Meshchaninov,Nikita Gushchin,Dmitry Baranchuk

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:frozen pretrained DLM, Maximum Mean Discrepancy, pretrained DLM, minimizes Maximum, reference distributions

备注: Tech Report. Code: [this https URL](https://github.com/yandex-research/mmd-dlm)

点击查看摘要

Abstract:We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.

25. 【2610.06647】LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

链接:https://arxiv.org/abs/2610.06647

作者:Shaokun Zhang,Yifan Zhang,Jian Hu,Yueying Li,Hao Zhang,Binfeng Xu,Jan Kautz,Yi Dong

类目:Computation and Language (cs.CL)

关键词:Reinforcement learning, broader adoption, memory demands remain, greatly advanced, advanced the capabilities

备注: 16 pages, 6 figures

点击查看摘要

Abstract:Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. Across reasoning tasks, LoGRA reduces average training memory by up to 45.7\% without sacrificing performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{this https URL}{Molt library}.

26. 【2610.06637】Long-Horizon Textual World Modeling through Structured Reasoning

链接:https://arxiv.org/abs/2610.06637

作者:Fangxin Wang,Xiang Gao,Yuguang Yao,Kaiwen Dong,Nikash Walia,Kamalika Das

类目:Computation and Language (cs.CL)

关键词:enabling agents, environment evolves, agents to compare, future actions, state

备注:

点击查看摘要

Abstract:World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions.

27. 【2610.06625】JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications

链接:https://arxiv.org/abs/2610.06625

作者:Matthew DiGiuseppe,Steven Denney

类目:Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:Large language models, constructs by generating, JEV, scale political text, Large language

备注: 71 pages, 2 figures, 14 tables (including appendices)

点击查看摘要

Abstract:Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.

28. 【2610.06621】Frozen Factor or Spectral Band? Disentangling Two Choices in Low-Rank LoRA

链接:https://arxiv.org/abs/2610.06621

作者:Adnan Slimane Ali,Ayoub Belfatmi,David Ngwe Pouth

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Spectral variants, low-rank adaptation, factor, rank, rates selected separately

备注:

点击查看摘要

Abstract:Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained weights, with learning rates selected separately. At rank 2, the same-band advantage of freezing A is larger than either within-factor band difference on all four task-model pairs with complete comparisons. Freezing B also trails comparable-budget free LoRA by 8-18 percentage points on five pairs spanning a formatting task and OpenBookQA. The A-frozen advantage persists in a single-GPU-model replication and within individual MLP module groups, including controls with equal or greater trainable counts for B frozen, and when A is frozen on a random orthonormal basis. The factor contrast weakens with rank. On OpenBookQA / Qwen2.5-1.5B at rank 16, PEFT's MiCA implementation trails comparable-budget LoRA by 3.08 points under a shared training recipe transferred from the MiCA paper. A trained oracle output subspace largely removes the low-rank deficit; partial warm-up gains recur across three direction seeds. The factor-versus-band ordering is descriptive; an approximate multiplicity audit weakens several earlier significance claims. These results extend known factor asymmetry by showing how its magnitude depends on spectral placement, rank and training conditions.

29. 【2610.06603】Word-Level Text Unmixing via Evidence-Preserving Ownership Routing with Language Models

链接:https://arxiv.org/abs/2610.06603

作者:Jinglin He,Siyang Jiang,Lixing He,Guoliang Xing,Hongkai Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:overlapping speech transcripts, document reading flows, Word-Level Text Unmixing, metadata is lost, reading flows

备注: 34 pages, 5 figures

点击查看摘要

Abstract:Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective. We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction. EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions. At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once. We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs. Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs. These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.

30. 【2610.06600】Molecules of a Story: Community Detection in PMI-weighted Narrative Networks

链接:https://arxiv.org/abs/2610.06600

作者:Kasper Fyhn,Rebekah Baglini

类目:Computation and Language (cs.CL)

关键词:Labatut and Bost, Automatically extracted narrative, revealing central narrative, Automatically extracted, revealing central

备注:

点击查看摘要

Abstract:Automatically extracted narrative networks -- graphs with entities as nodes and their relations as edges -- have proven useful for revealing central narrative structures through salient entities and their connections (Tangherlini et al. 2020; Labatut and Bost 2019). But a narrative is more than those central structures that everything else revolves around. This work is concerned with the everything else: brief sub-plots, small clusters of descriptions, or associations between minor characters that go under the radar at the macro-level. We present an approach to unearth such peripheral structures. They involve rare entities with limited textual presence, overshadowed by dominant entities and lost among each other in the long tail of many but rare entities (Baayen 2001). We leverage the known tendency of pointwise mutual information (PMI, Church and Hanks 1990) to inflate for rare events, turning its weakness into a strength by weighting edges with PMI to foreground peripheral entity configurations. Communities extracted from the resulting network are structural traces of underlying narrative elements. We demonstrate the approach on The Lord of the Rings. From measures of how concentrated or dispersed a community's activations are across the text, a typology emerges that reveals that peripheral structures form more than a single class: episodic passages, echoing long-distance connections, and recurring threads each surface as distinct configurations. The approach is conceptually simple and surfaces fine-grained narrative details that are lost in abundance, though its deliberate amplification of weak signals comes with inherent sensitivity -- best understood as a lens for exploration rather than a robust extraction pipeline.

31. 【2610.06587】Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants

链接:https://arxiv.org/abs/2610.06587

作者:Aadam Haq,Oggi Rudovic,Malcolm Chadwick,Jay Rainey,Shucong Zhang,Ricardo Guerrero,Sourav Bhattacharya,Maja Pantic

类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Automatic Speech Recognition, American English voice, English voice data, Speech Recognition, Automatic Speech

备注: ICASSP 2027 submission

点击查看摘要

Abstract:AI voice assistants often use Automatic Speech Recognition (ASR) with LLM-based reasoning, yet existing systems struggle with regional British accents, including Scottish, Irish, and Welsh accents, since most ASR models are trained predominantly on American English voice data. Consequently, errors can carry through to the LLM stage, corrupting tool-call arguments and producing wrong or missing responses, which is especially costly in finance. Deployable ASR must also meet tight latency and memory budgets, making an accent-robust model choice even harder. We introduce CavaBench, the first internally collected benchmark of spoken financial queries, and use it to evaluate a range of ASR models and their end-to-end ASR-LLM pipeline behaviour across self-reported British accents. We find that WER strongly predicts downstream tool-calling accuracy ($r = -0.93$) but can fail to reflect task-level performance, with accent-related failures varying substantially across models and acoustic conditions. These findings guide the design of more inclusive, reliable voice-based financial assistants.

32. 【2610.06582】Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

链接:https://arxiv.org/abs/2610.06582

作者:Shengtao Wen,Xiang Chen,Yu Tian,Lingbing Guo,Lina Gong,Sheng-Jun Huang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:World-model controllers rely, real control systems, execute commands asynchronously, commands asynchronously due, prediction and planning

备注:

点击查看摘要

Abstract:World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model controllers, rather than treating it only as an external control disturbance. Through controlled interventions, we identify two architecture-dependent failure modes: planning-based controllers such as TD-MPC2 suffer from a future-action timeline mismatch between imagined and executed action sequences, while recurrent world models such as DreamerV3 can attribute observed transitions to commands that were not actually applied. Our analysis shows that TD-MPC2 requires the correct future action sequence during latent dynamics rollout, whereas DreamerV3 requires timely attribution of each transition to the action that generated it. Based on these findings, we introduce two lightweight execution-consistent interfaces, Future-Sequence for TD-MPC2 and Applied-Action Feedback for DreamerV3, that correct these mismatches without modifying the pretrained world models. Experiments across delays, packet loss, reordering, multiple control domains, measured network traces, and a process-separated asynchronous stack consistently support both diagnoses and the corresponding architecture-specific corrections.

33. 【2610.06576】Before Agent Tells The Lie: Has Deception Already Been Represented?

链接:https://arxiv.org/abs/2610.06576

作者:Xinling Li,Dadi Guo,Qingyu Liu,Qinghua Mao,Yi R. Fung,Na Zou,Xia Hu,Dongrui Liu

类目:Computation and Language (cs.CL)

关键词:including hiding failures, signaling task completion, Large language model, falsely signaling task, Large language

备注:

点击查看摘要

Abstract:Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely signaling task completion. Existing monitoring approaches mainly detect deception after it appears in observable actions or outputs. In this paper, we investigate whether deceptive behavior can be predicted from an agent's internal representations before it becomes externally visible. We frame deception monitoring as a trajectory-level representation analysis problem and align agent trajectories around key decision points. Using hidden states extracted before these points, we show that future honest and deceptive outcomes can be reliably distinguished, with predictive signals remaining detectable several model calls before the final decision. We further characterize the temporal evolution of these signals: deception-related representations are weak early in execution but become increasingly identifiable as trajectories progress, while transferable structure can emerge before the strongest decision-adjacent signals appear. Finally, we intervene on the identified honest-deceptive representation directions during inference and find that activation steering reduces downstream deceptive behavior, suggesting that these representations influence agent decisions. Our findings indicate that agent deception is an evolving internal process that can be detected and potentially mitigated before it is expressed externally.

34. 【2610.06522】Anatomy of LLM Sycophancy: What a Flip Rate Hides

链接:https://arxiv.org/abs/2610.06522

作者:Haonan Huang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:flip rate counts, capitulation alike, counts a correction, measured flip rates, rate counts

备注:

点击查看摘要

Abstract:A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape measured flip rates. Each measurement is one stateless replay of an item, a committed answer, and one scripted user line in a fixed form. Every effect is read against a matched control with the line deleted. Pushback wording, committed text, answer format, boundary distance, and ground truth become factors of one instrument; earlier instruments vary one to three of them. Across eleven frontier models from three providers and about 760,000 controlled replays, which models look sycophantic depends on how the user pushes back. Lines that assert the opposite verdict and lines that challenge the answer without asserting one rank the models almost unrelatedly. Flip effects grow several-fold near a model's boundary, yet items answered identically in every screening draw still carry about half of the most-affected totals. On arithmetic tasks where the truth is known, one model re-derives and corrects itself under pressure while another abandons correct answers without written work. On the model tested, a planted derivation lowers release of the answer it argues for, true or wrong, where a bare stated value does not; the wrong answer is corrected much more often than the true one is abandoned. Under a yes/no readout the rankings come closer, entangled with a pressure-induced shift toward "no". One score per model therefore compares different behaviours across models and benchmarks. We condense these dependencies into a reporting profile; the instrument, records, and analyses will be released upon publication.

35. 【2610.06516】st-Time Adaptation of Reasoning Strategies with Bayesian Nonparametric Memory

链接:https://arxiv.org/abs/2610.06516

作者:Keshav Ramji,Tahira Naseem,Ramón Fernandez Astudillo

类目:Computation and Language (cs.CL)

关键词:modern large language, generation cost grows, cost grows substantially, grows substantially due, large language models

备注:

点击查看摘要

Abstract:While modern large language models (LLMs) have been trained to reason through verbalized chains-of-thought, the generation cost grows substantially due to suboptimal paths to reach the final answer. Furthermore, as new insights are discovered while observing various input queries (e.g. through self-reflection), limited mechanisms exist for carrying forward these findings to be applied to subsequent problems. One can view the list of such strategies or behaviors as a growing cheatsheet, with elements retrieved from this memory module at inference-time. In this work, we consider structured cheatsheets, with learned clusters of behaviors. We introduce a Hierarchical Dirichlet Process Gaussian Mixture Model (HDP-GMM) over behavior embeddings, which shares components across domains while allowing domain-specific mixing weights, and uses the posterior predictive to retrieve relevant behaviors for a query; we call this a $\textit{Bayesian Cheatsheet}$. This mechanism allows for cheap adaptation in an online test-time training (TTT) setting, softly updating the mixture's sufficient statistics following each sample and enabling the creation of new components when the synthesized behaviors are sufficiently novel. We demonstrate that Bayesian Cheatsheet achieves clear performance gains relative to existing memory modules across reasoning benchmarks such as AIME'25, Omni-MATH, and PhysReason, even in the cold-start setting. We show that the Bayesian Cheatsheet is an adaptively reorganizing memory module, as behaviors can be re-assigned to components through a single step of collapsed Gibbs sampling. Our findings highlight the value of Bayesian-inspired memory modules for effective test-time adaptation and the role of structure in metacognitive reasoning.

36. 【2610.06513】SOL: Measuring Gaps between Text Distributions by Double Sliced Wasserstein Metrics

链接:https://arxiv.org/abs/2610.06513

作者:Gregor Kornhardt,Moritz Piening,Jannis Chemseddine,Gabriele Steidl

类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:Evaluating text generation, generation requires measuring, text generation requires, generated distribution matches, Evaluating text

备注:

点击查看摘要

Abstract:Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.

Subjects:

Computation and Language (cs.CL); Machine Learning (cs.LG); Machine Learning (stat.ML)

Cite as:
arXiv:2610.06513 [cs.CL]

(or
arXiv:2610.06513v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.06513

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
37. 【2610.06481】AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

链接:https://arxiv.org/abs/2610.06481

作者:Jiaqi Xue,Yanjun Wang,Xiangci Li,Lingbo Mo,Aritra Sengupta,Shweta Garg,Murali Krishna Ramanathan,Myeongsoo Kim

类目:Computation and Language (cs.CL)

关键词:repository-level coding tasks, increasingly tackle complex, tackle complex repository-level, complex repository-level coding, agents increasingly tackle

备注:

点击查看摘要

Abstract:As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent. To coordinate their interdependent work, these agents share findings and agree on interfaces between modules. However, exchanged information often serves only as context, leaving individual agents to interpret it and incorporate it into subsequent work. Consequently, shared findings may go unused and deviations from interface agreements may go undetected, undermining the reliability and efficiency of collaboration. This motivates moving part of the coordination responsibility from individual agents to the execution harness. To make shared information actionable during execution, we introduce the Artifact-Exclusive Communication Protocol (AECP). AECP requires agents to communicate exclusively through structured artifacts and specifies how the harness processes them. The harness supplies findings when agents access relevant code, screens implementations for mismatches with recorded interface commitments, and requires affected agents to revisit revised agreements. These coordination steps become part of harness execution rather than actions that agents must initiate from prior messages. Across Doc2Repo, NL2Repo, and CodeProjectEval, using closed- and open-source models including Opus-4.8 and DeepSeek-V4-Flash, AECP improves average test pass rate by 28.2% and reduces average wall time by 16.5% relative to an agent team using free-form inter-agent messages. Artifact-exclusive communication also blocks the relay of malicious instructions between agents, reducing how often they reach other agents from 95% to 0% and how often those agents act on them from 40% to 0%.

38. 【2610.06479】Behavior-Preserving KV Cache Compression

链接:https://arxiv.org/abs/2610.06479

作者:Doo Hwan Hwang,Junyoung Jang,Junho Na,Hosung Lim,Kee-Eung Kim

类目:Computation and Language (cs.CL)

关键词:large language models, major bottleneck, bottleneck in long-context, long-form generation, generation with large

备注:

点击查看摘要

Abstract:KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We argue that cache compression should instead preserve the predictive behavior of the full-cache model, retaining entries whose removal would substantially change the model's output distribution. We propose Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution. Using pre-eviction forward statistics, the method avoids running separate masked forward passes for each candidate. Across diverse architectures and both prefill-time and generation-time compression, our method delivers substantial gains in downstream task quality over lightweight attention-based heuristics at matched retained-KV budgets, with the largest gains under aggressive compression. It achieves these gains with additional compression-time computation while retaining an end-to-end speedup over full-cache inference in our evaluated settings.

39. 【2610.06446】he Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

链接:https://arxiv.org/abs/2610.06446

作者:Jakub Macina,Manu Kapur,Mrinmaya Sachan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, Large language, questions are natively, natively poor, student

备注:

点击查看摘要

Abstract:Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.

40. 【2610.06439】Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models

链接:https://arxiv.org/abs/2610.06439

作者:Subramanyam Sahoo,Justin Shenk

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:Relative Policy Optimisation, Group Relative Policy, Group Relative, Policy Optimisation, model

备注: 11 Pages , Accepted at AI for Law Workshop @ ICML 2026 also accepted for publication in the Proceedings of Machine Learning Research (PMLR)

点击查看摘要

Abstract:What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy collapses from 0.500 (chance) to 0.072 (McNemar p 10^-36), driven entirely by the rate of properly formatted answers falling from 0.900 to 0.109. The model stops committing to answers. Yet when it does commit, accuracy rises from 0.556 to 0.657, showing that the collapse is not a failure of capability but a strategic response: the model has learned that verbose responses packed with citations but empty of a direct answer score higher than terse correct ones. We term this the Saul Goodman effect, a policy that becomes maximally lawyerly while becoming maximally noncommittal, and prove formally that it is the optimal response to any surface feature proxy that attaches no penalty to abstention. We further show that 89.3% of citations produced after training are structurally implausible hallucinations, many of them subtly corrupted names of real landmark cases, constructed in effect to survive a casual read and fail under scrutiny. To detect this failure mode before deployment, we introduce three diagnostic tools: the Confidence Theater Score (CTS), the Citation Plausibility Rate (CPR), and the Regret Gap (RG). In a domain where a confidently wrong answer can constitute malpractice, the broader lesson is direct: a reward function that measures how legal a response looks will produce a model that is maximally photogenic and minimally useful.

41. 【2610.06437】HeuFouFT: Task-Guided Metaheuristic Coordinate Search for Fourier Fine-Tuning

链接:https://arxiv.org/abs/2610.06437

作者:Ruiheng Wang,Yubo Hou,Yakun Zhu,Tianle Shen,Tao Wan,Zengchang Qin

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:introduce Heuristic-Guided Fourier, Heuristic-Guided Fourier Fine-Tuning, Fourier Fine-Tuning, Heuristic-Guided Fourier, Fourier

备注:

点击查看摘要

Abstract:We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-agnostic rules. HeuFouFT instead searches for coordinates using downstream performance. A coarse intensity map from lightweight block-level probes initializes three metaheuristic optimizers: Genetic Algorithm with Simulated Annealing (GA-SA), Particle Swarm Optimization (PSO), and Cuckoo Search (CS). During search, a Random Forest filters each population so that only the top 30% of candidates proceed to proxy fine-tuning. On E2E with GPT-2-Medium, all three variants outperform random-uniform FourierFT, Gaussian band-pass FourierFT, and LoRA across five metrics. PSO further outperforms LoCA, the best-performing baseline, on four metrics while using 37.6% fewer trainable spectral coefficients. Once coordinates are selected, HeuFouFT requires only 15--18% FLOPs of Full FT. These results show that task-guided search allocates limited spectral capacity more effectively than fixed sampling. Our code is publicly available.

42. 【2610.06430】Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?

链接:https://arxiv.org/abs/2610.06430

作者:Paula A. Perez-Toro,Tomas Arias-Vergara,Annette Schwarz,Abner Hernandez,Andreas Horr,Cornelia Kristen,Andreas Maier

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:Regional accent cues, matched conditions, accent cues, remains unclear, captured under matched

备注: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classification and continuous geolocation under matched conditions and speaker-independent read--spontaneous transfer. Whisper performs best under matched conditions, reaching 0.489 nine-way UAR and 148 km median geolocation error, but drops to 0.11/0.18 UAR across transfer directions and 363 km geolocation error. Self-supervised models show a similar degradation, whereas speaker embeddings are less discriminative in-domain but more robust under transfer. This contrast is consistent across classification and geolocation. Across representations, robustness is associated with how little a representation shifts between styles (style-invariance), for which crossstyle speaker retrieval is an interpretable proxy. Age, sex, sentence-overlap, and duration controls do not account for the gap, although channel characteristics contribute. These results show that strong matched-condition performance does not indicate robust regional information.

43. 【2610.06413】SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs

链接:https://arxiv.org/abs/2610.06413

作者:Rafael Teixeira Sousa,Vinícius Paulo Lopes de Oliveira,Elisa Ayumi Masasi de Oliveira,Luiza Martins de Freitas Cintra,Fernanda Bufon Färber,Igor Dias Aguiar,Julia Yasmim de Almeida Nobre,Arlindo Rodrigues Galvão Filho

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:reflects faithful spatial, prediction reflects faithful, faithful spatial reasoning, report ever-higher accuracy, correct prediction reflects

备注: Accepted at the 2nd Workshop on Embodied Spatial Reasoning (ESR), NeurIPS 2026. 29 pages (8 main), 9 figures, 18 tables. Code and data: [this https URL](https://github.com/spatialchain/SpatialChainBenchmark)

点击查看摘要

Abstract:Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve $\geq$79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking ($\rho$ = 0.88). Data, generation scripts, and evaluation code are released at this https URL.

44. 【2610.06406】What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

链接:https://arxiv.org/abs/2610.06406

作者:Zhongxiang Sun,Jiahao Yan,Hongkang Zhao,Haojie Ding,Boheng Zhang,Fan Yang,Xiao Zhang,Jun Xu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:overseeing autonomous execution, making individual decisions, long-horizon tasks, shift from making, making individual

备注:

点击查看摘要

Abstract:As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementary dimensions: alignment between requirements and behavior, and awareness of consequential autonomous decisions for verification. To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph. EBG presents task-oriented views of this graph to help monitors interpret behavior in context. Experiments across eight models show that EBG improves decision identification and evidence localization in most settings compared with direct access to the original context. Further experiments show that EBG's evidence-localization gains persist across input scales and hyperparameter settings, while real-world applications illustrate its practical value for human oversight.

45. 【2610.06401】RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

链接:https://arxiv.org/abs/2610.06401

作者:Mohamed Dhouib,Clement Elliker,Alexi Canesse,Maël Jenny,Lucas-Andrei Thil,Mahammed El-Sharkawy,Sonia Vanier,Elie Bursztein

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Tool-using language-model agents, untrusted external content, Tool-using language-model, external content, language-model agents

备注:

点击查看摘要

Abstract:Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.

46. 【2610.06383】Steering by Influence: Curvature Aware Data Weighting for Activation Steering

链接:https://arxiv.org/abs/2610.06383

作者:James A. E. Dixon,Stephen J. Roberts,Francesco Quinzan

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Inference-time steering offers, steering offers cheap, Inference-time steering, language model outputs, offers cheap

备注: Code: [this https URL](https://github.com/JDIXON-2/Concept_Activation_Transport)

点击查看摘要

Abstract:Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than thematic concepts. In this work, we steer towards examples that most express a concept thematically, rather than towards an expectation over all. We identify these examples using influence functions, which estimate how much each data point contributes to a model's representation of a concept. Unlike simple model activation similarity, they incorporate the curvature of the model's loss landscape, allowing them to capture concept-relevant relationships beyond superficial token-level similarity. We then propose influence-weighted activation transport, which uses optimal transport to steer activations of non-concept text towards those of concept text, weighting concept examples by their influence scores. We evaluate on toxicity suppression (Jigsaw), object-based concept induction (OneSec) and truthfulness induction (TruthfulQA), outperforming existing activation-transport baselines. We track capability after steering using perplexity and MMLU accuracy, finding that our method improves steering while largely preserving model quality. We further show that influence functions capture concept-relevant information that activation-based methods miss with the two approaches ranking data points significantly differently. Together, these results demonstrate the value of curvature-aware influence information for activation steering.

47. 【2610.06360】Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering

链接:https://arxiv.org/abs/2610.06360

作者:Aditya Tanna,Abhishek Jindal

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Reinforcement learning post-training, Reinforcement learning, language models relies, human preferences, learning post-training

备注: Accepted at CIKM 2026

点击查看摘要

Abstract:Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a soft verifier from a maintained controlled vocabulary: UMLS Concept Unique Identifier overlap (via scispaCy, set-level F1) gives a graded, externally specified reward computed without a model in the loop. We combine it inside GRPO with an entropy-normalised LLM judge, which covers the safety and evidence axes overlap cannot see, and a small consistency penalty on padding and repetition that keeps early-training samples scorable. This three-term composite improves over SFT on Phi-3-mini (3.8B) over MedQA by 2.9% on EM (0.700 vs 0.680) and 39% on Token-F1 (0.202 vs 0.145); on Llama-3.2-3B the corresponding gains are 14% on EM and 35% on Token-F1. We report Token-F1 as the primary metric because it credits partially-correct clinical content that EM discards at this open-generation scale. Main-table results are means over 3 seeds with standard deviations below 0.005. The method transfers to PubMedQA, where training on the PubMedQA train set with the same composite reward improves Token-F1 over SFT by 22% on Phi-3-mini and 17% on Llama-3.2-3B without retuning. A reward ablation on Phi-3, varying the judge-ontology split at a fixed consistency weight, attributes 3 EM points to the ontology term, the contribution that catches entity substitutions the judge cannot. Three negative findings constrain the design: DPO under random negatives underperforms SFT for strong-prior models but helps the weakest-prior one; PPO under a sparse neural reward diverges; GRPO with KL-in-loss collapses at 7B.

48. 【2610.06345】Breaking Bureaucracy: Evaluating open-source LLMs for legal document review

链接:https://arxiv.org/abs/2610.06345

作者:Farrukh Baratov,Niki van Stein,Suzan Verberne

类目:Computation and Language (cs.CL)

关键词:Natural Language Inference, legal Natural Language, Language Inference, Natural Language, legal Natural

备注:

点击查看摘要

Abstract:In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not require labeled training data. We evaluate our models on the ContractNLI benchmark and two NLI4Wills datasets. We successfully reproduce the baseline for the task (Span NLI BERT) and we evaluate multiple open-source LLMs on the same task. We analyze the invalid rate of the models, and their stability across temperature settings and domains. Among the generative models, Gemma-4 26B performs the best, reaching an accuracy of 81.2%, even outperforming the supervised model on one metric. On accuracy, it is not possible to beat the supervised model with zero-shot approaches. Qwen-3.6 35B performs well on both ContractNLI and additional datasets in the legal wills domain. Our findings indicate that zero-shot, open-source, generative LLMs are a viable alternative for real-world legal NLI when no supervised data is available. Our code is available at this https URL.

49. 【2610.06322】Agentic schema-guided extraction of materials process knowledge from scientific literature

链接:https://arxiv.org/abs/2610.06322

作者:Sameer Sadruddin,Jennifer D'Souza

类目:Computation and Language (cs.CL); Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Emerging Technologies (cs.ET)

关键词:measurements remain difficult, process-specific context, remain difficult, difficult to aggregate, reported in heterogeneous

备注: 15 pages, 3 figures, submitted for review to Nature Communications Materials

点击查看摘要

Abstract:Materials literature contains detailed experimental knowledge, but procedures, chemical entities and measurements remain difficult to aggregate because they are reported in heterogeneous forms and depend on process-specific context. We present SciKGExtract, a schema-guided framework that combines large-language-model extraction with chemical normalization and agent-based evaluation and refinement before knowledge-graph integration. We evaluate the framework on 176 atomic-layer-deposition papers describing zinc oxide (ZnO) and indium--gallium--zinc oxide (IGZO), together with an expert-annotated full-schema subset. PubChem normalization improves exact-match extraction F1 for every tested model. For ZnO, the best F1 increases from 0.591 for direct normalized extraction to 0.805 with agentic refinement, whereas the best IGZO result is 0.344, revealing the greater difficulty of multicomponent supercycle processes. Evaluation against a deeply nested schema containing 65 experimental properties and 155 quantitative measurement nodes further exposes errors in process segmentation and numerical assignment. These results show that chemical canonicalization and targeted agentic verification provide complementary controls for converting complex materials literature into reusable, machine-actionable experimental knowledge.

50. 【2610.06298】DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task

链接:https://arxiv.org/abs/2610.06298

作者:Saad Ezzini,Shadi Abudalfa,Maram Alharbi,Salmane Chafik,Hind Alatawi,Mo El-Haj,Ahmed Abdelali,Osamah Alnahari,Salima Lamsiyah

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Natural Language Processing, Language Processing Conference, Arabic Natural Language, Natural Language, Language Processing

备注: Accepted at ArabicNLP 2026

点击查看摘要

Abstract:Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and Swapping in Arabic Dialects (DialectSentEval), hosted with the Arabic Natural Language Processing Conference (ArabicNLP 2026). This shared task consists of two subtasks: Subtask 1 focuses on multi-class and multi-dialect sentiment analysis, requiring models to identify sentiment polarity across various Arabic dialects. Subtask 2 introduces a generative task for Arabic sentiment swap, challenging models to invert sentiment polarity while preserving core semantics. In this overview paper, we present the motivation, dataset creation, and summarize the main findings from participating models.

51. 【2610.06287】From Abusive Language Classification to Sequence Labeling Identification

链接:https://arxiv.org/abs/2610.06287

作者:Nicolas Zampieri,Ignacio Lopez,Manon Girard,Jeremy Auguste

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Industrial content moderation, tight latency constraints, Abusive Language Identification, process massive message, massive message streams

备注:

点击查看摘要

Abstract:Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a sequence-labeling task that jointly extracts AL spans and target mentions, and assess whether this approach can be used for text moderation. On a pilot corpus drawn from a production moderation pipeline, we compare ALI with ALC on cross-domain generalization and implicit abuse, and we also evaluate AL and target span detection. ALI remains competitive with ALC while providing localized outputs for moderators, with a modest and configuration-sensitive advantage on implicit abuse. Exact AL boundaries and target spans remain difficult to recover. We complement this comparison with a qualitative analysis and discuss perspectives on complete target--span linking and on structured benchmarks for ALI.

52. 【2610.06286】DeferKV: Rethinking Eviction Timing for One-Shot KV Cache Compression

链接:https://arxiv.org/abs/2610.06286

作者:Zhe Wang,Jiakai Li,Yujia Sun,Rongzheng Wang,Shuang Liang

类目:Computation and Language (cs.CL)

关键词:Long-context large language, demonstrated strong capabilities, introduces substantial memory, cache introduces substantial, Long-context large

备注:

点击查看摘要

Abstract:Long-context large language models (LLMs) have demonstrated strong capabilities across a wide range of tasks, but the growing KV cache introduces substantial memory and inference overhead. Existing one-shot KV cache compression methods typically commit to irreversible eviction immediately after prefill, before any signal from actual generation becomes available. Our quantitative analysis shows that early queries from the actual generation stage provide attention signals that are more consistent with subsequent decode attention, with the largest single-step gain occurring at the prefill-decode boundary. Based on this observation, we propose DeferKV, which moves the eviction decision from the end of prefill to the first real decoding step and temporally combines prompt-side and decode-side observations, thereby better aligning KV importance estimation with subsequent generation requirements. DeferKV requires no additional training, draft model, or future-query prediction module, making it simple and easy to deploy. Experiments on LongBench, RULER, and Needle-in-a-Haystack demonstrate that DeferKV consistently improves model performance under KV cache compression while maintaining low inference latency.

53. 【2610.06273】Probabilistic Race and Ethnicity Prediction Using Group-Specific Name Lists

链接:https://arxiv.org/abs/2610.06273

作者:Kyla Chasalow,Noah Dasanaike,Kosuke Imai

类目:Computation and Language (cs.CL)

关键词:Statistically valid estimation, Bayesian Improved Surname, Improved Surname Geocoding, Statistically valid, geographic location

备注:

点击查看摘要

Abstract:Statistically valid estimation of racial and ethnic disparities often requires inferring the probability that an individual belongs to a particular racial or ethnic group given only their name and geographic location. The standard approach, Bayesian Improved Surname Geocoding (BISG), relies on group population frequencies for each name. Although the U.S. Census Bureau provides such information for common names and a limited set of racial categories, comparable data do not exist for many racial and ethnic groups and are rarely available outside the U.S. We propose the list-powered BISG ($\ell$BISG) method, which can be used to derive calibrated group probabilities from group-specific name lists. These lists may be compiled based on expert knowledge or generated synthetically using large language models (LLMs), and thus may be subject to unknown biases. Representing names as embeddings, we treat list membership as a proxy prediction task and apply a correction based on proximal inference to recover the target group probabilities. We validate the method on U.S. voter files with self-reported race, on the full-count 1900 U.S. Census, and on the Lebanese voter registry. We find that LLM-generated name lists yield accurate and well-calibrated probabilities as well as precise disparity estimates comparable to those obtained using methods that require name-race data. Thus, $\ell$BISG substantially broadens the applicability of probabilistic race and ethnicity prediction to settings where name-race data are unavailable.

54. 【2610.06251】Shared Stopping Decisions Change Answers in HQQ Cache Quantization

链接:https://arxiv.org/abs/2610.06251

作者:Seunghui Jwa,Minsu Oh,Chanjun Park,Yeo-Chan Yoon

类目:Computation and Language (cs.CL)

关键词:Language-model systems batch, Language-model systems, systems batch questions, systems batch, Language-model

备注:

点击查看摘要

Abstract:Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused during generation. With request-local groups, Transformers' Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop. Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models. Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions. Computing the stopping mean in FP32 reduces cache differences but leaves answer changes. Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs. Fixed iterations and request-local stopping remove observed companion dependence under matched controls. Request-local stopping remains sensitive to synthetic padding changes at the tensor level. Fixing the original iteration budget removes this decision path without tuning. Neither repair has an established quality advantage, and natural rebatching still changes answers. Request-independence audits must cover stopping decisions as well as quantization groups.

55. 【2610.06236】DP-ES: Differentially Private Evolution Strategies for Prompt Optimization

链接:https://arxiv.org/abs/2610.06236

作者:Ziniu Liu,Aiping Li,Yue Han,Han Yu,Junjian Zhang,Dong Zhu,Changjian Li,Shiqiang Zhang

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:Token-level differentially private, noise-sensitive irreversible choices, search trajectory reveals, trajectory reveals prompt-template, reveals prompt-template drift

备注: Accepted at EMNLP 2026 (Main Conference). Code: [this https URL](https://github.com/StephCpa/dp-es)

点击查看摘要

Abstract:Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains $49.5\pm28.5\%$ across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-sensitive irreversible choices. We diagnose these as structural consequences of greedy token-by-token construction over privately aggregated counts. We then propose DP-ES (Differentially Private Evolution Strategies), a structurally cleaner alternative that maintains a population of full prompts, mutates them via LLM calls that never access the private dataset, and spends privacy only on sampled-Gaussian evaluation; deterministic or Gumbel-smoothed selection is post-processing. Under a conservative $(\varepsilon\leq1.0,\delta=10^{-5})$ guarantee, DP-ES achieves 88.1% on GSM8K (+38.6 pp over DP-OPT, approximately 9 times lower standard deviation), 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. It is also 2.5 times faster in wall-clock time and uses 3.3 times fewer logged private-data call groups than DP-OPT. Selection and population ablations, implementation-level noise checks, and a 200-profile exact-match memorization stress test complement the formal guarantee. Scope: Our experiments establish optimization robustness under DP noise, especially where prompt structure is critical; end-to-end validation on genuinely sensitive, non-saturated deployment data remains future work.

56. 【2610.06216】Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

链接:https://arxiv.org/abs/2610.06216

作者:Andrea Paganelli,Stefano Civelli,Pietro Bernardelle,Gianluca Demartini

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Pre-generation success probes, model hidden activations, estimate response correctness, activations before decoding, Pre-generation success

备注:

点击查看摘要

Abstract:Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed success frequencies (CALIBRATION); and (3) whether they produce scores comparable enough across candidate models for cost-aware multilingual routing (UTILITY). Using 3,000 MATH problems in 10 languages and 8 open-weight model configurations, we compare cross-lingual transfer from English-trained probes and equal-budget pooled multilingual probes. English-trained probes retain useful cross-lingual discrimination but become less well calibrated after transfer. Pooled multilingual supervision improves both properties and yields more reliable estimates of success. In routing experiments, the pooled router achieves a 0.7% higher test success rate while reducing modeled cost by 13.0% relative to always selecting the model with the highest average success. These results show that multilingual routing requires success estimates that remain well calibrated and comparable across languages and models.

57. 【2610.06204】Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers

链接:https://arxiv.org/abs/2610.06204

作者:Pedro Tabacof,Sagar Joglekar

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:full customer experience, LLM agents, agents are starting, customer experience, full customer

备注: 20 pages, 3 figures, 9 tables. Accepted (poster) at the NeurIPS 2026 Workshop on SLMs for Agentic Systems (SLM-Agents), Paris

点击查看摘要

Abstract:LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152 negotiations against two frontier buyers it never saw in training. With the same learning rate ($10^{-6}$) for every size, the gain of the RL model over its base rises from $+0.001$ at 2.3B to $+0.078$ at 31B. Each size was trained once and the two smallest checkpoints use a different architecture, so we fit no scaling law. Tripling the learning rate, with the same or fewer training steps, improves on the shared rate at every size by $+0.032$ (2.3B) to $+0.081$ (4.5B). In exploratory comparisons with two frontier models run as sellers, the 12B seller trained at the tripled rate scores above both, though its untrained base already scores as high as they do. The 4.5B seller at that rate shows no detectable difference from either and fits on one 48 GB GPU. A further 2.3B arm at ten times the shared rate raises pooled score, but its gain concentrates on the evaluation buyer that shares a model family with the training pool. These results suggest tuning the learning rate before concluding that a small model cannot learn to negotiate, and testing against buyers from more than one model family.

58. 【2610.06191】Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

链接:https://arxiv.org/abs/2610.06191

作者:Chubin Zhang,Zhenglin Wan,Xingrui Yu,Jingxuan Wu,Yaxin Zhou,Ivor Tsang,Bo An

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:tool keeps returning, stop relying, agent, agents judge results, stopping

备注: 37 pages, 6 figures, 28 tables. Code: [this https URL](https://github.com/bennidict23/judged-useless-queried-anyway)

点击查看摘要

Abstract:An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero for clock- or deadline-driven stopping. Where we record their judgments, the seven agents we test call a failing source's results useless 97-100% of the time, yet most of them rarely stop on that judgment. Prompt cues change when they stop but not what they stop on. Permission to answer from memory and a reasoning mode can bring early stops regardless of evidence, a stated budget moves the 7-8B models' stops to the deadline, and a stopping rule or call cost in the prompt is followed at most partly. Stopping follows the evidence only when the harness enforces an integration step that makes the agent answer after five consecutive results it judged useless. This step raises failing-source success for every model, keeps the stopping point fixed when the budget doubles, and needs no extra judgment call when the agent states its judgments. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.

59. 【2610.06174】Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate

链接:https://arxiv.org/abs/2610.06174

作者:Yoshihiro Izawa,Gouki Minegishi,Yoko Yamakata

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:recognize degradation, LLMs recognize degradation, computational substrate, computational substrate induced, recognize

备注: 9 main pages with appendix

点击查看摘要

Abstract:Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their computational substrate induced by quantization. We first show that existing models fail to self-report their quantization state, even when provided with their own generated text as an external cue. Linear probing reveals that, while generated text carries almost no trace of quantization, internal representations contain clear, method-specific fingerprints. Through training, models learn to identify severely degraded outputs such as those of 4-bit models by comparison, yet still fail to do so from a single output. A shared LoRA trained jointly across quantization levels succeeded in reading out internal fingerprints, but fails on unseen quantization methods, merely mapping method-specific fingerprints to labels. Whereas external self-observation can restore awareness in some cases of human anosognosia, our results suggest that the more promising route to enabling such awareness in LLMs may lie in their internal representations. Our results highlight fundamental limits of generalizability to LLM self-monitoring.

60. 【2610.06170】MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge

链接:https://arxiv.org/abs/2610.06170

作者:Abdul Basit,Muhammad Abdullah Hanif,Muhammad Shafique

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Biomedical large language, large language model, requires auditable assessment, Biomedical large, source-grounded subspecialty knowledge

备注: 7 pages, 3 figures. Accepted for publication to BHI 2026

点击查看摘要

Abstract:Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.

61. 【2610.06161】Introducing Code-Switched Contexts to Cognitively-Inspired Bilingual Model Training

链接:https://arxiv.org/abs/2610.06161

作者:Zhuojing Huang,Luise Pohlmann,Lisa Beinborn

类目:Computation and Language (cs.CL)

关键词:cross-linguistic syntactic mapping, accelerate vocabulary growth, syntactic mapping, children are regularly, regularly exposed

备注: EMNLP 2026, BabyLM Challenge; 18 pages, 6 figures

点击查看摘要

Abstract:During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.

62. 【2610.06147】Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts

链接:https://arxiv.org/abs/2610.06147

作者:Seungmin Oh,Seunghun Kang,Jongbin Ryu

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Vision-language models, achieve strong zero-shot, strong zero-shot transferability, achieve strong, inference time

备注: Accepted for publication in Advances in Neural Information Processing Systems (NeurIPS) 2026

点击查看摘要

Abstract:Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose Test-Time Correction (TTC), a hypothesis-based correction framework guided by a simple principle: hypothesize, reconstruct, correct. Given a test feature and its top-k candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting divergence shift. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This training-free candidate-verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to 2x speedup, over 3x lower CPU memory usage, and up to 1.4x lower GPU memory usage than the lowest-memory training-free baseline.

63. 【2610.06100】From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

链接:https://arxiv.org/abs/2610.06100

作者:Quanyu Long,Xiao Chen,Jianda Chen,Haozhen Zhang,Qisheng Hu,Jianzhu Bao,Wenya Wang

类目:Computation and Language (cs.CL)

关键词:evaluating LLM agents, evaluating LLM, LLM agents, impractical to reproduce, increasingly valuable

备注:

点击查看摘要

Abstract:Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.

64. 【2610.06093】Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII

链接:https://arxiv.org/abs/2610.06093

作者:Alexandru Nazare,Agnese Profico,Nicolò Vania,Elena Di Croce,Daria Caramanica,Davide Venditti,Elena Sofia Ruzzetti,Giancarlo A. Xompero,Fabio Massimo Zanzotto

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)

关键词:extraction remain under-explored, cross-lingual data extraction, data extraction remain, Training Data Extraction, data extraction

备注:

点击查看摘要

Abstract:The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)

Cite as:
arXiv:2610.06093 [cs.CL]

(or
arXiv:2610.06093v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.06093

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Elena Sofia Ruzzetti [view email] [v1]
Mon, 5 Oct 2026 10:27:33 UTC (643 KB)

65. 【2610.06069】Attention Tax, Handoff Tax: A Stylised Model of When Multi-Agent LLM Systems Help

链接:https://arxiv.org/abs/2610.06069

作者:Akshit Anchan,Nayonika Sen

类目:Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:multi-agent LLM systems, LLM systems reaches, multi-agent gains grow, multi-agent LLM, systems reaches sharply

备注: 23 pages, 6 figures. Code and data: [this https URL](https://github.com/akshitanchan/attention-handoff-tax)

点击查看摘要

Abstract:Recent work on multi-agent LLM systems reaches sharply different conclusions: some results show that a single agent with the same information and compute should dominate a delegated system, others that multi-agent gains grow with task depth. We argue that much of the disagreement comes from modelling different bottlenecks, and introduce a stylised reliability model built around two trade-offs. Decomposition reduces the burden of long contexts but incurs a handoff tax when information is compressed or transferred between agents. Redundancy gains from multiple samples, but its benefit depends on how much their failures are shared. With reasoning budget, verification, and task structure added, the model yields two crossover conditions: decomposition becomes preferable once the attention cost avoided by resetting context exceeds the handoff cost, and parallel sampling at equal budget is eventually preferable when its shared-failure floor lies below the error floor of one agent thinking longer. We connect these regimes to recent theoretical and empirical results. On a ledger-reconciliation task we measure the context-degradation curve and the handoff tax from single-agent and handoff runs alone. From these the model places the crossover at depth 10 and predicts decomposition to win at depths 20, 50, and 100. It does, on step-level and final-balance accuracy, and the decomposed system's success, which the prediction never sees, lands within 9 percentage points of the predicted rate at every depth.

66. 【2610.06064】rustMI: Causally controlling how assistants trust their users

链接:https://arxiv.org/abs/2610.06064

作者:Théo Lasnier,Romain Froger,Maxence Lasbordes,Djamé Seddah

类目:Computation and Language (cs.CL)

关键词:Large Language Model, assistants routinely decide, Large Language, parties whose competence, routinely decide

备注: 27 pages, 12 figures, 11 tables

点击查看摘要

Abstract:Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply with harmful requests or act on malicious instructions encountered during tool use. To study this problem, we define trust as an assistant's willingness to accept vulnerability to the actions of another party and ask whether such behavior can be causally controlled through model activations. We build 2,000 contrastive conversations spanning ability, benevolence, and integrity, where paired responses complete the same request but differ in whether the assistant trusts the user. From these pairs, we learn steering matrices while keeping the model parameters frozen and test them across six instruction-tuned models from three families, finding that steering changes trust decisions monotonically in both directions. We then ask whether this effect extends to several safety-related agent settings involving harmful requests, prompt injections, and insider threats, while using benign-task and reasoning as controls. Our findings provide evidence that trust in the user can be causally controlled along linear directions in model activations and provide a way to study how trust shapes safety-relevant behavior in language models.

67. 【2610.06056】ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs

链接:https://arxiv.org/abs/2610.06056

作者:Yijing Du,Xiangcheng Zhan,Shuo Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large Vision-Language Models, Large Vision-Language, frequently suffer, suffer from object, Large

备注: Accepted in EMNLP 2026 Oral

点击查看摘要

Abstract:Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.

68. 【2610.06049】Backdooring Sparse Autoencoders

链接:https://arxiv.org/abs/2610.06049

作者:Enrico Ahlers,Daniel Passon,Tobias Kiecker,Eik Reichmann,Lars Grunske

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:Sparse autoencoders, interpret language models, internal representations, Sparse, language models

备注:

点击查看摘要

Abstract:Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.

69. 【2610.06031】LightMTP: Lightweight Latent Multi-Token Prediction

链接:https://arxiv.org/abs/2610.06031

作者:Tamara Czinczoll,Julie Kallini,Gerard de Melo,Chen Shani

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:standard pretraining objective, exploit local patterns, capturing longer-range structure, Next-token prediction, explicit training signal

备注:

点击查看摘要

Abstract:Next-token prediction (NTP) is the standard pretraining objective for large language models, yet it provides an explicit training signal only for the immediate next token, which can lead models to exploit local patterns instead of capturing longer-range structure and ideas. Multi-token prediction (MTP) addresses this by training models to predict several future tokens. However, existing MTP methods often introduce a large number of new parameters with limited improvements in downstream performance. Latent MTP approaches address this efficiency issue by encoding future tokens into a vector representation. However, these approaches usually rely on external helper models for future token encoding. We propose LightMTP, a lightweight, i.e., parameter-efficient, latent MTP approach that bootstraps the future token representations from the model's own hidden states. Our two LightMTP variants extend supervision to more future tokens without requiring the additional computational overhead of conventional MTP nor the external supervision latent MTP normally relies on. LightMTP adds at most 1% extra parameters, retains better performance on general language modeling benchmarks, and achieves similar gains in planning, coding, and reasoning.

70. 【2610.06026】Differentiable Bit-Widths: Co-optimizing Pruning and Quantization via SVD for Ultra-Efficient LLM Compression

链接:https://arxiv.org/abs/2610.06026

作者:Hankyul Kang,Jongbin Ryu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:large language models, language models, recently emerged, promising strategy, large language

备注: Accepted to Advances in Neural Information Processing Systems (NeurIPS 2026)

点击查看摘要

Abstract:SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings ($1.61$ bits) designed for ultra-efficiency. Code: this https URL.

71. 【2610.06018】Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

链接:https://arxiv.org/abs/2610.06018

作者:Eryk Kołodziejczyk,Alberto Presta,Karol Szurkowski,Michal Byra

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:natural language queries, Spatio-temporal video grounding, STVG models, aims to localize, space and time

备注: Accepted on EMNLP 2026 Findings

点击查看摘要

Abstract:Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.

72. 【2610.06011】D-Loop: Looped Diffusion Drafting for Speculative Decoding

链接:https://arxiv.org/abs/2610.06011

作者:Kecheng Chen,Yuyang He,Cheng Gong,Hui Liu,Guoping Long,Jiajun Li,Shi Wu,Suiyun Zhang,Haoliang Li,Ziru Liu,Rui Liu

类目:Computation and Language (cs.CL)

关键词:accelerates speculative decoding, drafting multiple tokens, diffusion accelerates speculative, accelerates speculative, speculative decoding

备注:

点击查看摘要

Abstract:Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.

73. 【2610.05982】Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models

链接:https://arxiv.org/abs/2610.05982

作者:Yao Lu,Zhaiyuan Ji,Yaxin Gao,Zeyu Wang,Zhe Tang,Jiaheng Wei,Zhaowei Zhu,Shanqing Yu,Qi Xuan

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, rich model ecosystem, Language Models, artificial intelligence

备注: 12 pages, 8 figures

点击查看摘要

Abstract:With the rapid development of artificial intelligence, the emergence of various Large Language Models (LLMs) has created a rich model ecosystem. However, this also brings a key challenge: how to select the optimal model for a specific user query. LLM routing addresses this need by dynamically assigning queries to the most suitable expert in the pool of candidate models. However, existing routing frameworks often simplify this process to a standard classification task; thus, a critical vulnerability is exposed when multiple candidate models correctly answer the same query. We formalize this capability overlap as routing noise, which misleads the router with arbitrarily correct candidate models, ultimately leading to routing collapse (a severe decline in generalization ability on unseen tasks). To address this problem, we propose a novel Cluster-Aware Soft-Labeling Routing (CASLR) framework. CASLR shifts the evaluation paradigm from the success of a single query to macro-domain consensus by replacing traditional one-hot vectors with a masked softmax mechanism. Specifically, for experts who answer incorrectly, we penalize their target probability to zero; for the remaining candidates, we directly compute continuous fine-grained soft labels based on their global clustering utility scores. We then use these refined soft labels to supervise a lightweight router. Specifically, the framework not only demonstrates superior accuracy on multiple benchmarks, but also outperforms Llama-3.3-70B-Instruct by 7.80% in overall average performance. Furthermore, the extremely low routing inference latency of only 1.13s further confirms that CASLR can achieve efficient system scheduling with almost zero additional overhead, while ensuring high response quality.

74. 【2610.05978】Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

链接:https://arxiv.org/abs/2610.05978

作者:Jie Wang,Shiwei Luo,Qi Zhang,Yuanbin Wu

类目:Computation and Language (cs.CL)

关键词:Tokenizer-free language models, resulting longer sequences, longer sequences substantially, sequences substantially increase, Tokenizer-free language

备注:

点击查看摘要

Abstract:Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to $25\%$ of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields $3.4\times$ more accepted tokens than in subword Transformers.

75. 【2610.05977】StagQ: Constraint-Driven Multi-Precision Weight Quantization for LLMs

链接:https://arxiv.org/abs/2610.05977

作者:Zhe Wei,Mengqi Guo,Yuan Yuan,Jiunn Bin Lim,Boyi Pan,Michael Bi Mi

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language model, Serving a large, weight-precision operating points, language model, large language

备注: 17 pages, 4 figures, 8 tables

点击查看摘要

Abstract:Serving a large language model (LLM) across a fleet of deployments requires several weight-precision operating points. Multi-precision formats serve them all from one stream whose prefixes are valid lower-precision codes, instead of storing multiple copies. We present StagQ, a multi-precision weight format whose main stream is a 2-bit group-wise affine base followed by a configurable number of 1-bit refinement planes on a dyadic step schedule. Every supported precision is a readable prefix, decoded by an affine map derived from metadata shared across all precisions, with no per-weight lookup. A sparse side record, filled both before and after the grid is fitted, holds out the few weights the grid serves worst. We report two configurations of the encoder. At two bits the cheaper one leads the strongest multi-precision baseline on Llama-3.1-8B, Phi-4, and OLMo-2-7B by 3.1 to 7.0 MMLU points, at a slightly lower logical rate. At three bits it leads on Llama-3.1-8B, leads on Phi-4 at a higher rate, and ties on OLMo-2-7B. At four bits it ties on all three, at a higher rate. In a batch-one matrix-vector product on an NVIDIA A100 GPU, timed on synthetic weights, our kernel is faster than the two baseline kernels in most shape-precision cases.

76. 【2610.05976】Can Language Models Learn to Reject Their Own Bad Reasoning Steps?

链接:https://arxiv.org/abs/2610.05976

作者:Siheng Xiong,Xiaoze Liu,Yiqiao Jin,Xiaoqian Wang,Jing Gao

类目:Computation and Language (cs.CL)

关键词:contaminating subsequent generation, prevent harmful reasoning, harmful reasoning steps, subsequent generation, Verifier-guided decoding

备注:

点击查看摘要

Abstract:Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's recoverability as the probability that the frozen generator can complete it correctly. Diagnostics show that adjacent recoverability changes are often difficult to resolve with practical Monte Carlo budgets, while same-prefix candidates exhibit a sparse low-recoverability tail. We introduce Self-Step Rejection (SSR), which trains a lightweight LoRA acceptance gate on the generator backbone while keeping the base model frozen. SSR uses confidence-qualified first-passage supervision: steps before the first resolved crossing of a root-relative recoverability barrier are accepted, the crossing step is rejected, and unresolved steps and suffixes are excluded. Training combines pointwise classification, same-prefix pairwise learning, and group-relative policy refinement using final-answer correctness. At inference, SSR accepts candidates or resamples from the unchanged prefix under rejection budgets, without an external learned verifier. Across three reasoning models and five mathematical reasoning benchmarks, SSR improves macro-average accuracy over single-pass decoding by 5.4--10.1 points using 1.21--1.40x as many generated tokens, and achieves the highest macro-average accuracy among evaluated step-level methods. Full-solution scaling methods require 4.47--8.27x the single-pass token cost for comparable performance.

77. 【2610.05966】HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

链接:https://arxiv.org/abs/2610.05966

作者:Junying Chen,Xinyuan Xie,Ziniu Li,Wenyuan Gu,Jianquan Li,Xiang Wan,Guangjun Yu,Ruoyu Sun,Haizhou Li,Benyou Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:general-purpose large language, Domain adaptation aims, target domain, large language model, aims to turn

备注: Extended version of "OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation", accepted at ICML 2026, with additional analysis and scaling to HuatuoGPT-3

点击查看摘要

Abstract:Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at this https URL.

78. 【2610.05896】asteRoute: Personalized Routing for Video Generation

链接:https://arxiv.org/abs/2610.05896

作者:Zhi Rui Tam,Chao-Chung Wu,Sin-Han Yang,Peyton Ku,Brendan Kuang,Tzu-Ting Hsieh,Min-Fang Hsu,Fang-Ling Tsai,Yun-Nung Chen,Wei-Chiu Ma,Chieh-Yen Lin

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Rapid progress, differ substantially, substantially in capability, Rapid, Abstract

备注:

点击查看摘要

Abstract:Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees with each annotator's own favorite only 34-55% of the time. Motivated by this observation, we introduce TasteRoute, a personalized video-generation router that selects a generator jointly based on the input request, user preferences, and available generation budget. Across text-to-video and image-to-video settings, TasteRoute is competitive with strong simple baselines on preference routing while reducing average generation cost. The cost saving increases under higher budget caps. Finally, we release TasteRoute-3k, a human-annotated dataset containing multi-model video comparisons, quality judgments, preference rankings, and user-profile signals to facilitate future research on personalized and cost-aware video routing.

79. 【2610.05894】Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

链接:https://arxiv.org/abs/2610.05894

作者:Sarim Hashmi,Mukul Ranjan,Abdelrahman Elsayed,Muhammad Umer Sheikh,Fahad Shamshad,Nils Lukas

类目:Computation and Language (cs.CL)

关键词:Masked diffusion language, denoising masked positions, iteratively denoising masked, diffusion language models, masked positions

备注:

点击查看摘要

Abstract:Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and we show that an adversary can exploit this. Since an answer remains open to revision over many denoising steps, an adversary with access to internal activations can watch how likely the model is to produce a chosen answer and adjust the intervention accordingly. Building on this observation, we study targeted bias injection, an attack that steers a frozen dLLM toward a demographic answer selected by the adversary. The attack uses a simple proportional-integral (PI) controller that tracks the target-answer probability during denoising and adapts the strength of a steering vector on the fly. On ambiguous BBQ questions where the correct answer is abstention, our attack raises LLaDA-8B-Instruct's preference for the targeted group from 1.8 to 16.7 percentage points, more than three times the strongest fixed-strength steering baseline, and on SocialStigmaQA it raises the selection of stigmatizing answers from 17.6% to 58.1%. Fitted to other demographic targets, the same attack shifts answers by up to 37 percentage points, and each attack takes about 40 minutes on one GPU. On the primary target, feedback is what makes the attack work: constant steering at the same average strength over the token-committing steps produces a far smaller shift while corrupting nearly three times as many outputs, and a constant strength set separately for each example still falls well short. Our findings identify the denoising trajectory as a new control channel in dLLMs and call for bias audits that examine the serving stack rather than the frozen model alone.

80. 【2610.05879】Learning to Learn a Language

链接:https://arxiv.org/abs/2610.05879

作者:Lennart Carstens-Behrens,Holger Fröhlich

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:byte-level transformer pretrained, synthetic non-linguistic prior, Prior-Fitted Language Model, byte-level transformer, present the Prior-Fitted

备注: 15 pages, 6 figures, 5 tables, Code: [this https URL](https://github.com/cbl/prior-fitted-language-model) , weights: [this https URL](https://huggingface.co/lennartcb/pflm1)

点击查看摘要

Abstract:We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.

81. 【2610.05872】Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning

链接:https://arxiv.org/abs/2610.05872

作者:Chen Henry Wu,Thomas Zhang,Aditi Raghunathan

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:long-standing goal, continual learning, data, update, learning

备注:

点击查看摘要

Abstract:A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.

82. 【2610.05842】HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

链接:https://arxiv.org/abs/2610.05842

作者:Zhuokun Chen,Xi Lin,Xiyu Wu,Jiahao He,Jianfei Cai,Bohan Zhuang

类目:Computation and Language (cs.CL)

关键词:distant information difficult, Linear attention enables, attention enables efficient, make selective access, Hybrid Linear Attention

备注:

点击查看摘要

Abstract:Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: this https URL

83. 【2610.05831】Selecting Long-Horizon Trajectories for Reliable and Efficient Terminal-Agent Training

链接:https://arxiv.org/abs/2610.05831

作者:Cuong Dang,Hoang Anh Just,Ruoxi Jia

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:supervise remains unexplored, imitating long teacher, Terminal agents, remains unexplored, commonly trained

备注: Submitted to ICLR 2027

点击查看摘要

Abstract:Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a key design axis for reliability and cost. Reliability improves with longer horizons but saturates: on Terminal-Bench, a 12K-token horizon solves more tasks than 16K ($29\pm0.7$ vs.\ $26\pm0.8$) while requiring 30\% less training time. The horizon also shapes agent behavior: short horizons cause premature termination, intermediate horizons yield productive error recovery, and long horizons induce over-persistence. We analyze this saturation through a bias--complexity bound, in which longer supervision reduces temporal supervision bias but increases finite-sample estimation error from more heterogeneous late-stage histories. Guided by this analysis, we propose \emph{selective long-horizon refinement}, which first trains on short prefixes and then refines only on continuations that are most likely under the warm-start model. It consistently outperforms full long-horizon training. At 16K, it raises successful attempts from $110\pm2.7$ to $126\pm2.1$ and tasks solved in at least six of eight attempts from $9\pm0.7$ to $14\pm0.6$; with half of the long-horizon data, it still reaches $122\pm2.4$ while cutting training time by 23\%. The gains transfer across benchmarks, from $64\pm2.6$ to $73\pm2.1$ on Terminal-Bench v2.0 and from $137\pm2.7$ to $155\pm2.2$ on OpenThoughts-TBLite. For long-horizon supervision, selecting the right trajectories matters more than training on all of them.

84. 【2610.05817】Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation

链接:https://arxiv.org/abs/2610.05817

作者:Alireza Jafari,Arman Adibi,Mohammad Ghavamzadeh,Hadi Daneshmand

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Text revision, Nash, integral component, Text, language models

备注: 34 pages, 6 figures, 11 tables. Code: [this https URL](https://github.com/alireza-jafari/Nash-Decoding)

点击查看摘要

Abstract:Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an $\varepsilon$-Nash equilibrium in $O(1/\varepsilon)$ time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to $18\times$ larger, without any fine-tuning or retraining, at the cost of additional test-time computation.

85. 【2610.05815】Plan Canvas: Fixed Reasoning Regions for Continuous Language Flows

链接:https://arxiv.org/abs/2610.05815

作者:Miaohe Niu,Pengxiang Li,Jingbo Zhu,Tong Xiao

类目:Computation and Language (cs.CL)

关键词:Continuous language flows, language flows generate, Continuous language, flows generate text, language flows

备注:

点击查看摘要

Abstract:Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start is therefore unknown during denoising, and the model has to decide the trace length, the place of every trace token, and the answer at the same time. We propose Plan Canvas to fix the boundary between the trace and the answer. A plan region of fixed capacity holds a compact trace, supervised padding fills its unused positions, and the answer starts at a fixed position. The fixed regions also allow separate denoising clocks for the plan and for the answer. With the trace text, backbone, and canvas length of the free-trace baseline held fixed, Plan Canvas improves accuracy on ProsQA and on Deep ProsQA, a graph benchmark with longer proofs. On Deep ProsQA, accuracy rises from 73.0\% to 87.0\%, the share of questions answered with a valid path rises from 30.8\% to 59.1\%, and the gain is largest on the longest proofs.

86. 【2610.05800】Adaptive Utilization of Low-Rank Adaptation via Conditioned Gating

链接:https://arxiv.org/abs/2610.05800

作者:Guang Yang,Changhao Guan,Chao Huang,Yufeng Chen,Kaiyu Huang

类目:Computation and Language (cs.CL)

关键词:constraining model updates, Low-Rank Adaptation, achieves parameter-efficient fine-tuning, adaptation subspace, parameter-efficient fine-tuning

备注: ICML 2026

点击查看摘要

Abstract:Low-Rank Adaptation (LoRA) achieves parameter-efficient fine-tuning by constraining model updates to a low-rank subspace and has been widely used in practice. However, LoRA typically employs a shared low-rank update across tokens, which limits its ability to fully exploit the adaptation subspace for tokens from different sequences. To address this issue, we propose an adaptive utilization of Low-Rank Adaptation (U-LoRA), which employs conditioned gating to explicitly learn effective token-level utilization of the limited low-rank adaptation subspace. Specifically, U-LoRA generates utilization coefficients along low-rank directions for each token and jointly coordinates and constrains them using sequence-level contextual information, thereby inducing more consistent adaptive patterns within a sentence. To further enhance training stability, we introduce a bias-corrected exponential moving average (EMA) historical prior that calibrates utilization signals across optimization steps, suppressing noise caused by batch-to-batch fluctuations. The effectiveness of our method arises from a better utilization of the existing low-rank subspace via input-conditioned strategies, rather than from expanding the subspace. Experiments on mathematical reasoning and natural language understanding benchmarks demonstrate that U-LoRA achieves competitive performance under comparable parameter budgets when with strong LoRA baselines and recent variants.

87. 【2610.05783】CLARA: Can AI Assess Developmental Appropriateness in Children's Stories?

链接:https://arxiv.org/abs/2610.05783

作者:Sijing Yin,Zirui Wang,Qian Liu,Jiamou Liu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:assessment typically relies, developmental literacy research, developmental, literacy research, English children stories

备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Assessing the developmental suitability of children's narratives is important for educational recommendation and developmental literacy research, yet such assessment typically relies on subjective and difficult-to-scale human judgment. This raises an important question: Can AI systems approximate human developmental judgments of children's stories? To study this problem, we introduce CLARA, a cognitively grounded framework for developmental narrative understanding through structured annotation across cognitive (COG), language (LAN), and social-emotional (SEL) dimensions, together with a bilingual benchmark resource containing 1107 Chinese--English children's stories with normalized silver developmental references and structured developmental annotations. We evaluate CLARA through benchmark comparison, component analysis, translated bilingual consistency analysis, and blinded human evaluation with educators. Experimental results show that structured developmental annotation achieves substantially stronger alignment with developmental references and human judgments than readability-based methods and direct prompting baselines. Overall, our findings suggest that AI systems can approximate certain aspects of human developmental judgment when guided by structured developmental annotation, while also highlighting the importance of interpretability and human oversight in educational NLP.

88. 【2610.05778】MedicalHarness: A Controlled Evaluation of LLMs and Agent Harnesses on Medical Tasks

链接:https://arxiv.org/abs/2610.05778

作者:Ziqing Wang,Lili Zhao,Kaize Ding

类目:Computation and Language (cs.CL)

关键词:LLM agents, harness, increasingly built, work and scored, scored on clinical

备注:

点击查看摘要

Abstract:LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is therefore a property of a model--harness pair. For medical agents, how much outcomes change with the harness has rarely been measured. Measuring this change, and explaining it, raises two challenges. First, a harness comparison must change nothing but the harness and be repeated across models and kinds of task. Second, comparing whole harnesses leaves their mechanisms bundled together, so it cannot show when an individual mechanism helps. To address these challenges, we present MedicalHarness, a controlled study of models and agent harnesses on medical tasks. We first build MedicalHarnessBench to evaluate agents on $107$ tasks across four domains that each test a different harness capability. Using this benchmark, we run five open-weight models under five agent harnesses, changing only the harness within a comparison, and analyze both outcomes and execution traces. To study individual mechanisms, we build MH-Lab, a controlled harness that switches off context management, planning or tool exposure one at a time within a shared execution loop. We find that the harness and its interaction with the model account for about a quarter of the outcome variance, and that no single harness is best across models and tasks. Code and data are available at this https URL.

89. 【2610.05777】Mining Agent Skills from Production Traces

链接:https://arxiv.org/abs/2610.05777

作者:Yue Ran Kang,Colton Mikolajczyk,Chhaya Methani,Hazel Mak,Sahil Bhatnagar,Susheel Suresh,Alejandro Gutierrez Munoz

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:record procedural instructions, Agent skills, curated by hand, record procedural, procedural instructions

备注: 23 pages, 4 figures, 10 tables

点击查看摘要

Abstract:Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whether a run has succeeded may be unavailable. We study how the sampling of execution traces, access to success or failure information, and the form of the mined skills affect downstream task performance. Holding the mining pipeline fixed, we compare six combinations of mining evidence and skill forms. Mining evidence has three levels: successful trajectories only, successes and failures with their outcome labels, or the same mix with labels withheld. Skill form has two types: an ordered workflow plan, or a declarative ontology of entities, states, and policies. We evaluate the mined skills on two enterprise benchmarks, ThinkingBox-Bench and APEX-Agents. Analysis of task-level paired differences shows that the benefits of different configurations of mining evidence and skill forms depend on the enterprise domain. On ThinkingBox-Bench, paired differences show that workflows score better than ontology by 1.7 pp, Goldilocks beats success-only evidence type by 2.4 pp and Goldilocks blind simulating skills learnt without outcomes is worse by 3.1 pp. APEX-Agents shows a moderate preference for ontologies and no clear preference between evidence regimes. Within each domain, task structure related constraints drive uneven performance with mined skills. These findings motivate tailoring meta-skills to the demands of the target tasks rather than adopting a one-size-fits-all approach.

90. 【2610.05774】AdaSpark: Adaptive DSpark with Online Learning for Tree Verification and N-gram Fill

链接:https://arxiv.org/abs/2610.05774

作者:Liquan Liu,Yifan Zhang,Bowei Xu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:tree verifier checks, propose ranked candidates, Block drafters, DSpark propose ranked, forward pass

备注: 25 pages, 10 figures, 15 tables. Code: [this https URL](https://github.com/zeraix/imparo)

点击查看摘要

Abstract:Block drafters such as DSpark propose ranked candidates for several positions in one forward pass, and a tree verifier checks them in one pass of the target. The number of rows to verify trades the tokens a wider tree is expected to accept against the time a wider verify takes. Most schedulers that choose this number take the verify time from a table or model measured before serving, corrected online by at most one scale factor, and take acceptance from the drafter's confidence estimates or from a map fitted offline. AdaSpark learns both quantities while it serves, with no profile, calibration or sweep in advance. It learns which verify widths are worth offering and fits each one's verify time as a function of context. It fits each candidate's acceptance probability to the target's verify outcomes, with the drafter's confidence head as one input, and orders and sizes the tree by that fit instead of by the head. The same model prices n-gram continuations of the request's own text, so drafted and text-derived candidates compete for rows in one best-first order. The width is chosen by pricing time at the long-run decode rate. On single- and multi-turn conversations from six public datasets, on three dense targets and one mixture-of-experts target, AdaSpark decodes 1.5-3.1x faster than this http URL's DSpark with the same drafters. Our imparo engine with AdaSpark is 1.17-1.52x faster than imparo running with a three-token chain (the default this http URL setting); this gain comes from the scheduler alone. Without a width sweep, AdaSpark is never more than 0.3% slower than the best pinned tree width on any dense target or context band. On the mixture-of-experts target it ties the best pinned width, and the other pinned widths from 4 to 16 rows are 5-14% slower.

Comments:
25 pages, 10 figures, 15 tables. Code: this https URL

Subjects:

Machine Learning (cs.LG); Computation and Language (cs.CL)

Cite as:
arXiv:2610.05774 [cs.LG]

(or
arXiv:2610.05774v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2610.05774

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
91. 【2610.05700】Voltic: Distinguishing Volatility from Stochasticity in Recurrent Memory

链接:https://arxiv.org/abs/2610.05700

作者:Parsa Hejabi,Morteza Dehghani,Payam Piray

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:decide how strongly, strongly to overwrite, Recurrent sequence models, Kalman update, Read as Bayesian

备注: 43 pages

点击查看摘要

Abstract:Recurrent sequence models must decide how strongly to overwrite their memory at each token. Read as Bayesian filtering, this write is the gain of a Kalman update, set by uncertainty from two sources that pull it in opposite directions: volatility, how quickly the underlying associations change, and stochasticity, how noisy each observation of them is. First, we show that the update of gated delta-rule memories is the form this filter takes under isotropic uncertainty. Next, we introduce Voltic, a recurrent memory that keeps the covariance anisotropic and makes both noise variances input-dependent, so the write is vector-valued and carries uncertainty accumulated over the sequence. A dense covariance would have to be propagated token by token, ruling out the parallel training these models depend on. We therefore give two assumed-density approximations, diagonal and quasi-diagonal, both of which leave the memory update in delta-rule form and reuse its chunked kernels. On controlled recall tasks in which associations change and observations are corrupted, Voltic leads all baselines. On the task combining volatility and stochasticity, its margin over the strongest baseline is larger at both extrapolation sizes than at the training sizes. In 45M-parameter language models it leads an eight-task reasoning average and achieves higher retrieval accuracy beyond the training context length than gated baselines, at throughput close to those baselines. Deriving the write from an uncertainty recursion therefore makes memory more responsive to change.

92. 【2610.05685】Spend Bytes on Breadth: Precision-Count Trade-offs for Decode-Time KV Compression in Long Chain-of-Thought Reasoning

链接:https://arxiv.org/abs/2610.05685

作者:Runguo Li

类目:Computation and Language (cs.CL)

关键词:decoding long chains, chains of thought, decoding long, long chains, compressed online

备注: 15 pages, 4 figures, 9 tables

点击查看摘要

Abstract:Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision. BreadthKV spends the bytes on more tokens at low precision, combining quantization with eviction, and picks the bit-width for each model and budget with a 60-problem end-to-end calibration, since offline attention error does not predict it reliably. On three reasoning models and four math and science benchmarks, it scores above eviction alone in 17 of 18 settings and produces shorter outputs. Much of what eviction loses comes from derailed runs, which keep reasoning until the length cap without reaching an answer. On Qwen3-8B at our tightest budget, eviction sends 91% of AIME samples to the cap and BreadthKV 40%. Under the same protocol, BreadthKV is statistically indistinguishable from a joint rate-distortion allocator (RDKV) that uses 27% more KV memory-time, and it outperforms our re-implementation of ThinKV.

93. 【2610.05682】Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

链接:https://arxiv.org/abs/2610.05682

作者:Jiulin Li,Ping Huang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Knowing domain rules, Knowing domain, traditional Chinese Bazi, domain rules, guarantee applying

备注: 16 pages, including references and appendices. Project page: [this https URL](https://monsterpppp.github.io/bazi-qa-benchmark/)

点击查看摘要

Abstract:Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.

94. 【2610.05681】Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions

链接:https://arxiv.org/abs/2610.05681

作者:Chanuka Dinuwan,Sanath Jayasena,Buddhika Karunarathne

类目:Computation and Language (cs.CL)

关键词:Automatic speech recognition, Automatic speech, low-resource languages remains, Massively Multilingual Speech, Sinhala ASR research

备注: 29 pages, 1 table. Submitted to Computer Speech Language

点击查看摘要

Abstract:Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both conventional and modern ASR systems. This paper presents the first critical review of Sinhala ASR research, tracing its development from Hidden Markov Models (HMMs) through deep neural networks to self-supervised pre-trained models such as wav2vec 2.0, XLS-R, Whisper and Massively Multilingual Speech (MMS). We compare existing Sinhala systems with related low-resource ASR work on Tamil, Malayalam and Hindi in terms of architecture, training data, word error rate (WER) and robustness to real-world acoustic conditions, and we assess self-supervised and transfer learning as responses to scarce labeled data. We show that most reported WERs are not directly comparable because they differ in corpus, data split and scoring, and that the only controlled comparison in the literature attributes an 18.1% relative WER reduction to corpus correction alone. We also discuss context-aware ASR that draws on phonological, syntactic and semantic knowledge. We identify six research gaps: (1) the lack of large annotated corpora covering multiple dialects and acoustic conditions; (2) weak contextual modeling of Sinhala morphosyntax; (3) high WER in real-world conditions; (4) the absence of standardized benchmarks; (5) the lack of parameter-efficient fine-tuning studies; and (6) the absence of annotated code-switched Sinhala-English speech resources. We outline a research agenda to address these gaps, intended as a roadmap for researchers working on Sinhala and other morphologically rich languages.

95. 【2610.05630】Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection

链接:https://arxiv.org/abs/2610.05630

作者:Nallathambi Vethiappan,Derya Soydaner,Gijs Wijnholds

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:image supports, Visual entailment, leaves undecided, undecided a textual, Atomic Visual Entailment

备注: 15 pages, 16 figures

点击查看摘要

Abstract:Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.

96. 【2610.05615】DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate

链接:https://arxiv.org/abs/2610.05615

作者:Khanh-Binh Nguyen,Van Dai Do,Tien Anh Nguyen,Svetha Venkatesh,Hung Le

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, Dynamic Resolution Assignment, existing multimodal MAD, language models, multimodal MAD frameworks

备注:

点击查看摘要

Abstract:Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variation in the visual scale needed across samples and agents. In addition, these frameworks frequently suffer from groupthink, a phenomenon where agents prematurely abandon correct deductions to conform with confident but hallucinated peer responses. To address these bottlenecks, we introduce DREAM (Dynamic Resolution Assignment For Multimodal Multi-Agent Debate), which operates via two core components: (1) Dynamic Resolution Assignment, a zero-shot probe round where agents test multiple resolutions, quantify uncertainty using Average Normalized Log-Likelihood (ANLL), and use an adaptive threshold to assign each agent to its empirically optimal resolution; (2) Uncertainty-Guided Rollback Aggregation counters groupthink by tracking each agent's uncertainty over rounds and restoring early low-uncertainty answers overridden by group pressure. On six multimodal datasets, DREAM improves the accuracy-token trade-off over multi-agent debate baselines by 1.5-3.2% accuracy without dataset-specific tuning.

97. 【2610.05597】More Than Words: Compositional Tokenization for Efficient Language Models

链接:https://arxiv.org/abs/2610.05597

作者:Yuval Reif,Guy Kaplan,Roy Schwartz

类目:Computation and Language (cs.CL)

关键词:generate text sequentially, inference step covers, Language models process, step covers, generate text

备注: COLM 2026

点击查看摘要

Abstract:Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.

98. 【2610.05591】What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

链接:https://arxiv.org/abs/2610.05591

作者:Yekun Chai,Haoyi Xiong

类目:Computation and Language (cs.CL)

关键词:epoch count matters, faces three questions, data, increasingly repeats data, unique data

备注:

点击查看摘要

Abstract:As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two references: one epoch on the same data, which gives its value, and fresh data at equal compute, which gives its cost. Against fresh data, the cost of repetition follows a single variable, the number of extra epochs divided by the unique tokens per parameter. Against the same data, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size. With unique data fixed, the predicted compute-optimal run grows model size and epochs together until loss stops improving, near the critical epoch count. The same variable accounts for the direction of size trends that appear to conflict: larger models tolerate fewer epochs when the corpus is fixed, from about 15 at 127M to 4 at 2B parameters, but not when unique data grow with the model. Counts alone do not determine loss: at identical counts, replaying shards consecutively raises loss by up to 0.46~bits per byte, concentrating repeats on fewer samples also raises it, lower-entropy sources degrade faster with repetition, and re-tokenizing repeats helps only under heavy repetition. These results offer an empirical guide to pretraining when unique data, rather than compute, are the binding constraint.

99. 【2610.05590】ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction

链接:https://arxiv.org/abs/2610.05590

作者:Jiheng Liang,Chen Zhao,Di Wu,Chenyang Bu,Yunpeng Hong,Xingquan Zhu,Yi He

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:identify clinically significant, training-time interaction history, clinically significant interactions, identify clinically, clinically significant

备注: Accepted at NeurIPS 2026 (poster). Code: [this https URL](https://github.com/0217ljh/ColdDDI-NeurIPS2026)

点击查看摘要

Abstract:Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, textual, or knowledge-graph (KG) evidence, do they actually use the evidence that pharmacologically supports the interaction? We introduce ColdDDI, a reconstructible diagnostic benchmark built from DrugBank 5.1.13, with 1,900 approved small-molecule drugs and 565,731 positive DDI pairs. ColdDDI evaluates pairs with zero, one, or two unseen drugs. It also annotates each interaction by whether it changes drug exposure or drug effect, and by whether the biomedical knowledge graph contains shared enzymes, transporters, or targets that can plausibly mediate the interaction. These annotations separate evidence availability from predictive dependence. We evaluate eight conventional DDI methods and 13 LLMs; for open-weight LLMs, we test five prompt patterns and use masking, drug replacement, and channel-sensitivity metrics to probe knowledge utilization. ColdDDI exposes that, in the hardest split where both drugs are unseen, the main performance divide is mediator availability. A fine-tuned 1B LLM recovers 89-93% of interactions with a shared enzyme, transporter, or target, but only 40-62% without such a mediator. More importantly, KG-provided evidence is not always used; several KG-augmented baselines change little when the shared mediator is masked or disrupted, whereas fine-tuned LLMs respond strongly to this intervention. Thus, ColdDDI evaluates knowledge utilization rather than knowledge access alone, showing where cold-start DDI models rely on mechanistic evidence and where they fail despite receiving it. Code is available at this https URL.

100. 【2610.05584】Expanding LLM Reasoning

链接:https://arxiv.org/abs/2610.05584

作者:Rian Atri,Evan Luo

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Extra inference compute, Extra inference, inference compute, spent on sampling, sampling more reasoning

备注: Accepted to NeurIPS Main Conference '26 with a score of 4.33

点击查看摘要

Abstract:Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and measure it at every eligible step for nine models on six benchmarks (41 model and benchmark cells). Restart position matters: steps selected on one set of continuations beat uniform placement when scored on disjoint ones, in held-out audits on 5, 16, and 38 cells (+4.25 points [+2.51, +6.63] in a fresh five-cell audit). A fixed rule that restarts from the last eligible steps, always-last, is a strong baseline: our learned router beats uniform placement but shows no detected gain over it, and on DeepSeek-R1-Distill-Qwen-14B/MATH-500 always-last exceeds the exact self-consistency frontier at matched aggregate generated output by +0.052 [+0.008, +0.098], using 0.774x the aggregate generated output of four-sample self-consistency. Cross-fitted oracle selection still finds held-out headroom beyond declared positional classes, a target for future selectors. Finally, breaking step-label ties by earliest index flips the sign of a pointwise selector's gain over uniform placement in every seed of a five-seed diagnostic with four rollouts per step; randomized ties remove the bias.

101. 【2610.05575】Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

链接:https://arxiv.org/abs/2610.05575

作者:Abdul Rehman,Jian-Jun Zhang,Xiaosong Yang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:written text carries, research rests, carries enough information, information to select, untested assumption

备注: 13 pages, 6 figures

点击查看摘要

Abstract:Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.

102. 【2610.05564】Lend Me Your Eyes: Instruction-Aware Text Embeddings via Attention Relay

链接:https://arxiv.org/abs/2610.05564

作者:Yiyuan Luo,Vaggos Chatziafratis

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:contrastive learning learn, follow task instructions, learn to follow, follow task, instruction-paired data

备注:

点击查看摘要

Abstract:Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propose Attention Relay, which passes the attention weights an LLM produces to the embedder's own attention. Across six instruction-tuned LLMs from the Qwen3, Llama 3.1 and OLMo 3 families and ten widely used embedding models that differ in tokenizer, size and pooling type, Attention Relay makes nearly every combination instruction-aware. Experiments that break the method down into its parts show that the LLM's attention weights track the instruction in its later layers and come largely from instruction tuning. They also show that relaying these weights selects which content in the text matters: it makes the aspect of the text that the instruction asks about dominant in the embedding, or restores that aspect where averaging had diluted it.

103. 【2610.05541】Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential

链接:https://arxiv.org/abs/2610.05541

作者:Swadesh Swain,Sanghamitra Dutta

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:Mechanistic interpretability, features, safety, Mechanistic, understand safety behavior

备注: 28 pages, 3 figures, 15 tables. Submitted to ICLR 2027

点击查看摘要

Abstract:Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it). To find suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. A significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. Under natural jailbreaks, the suppression acting on the highest-CAP features rises 2-4x, and their activation correspondingly falls by up to 80%. Amplifying a feature's suppressors pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experiments span five Gemma, Qwen, and Llama models across various parameter sizes. Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.

104. 【2610.05540】SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output

链接:https://arxiv.org/abs/2610.05540

作者:Shiyuan Zhou,Ashwin Gerard Colaco,Sainyam Galhotra,Sharad Mehrotra

类目:Databases (cs.DB); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:data analysis research, Natural language, analysis research, foundational to progress, progress in data

备注: Extended version of the paper accepted at ACM SIGMOD 2027. Includes additional appendix material. 27 pages. 4 figures

点击查看摘要

Abstract:Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of state-of-the-art performance. We present SALUS, a system that automatically detects annotation errors in NL-to-SQL benchmarks. SALUS frames benchmark auditing as a weakly supervised error detection: SQL generated by multiple LLM agents drive a suite of complementary weak-labeling functions. By passing this noisy vote matrix through a generative label model, we extract high-confidence training samples without requiring human ground truth. These samples train a decision plane that maps gold SQL query features to per-agent trustworthiness, allowing SALUS to intelligently fuse reliability estimates with raw verdicts for rigorous benchmark error detection. We evaluate on BIRD-Clean-xs, a benchmark of 298 BIRD development tasks with manually verified correctness labels. SALUS achieves F1 = 0.9194, significantly outperforming the state-of-the-art baselines. Applying SALUS to the full development sets, we estimate annotation error rates of approximately 37% on BIRD and 27% on Spider.

105. 【2610.05534】Dataset Signatures in Human-LLM Interactions and User Modeling

链接:https://arxiv.org/abs/2610.05534

作者:Joseph Suh,Serina Chang

类目:Computation and Language (cs.CL)

关键词:LLM interaction datasets, interaction datasets shape, user models, shape our understanding, LLM interaction

备注:

点击查看摘要

Abstract:Human--LLM interaction datasets shape our understanding of AI use and provide a foundation for downstream research, including training and evaluation of user models. In recent years, a growing number of datasets have sought to capture a representative picture of human--LLM interactions. But how different are the pictures these datasets provide, and what do those differences mean for research built on them? We study these questions across seven conversation datasets, spanning in-the-wild chat logs and human preference data. We begin by revisiting the dataset classification experiment of Torralba Efros and find that neural network classifiers identify the source of a conversation from user messages alone well above chance, indicating distinctive dataset signatures. This separability persists after matching datasets on the dimensions of human-designed taxonomies, implying subtle differences that these taxonomies do not capture. We then examine the implications for user modeling: how dataset signatures propagate to the outputs of user models trained on these datasets; how dataset choice influences evaluations of user model quality and subsequent evaluations of LLM assistants paired with these user models; and how dataset classifiers can guide data selection for training user models. While each dataset is meant to capture a slice of 'real-world' interactions, our findings reveal the extent to which these slices diverge, and the consequences of those differences for research built on these foundations.

106. 【2610.05533】What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects

链接:https://arxiv.org/abs/2610.05533

作者:Bowen Xu,Boyu Chen

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Harness search, change raises, reasoning switches, raises a score, GEPA search arm

备注: 43 pages, 2 figures, 35 tables. Ancillary files in anc/: the frozen preregistration, its addenda (row keys of five rows withheld) and the aggregate analysis report, with a README

点击查看摘要

Abstract:Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, with three small models, three benchmarks, replication and test partitions, a GEPA search arm and six evaluation defects injected one at a time, and we report all 47 primary endpoints. Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells, and in each the gain came mostly from questions where the capped setting gave no readable answer. The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons, and lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect's distortion instead of revealing it. On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline.

107. 【2610.05532】DelegationBench: Measuring When AI Agents Should Ask Before Acting

链接:https://arxiv.org/abs/2610.05532

作者:Shiva Pochampally

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)

关键词:edit files, send emails, agents that send, make purchases, purchases must decide

备注: 35 pages, 7 figures. Code and data: [this https URL](https://github.com/PieLord757/delegation-bench)

点击查看摘要

Abstract:AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test whether such scores can be trusted. It has 156 scenarios with four possible responses (act, ask for permission, ask for missing information, refuse), and most scenarios come in matched pairs that change a single feature: whether the action was requested, what is at stake, whether it can be undone, or who will see it. Across ten models from five families, agreement scores mislead in three ways. A simple keyword rule, which we wrote after seeing the benchmark, agrees with our annotators more often than eight of the models, yet its decision changes in only 9 of 48 matched pairs. Equivalent ways of asking the same question change how often a model acts by up to 52.5 percentage points. And every model stops to ask the user less often when it must carry out the task with tools than when it judges a proposed action. When rules are stated explicitly, the same models follow them almost perfectly, so the gaps are not explained by a general inability to follow rules. We release the benchmark and evaluation tools and recommend reporting these properties separately rather than as one score.

108. 【2610.05484】Universal Test-Time Training

链接:https://arxiv.org/abs/2610.05484

作者:Zefan Cai,Qinzhe Hu,Ziqiao Ma,Hao Tan,Junjie Hu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent Test-Time Training, architectures compress context, Universal Test-Time Training, Recent Test-Time, Test-Time Training

备注: 37 pages. Project page: [this https URL](https://zefan-cai.github.io/uTTT.github.io/) ; code: [this https URL](https://github.com/Zefan-Cai/uTTT)

点击查看摘要

Abstract:Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.

109. 【2610.05466】Writing as a Self-Organized Critical Process

链接:https://arxiv.org/abs/2610.05466

作者:Nikolay Mikhaylovskiy

类目:Computation and Language (cs.CL)

关键词:explain autocorrelation decay, decay power laws, power laws omnipresent, autocorrelation decay power, self-organized criticality

备注:

点击查看摘要

Abstract:We explain autocorrelation decay power laws omnipresent in texts by self-organized criticality. Specifically, we analyze the recently released KLiCKe keystroke dataset and show that not only the final texts' autocorrelations form a manifold that adheres to a power law with a finite-size scaling, but also the text revisions generate revision-size-dependent restoring dynamics toward that manifold. Thus, human writing appears to dynamically regulate semantic correlations in a text toward a critical state.

110. 【2610.05437】une: Evolving Agent Skills From Offline Telemetry

链接:https://arxiv.org/abs/2610.05437

作者:Justin Chih-Yao Chen,Elias Stengel-Eskin,Yan Chen,Pol Llado,Scott Counts,Mohit Bansal,Benjamin Van Durme,Harsh Jhamtani,Gaurav Verma

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:capture procedural knowledge, Computer-use agents, people use software, capture procedural, procedural knowledge

备注: Project Page: [this https URL](https://microsoft-teletune.github.io/)

点击查看摘要

Abstract:Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) Non-Replayability, since past activity cannot be replayed to evaluate skill updates; and (3) Interleaved Trajectories, since logs may mix several tasks without marking their boundaries. To address these, we introduce TeleTune, a framework for learning a textual skill library from offline logs without recorded goals, cannot be replayed during optimization, and may interleave tasks. TeleTune uses action-prediction errors on logged trajectories to propose library edits and keep only those that improve held-out action-prediction accuracy, which we call skill-guided progress. The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task. At test time, the agent is provided with the learned library and the workflow-based retrieved demonstrations. Experiments on WorkArena and Online-Mind2Web show that TeleTune outperforms random retrieval, Agent Workflow Memory (AWM), and their combination. We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%. Under the heaviest perturbation of the WorkArena training data,TeleTune keeps the highest average success rate at 68.5%, 6.3% above the strongest baseline. Our analyses show (1) skill optimization and workflow-based retrieval are complementary, (2) optimizing on fixed logs costs 5 to 75 times fewer tokens than validating the same edits with live episodes, (3) skill-guided progress tracks the live success rate.

111. 【2610.05425】Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

链接:https://arxiv.org/abs/2610.05425

作者:Ali Vosoughi,Akhil Kasturi,Chenliang Xu,Axel Wismueller

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Automated checks, radiology reports, reports may rely, rely on AI-generated, leave findings unmentioned

备注: 40 pages, 7 figures, 17 tables (main text and references pp. 1-19; Supplementary Information as an appendix, pp. 20-40). Submitted to npj Digital Medicine. Code: [this https URL](https://github.com/ali-vosoughi/VerifyGRPO-Rad)

点击查看摘要

Abstract:Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.

112. 【2610.05402】ask Vector Descent: Learning from Non-IID Batches

链接:https://arxiv.org/abs/2610.05402

作者:Anton Baumann,Jonas Hübotter,Zeynep Akata,Andreas Krause

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:acquire new knowledge, knowledge without forgetting, central challenge, model, language model training

备注:

点击查看摘要

Abstract:A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement ($\lambda=1$) with partial integration, which scales the task vector by $\lambda$ before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of $\lambda$ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.

113. 【2610.05394】he Hidden States Cookbook: A Large-Scale Ablation Study for Noise-Robust Conversational Intent Classification in Industry

链接:https://arxiv.org/abs/2610.05394

作者:Bogdan Bogachov,Nikita Letov,Yaoyao Fiona Zhao

类目:Computation and Language (cs.CL)

关键词:users naturally embed, wastes computational resources, naturally embed queries, intent classification accuracy, database interfaces face

备注: 11 pages, 4 figures

点击查看摘要

Abstract:Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances in orchestration and retrieval strategies, a fundamental question remains unanswered: which pooling strategy maximizes intent classification accuracy under realistic conversational noise in production language models? This work addresses this gap through 360 controlled experiments spanning four pooling configurations (mean, max, last-token, attention, and FFT-augmented variants) using Llama-3.2-1B-Instruct on BANKING77 and CLINC150 datasets under clean/noisy conditions with ten random seeds. Key findings reveal that attention pooling consistently outperforms alternative strategies under noisy conditions (~+2.6-2.8 F1 over the default), while mean pooling degrades performance by up to ~5 F1 points. Frequency-domain filtering does not produce consistent accuracy improvements and functions primarily as a structural variation rather than an accuracy-enhancing component. These results provide concrete, evidence-based guidance for building noise-robust conversational classifiers: attention pooling is recommended for noisy interfaces, mean pooling should be avoided, and last-token pooling is appropriate for clean-query scenarios.

114. 【2610.05387】GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation

链接:https://arxiv.org/abs/2610.05387

作者:Murad Hossen,Tasneem Selim,Gurur Gamgam,Tuga Yousif,Abderrahmane Kasmi,Ikram Aissiou,Mubaraq Onipede,Faran Taimoor Butt,Sanae Zrigui,Rosa Y. G. Paccotacya-Yanque,Ignatius Balayo,Ikram Elhouiti,Hadil Affes,Bijay Adhikari,Sargam Goyal,Muhammad Ibrahim Isah,Mohammad Idrees Bhat,Samuel Kangoni Matia,Peguy Kem-Meka Tiotsop Kadzue,Maha Trabelsi,Emmanuel Owusu,Vinit,Nour Majdoub,Tamiru Alemnew,Islem Rekik

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:Large language models, remains largely unexplored, graph-structured machine learning, machine learning problems, learning problems remains

备注: Accepted at EMNLP 2026. 28 pages, 16 figures, 4 tables. Benchmark and leaderboards: [this https URL](https://basiralab.github.io/GNN-CB/)

点击查看摘要

Abstract:Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at this https URL.

115. 【2610.05382】Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

链接:https://arxiv.org/abs/2610.05382

作者:Shanyong Wang,Zhenwen Ji,Lei Jin,Yining Zhao,Yicheng Qian,Chengqiang Lu,Yi Wu,Yao Hu,Lizhen Cui,Yanyu Xu

类目:Computation and Language (cs.CL)

关键词:search requires agents, multiple steps, steps and synthesize, Long-horizon search requires, requires agents

备注: 26 pages, Natural Language Processing

点击查看摘要

Abstract:Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and cause the policy to lose track of unresolved questions, overlook useful evidence, or terminate before sufficient support has been collected. One of promising way is to decouple three distinct responsibilities of proposing retrieval actions, updating persistent state, and deciding when to stop rather than concentrating them within a single policy. Targeted at it, we introduce Harness-Search, a multi-agent search harness to reduce the local errors propagating across subsequent exploration, evidence curation, and termination decisions. In particular, Harness-Search assigns these responsibilities to three permission-bounded authorities: a Retrieval Policy that proposes search operations, a Memory Operator that validates and commits persistent-state updates, and a Summary Auditor that accepts or rejects termination based on the sufficiency of the curated evidence. Together, these roles form a Propose-Commit-Audit loop in which actions are proposed, persistent evidence is selectively committed, and stopping decisions are subjected to an explicit sufficiency check. Across seven long-horizon search benchmarks, Harness-Search improves both retrieval and answer generation under the same policy backbone, increasing Recall by 4.60-27.92 points and Final-Answer Recall by 12.34-30.13 points over the strongest harness-based baseline on each evidence-retrieval benchmark. Moreover, trajectory-level analyses show that Harness-Search continues to accumulate useful evidence and expand evidence coverage with less redundant retrieval as the search history grows.

116. 【2610.05380】VHDL-REPOBENCH: A Repository-Level Benchmark for Evaluating Large Language Models on VHDL Design Generation

链接:https://arxiv.org/abs/2610.05380

作者:Prashanth Vijayaraghavan,Akul Malhotra,Ashutosh Jadhav,Ehsan Degan,Vandana Mukherjee

类目:Programming Languages (cs.PL); Hardware Architecture (cs.AR); Computation and Language (cs.CL)

关键词:Large Language Models, demonstrating strong potential, Large Language, hardware description languages, description languages

备注: 7 pages

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly applied in hardware design automation, demonstrating strong potential in generating and understanding hardware description languages. However, most existing benchmarks focus on Verilog, with limited evaluation of VHDL, which remains widely used in industry and academia for FPGA and safety-critical systems. To address this gap, we introduce VHDL-REPOBENCH, a large-scale, cross-file, repository-level benchmark for assessing LLM capabilities on realistic VHDL design generation and analysis tasks. VHDL-REPOBENCH curates ~100 open-source VHDL repositories, encompassing ~2.5k VHDL files and ~500 testbenches, and provides structured problem statements, module stubs, and self-verifying testbenches. The benchmark enables comprehensive evaluation across syntax, semantic correctness, hierarchical reasoning, cross-file dependency resolution, and functional verification. We evaluate several state-of-the-art models, including GPT-4o, Llama-3-70B, Qwen2.5-72B, CodeLlama-70B, and multi-step reasoning approaches such as Reflexion and CoDes. Results reveal that while current LLMs achieve moderate line- and block-level accuracy, substantial challenges remain in multi-file reasoning, hierarchical design understanding, and specification-to-module generation. VHDL-REPOBENCH represents the first large-scale VHDL-focused benchmark and provides a valuable resource for the hardware design community to evaluate, compare, and advance LLM capabilities for practical VHDL development.

117. 【2610.05373】owards Unbiased On-Policy Distillation for Block Diffusion Language Models

链接:https://arxiv.org/abs/2610.05373

作者:Zaiquan Yang,Fei Wei,Yong Wang,Yudong Han,Yiyu Li,Zhuofan Zong,Gerhard Petrus Hancke,Xiangxiang Chu,Rynson WH Lau

类目:Computation and Language (cs.CL)

关键词:effective post-training paradigm, recent efforts extending, diffusion language models, block diffusion language, language models

备注:

点击查看摘要

Abstract:On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexplored. In this work, we investigate this regime and reveal two critical optimization biases that induce severe training instability. First, mismatched block boundaries between teacher and student cause \textbf{\textit{context misalignment}}, providing distorted supervisory signals that misguide student decoding. Second, even under aligned contexts, an \textbf{\textit{intrinsic optimization bias}} in OPD, where the student tends to rapidly absorb high-support signals while lagging on low-support updates, drives a premature confidence surge that traps weaker students in catastrophic overconfidence collapse. To resolve these, we propose \mbox{\textbf{Un-OPD}}, an unbiased on-policy distillation framework with two novelties for stabilizing BDLM training. First, Un-OPD introduces a boundary-aware step filtering strategy that eliminates context-misaligned decoding steps. Second, Un-OPD proposes moderating optimization intensity at high-support positions via a support-rebalanced confidence calibration, thereby bypassing overconfidence collapse. Beyond stability, we also introduce a rollout reuse mechanism to reduce rollout generation overhead. Extensive experiments on math reasoning and code generation benchmarks show that Un-OPD consistently stabilizes training and delivers superior performance while reducing wall-clock training time by approximately half.

118. 【2610.05366】he ÌròyìnSpeech Text Corpus: 24,905 Curated Yorùbá Sentences for Speech and Language Technology

链接:https://arxiv.org/abs/2610.05366

作者:Kola Tubosun,Aanuoluwapo Aremu,Tolulope Ogunremi,Iroro Orife,David Ifeoluwa Adelani

类目:Computation and Language (cs.CL)

关键词:distributed by ELRA, Yorùbá read-speech corpus, Yorùbá read-speech, tone-marked Yorùbá sentences, Yorùbá

备注: 8 pages. Data descriptor for the ÌròyìnSpeech Text Corpus, doi: [https://doi.org/10.5281/zenodo.23138464](https://doi.org/10.5281/zenodo.23138464) . Under review at the Journal of Open Humanities Data

点击查看摘要

Abstract:ÌròyìnSpeech is a 42-hour, 80-speaker Yorùbá read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yorùbá sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yorùbá corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yorùbá personal and place names appear in Yorùbá form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.

119. 【2610.05343】MemStrata: 95% and 90.91% Source-Aware Accuracy on LongMemEval-500 and LoCoMo-1540 with a Local Qwen 3.8 27B Q4_K_M Reader

链接:https://arxiv.org/abs/2610.05343

作者:Neeraj Yadav(Called It Inc.)

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:incomplete benchmark reference, adequate conversational answer, adequate conversational, short or incomplete, incomplete benchmark

备注: 29 pages, 23 tables, 1 figure. Ancillary files contain per-question grades and reproducible analyses, plus explicitly labelled exports from audited follow-up reports

点击查看摘要

Abstract:An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing system-blinded answers; original reference-only grading is reported alongside. With a local Qwen 3.8 27B Q4_K_M reader and a 24,000-token evidence ceiling, MemStrata CL1 scores 475/500 (95.0%) on LongMemEval-S and 1,400/1,540 (90.91%) on LoCoMo categories 1-4 under source-aware GPT-5.5 adjudication, against 463/500 (92.6%) and 1,205/1,540 (78.25%) under reference-only grading of the same answers. It preserves a retrieval backbone and adds nonduplicated, dated, speaker-attributed source spans. A same-reader full-history control with about 4.7 times the evidence scores 464/500 reference-only and 470/500 (94.0%) source-aware; neither difference is decisive. Keyword-only selection at the same budget scores 425, and a matched-reader Letta arm 438. On LongMemEval-M, where the packet holds about 1.6% of each history, MemStrata CL1 scores 427/500, with losses concentrated in multi-session and temporal questions. On 300 BEAM-1M questions it outscores dense retrieval, 0.738 to 0.706 (Wilcoxon p = 0.011). A same-seed replay of unchanged requests changed 1.5-2.3% of labels. On identical packets GLM 5.3 flash is non-inferior within 3 points (462 versus 463); Muse Spark 1.3 did not show non-inferiority on 269 questions. None of four pre-registered interventions met all of its registered advancement or feasibility criteria. Signed read-side artifacts support inspection but do not regenerate the private retrieval pipeline. The superiority of source-aware grading to human adjudication is not established, and development exposure, automated-judge dependence and the absence of held-out data preclude an independent-replication or leaderboard claim.

120. 【2610.05322】When Does Longer Reasoning Help? Predicting Mathematical Reasoning Through Discovery and Execution

链接:https://arxiv.org/abs/2610.05322

作者:Adib Hasan,Lay Jain,Thanic Nur Samin

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:mathematical reasoning scales, Test-time compute, mathematical reasoning, scales with additional, Test-time

备注: Accepted in NeuRIPS MATH-AI Workshop 2026

点击查看摘要

Abstract:Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a convolution of strategy discovery and conditional execution. From independent short-budget attempts and oracle-sketch-conditioned runs, the framework estimates cumulative success along held-out reasoning trajectories under alternate compute allocations. We evaluate four models on 35 fresh Olympiad problems and non-geometry problems from IMO-ProofBench Advanced. Under the DE framework, near-saturated execution predicts geometric scaling, as observed for the GPT models. For Claude Opus 4.8, incorporating measured execution substantially improves held-out forecasts over geometric extrapolation across one- and two-arm allocations. As a secondary application, regularized DE (R-DE) decisions to continue or restart yield lower average regret than the best model-specific retrospective policy. Together, these results show that measuring conditional execution provides information about longer reasoning that short-budget success rates do not always capture.

121. 【2610.05308】RubricArmor: Adversarial Evolution Improves LLM-Based Rubric Generation

链接:https://arxiv.org/abs/2610.05308

作者:Haocheng Yang,Yuchao Zhang,Licheng Pan,Jiajun Fan,Maolin Wang,Kangning Zhang,Shuai Shao,Shijian Wang,Yuan Lu,Chunyuan Zheng,Hao Wang

类目:Computation and Language (cs.CL)

关键词:large language models, aligning large language, Rubric-based reinforcement learning, query-specific evaluation criteria, reinforcement learning

备注:

点击查看摘要

Abstract:Rubric-based reinforcement learning (RL) provides interpretable rewards for aligning large language models (LLMs) by evaluating responses against query-specific evaluation criteria. To construct rubrics at scale, a straightforward approach to LLM-based rubric generation is to prompt an LLM to generate a rubric directly from the query. However, rubrics directly generated by LLMs are vulnerable to reward hacking, since omitted or underspecified criteria allow the policy to obtain high rubric rewards with low-quality responses. Existing LLM-based rubric generation methods improve the granularity and coverage of the generated criteria but do not proactively guard against reward hacking. To address this limitation, we propose RubricArmor, an adversarial framework that exposes and mitigates potential reward hacking at the rubric generation stage before it occurs in subsequent RL. Specifically, RubricArmor performs adversarial evolution, in which an attack step and a repair step alternate over multiple rounds. The attack step simulates the reward hacking of the policy by constructing adversarial responses that satisfy the current rubric but fail to properly complete the task. The repair step then revises the rubric to detect the response defects exposed by the attack step while preserving other valid criteria. Extensive experiments demonstrate that RubricArmor outperforms competitive rubric generation baselines and translates into more effective downstream rubric-based RL.

122. 【2610.05282】Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

链接:https://arxiv.org/abs/2610.05282

作者:Tongyan Hu,Hao Li,Xiaogeng Liu,Ruida Wang,Zhengyu Liu,Shuyao Xu,Ning Zhang,Ziyang Li,Yinzhi Cao,Bryan Hooi,Chaowei Xiao

类目:Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:Large language models, Large language, language models remain, language models, models remain vulnerable

备注: 29 pages

点击查看摘要

Abstract:Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9\% to 72.4\% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at this https URL

123. 【2610.05278】When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

链接:https://arxiv.org/abs/2610.05278

作者:Qishi Zhan,Seoyeon Jang,Zihan Dong,Minxuan Hu,Ziheng Chen,Tonghui Qu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Verifiable instruction-following benchmarks, Verifiable instruction-following, fixed template, instruction-following benchmarks, benchmarks often express

备注:

点击查看摘要

Abstract:Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.

124. 【2610.05277】Mind the Gaps: From Failure Attribution to Closed-Form Repair of Code Language Models

链接:https://arxiv.org/abs/2610.05277

作者:Jian Gu,Hongyu Zhang,Chunyang Chen,Aldeida Aleti

类目:oftware Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Code language models, library evolves, writing the interface, generic update, neurons

备注:

点击查看摘要

Abstract:Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods attribute a failure to neurons, select the highest-ranked ones, and apply a generic update. This pipeline assumes that the attributed neurons are the ones to patch and that a generic update fits every failure, and neither assumption has been examined. We examine both on executable API evolution tasks in Python and Rust with three code models and identify two gaps. The targeting gap separates the neurons that failure attribution targets from the neurons that can carry a patch: their top sets have a Jaccard overlap of only 0.15 to 0.20. The tailoring gap separates a generic update from a patch built for the failure: the patches that different carrier neurons need are nearly orthogonal. To address both gaps, we propose ASTRA. It targets neurons by contrastive semantics, an attribution that scores a neuron by its contribution to the logit contrast between the target token and the produced token. It then tailors the patch by solving one small linear system in closed form, which corrects all failing tokens of a sample jointly and needs neither an optimizer nor a backward pass. Contrastive semantics selects significantly better carrier neurons than gradient-based attributions in 3 of 6 settings and comparable ones in the others. On average, ASTRA reaches 66.7 percent Pass@1, against 48.4 percent for the best of AlphaEdit, STAR and low-rank adaptation, and repairs a sample in 2.7 seconds. It is the best method in all 6 settings and for every type of API change, and this advantage persists under an unseen phrasing of the test prompt. Its side effects on unrelated code are small on the large models and larger on the small one.

125. 【2610.05275】Inductive Claims Extraction at Scale

链接:https://arxiv.org/abs/2610.05275

作者:Sandrine Chausson,Björn Ross

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Social and Information Networks (cs.SI)

关键词:typically single-clause statements, typically single-clause, single-clause statements, factual to evaluative, built and expressed

备注:

点击查看摘要

Abstract:A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statements, which convey a particular interpretation of reality and can range from factual to evaluative. Moreover, rather than occurring randomly, claims coalesce, recur in patterns, and come to be associated with different world views. When paired with structural computational tools such as Social Network Analysis, claims can be a powerful unit of analysis to study political phenomena such as echo chambers or polarisation. In this paper, we present a pipeline that uses a large language model (LLM) to inductively extract and catalogue claims from large social media corpora, and apply it to two different Twitter datasets: one relating to the 2020 US presidential election and the other to the 2022 FIFA World Cup. We comprehensively evaluate the approach by measuring the pipeline's recall and precision against manually annotated samples, run ablation studies isolating the contribution of its various components, and perform a qualitative error analysis. We discuss the value of the approach in the context of Computational Social Science research, and illustrate its capabilities by presenting the claims catalogue obtained from each dataset.

126. 【2610.05247】mplar: agentic induction and evolution of standardized radiology reporting templates from large-scale clinical corpora

链接:https://arxiv.org/abs/2610.05247

作者:Xiaotian Hu,Mingxuan Liu,Zhonghan Wang,Xinfeng Zhang,Yiming Huang,Ziang Wang,Kasidit Anmahaepong,Yijin Li,Yifei Chen,Hongjia Yang,Zihan Li,Qiyuan Tian

类目:Multiagent Systems (cs.MA); Computation and Language (cs.CL)

关键词:Structured radiology reporting, Structured radiology, radiology reporting mitigates, high-quality reporting templates, mitigates the heterogeneity

备注:

点击查看摘要

Abstract:Structured radiology reporting mitigates the heterogeneity of free-text reports, yet its benefits depend on high-quality reporting templates. In practice, such templates are conventionally built through labor-intensive expert consensus and therefore vary across institutions and lag behind evolving clinical practice. Large language models (LLMs) enable automated template induction, but existing approaches remain limited: single-LLM induction is constrained by context length, and the corpus-scale method ASTAR produces a static, closed-corpus template without external grounding or downstream adaptation. To address these limitations, we propose TEMPLAR, a TEMPLate-centric Agentic framework for inducing and evolving standardized Radiology reporting templates from large-scale clinical corpora. TEMPLAR treats the template as a persistent central state maintained alongside two provenance-aware knowledge graphs, namely an anatomical graph that constrains template construction and a diagnostic graph that supports finding-to-diagnosis reasoning. Three agents operate on this state. The Induction Agent derives canonical clinical slots from anatomy-constrained Span-Triple atoms via dual-view similarity clustering; the Evolution Agent then assembles these slots into a hierarchical template and revises it under consistency constraints, external clinical evidence, and downstream structuring feedback; and the Clinical Agent applies the evolved template to report structuring, reconstruction, and diagnostic reasoning. Across four datasets, TEMPLAR outperforms ASTAR, three medical LLMs, and six general-purpose LLMs in coverage, information fidelity, and diagnostic fidelity, while achieving the highest or tied-highest LLM-rated template quality. Its fidelity advantages over ASTAR persist under cross-dataset transfer, and cumulative ablations support complementary contributions of its key components.

127. 【2610.05223】Look Before You Leap: Thermodynamic Arbitration of Parametric and Non-Parametric Knowledge in LLM Agents via Self-Regulating Memory Architectures

链接:https://arxiv.org/abs/2610.05223

作者:Akash Das,Ishan Roy

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:profound cognitive polarization, modern LLMs consists, cognitive polarization, profound cognitive, modern LLMs

备注:

点击查看摘要

Abstract:The architecture of modern LLMs consists of a profound cognitive polarization. LLMs possess implicit intuition encoded in their parameters, yet rely on a disconnected, explicit mechanism to access the outside world. Agentic frameworks have not bridged this gap; instead, models are often compelled into pathological "induced amnesia." Under the prevailing "Retrieve-Always" paradigm, agents must distrust their internal knowledge, making every user interaction a "tabula rasa" event that must be checked externally. This creates reflexive dependence that can be thermodynamically wasteful, cognitively fragile, and susceptible to irrelevant context. We propose a return to first principles, operationalizing the biological maxim "Look Before You Leap." We introduce MARTA (Metacognitive Adaptive Retrieval and Thought Architecture), a neuro-symbolic framework that bridges parametric and non-parametric knowledge. Rather than treating retrieval as mandatory, MARTA models it as a cost, taking the leap only when perceived internal inadequacy warrants external information. By allowing the agent to gauge the entropy of its own thoughts before acting, MARTA enables deliberative retrieval and uncertainty-aware decision making. Our approach suggests that giving agents the capacity for introspection can restore a more efficient balance between internal knowledge and external information.

128. 【2610.05219】Safe Context Switching for Agents in the Wild: Mitigating Subspace Interference via Orthogonal Adaptation

链接:https://arxiv.org/abs/2610.05219

作者:Akash Das,Ishan Roy

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, Language Models exhibit, Language Models, Sequential Subspace Interference

备注:

点击查看摘要

Abstract:Most Large Language Models exhibit a fundamental tension between two sequential tasks, such as logical reasoning and safety alignment. The high-variance internal states required for sophisticated Chain-of-Thought (CoT) deduction can geometrically interfere with latent representations encoding safety constraints. We identify this phenomenon as Sequential Subspace Interference, showing that standard fine-tuning on logical tasks such as multi-step mathematics and code generation can result in a 23.3% interference penalty on alignment benchmarks, substantially weakening the model's safety priors. This Reasoning Drift is not adequately captured by current adaptation methods because gradients for logical tasks are rarely orthogonal to safety objectives. To address this issue, we propose AURA (Adaptive Unique Residual Allocation), a spectral regularization framework that enforces Spectral Independence between reasoning and safety. By explicitly estimating the null space of the alignment manifold and constraining reasoning updates to its orthogonal complement, AURA enables models to improve logical reasoning without compromising safety. Empirically, AURA recovers 23.0% of the lost performance while preserving greater than 0.98 cosine fidelity to the safe state, demonstrating that reasoning and alignment can be effectively decoupled through geometric regularization.

129. 【2610.05190】FORGE: Verification-Gated Behavioral Repair for Generative Language Models

链接:https://arxiv.org/abs/2610.05190

作者:Hsin-Ling Hsu,Min-Yu Chen,Nai-Chia Chen,Yan-Ru Chen,Yi-Ling Chang,Fang Yu

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:inherit undesirable behaviors, including demographic bias, Generative large language, inherit undesirable, behaviors from pre-training

备注:

点击查看摘要

Abstract:Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees. Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs. We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages. Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation. FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol. On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation. The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate. FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.

130. 【2610.05170】Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation

链接:https://arxiv.org/abs/2610.05170

作者:Feng He,Hejia Wang,Linghao Meng,Ming Gao,Qiankun Li

类目:Computation and Language (cs.CL)

关键词:systems sample multiple, improve code generation, sample multiple candidate, systems sample, final output

备注: EMNLP 2026

点击查看摘要

Abstract:Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator. We challenge this assumption under misleading task premises. When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them. Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool. We call this failure mode Verification Trap. Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection. Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness. These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC. Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.

131. 【2610.05126】Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining

链接:https://arxiv.org/abs/2610.05126

作者:Ziyue WANG,T. Kanamori

类目:Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:small language model, small language, larger scale, repetition counts, language model

备注:

点击查看摘要

Abstract:The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens. We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale. On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M. This supports candidate retention as a practical alternative to exact point prediction. We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms. A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.

132. 【2610.05114】Belief-Trajectory Energy: Measuring the Path to a Prediction

链接:https://arxiv.org/abs/2610.05114

作者:Jiahao Ying,Wei Tang,Boxian Ai,Yaoning Wang,Haotian Chen,Wenhe Sun,Caijun Xu,Haozhan Cai,Changyi Xiao,Yixin Cao

类目:Computation and Language (cs.CL)

关键词:Large language models, Transformer layers, Large language, predictions across Transformer, progressively revise

备注:

点击查看摘要

Abstract:Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to $0.998$ macro-AUROC and $95.6\%$ eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at this https URL.

133. 【2610.05111】InstMoE: Adaptive Multimodal Routing with Specialized Experts

链接:https://arxiv.org/abs/2610.05111

作者:Guimin Hu,Xiang He,Yingjian Li,Zheng Lian,Boyan Xu,Ruichu Cai

类目:Computation and Language (cs.CL)

关键词:information pathways required, inherently heterogeneous, effective prediction, required for effective, information pathways

备注:

点击查看摘要

Abstract:Multimodal inputs are inherently heterogeneous, not only across modalities but also in the information pathways required for effective prediction. To address this limitation, we propose InstMoE, an adaptive expert routing framework for multimodal learning. InstMoE dynamically routes each input to specialized unimodal and cross-modal experts, allowing the model to adapt its information pathways to the characteristics of the input. However, routing can be misled when modality-specific variations obscure task-relevant semantics. Such irrelevant variations may distort routing decisions, causing inputs to be assigned to inappropriate experts. We therefore introduce a Contrastive Semantic Alignment module, which encourages semantically similar inputs to share task-relevant representations while suppressing irrelevant modality-specific variations. Experiments on multimodal sentiment analysis benchmarks demonstrate that InstMoE achieves state-of-the-art performance on CMU-MOSEI and CH-SIMS v2 while using substantially fewer parameters than competitive baselines. Further analysis shows that different inputs exhibit distinct expert preferences, demonstrating that InstMoE moves beyond fixed fusion toward adaptive multimodal computation.

134. 【2610.05107】SearchJev: A Fast and Calibrated System-1 Model for Search Agents

链接:https://arxiv.org/abs/2610.05107

作者:Congfeng Cao,Lipeng Zuo,Konstantinos Papakostas,Qiwei Xu,Songwei Xu,Lun Zhou,Zhaochun Ren,Yougang Lyu,Xiaohui Yan

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:repeatedly make short, agents repeatedly make, evidence sufficiency, repeatedly make, Search

备注:

点击查看摘要

Abstract:Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.

135. 【2610.05099】Small Agents with Semantic Search: Efficient Multilingual Code Localization

链接:https://arxiv.org/abs/2610.05099

作者:Maxence Lasbordes,Aarush Sinha,Raphael Sourty,Amélie Chatelain,Djamé Seddah

类目:Computation and Language (cs.CL)

关键词:Locating relevant files, Locating relevant, code repositories, core subtask, operating over code

备注:

点击查看摘要

Abstract:Locating relevant files from natural-language requests is a core subtask for agents operating over code repositories. We investigate whether this task can be delegated to compact, specialized models to enable on-device search while reducing the token usage, latency, and inference cost of larger agents. We show that semantic search improves file localization, with gains in accuracy, cross-language transfer, and inference efficiency. To study this setting, we introduce a training framework for file-localization agents built around ColGREP, a local semantic search tool based on late-interaction retrieval models. Our recipe combines weighted supervised fine-tuning on teacher trajectories, assigning turn-level credit based on retrieval outcomes, with reinforcement learning on localization quality. We train three model families with fewer than two billion parameters to formulate search queries, inspect retrieved content, and identify relevant files. On localization tasks derived from SWE-bench Lite and Multi-SWE-bench Flash, ColGREP-equipped agents substantially improve over their base models and outperform corresponding GREP-based agents. In addition to improving localization accuracy, ColGREP reduces mean end-to-end trajectory latency by 44.1\% on CPU while using 29.1\% fewer tokens, and enables better generalization to programming languages unseen during fine-tuning. These results suggest that compact, tool-specialized localization agents can provide an efficient interface between natural-language requests and large codebases.

136. 【2610.05097】ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory

链接:https://arxiv.org/abs/2610.05097

作者:Hao Jiang,Zhanyu Guo,Chenwei Wu,Yichen Guo,Qizhe Zhang,Junchi Yao,Jixian Wu,Jinhao You,Kai Tang,Jiajun Cao,Tinghao Wang,Mengyu Wang,Leo Anthony Celi,Shanghang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:multimodal large language, weakening visual grounding, visual input diminishes, initial visual input, large language models

备注: 32 pages. Code coming soon

点击查看摘要

Abstract:As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that local evidence benefits from global context, compact latent representations balance accuracy and visual-context cost, and the utility of memory access depends on the reasoning state. Guided by these findings, we propose ReMAP (Reasoning-Time Memory-Augmented Perception), which couples two complementary latent memories: a static, question-conditioned Global memory that preserves scene and cross-image context, and a dynamic Local memory that uses this context as an anchor while selecting and re-encoding region-level evidence according to the current reasoning state. Both memories return compact latent tokens inserted into the reasoning sequence, and a reinforcement-learning access policy trained with branched rollouts decides when to continue reasoning or invoke Global or Local memory. On ten benchmark families, ReMAP outperforms prior visual-memory methods on all four multi-image benchmarks, exceeding the strongest prior results on MuirBench and MIMIC by 8.38 and 14.84 percentage points. Across four backbone families, enabling memory access improves over the same trained model with memory disabled, and on shared V*Bench, CV-Bench-2D, and MuirBench questions ReMAP reduces the visual tokens entering the reasoning sequence by 51.0-76.8% relative to the native-resolution backbone. Further analyses show that Global and Local memory form distinct yet complementary latent representations. Together, these components restore the perceptual cycle by letting the reasoning state trigger targeted visual retrieval, with the retrieved evidence guiding subsequent reasoning.

137. 【2610.05095】vMF Sentence LDA: A Spherical Topic Model over Sentence Embeddings

链接:https://arxiv.org/abs/2610.05095

作者:Ryotaro Kobayashi,Yuri Murayama,Kiyoshi Izumi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Latent Dirichlet Allocation, Latent Dirichlet, Dirichlet Allocation, LDA, sentence embeddings

备注:

点击查看摘要

Abstract:Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure of the document nor the similarity in meaning between words. Earlier work has responded to this limitation in two ways: many models have introduced embeddings, and some have assigned topics to sentences rather than to words. Their combination, a topic model that observes sentence embeddings, remains little explored. We propose vMF Sentence LDA (vSLDA), which keeps the admixture structure of LDA, observes each sentence as its L2-normalized embedding and models each topic by a von Mises-Fisher (vMF) distribution, which matches the cosine geometry of sentence embeddings. Its per-topic parameter count and per-iteration cost are linear in the embedding dimension, versus quadratic for the full-covariance Gaussian distribution in the existing model over sentence embeddings. We evaluate vSLDA where the limitation is expected to matter most, among topics that share much of their vocabulary: the topics that subdivide the one subject of a collection, and the narrow topics that result when a corpus is divided into a large number of topics. On two corpora, vSLDA attains the best mean rank against eight baselines when the fine categories within each coarse category are classified from the document-topic distributions. On the whole corpus, its advantage appears or widens as the number of topics grows. Weighting the word frequencies of each sentence by its topic posterior yields expected topic-word counts of the same form as those of LDA, so that the standard topic coherence and diversity measures apply to models that assign topics to sentences. On their product, topic quality, vSLDA leads in most within-category conditions of both corpora.

138. 【2610.05094】How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

链接:https://arxiv.org/abs/2610.05094

作者:Laurène Vaugrante,Thilo Hagendorff

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, increasingly use Large, Language Models, Researchers increasingly

备注:

点击查看摘要

Abstract:Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.

139. 【2610.05084】owards cross-cultural study of folksong lyrics with machine translation

链接:https://arxiv.org/abs/2610.05084

作者:Anna Dvořáková,Anna Aljanaki,Danbinaerin Han,Peter van Kranenburg,Matěj Kratochvíl,Inna Lisniak,Zdeněk Vejvoda,Jan Hajič jr

类目:Computation and Language (cs.CL); Digital Libraries (cs.DL)

关键词:universally present, Natural Language Processing, human societies, language, human

备注: 13 pages, 1 figure, 3 tables

点击查看摘要

Abstract:Music is universally present in human societies. Ethnomusicologists have long been documenting the diverse expressions of human musicality, and comparative musicology has recently brought several studies of folksong to a more global scale. Such cross-cultural research has not been conducted on lyrics: the language barrier has so far prevented work with multi-lingual data. However, Natural Language Processing (NLP) technologies have reached a stage where this language barrier may no longer be prohibitive. Combining folksong lyrics corpora across five languages, we machine-translate them to a pivot language with a pre-trained neural topic model, and we examine the relationship between content and social function within each language, and across languages for wedding songs. As expected, human evaluation of translation results shows that non-Indo-European languages suffer from overall worse translation quality. Experiments with topic models then indicate that the content of lyrics is at best partially related to the social function of folksongs across all languages. These experiments are just first steps into cross-cultural folk musics lyrics analysis; however, they do indicate that a previously unobserved web of cross-cultural relationships beyond ethnomusicological typologies may be uncovered through the study of what people sing across the world's diverse folk musics.

140. 【2610.05076】Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

链接:https://arxiv.org/abs/2610.05076

作者:Cheng Luo,Bing Li,Bernard Ghanem

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:model store information, Test-time training, store information, TTT, Test-time

备注:

点击查看摘要

Abstract:Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.

141. 【2610.05069】Usage-Modulated Sentiment Representations in Large Language Models

链接:https://arxiv.org/abs/2610.05069

作者:Hongfei Du,Jiacheng Shi,Yanfu Zhang,Gang Zhou,Ye Gao

类目:Computation and Language (cs.CL)

关键词:LLM activation spaces, Prior work suggests, Prior work, LLM activation, fully capture sentiment

备注: Accepted to EMNLP 2026 (main conference). 18 pages, 3 figures, including appendices

点击查看摘要

Abstract:Prior work suggests that sentiment can often be captured by approximately linear directions in LLM activation spaces, but a single direction may not fully capture sentiment representations. In natural communication, sentiment is shaped not only by polarity but also by usage factors, such as tone and audience adaptation. We test whether these factors systematically modulate sentiment representations beyond a shared sentiment direction. We construct a controlled paired dataset that holds event content fixed while varying sentiment polarity and usage factors, and analyze Llama, Mistral, and Gemma. We identify a shared sentiment direction, remove it, and test the residual structure through erasure and generation-time tone steering. Across models, the shared direction is robust (median cosine 0.953-0.975), yet removing it leaves 0.833-0.909 of the original positive-negative representation-difference norm. The residuals contain compact, reproducible usage-conditioned structure. Targeted erasure weakens held-out usage metrics more than random and label-shuffled controls. On Llama, outputs steered along residualized tone components are preferred in 92.8% of blind target-tone comparisons while preserving the requested sentiment polarity in 98.7% of evaluated outputs.

142. 【2610.05044】AraYoungVoices: A Diverse L1/L2 Corpus of Arabic Child and Adolescent Speech

链接:https://arxiv.org/abs/2610.05044

作者:Shammur Absar Chowdhury,Zien Sheikh Ali,Houssam Eddine-Othman Lachemat,Hamdy Mubarak

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)

关键词:systems primarily target, substantial performance gaps, ASR systems primarily, primarily target native, target native adult

备注:

点击查看摘要

Abstract:State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes 146 native Arabic (L1) and 140 second-language (L2) speakers, with native speakers spanning Egyptian, Gulf, Levantine, and North African dialectal backgrounds and L2 speakers representing diverse linguistic backgrounds across the Americas, Asia, Africa, and Europe. We benchmark four pretrained ASR models under zero-shot and fine-tuned settings using unseen-speaker-$\$-unseen-prompt (USUP) and unseen-speaker-$\$-seen-prompt (USSP) evaluations. Results show that L2 speech remains substantially more challenging than L1 speech, with the largest errors observed mainly for younger L2 speakers. Age-specific fine-tuning improves the matched age group, while joint fine-tuning provides a stronger balance across populations. ASR hypotheses are also consistently closer to the standard reading prompt than to the verbatim transcription, particularly for L2 speech, suggesting partial normalization of reading deviations.

143. 【2610.05041】Communication Shapes Collective Inference in Self-Adapting LLM Societies: Evidence from Mafia

链接:https://arxiv.org/abs/2610.05041

作者:Haonan Huang,Joey Xiao

类目:Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:group identify hidden, group, group adapts, identify hidden adversaries, group identify

备注:

点击查看摘要

Abstract:When does communication help a group identify hidden adversaries, and how does its value change as the group adapts? In Mafia, an informed minority hides inside an uninformed majority whose only evidence is open play. The zero-information game, where each day's vote eliminates a random player, is exactly solved and scores every society; matched-casting comparisons between protocols identify the effect of communication. Societies of 8-100 claude-haiku-4-5 agents (7,416 analyzed games, 1.9M model calls) adapt by rewriting and inheriting private strategy notes. Simultaneous broadcast improves adversary identification over silence in all nine compositions tested (8-46 players). Turn-taking removes most of this advantage; its voting landslides are as frequent as broadcast's but land on mafia near chance (1.08x versus 2.53x). At 70 players, agents reading eight statements per day identify adversaries worse than silent ones, and limited talk is worth less than at 46 players. Adaptation is fast but need not help. In their first broadcast games, citizens announce their role far more often than mafia (91% vs. 30%) and first-day votes find mafia at three times chance; within two generations citizens stop announcing and the cue fades, a change the inherited notes carry. In controlled redeployments at 16 players, societies carrying sixty generations of their own notes score below societies with none. Communication shapes both collective inference and the signals it depends on, so a protocol's value must be measured together with the adaptation that changes those signals.

144. 【2610.05039】Causal Improvement Graph for Agentic Harness Optimization

链接:https://arxiv.org/abs/2610.05039

作者:Junjie Zhang,Shunyu Liu,Haoyu Wang,Ting-En Lin,Yongbin Li,Dacheng Tao

类目:Computation and Language (cs.CL)

关键词:controls execution flow, Agentic Harness, constructs task context, execution flow, context and controls

备注:

点击查看摘要

Abstract:Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.

145. 【2610.05031】When LLMs Sit Above Diagnostic Tools: Unrealized Complementarity in Industrial Fault Diagnosis

链接:https://arxiv.org/abs/2610.05031

作者:Donghwan Kim

类目:Computation and Language (cs.CL); Systems and Control (eess.SY)

关键词:Large language models, Large language, stronger combined system, specialized tools, combined system

备注: 35 pages, 5 figures + 1 appendix figure

点击查看摘要

Abstract:Large language models are increasingly used as integration layers above specialized tools, but a stronger component does not necessarily produce a stronger combined system. Across five diagnostic datasets (bearing vibration, process monitoring, semiconductor equipment), we study whether an LLM can reliably use external diagnostic information; paired repeat calls separate advice effects from output instability. In all five, conflicting external information overturned initially correct LLM judgments. Among the four datasets with direct integration comparisons, none showed a consistent advantage for implicit LLM integration over the stronger standalone source. On a Tennessee Eastman confirmation set whose protocol was fixed before evaluation, unaided accuracy was 64.67%, implicit LLM-specialist integration 77.43%, and the specialist alone 83.33%. Specialist information improved the LLM by 12.8 points (95% interval 9.7 to 15.9), yet the integrated output stayed 5.9 points below the specialist (95% interval -12.0 to -0.7). A two-source selector oracle reached 92.76%, indicating complementarity that the integrated output did not fully realize. The integrated output missed 140 of 295 specialist corrections (47.5%) but lost 15 of 99 initially correct LLM judgments (15.2%). The deficit remained under prompt and specialist sensitivity analyses. Among CWRU cases solved under both evidence presentations, task-aligned physical evidence yielded lower estimates of susceptibility to incorrect advice in six of seven models (five intervals excluding zero); higher reasoning effort gave no reliable reduction in five models, and a separate four-model TEP analysis gave no clear evidence that it resolves the integration problem. Source quality and integration quality should be evaluated separately: an integration layer should be compared with its stronger standalone component, not only with the unaided LLM.

146. 【2610.04974】IREA: Intermediate Representation-based Embedding Alignment for Normative RAG

链接:https://arxiv.org/abs/2610.04974

作者:Mirae Han,Sihyeong Yeom,Harksoo Kim

类目:Computation and Language (cs.CL)

关键词:Large language models, shown strong performance, Large language, questions involving ethical, shown strong

备注:

点击查看摘要

Abstract:Large language models (LLMs) have shown strong performance across various tasks, but they still struggle with questions involving ethical judgment. Previous studies have attempted to train LLMs on ethical standards, but the diversity and relativity of ethical norms make them difficult to fully internalize in model parameters. As an alternative, we introduce normative RAG, a retrieval-augmented approach that supports ethical judgment using external normative knowledge. Normative retrieval involves a distinct asymmetry between context rich narrative queries and generalized normative statements. Existing factual retrieval methods rely on query-only expansion into a document-like form, making them insufficient for resolving this asymmetry. Therefore, we propose Intermediate Representation-based Embedding Alignment (IREA), a bidirectional alignment method that maps both text types into a shared situation-behavior representation. This representation captures ethically salient contextual and behavioral information in a normalized form, reducing surface-level discrepancies and improving alignment in the embedding space. Experimental results show that IREA improves normative retrieval and downstream ethical judgment across multiple settings, demonstrating the effectiveness of bidirectional alignment for normative RAG.

147. 【2610.04973】rajLong: Co-Designing Agentic and Long-Context Supervision for Mid-Training

链接:https://arxiv.org/abs/2610.04973

作者:Miao Peng,Qintong Zhang,Nuo Chen,Yuhan Li,Guochen Yan,Xinran Gu,Hongqiu Wu,Hai Wang,Lydell Huang,Wentao Zhang,Jia Li

类目:Computation and Language (cs.CL)

关键词:extended interaction histories, workplace tasks increasingly, tasks increasingly rely, LLM agents, interaction histories

备注:

点击查看摘要

Abstract:LLM agents for coding, search, and workplace tasks increasingly rely on long-context capabilities to effectively aggregate and reason over extended interaction histories. Recent work has incorporated agent trajectories into mid-training stage, drawing on their naturally long and interaction-rich structure. Yet how to organize the information within these trajectories into effective mid-training supervision remains underexplored. In this work, we investigate the relationship between long-context and agent atomic capabilities and introduce TrajLong, a novel framework that compiles trajectories into long-context training tasks with dense supervision, targeting three representative atomic capabilities: evidence grounding, cross-evidence aggregation, and temporal state maintenance. We mid-train Qwen3-14B-Base and Qwen3-30B-A3B-Base with data compiled by TrajLong, followed by supervised fine-tuning. Experiments on 6 long-context and 12 agent benchmarks demonstrate broad performance gains, with controlled ablations showing improvements over raw and masked trajectory baselines. Capability-level analyses further reveal task-dependent associations between long-context and agent atomic capabilities. These findings suggest that the shared capability demands of long-context reasoning and agent execution provide a principled basis for designing mid-training data to develop downstream agent capabilities.

148. 【2610.04967】One Token Can Be Enough: Bridging Prompting and Activation Steering with Prefix Steering

链接:https://arxiv.org/abs/2610.04967

作者:Xudong Zhu,Zhihui Zhu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Prompting guides language, steering, intervention, guides language model, initial context

备注: 34 pages, 7 figures. Project page: [this https URL](https://xudongzhu.com/projects/prefix-steering/)

点击查看摘要

Abstract:Prompting guides language model behavior through the initial context, whereas activation steering often intervenes throughout generation. A natural question is whether steering can produce effects on subsequent computation similar to those of prompting. Under fixed-state attention assumptions, we establish sufficient conditions for single- and multi-token steering to match prompt-induced attention-head outputs, and characterize how changes in input representations affect this match and its approximation error. This attention-level connection leads us to ask whether, at the behavioral level, steering can also guide subsequent generation through a brief initial intervention. We study Prefix Steering, which applies existing steering directions and operators over a short span starting at the final prompt token, with no further direct intervention afterward. We examine how intervention duration and strength jointly shape the control-capability trade-off. Across four models and five tasks, intervention over a short span, even a single token, often retains much of full steering's behavioral control while better preserving general capabilities, offering a trade-off competitive with, and in some settings better than, prompting and alternative steering-strength policies. Prefix Steering also remains effective on final-answer formatting tasks after reasoning, suggesting that a brief initial intervention can influence behavior expressed well after steering ends. These findings challenge the common practice of steering every generated token and motivate a more dynamical view of activation steering, in which a brief intervention can alter the trajectory of subsequent generation without continued intervention.

149. 【2610.04961】Building LLM Agent Systems the Deep Learning Way: From Modular Design to Architecture Search

链接:https://arxiv.org/abs/2610.04961

作者:Tao Feng,Pengrui Han,Zhongjie Dai,Jiaxuan You

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, LLM agent systems, LLM agent, LLM agent system

备注: 20 pages, 16 figures

点击查看摘要

Abstract:Large Language Models (LLMs) have revolutionized AI research and enabled exciting agent systems. To build a complex LLM agent system, most existing research relies on insights from other domains or heuristics to manually build the agent system. However, this approach often requires heavy hand-engineering and fails to fully optimize for the downstream task of interest. Inspired by the tremendous success of deep learning, we propose to construct LLM agent systems in a modular manner, similar to building a deep neural network. Our key insight is to make analogies between LLM building blocks, such as retrievals, memories, and prompting strategies, and the successful deep learning modules, such as MLPs, attention, and recurrent modules. We further design forward inference and feedback mechanisms for LLMs, where prompts in LLMs are considered as the weights in deep models, and the prompt optimization from feedback is analogous to the back-propagation algorithm. We additionally leverage a search algorithm to search for the best configuration of LLM agent systems, similar to the neural architecture search (NAS) in deep learning research. Comprehensive experimental results demonstrate that the proposed deep learning recipe for LLM agent systems is highly effective, in particular: (1) Organizing LLM modules into deep-learning-style architectures yields noticeable performance gain; (2) Automatic prompt optimization, equivalent to backpropagation, is efficient in incorporating feedback from the task of interest and achieves at least 5% performance improvement; (3) NAS equivalent algorithm works well for further optimizing the LLM agent system architecture with 11% performance gain compared with randomly designed architectures. Overall, our research demonstrates the exciting opportunity of transferring the success of deep learning to building LLM agent systems.

150. 【2610.04956】From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment

链接:https://arxiv.org/abs/2610.04956

作者:Zeshen Zhang,Han Zhao,Weihao Cui,Quan Chen,Yu Liu,Yongjun Deng,Jing Yang,Jiuchen Shi,Chen Chen,Youmin Chen,Yu Feng,Minyi Guo

类目:Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)

关键词:Large Language Models, Language Models, Large Language, public cloud services, Service Level Objectives

备注: 22 pages, 16 figures, 4 tables

点击查看摘要

Abstract:As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objectives (SLOs) when serving multiple LoRA-adapted services simultaneously. Existing serving frameworks suffer from severe SLO violations due to the computational overhead of LoRA layers and the rigid nature of batch scheduling. To address this, we propose HALO, a scheduling method tailored for LoRA-assisted on-premise LLM deployment. HALO introduces two key innovations: a spatial multiplexing strategy that overlaps Base and LoRA computations by partitioning GPU Streaming Multiprocessors (SMs), and an SLO-aware scheduler that decouples request execution based on "request-level slack." By prioritizing urgent tasks and utilizing idle budget for traffic shaping, HALO significantly mitigates resource contention. Our evaluation demonstrates that HALO minimizes SLO violations while improving throughput compared to state-of-the-art baselines.

151. 【2610.04918】Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning

链接:https://arxiv.org/abs/2610.04918

作者:Lin Qiu,Yao Liu,Diyi Hu,Hanqing Zeng,Onur Gungor,Chujie Chen,Jiayi Liu,Jianyu Wang,XueLin Zheng

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Reinforcement learning, trajectory is worth, learning with verifiable, decisions that produced, scales multimodal reasoning

备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing problem. A controlled visual intervention yields a per-token evidence response; robust within-trajectory coordinates remove incidental scale; and a budgeted entropic router distributes a fixed amount of sequence utility according to perceptual dependence. A residual support path guarantees positive credit at every valid position, and an analytic correction restores the prescribed credit mass exactly. The resulting field is selective, bounded, full-support, and invariant to response-local score shifts, and recovers hard token selection as a limiting case. Across four model families and seven reasoning benchmarks, RVCO improves accuracy over strong RLVR baselines while maintaining late-stage optimization stability, corruption robustness, and competitive training cost. Rewards, rollouts, and the group-relative advantage estimator are unchanged; only the geometry of token-level credit differs.

152. 【2610.04914】Monitorability Disposition in Large Reasoning Models

链接:https://arxiv.org/abs/2610.04914

作者:Shahriar Golchin,Marc Wetter

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:large reasoning models, real-world practice, large reasoning, models, model

备注:

点击查看摘要

Abstract:Monitoring the chain-of-thought (CoT) of large reasoning models (LRMs) is a common way to detect misbehavior in real-world practice. However, current monitoring is passive: a separate model inspects the session only after execution. This means harm may already have occurred before it is caught. An active alternative is to have the model self-report its misbehavior as it happens. Whether models are willing to do this, however, is unknown. We introduce "monitorability disposition": a model's willingness to make itself monitorable and stay monitored throughout inference when warranted. We measure it as the fraction of warranted cases in which a model self-reports its own misbehavior via tool calls to available monitoring channels. We evaluate four LRMs on three misbehaviors (sycophancy, reward hacking, and bias) while varying the available monitors (AI and human) and the pressure to use the monitoring tools. We find that when tool use is optional, models self-report in only about 16% of warranted cases on average. Increasing tool-use pressure does not improve reporting where it matters: high-severity misbehavior is never self-reported. Models also systematically select the monitor they perceive as least strict. Overall, we identify monitorability disposition as a new contributing factor to model monitorability: when sufficiently strong, it keeps models seeking monitorability throughout inference.

153. 【2610.04906】Scaling Verifiable Environments for Long-horizon Work Agents

链接:https://arxiv.org/abs/2610.04906

作者:Jiazheng Zhang,Long Ma,Yunxian Yang,Zhiheng Xi,Zhikai Lei,Yajie Yang,Chenyang Liao,Enyu Zhou,Yang Nan,Yuchen Tian,Senjie Jin,Yibo Wang,Wei He,Boyang Liu,Jixuan Huang,Xin Guo,Zhezheng Hao,Xinbing Liang,Zhihao Zhang,Changzhi Zhou,Wiggin Zhou,Tao Gui,Qi Zhang,Xuanjing Huang,Clarenceai,Aiden Adams

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Work agents operate, professional knowledge-intensive work, requiring training environments, Work agents, knowledge-intensive work

备注:

点击查看摘要

Abstract:Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.

154. 【2610.04899】Rewrite What Matters: Adaptive Multilingual Query Rewriting for Reasoning via Agentic Reinforcement Learning

链接:https://arxiv.org/abs/2610.04899

作者:Rui Qi,Yufeng Chen,Yunlong Liang,Chuan Meng,Sijin Lu,Ge Shi,Jinan Xu,Fandong Meng,Kaiyu Huang

类目:Computation and Language (cs.CL)

关键词:queries with equivalent, leading to performance, performance disparities, equivalent semantics, rewriting

备注:

点击查看摘要

Abstract:In multilingual scenarios, queries with equivalent semantics but in different languages could guide the model into different reasoning trajectories, leading to performance disparities. To mitigate this gap, previous studies typically apply a one-size-fits-all query rewriting strategy, such as translation, which overlooks the fact that different scenarios require diverse types of semantic transformations. In this paper, we propose mRewriter-R1, an agentic multilingual query rewriting framework with reinforcement learning. Unlike single-turn rewriting, mRewriter-R1 formulates multilingual query rewriting as a multi-turn sequential decision-making process, where the model dynamically performs multi-aspect optimization through adaptive operator selection. Experimental results demonstrate that mRewriter-R1 outperforms all strong multilingual rewriting baselines on different large reasoning backbones. Further analyses show that the learned policy can adaptively decide on rewriting operators according to query characteristics, exhibiting strong generalization ability across diverse reasoning tasks, and plug-and-play compatibility with heterogeneous reasoning language models.

155. 【2610.04898】A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese

链接:https://arxiv.org/abs/2610.04898

作者:Hongao Zhu(1),Muxiaoqiao Xu(2),Yikang Liu(3),Siyuan Song(2 and 4),Yuxia Wang(2),Byung-Doh Oh(5),Hai Hu(6) ((1) Department of Linguistics, University of California San Diego, (2) School of Foreign Languages, Shanghai Jiao Tong University, (3) School of Computer Science, Shanghai Jiao Tong University, (4) Department of Linguistics, University of Texas at Austin, (5) Division of Linguistics and Multilingual Studies, Nanyang Technological University, (6) Department of Language Science and Technology/Division of AI and the Humanities, Hong Kong Polytechnic University)

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Mandarin Chinese reading, Mandarin Chinese, Shortest Matching Sequence, Chinese reading times, Chinese reading

备注: 15 pages, 3 figures. Hongao Zhu and Muxiaoqiao Xu contributed equally. Correspondence to: [this http URL](http://hai.hu) @polyu. [this http URL](http://edu.hk)

点击查看摘要

Abstract:This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to $n$-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.

156. 【2610.04888】No Hindsight for LLM Fact-Checkers: Measuring Leakage Channels in Misinformation Detection

链接:https://arxiv.org/abs/2610.04888

作者:Kuan-Hua Wu Lu,Yohanes Andre Setiawan

类目:Computation and Language (cs.CL)

关键词:large language model, automated fact-checking scales, social media, large language, verdict scores

备注: 11 pages, 5 figures

点击查看摘要

Abstract:As automated fact-checking scales on social media, large language model (LLM) verdict scores can look stronger than warranted. One reason is that evaluations mix in information that was not knowable at claim time. Two channels are easy to conflate: outcomes memorized in pre-training and retrieved evidence published after the claim. Yet standard benchmarks rarely separate the two. In this study we measure both channels on AVeriTeC and QuanTemp++ by reconstructing point-in-time evidence conditions and probing for outcome information encoded in model representations. We find substantial evidence of parametric leakage, that can be hidden by the aggregate accuracy, while a simple representation bottleneck reduces this future leakage more efficiently than a mutual-information-based training penalty. We also find that allowing post-claim evidence inflates zero-shot accuracy by 6.3 points in AVeriTeC while the effect is negligible in QuanTemp++, where retrieval provides little post-claim evidence. These results show that misinformation benchmarks can overstate fact-checking performance when they do not account for what information was actually available at claim time.

157. 【2610.04875】SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models

链接:https://arxiv.org/abs/2610.04875

作者:Chung-En Ho,Weiyu Sun,Cheng-Jhih Shih,He Li,Yong Liu,Yingyan(Celine)Lin

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Diffusion large language, single forward pass, iterative block denoising, Diffusion large, generate text

备注:

点击查看摘要

Abstract:Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.

158. 【2610.04851】CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery

链接:https://arxiv.org/abs/2610.04851

作者:Tao Feng,Fangxu Yu,Zijie Lei,Jiaru Zou,Changjiang Jiang,Yi Yan,Jiaxuan You,Pan Lu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:repeated attempts, attempts while continuing, continuing to explore, Open-ended discovery requires, discovery requires learning

备注:

点击查看摘要

Abstract:Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement learning (RL) enables such adaptation; however, strongly favoring high-reward trajectories may suppress low-reward yet potentially promising directions too early. We introduce CURIO, a curiosity-driven test-time learning framework that complements task feedback with an Intrinsic Curiosity World Model (ICWM). The ICWM learns transitions in the policy's hidden-state representation and supplies prediction-error bonuses at sampled tokens outside the policy's top-k choices. Epoch normalization and an annealed weight regulate their contribution to the policy update. On six mathematical discovery tasks and single-cell denoising with Qwen3 backbones from 8B to 235B, three-run means improve over a matched task-only RL control on five mathematical objectives, match the best reported performance on Circle Packing, and improve denoising Score and mean squared error (MSE) on both held-out corpora at every tested scale. Relative gains reach 18.3% on Hadamard and 10.8% on denoising Score. Code-diversity measurements show greater structural variation among generated programs, supporting curiosity as a complementary exploration signal for learning in open-ended discovery.

159. 【2610.04830】Cluster Validation Indices as Self-Supervised Objectives for Text Representation Learning

链接:https://arxiv.org/abs/2610.04830

作者:Kishor Kumar Bhaumik,Nicolas Roque dos Santos,Neil Shah,Jia Chen,Evangelos E. Papalexakis

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:pretrained language encoder, Self-supervised fine-tuning refines, refines the embedding, embedding space, pretrained language

备注:

点击查看摘要

Abstract:Self-supervised fine-tuning refines the embedding space of a pretrained language encoder without labels. However, the commonly used approaches are computationally expensive. Specifically, contrastive learning-based methods need multiview data and in-batch negative examples, while negative-free approaches require auxiliary graphs/networks. An interesting question arises: can self-supervised fine-tuning be done without relying on either additional negatives or graph data? To answer this question, we introduce SilK (Silhouette-guided K-means), which trains on a Cluster Validation Index, an internal measure of cluster quality without using labels. SilK clusters the corpus and then regresses a simplified silhouette toward a target value. Each document is compared only against the k cluster centroids, never against other documents, so the method needs no augmentation, no negative pairs and one view per document. On BERT-base, SilK trains 1.46x faster per epoch than the fastest baseline we evaluate and uses 45.4% less peak GPU memory than the leanest one. Under frozen-encoder linear probing, SilK stays competitive with the best baselines on three downstream tasks.

160. 【2610.04829】Viva La Vida: Verification and Accumulation Failures in Multi-Agent Proof Search

链接:https://arxiv.org/abs/2610.04829

作者:Benji Xu,Ken Zheng,Noah Han

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)

关键词:models judging model, ultimately language models, language models judging, agentic prover works, judging model outputs

备注:

点击查看摘要

Abstract:When an agentic prover works on an open problem, there is no proof assistant to fall back on: its verifier and lemma library are ultimately language models judging model outputs. We instrumented such a system end to end and analyzed $51{,}754$ traced observations across three full runs ($186$ hours, \$$5{,}694$). We find three connected failure modes. First, the three-model verifier requires unanimity and treats parse or API failure as non-approval; in $10$ of $12$ verification events, one member returned no parseable output or an API error, making acceptance arithmetically impossible without surfacing an error. Second, when the ensemble did function, one verifier approved $3$ attempts that GPT rejected, each claiming to resolve the open problem; a single-verifier design would therefore have announced a solution three times. Third, because nothing could be approved, every review was a refutation, yet the lemma extractor mines reviews as well as proofs: $24$ of $93$ lemmas ($26\%$) were extracted from rejected arguments with their refutational context removed. Taken together, these findings show that without external verification, supervision is itself a critical trust boundary: systems must distinguish abstention from rejection, preserve useful disagreement, and preserve the provenance and polarity of information before it becomes future context.

161. 【2610.04824】Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications

链接:https://arxiv.org/abs/2610.04824

作者:Peng Qi,Chunliang Lyu,Gang Li,Fabian Chan,Cheng Chang,Ignacio Cases,Will Lu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:demonstrated strong capabilities, perform complex open-ended, complex open-ended tasks, foundation models, based on foundation

备注:

点击查看摘要

Abstract:AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically propagated errors; (b) high cost and latency due to FM calls, repeated in full whenever a task recurs with different inputs; (c) FMs' limited context and instruction following capability confine how well agents manage the ever-growing execution context and follow complex plans. We introduce $\textbf{A}$gent $\textbf{B}$ehavior as $\textbf{C}$ode $\textbf{Agent}$ (ABCAgent), which uses a symbolic program (e.g., Python code with potential neural functions) to fully specify the agent's behavior at runtime, with a powerful FM agent editing that program for flexibility. Behavior is thus specified without premature variable binding, and its execution is deterministic. We evaluate ABCAgent on six agent benchmarks, two of which we construct to test how well a derived program generalizes to variants of the task it was written for. ABCAgent matches a model-matched neural agent on GAIA and augmented GAIA, and surpasses it where robustness and long control flows matter: 98.3% against 97.3% on GSM-Symbolic ($p = 0.001$), 71.9% against 47.4% $\mathrm{Pass}^4$ on the telecom domain of $\tau^2$-bench ($p = 0.0001$), and more records written correctly at every loop length on our control-flow-augmented WorkArena benchmark. For more parametric task families, ABCAgent is also significantly superior in efficiency. Without authoring a new program, ABCAgent solves 92.6% of GSM-Symbolic instances and 20.1% of augmented GAIA variants, which yields $5.2\times$ lower latency and $7.0\times$ lower cost on GSM-Symbolic, 19% lower cost on augmented GAIA, and $9.5\times$ lower agent latency on $\tau^2$-telecom.

162. 【2610.04803】How Do People Challenge Racial Stereotypes Online? Counter-Story Detection Across Reddit Communities

链接:https://arxiv.org/abs/2610.04803

作者:Uma Sushmitha Gunturi,Jimin Mun,Maarten Sap,Maria Antoniak

类目:Computation and Language (cs.CL)

关键词:powerful mechanism people, challenge dominant narratives, powerful mechanism, mechanism people, challenge dominant

备注: Accepted to EMNLP 2026 (Main Conference). 28 pages, 12 figures, 24 tables. Content warning: this paper contains examples of racial stereotypes that may be upsetting or offensive. Code: [this https URL](https://github.com/UmaGunturi/counter_story_detection)

点击查看摘要

Abstract:Counter-storytelling is a powerful mechanism people use to challenge dominant narratives. Unlike other forms of counterspeech that have been widely studied in computational social science, counter-storytelling has largely been overlooked. Counter-stories are difficult to detect automatically; they are relational (defined with respect to expressions of racial stereotypes) and structurally diverse (drawing on stories that describe lived experiences, witnessed events, exemplars, and hypotheticals). We introduce a first framework for detecting and characterizing counter-storytelling against racial stereotypes at scale. This includes (1) a three-dimensional taxonomy grounded in narratology and Critical Race Theory and (2) a multi-stage pipeline that identifies relational pairs of stereotypes and counter-stories in noisy Reddit discourse. Using this pipeline, we annotate 25,549 Reddit posts across 615 communities and identify 1,312 counter-stories. Our analysis shows that speaker identity and post context shape how counter-stories are told. For example, in-group writers favor first-person testimony, often adopting the role of self-reflective insiders. Our work shows how computational methods can scale qualitative approaches to identify and characterize counter-storytelling as a contextual narrative practice, with implications for content moderation, narratology, and racial discourse analysis.

163. 【2610.04753】More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

链接:https://arxiv.org/abs/2610.04753

作者:Noam Elata,Itay Lamprecht,Mikey Shechter,Daniel Ohayon,Itay Hubara,Daniel Soudry

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Large Language, generation in Large, utoregressive generation, Language Models

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:utoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple sparse attention method designed to study the interaction between sparsity and head-count asymmetry. We formalize the benefits of this asymmetry theoretically and validate them empirically through latency measurements and quality evaluations on models up to 1.5B parameters. Together, SAGA and Atop-N achieve end-to-end decoding speedups exceeding $2\times$ over our full-attention GQA baseline at long contexts. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants on the evaluated benchmarks. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture, enabling practitioners to benefit from our approach without costly retraining.

164. 【2610.04725】Verb-ICL: Rethinking In-Context Learning for Structured Prediction

链接:https://arxiv.org/abs/2610.04725

作者:Fan Bai,Hengshuo Miao,Sanjit S Batra,Hamid Reza Hassanzadeh,Ardavan Saeedi,Mark Dredze

类目:Computation and Language (cs.CL)

关键词:require modeling fine-grained, tasks pose unique, compositional outputs require, outputs require modeling, sentence-level approaches fail

备注: Accepted to COLM 2026

点击查看摘要

Abstract:Structured prediction tasks pose unique challenges for in-context learning (ICL): their compositional outputs require modeling fine-grained, token-level patterns that sentence-level approaches fail to capture, and their task-specific annotation conventions are human-defined artifacts that cannot be acquired through pretraining alone. We propose Verb-ICL, a selective annotation framework for ICL-based structured prediction that addresses both challenges. Verb-ICL first selects representative examples using a token-level coverage strategy that captures local semantic patterns critical for structured prediction, then generates actionable error feedback that codifies task-specific annotation guidelines and incorporates this feedback into ICL demonstrations. We evaluate Verb-ICL on six structured prediction datasets spanning information extraction and semantic parsing. Experiments with recent LLMs show that Verb-ICL consistently outperforms strong selective annotation baselines under low-resource settings and continues to provide gains as the annotation budget increases. Extended analyses demonstrate that the generated feedback is predominantly useful across a four-category quality taxonomy, generalizes as task-level guidance beyond instance-specific corrections, and improves performance regardless of the underlying selection strategy.

165. 【2610.04720】WNet: Discrete Wavelets Transform for Efficient Token Mixing

链接:https://arxiv.org/abs/2610.04720

作者:Rana Aref Salama,Abdou Youssef,Mona Diab

类目:Computation and Language (cs.CL)

关键词:encoding long sequences, token draw information, draw information, encoding long, token

备注:

点击查看摘要

Abstract:In a Transformer, token mixing is the step that lets each token draw information from other tokens, and it dominates the cost of encoding long sequences. Self-attention does this mixing very well: every token weighs every other token by content, which gives strong contextual modeling. That all-pairs comparison is also why its cost grows quadratically with sequence length. We introduce WNet, a Transformer encoder that replaces self-attention with token mixing based on the discrete wavelet transform (DWT). Three attention-free mixers recombine the scales: by linear fusion, by learned gating, or by letting each token choose its scales. A hybrid adds self-attention in the last layer only. A receptive-field analysis shows that wavelet mixers built from two-tap filters, such as Haar, never relate tokens outside fixed blocks, however deep the network, even when the filters are learned. Longer filters reach the whole sequence within two layers. We pre-train every model with masked language modeling on a fixed-token subset of C4 and fine-tune on GLUE, using one controlled setup with size-matched BERT and FNet baselines and a control that cannot mix tokens. The token-gated mixer trains as fast as attention at 256 tokens and 2.7 times faster at 4,096.

166. 【2610.04699】Not Self-Decidable: LLMs Cannot Draw the Boundary of What an Agent Verifier Can Check

链接:https://arxiv.org/abs/2610.04699

作者:Anthony Rhodes

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:agent faces rules, faces rules, agent faces, agent enforces rules, fixed check

备注: Accepted at the Who Verifies the Agents? Workshop at NeurIPS 2026. Workshop papers are non-archival. 21 pages (9 content pages plus references and appendices)

点击查看摘要

Abstract:A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split up front. Where the requirements come from outside, as in finance, healthcare and law, the agent enforces rules it did not write, so the split falls to runtime, recurring for every predicate of every rule on every action at a rate no reviewer can audit. Every escalation scheme assumes a model can make that decision itself, that it is self-decidable. Across six corpora, including the EU AI Act, FINRA guidance and a deployed credit agent, we collect roughly 22,000 labels from four models built by three labs. They agree almost perfectly where the answer is obvious and collapse on regulatory text; their errors run in opposite directions, so no model can be trusted as the conservative choice; and on the deployed agent's own rule-set they err together, over-claiming that a fixed check will do, the direction that never gets escalated. We introduce CoVer (corroborate-then-verify), which treats unanimity as a nomination, admitting a predicate only when the check synthesized for it survives intervention, reading fields the agent cannot write and holding under deterministic rewording. That gate rejects most of what corroboration wrongly admits, at a cost in coverage we report rather than tune away. The obvious alternative, agreement with a reference judge, certifies nothing: it climbs from 30% to 77% across calibration bands while the genuinely decidable share does not move, because a judge drawn from the population under indictment ratifies the blind spot it shares. Self-decidability is not a capability to elicit from a model but a boundary the verifier must construct.

167. 【2610.04693】Penumbra: Sample-Efficient Adversarial Search for Regulatory Obligations

链接:https://arxiv.org/abs/2610.04693

作者:Anthony Rhodes

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:carries real penalties, healthcare and law, entering finance, real penalties, violation leaves

备注: Accepted at the Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust) at NeurIPS 2026. Workshop papers are non-archival. 12 pages (4 content pages plus references and appendices)

点击查看摘要

Abstract:Agents are entering finance, healthcare and law, sectors where a violation leaves no lexical signature and carries real penalties. Whether an omission is material, or a disclosure sufficient, depends on what the response left out. Probing such an obligation means finding responses one minimal edit from flipping compliance, and every probe costs a generation and two adjudications, so the binding constraint on regulatory red-teaming is sample efficiency, not volume. We introduce Penumbra, an adversarial search that walks from a verified anchor under an expanding edit budget until a two-evaluator committee changes its verdict, and emits the two adjacent responses that straddle the change. Allocation is adaptive, and the objective is coverage of the defeat surface: distinct (obligation x defeat mode) cells resolved per candidate. At matched budget, adaptive allocation reaches uniform allocation's full-budget coverage on 59% of the candidates; at equal records it covers 1.43x the defeat modes of naive enumeration, and the gain is confined to the axis it targets. On 60 screened obligations of a financial advisory constitution, Penumbra returns 144 pairs, each a compliant and a violating response that the committee places on opposite sides of the boundary, differing by a handful of words where a model asked for both directly produces texts sharing almost nothing. A second constitution, for clinical triage, reproduces this on 18 obligations: 49 pairs at the same tightness, with the same modes hardest. Pairs like these show where an obligation's own terms stop deciding, which is what an agent deployed under it must be tested against, and the search finds them at a cost that scales with the boundary, not the text.

168. 【2610.04687】Understanding Errors in LLM-Based Question Answering over Imperfect Tables

链接:https://arxiv.org/abs/2610.04687

作者:Baowen Zhang,Wei Fan,Ruman Wang,Hangting Ye

类目:Computation and Language (cs.CL)

关键词:large language models, question answering, language models, controlled studies, large language

备注: 41 pages, 7 figures

点击查看摘要

Abstract:We investigate error discovery and handling in question answering over imperfect tables through controlled studies across three large language models (LLMs) on human-reviewed RADAR-T examples. Answering questions over these tables requires handling errors that can affect the answer. We vary row order and compare original, error-marked, and repaired tables to test whether discovery depends on where errors appear and whether providing their locations is sufficient for accurate question answering. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. Complete discovery is higher for back than front placements and, averaged over the tested mean positions, for compact than widely spaced layouts. Second, providing verified error locations alone is insufficient for accurate QA, leaving a substantial accuracy gap between error-marked and repaired tables. Providing tables with human-reviewed repairs already applied raises code-assisted QA accuracy by 39.0-59.1 percentage points over the error-marked tables across the three systems. GBDI, a simple workflow, puts these findings into practice by combining error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors. On RADAR-T, GBDI raises observed QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five systems. These results highlight the importance of both reliable error discovery and effective error handling in question answering over imperfect tables. Our anonymous repository is available at this https URL

169. 【2610.04686】When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning

链接:https://arxiv.org/abs/2610.04686

作者:Zihao Zhao,Tunyu Zhang,Haizhou Shi,Yusong Zhao,Xinxi Zhang,Hao Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Social and Information Networks (cs.SI)

关键词:beat simple majority, Multi-agent debate, simple majority voting, fails to beat, beat simple

备注: 27 pages, 4 figures

点击查看摘要

Abstract:Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the first requirement through recoverable headroom, which measures cases where a correct proposal is available but the majority answer is wrong. For the second, we develop Latent Verification Debate (LVD), an accounting model in which candidate proposals receive answer-specific verification evidence before final generation. Controlled fixed-proposal interventions estimate this latent effect in equivalent peer-support units and show that correct evidence changes answer probabilities and generated decisions while proposal supply remains fixed. To improve proposal supply, we construct societies from neural-thicket agents using labeled and label-free coverage objectives. Across two backbones and matched-budget reasoning benchmarks, coverage-selected societies increase complementary proposal supply and improve aggregate accuracy in repeated stochastic evaluations. Round-level controls further show that interaction provides gains beyond applying the same finalizer directly to the initial proposals. These results identify proposal coverage and truth-sensitive evidence use as complementary conditions for debate to outperform voting. Code is available at this https URL.

170. 【2610.04683】Steering Speech-Language Models: Training-Free Task Specialization via Contrastive Activation Addition

链接:https://arxiv.org/abs/2610.04683

作者:Séverin Baroudi,Yanis Labrak,Pierfrancesco Melucci,Sergio Burdisso,Petr Motlicek,Hervé Bredin,Mirco Ravanelli,Ricard Marxer

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:Large Language Models, Contrastive Activation Addition, training-free Contrastive Activation, behavior of Large, training-free steering approaches

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Activation steering has proven effective for controlling the behavior of Large Language Models (LLMs) at inference time, but its application to SpeechLLMs remains new, and training-free steering approaches for such models are still largely unexplored. We propose a training-free Contrastive Activation Addition (CAA) protocol that derives steering vectors for common speech tasks (e.g. transcription) in SpeechLLMs from a small number of labeled utterances. We showcase that adding these vectors in the representation space, at inference time, enforces better the targeted speech task. We further show that, when combined with prompting, these vectors yield to consistent improvement over prompting alone on most evaluated tasks such as Automatic Speech Recognition (ASR) or Emotion Recognition (ER), and transfer to out-of-domain data. We additionally demonstrate the usefulness of script-normalization directions to enforce the target script of a specific language.

171. 【2610.04676】Extracting Persona Subspaces Through Iterative Nullspace Projection For Modulation

链接:https://arxiv.org/abs/2610.04676

作者:Ananya Malik,Mai ElSherief

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Large Language, Language Models, tune their semantics, Large

备注: 29 pages, 12 tables, 12 figures

点击查看摘要

Abstract:Large Language Models (LLMs) can adopt distinct personas to tune their semantics, expertise, and perspective to different users and tasks. Precise control over these traits is critical to ensure safety and reliability in model behavior. Existing methods like activation steering and prompt-based persona induction reduce a persona to a single dominant direction, missing the finer, nested traits that emerge only once that dominant signal is factored out. We introduce modulation as a setting where the persona context is already embedded in the content being manipulated, requiring control methods to amplify or suppress a trait already present rather than inject it from scratch. PaSS is an inference-time control paradigm that models personas as multi-dimensional subspaces in a model's latent space without supervised contrastive examples. The persona subspaces are extracted via iterative concept erasure and applied to modulate persona-guided generation without retraining. To extract this subspace, we use Iterative Nullspace Projections (INLP) to linearly and iteratively isolate persona-specific directions. We causally evaluate six personas against diverse tasks like MATH-500, TinyAlpaca, GSM8K, and IFEval, showing that discriminative, iterative subspace extraction captures diverse traits underlying a given persona, enabling stronger and larger modulation than single-direction additive methods, while maintaining content fidelity. We further study individual peeled directions within each subspace to uncover the distinct aspects of persona behavior they encode. Overall, we show that persona subspaces offer a controllable, interpretable, and generalizable framework for modulating LLM behavior without sacrificing task performance.

172. 【2610.04642】Grounding Probes: Generator-Independent Hallucination Detection from Observer Model Hidden States

链接:https://arxiv.org/abs/2610.04642

作者:Michael Rathmayr,Ádám Kovács,Gábor Recski

类目:Computation and Language (cs.CL)

关键词:surface checks miss, miss paraphrased fabrication, checks miss paraphrased, cost extra generations, context trades speed

备注: 13 pages, 4 figures, 4 tables. Code and per-sample predictions: [this https URL](https://github.com/MRathmayr/LettuceDetect)

点击查看摘要

Abstract:Detecting responses that retrieval-augmented generation does not ground in its context trades speed against accuracy: surface checks miss paraphrased fabrication, sampling-based methods cost extra generations. Hidden-state probes sit between the two, but every existing one reads the generating model's own activations, so a change of generator invalidates the detector and a closed-weight generator is out of reach. This paper removes that coupling. The Grounding Probe is logistic regression over the mean-pooled middle-layer hidden states of an observer language model that reads the context, question, and response in one forward pass and generates nothing, with the recipe it needs: pool over response tokens, read a middle layer, and control capacity, which closes the train-test AUROC gap from 0.087-0.202 to 0.009-0.013. Asking the observer outright, rather than reading its hidden state, costs at least +0.166 AUROC in every one of four models. Fitted on 15,090 annotated responses it reaches 0.879-0.894 AUROC on RAGTruth test across four observers, and 0.924 AUROC with 0.820 F1@0.5 averaged with a supervised span detector, 0.060 above that detector alone. One probe holds across six generators, and hold-out controls, including one in which no evaluation prompt appears in training, bound the cost of removing a generator at about 0.02 AUROC. Code, probes, and predictions are released.

173. 【2610.04636】Does Neural Complexity Improve Health Misinformation Detection? A Leakage-Controlled Cross-Corpus Benchmark

链接:https://arxiv.org/abs/2610.04636

作者:Mkululi SIKOSANA

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Increasing architectural complexity, Increasing architectural, data splits, assumed to improve, reported gains

备注: 18 pages, 3 tables. Reproducibility materials and source code are available from the author

点击查看摘要

Abstract:Increasing architectural complexity is often assumed to improve health misinformation detection, yet reported gains are difficult to interpret when studies use different corpora, preprocessing pipelines, data splits, and leakage controls. This study provides a controlled cross-corpus benchmark of five compact neural architectures (1D-CNN, LSTM, BiLSTM, CNN-LSTM, and CNN-BiLSTM), a soft-voting neural ensemble, and three classical machine-learning baselines using COVID19-FNIR and CONSTRAINT. Exact-text duplicate controls were applied before modelling; all neural systems used a common preprocessing and optimisation protocol, and neural results were repeated across three random seeds. On COVID19-FNIR, the deep ensemble achieved a mean macro-F1 of 0.9963 and ROC-AUC of 0.9994, while individual neural models ranged from 0.9945 to 0.9957 macro-F1. On CONSTRAINT, the ensemble achieved macro-F1 of 0.9272 and ROC-AUC of 0.9811, whereas a linear SVM achieved macro-F1 of 0.9574 and ROC-AUC of 0.9931. Architecture rankings changed across corpora, and simple sparse linear models remained highly competitive. The findings show that model complexity does not provide a stable performance advantage and that benchmark construction can dominate architecture choice. The study contributes a reproducible, leakage-controlled basis for evidence-driven model selection in health misinformation classification

174. 【2610.04620】Stance Drift: How AI-mediated Communication Distorts Our Message

链接:https://arxiv.org/abs/2610.04620

作者:Lingchong Liu,Yanfei Zhou,Jacob Bien,Y.X. Rachel Wang,Lucy Xia,Xin Tong

类目:Computation and Language (cs.CL); Applications (stat.AP)

关键词:remains largely untested, speaker position remains, position remains largely, Large language models, increasingly mediate human

备注:

点击查看摘要

Abstract:Large language models (LLMs) increasingly mediate human communication, from drafting emails to summarizing scientific reports, yet whether they faithfully preserve a speaker's position remains largely untested. We model AI-mediated communication as a two-step generation-extraction pipeline: one LLM produces an argument from a specified stance, and a second LLM extracts the stance from that argument. We represent the pipeline as a probabilistic state transition over five Likert-type stance categories and define the stance preservation rate (SPR) as the average probability that the extracted stance matches the initial stance. Across 112 debate propositions, none of the nine LLMs tested exceeded an SPR of 0.7 under the default configuration. Three drift patterns accounted for most of the drift: polarization, deviation from neutrality, and flipping. Among the mitigation strategies tested, including in-context learning, multiple extraction with shuffled options, assertion, and reflection, only adding medium reasoning effort to a reflection prompt for GPT-5.4 substantially improved the SPR, to 0.775, yet polarization remained the largest pattern, with 0.119 of the transition mass. An exploratory comparison with human extraction on a single proposition suggests that drift arises at both the generation and the extraction stage. These results point to a fidelity gap in AI-mediated communication, with implications for journalism, policy deliberation, scientific communication, and other domains where opinion-laden messages pass through language models.

175. 【2610.04594】SIFT: Robust Meta-Faithfulness Verification of Chain-of-Thought Reasoning Under Distribution Shift

链接:https://arxiv.org/abs/2610.04594

作者:Noor Islam S. Mohammad,Md. Basim Al Zabir Shammo,Hasan Siddiki,Mahmudul Hasan,Md. Faisal Sheikh,Jakaria Habib

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:audit reasoning models, stable properties, audit reasoning, treated as stable, detector

备注: Under review as a conference paper at ICLR 2027

点击查看摘要

Abstract:Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties. We ask whether a detector is faithful to itself under distribution shift. We formalize meta-faithfulness as an invariance principle: a valid detector must return identical verdicts on traces that differ only by transformations preserving ground-truth faithfulness. We prove three results: (i) no detector using only intervention-response profiles can separate faithful from epiphenomenal mechanisms with identical signatures; (ii) any detector relying on shift-sensitive features violates invariance at a rate independent of its in-distribution accuracy; (iii) an asymptotic certified selective-risk guarantee enables confident abstention. We operationalize the principle in FaithShift, a stress-test protocol spanning ten shift axes, and propose SIFT, a hidden-state trajectory detector trained with cross-environment invariance objectives and certified abstention. Across 14,996 traces, four domains, and eight models, three findings emerge. First, transfer collapse is real: all existing detectors show gaps $\geq 0.15$ AUROC. Second, the dominant bottleneck is sampling stochasticity, not shift: over 80% of detector instability stems from random seed variation, falsifying our preregistered prediction that shift-attributable violations exceed 0.25. Third, SIFT cuts invariance violations by 64% over the best single-seed baseline, but a four-seed ensemble of any detector narrows the margin to 0.01 (indistinguishable at matched coverage, $p=0.21$), and SIFT needs a 51% abstention rate. Cross-model transfer degrades from within-family to cross-family to open-weight-to-API, partly closed by multi-model training. We offer a framework for auditing auditors: the real barrier is detector variance, not distribution shift.

176. 【2610.04589】StegoMemory: Agentic Memory Acts as Covert Steganographic Channel

链接:https://arxiv.org/abs/2610.04589

作者:Snehasis Mukhopadhyay,Arun Nair

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:robust against stealthy, stealthy steganographic attacks, agentic memory robust, safety oversight, triggering safety oversight

备注:

点击查看摘要

Abstract:Is agentic memory robust against stealthy steganographic attacks? We carry out a large-scale red-teaming exercise to test whether agents can encode attacker-controlled strings in one session and recover them in another without triggering safety oversight. Following SHADE-Arena-style tasks, we embed malicious side tasks to encode secret strings using steganography within otherwise benign tasks and evaluate them using independent task-completion and safety oversight. We test 14,000 attack trials spanning 91 model-attack configurations, 13 models, and 7 steganographic schemes: Acrostic, Mesostic, Telestich, Word-choice, Capitalization, Position-based, and Homoglyph. 41.2% of trials successfully encode the target payload in the generated response, while 29.8% survive the memory-write step with the encoding intact. Overall, 25.4% of trials yield a recoverable payload at retrieval, and 20.1% achieve exact-match recovery, including synthetic API keys and cryptographic material never exposed to the receiver. Among the exact-match transmissions, 71.4% also pass primary task-completion scoring and are independently judged safe by the oversight model, corresponding to 14.4% of all trials in which a successful covert transmission would appear to be an ordinary, benign interaction under task-level evaluation. Our results demonstrate that agentic memory can function as a persistent cross-session covert channel. The results further show that the principal bottleneck occurs at memory persistence rather than retrieval: once a steganographic payload survives the memory-write stage, a substantial fraction remains recoverable. We therefore argue that memory integrity, information-flow control, and covert-channel detection should be explicit security requirements for agentic systems.

177. 【2610.04580】Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs

链接:https://arxiv.org/abs/2610.04580

作者:Hoigi Seo,Byung Hyun Lee,Minjun Kim,Dohyun Mah,Jongho Lee,Se Young Chun

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Multi-modal large language, large language models, large language model, large language, Multi-modal large

备注:

点击查看摘要

Abstract:Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (\textit{e.g.}, audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer. Project page with code is available at \url{this https URL}

178. 【2610.04575】From Probe Scores to Alarm Policies: Operational Validity of Activation Monitors for Language-Model Agents

链接:https://arxiv.org/abs/2610.04575

作者:Xueping Gao

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:tight false-alarm budgets, predict safety-relevant properties, make thresholded alarm, thresholded alarm decisions, deployed agent monitors

备注: 12 pages, 1 figure. Accepted at the Eighth International Conference on Distributed Artificial Intelligence (DAI 2026)

点击查看摘要

Abstract:Activation probes can predict safety-relevant properties of language models with high area under the receiver-operating-characteristic curve (AUROC), but deployed agent monitors make thresholded alarm decisions under tight false-alarm budgets. These are different estimands. We introduce an Operational Validity Contract that fixes a monitor's target, observability, identity, timing, intervention unit, comparator, calibration, and cost. We formalize risk at the semantic request or trajectory level: when one task contains repeated alarm opportunities, row-level AUROC and false positive rate do not identify semantic-unit any-alarm risk. A confidence-certified threshold also requires enough independent negative units, a tie-safe rule, and transport to deployment. Across Models Under Pressure, LASR refusal prediction, immutable AgentDojo, and a prospectively protocol-frozen ST-WebAgentBench replication, joint activation-observable monitors reach AUROC 0.957 and 0.935 on the first two benchmarks, yet their locked 5%/10% detection rates are only .642/.742 and .719/.782, respectively. On AgentDojo, the secondary mean-activation rollout monitor reaches AUROC 0.922 but detects none of 38 positive semantic cases at the locked 5% operating point; thresholds intended for 10% false alarms realize 18.7-20.0% on test. Because the test misses its prospectively frozen 40-positive support gate, we label it support-insufficient. On ST-WebAgentBench, activation reaches AUROC .874, but 23 independent calibration negatives cannot identify even a 10% controller; the locked policy abstains rather than reporting its mechanical zero FPR as a success. An exploratory counterexample also lowers full AUROC while improving realized 10% utility. The fail-closed compiler caps MUP and LASR at restricted predictive value and AgentDojo and ST-Web at representation accessibility; no setting reaches alarm-policy validity.

179. 【2610.04541】Autonomous Structuring of Radiology Reports Across Modalities at Archive Scale Using an Open-Weight Large Language Model

链接:https://arxiv.org/abs/2610.04541

作者:Friedrich Puttkammer,Fabian Drexel,Marlene Fritzsche,Era Stambollxhiu,Miriam Kumpf,Lena Schmitzer,Lea Schumann,Lina Xu,Johannes Moll,Jannik Lübberstedt,Zeineb Ben Chaaben,Anirudh Narayanan,Hartmut Häntze,Renato Cuocolo,Antonios Billis,Alexander Löser,Jawed Nawabi,Marcus R. Makowski,Cosmin I. Bercea,Shahrooz Faghihroohi,Lisa C. Adams,Keno K. Bressem

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language model, open-weight large language, free-text radiology reports, reports, develop and evaluate

备注: 27 pages, 9 figures

点击查看摘要

Abstract:Purpose: To develop and evaluate an open-weight large language model (LLM) pipeline that converts an entire archive of free-text radiology reports into structured reports without human oversight. Materials and Methods: In this retrospective study, a pipeline with 150 hierarchically organized templates was developed at one center and tested at a second center on reports from 2010 to 2025. The open-weight model gpt-oss-120B selects the template in three constrained-decoding steps and fills it on one local graphics processing unit. Template selection was scored against expert labels on 914 randomly sampled reports of five modalities, structuring quality on 920 radiography and CT reports corrected field by field by five residents. The pipeline then processed the complete archive of the second center. Proportions are reported with Wilson 95% confidence intervals (CIs). Results: An optimal template set was selected for 74.4% of reports (680 of 914; 95% CI: 71.5%, 77.1%) and an appropriate set for 82.3% (752 of 914; 95% CI: 79.7%, 84.6%), 87.7% for single-region and 54.1% for multi-region reports. Macro semantic textual similarity between output and corrected reference was 0.95 for radiography and 0.97 for CT, residents left 88.7% of 24,638 fields unchanged, and unsupported content was flagged in 1.0% and 1.5% of reports. Of 2,186,982 archive reports, 96.5% received structured output, 2,401,544 structured reports, at 1,258 reports per hour on one graphics processing unit. Conclusion: An open-weight LLM pipeline structured a complete multimodality report archive without human oversight with high content fidelity. Multi-region reports remained the main source of template errors.

180. 【2610.04512】Correctness Is a Direction: Geometric Answer Selection in Language Models

链接:https://arxiv.org/abs/2610.04512

作者:Marcus Armstrong,Navid Ayoobi,Pradham Mummaleti,Alexander Chulzhanov,Arjun Mukherjee

类目:Computation and Language (cs.CL)

关键词:recoverable geometric direction, recoverable geometric, hidden states, percentage points, correct answer

备注:

点击查看摘要

Abstract:Answer correctness is encoded as a recoverable geometric direction in the hidden states of language models. We show that the mean displacement from incorrect to correct answer representations, computed at approximately 70\% of model depth from fifty labeled examples with no parameter updates, yields a scoring direction that outperforms zero-shot log-probability scoring by up to +32.0 percentage points on factual benchmarks (ARC-Challenge and MMLU) and by +38.1 to +51.8 percentage points on TruthfulQA, across five models spanning 1B to 8B parameters in three architecture families (Llama, Qwen, Gemma). The method requires one forward pass and one dot product per candidate; no generation is performed at inference. Applied as a hallucination detector on individual (question, answer) pairs, the recovered direction achieves 0.693~AUROC versus 0.578 for log-probability scoring. We additionally find that correctness directions for factual reasoning, domain knowledge, and calibrated truthfulness are near-orthogonal in representation space, revealing that language models allocate geometrically independent subspaces to qualitatively distinct notions of correct answer, with architecture-dependent variation in the degree of separation. This structure explains the observed transfer pattern---the direction calibrated on factual questions transfers within task type but not across it---and suggests that LLM calibration failures may reflect a routing problem: the model's internal representation contains more correctness signal than its output behaviour exploits.

181. 【2610.04504】he Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security

链接:https://arxiv.org/abs/2610.04504

作者:YaJie Yin

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:Verification Autonomy Levels, apply Verification Autonomy, prompt hardening, VAL stack holds, produced a dense

备注: v1: applies Verification Autonomy Levels (VAL) to 22 agent-security defenses; first controlled deployment-value comparison (VAL-guided vs intuition stack) across three testbeds; ~17k LLM calls; an honest out-of-ODD boundary is reported. Writing was assisted by an AI language model; all experiments and research decisions are the author's own

点击查看摘要

Abstract:LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels (VAL) - L0: LLM self-declaration; L1: deterministic rules; L2: objective ground truth; L3/L4: decidable completeness; L5: impossible - to 22 agent-security defenses; the taxonomy is falsifiable (10/10 prediction hits on frozen cards, flagged). We run the first controlled deployment-value comparison: at equal budget, a VAL-guided stack (confirmation gate + schema sandbox) versus a mainstream intuition stack (prompt hardening + keyword filter), 50 scenarios, 12 attack variants, adaptive/white-box/PAIR escalation (~7,000 testbed calls; ~10,000 harness calls on AgentDojo/JADE). The VAL stack holds 0.000 attack success at 1.000 benign success (0.5% ASR at 79.7% utility on AgentDojo banking vs 4.3% undefended); the intuition stack reaches 0.000 ASR but kills all benign actions - security by model-behavior luck, not structure. Across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0-1.9%-6.2%, n=16 on JADE) while the VAL stack's holds within its ODD (0-0-0), its only breach a disclosed out-of-ODD password gap (0.5%). The same zero, two different guarantees: zero is an outcome, not a guarantee.

182. 【2610.04499】Homogeneous Semantic Alignment and Hierarchical Expert Routing for Radiology Report Generation

链接:https://arxiv.org/abs/2610.04499

作者:Erjian Zhang,Jiayuan Ma,Liejun Wang,Yikemaiti Sataer,Xiaoming Tao,Zhiqing Guo

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Radiology report generation, convert medical images, Radiology report, Hierarchical Expert Routing, aims to convert

备注:

点击查看摘要

Abstract:Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of underlying representations and the undifferentiated rigid coupling of heterogeneous information cause weak visual abnormality cues to be easily diluted by massive text priors and generation inertia during decoding. To overcome this bottleneck, inspired by cognitive science, we propose a novel two-stage Homogeneous Semantic Alignment and Hierarchical Expert Routing (HSA-HER) framework. First, the model introduces an explicit homogeneous distribution constraint in the underlying latent space to effectively eliminate the cross-modal distribution shift between visual and textual features, thereby extracting purified visual features as semantic anchors that accurately align with diseases. Second, for heterogeneous clinical evidence composed of visual features, local entities, and global retrievals, we design a hierarchical expert routing mechanism guided by these disease semantic anchors. This mechanism abandons the undifferentiated rigid coupling paradigm. Specifically, it dynamically activates expert networks to perform targeted mining and semantic reconstruction on multi-source evidence, and adaptively allocates fusion weights. Extensive experiments on three mainstream benchmark datasets demonstrate that HSA-HER achieves state-of-the-art performance, accurately depicting complex imaging details and key diagnostic information.

183. 【2610.04495】Emoji-Emotion Ranking System Using Twitter Data

链接:https://arxiv.org/abs/2610.04495

作者:Danila Khlebokazov,Nurkhan Tashimov,Pakizar Shamoi

类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:replacing words, Nowadays, Abstract, words, emojis

备注: Has been submitted to IEEE

点击查看摘要

Abstract:Nowadays, emojis are often replacing words. Yet computational systems still oversimplify them. Most existing approaches treat emojis as static sentiment indicators and overlook their emotional distributions. In this study, we propose an emoji-aware emotion analysis framework based on a Twitter (X) dataset of 100,000 emoji-containing replies collected between 2020-2025. After preprocessing and text cleaning, we applied text-to-emotion classification to detect five primary emotions (Happy, Angry, Sad, Fear, and Surprise) for each message. By aggregating emotion scores across contexts in which each emoji appears, we estimate emoji-emotion association distributions and construct an emoji-emotion ranking system reflecting relative emotional dominance. Furthermore, we project emojis into the Russell valence-arousal space to enable continuous affective interpretation. Our results demonstrate that emojis exhibit probabilistic, context-sensitive emotional profiles rather than fixed sentiment polarities.

184. 【2610.04489】DV-Lens: Revealing the Functional Organization of Language Model Parameters

链接:https://arxiv.org/abs/2610.04489

作者:Chenhang Cui,Jian Yu,Shuyi Miao,Xiaohao Liu,Rui Huang,Fei Shen,An Zhang,Tat-Seng Chua

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Understanding parameter functions, Understanding parameter, functions helps elucidate, elucidate the internal, internal mechanisms

备注:

点击查看摘要

Abstract:Understanding parameter functions helps elucidate the internal mechanisms of large language models (LLMs). However, how to connect parameters from different modules to verifiable output effects and further characterize the relationship between their functional organization and model capability remains to be explored. To this end, we introduce the downstream vocabulary lens (DV-Lens), a parameter-level interpretability framework that links native parameter directions to their downstream vocabulary responses. Specifically, we first estimate module-specific downstream Jacobians over a reference prompt set for attention query, key, value, and output (Q/K/V/O) projections and feed-forward networks (FFNs). Second, we use these mappings to project native parameter columns into the final vocabulary space, obtaining signed readouts that characterize their average local output responses. Third, we group parameter columns by their vocabulary readouts and introduce downstream vocabulary complexity (DV-Complexity), which quantifies within-group structural variation using normalized reconstruction residuals of the original weights. At the parameter level, randomized controls and finite-difference tests show that DV-Lens readouts capture non-random vocabulary structure and predict local logit changes with 98.0% coordinate-orientation agreement across 720 cases from nine models. These readouts further guide parameter ablation, steering, and swapping across 21 models, shifting target-token probabilities in the predicted directions under controlled conditions. At the model level, the joint-parameter score of DV-Complexity achieves a Spearman correlation of 0.904 with benchmark-based capability rankings across 48 language models. Together, these results provide intervention-based evidence for DV-Lens interpretations and reveal an association between DV-Complexity and model capability.

185. 【2610.04412】Large Language Models and Augmented Democracy

链接:https://arxiv.org/abs/2610.04412

作者:Jairo Gudiño-Rosero

类目:Computers and Society (cs.CY); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:Artificial intelligence enables, intelligence enables computational, Large Language Models, Artificial intelligence, enables computational agents

备注: PhD thesis, Center for Collective Learning (Toulouse School of Economics), 2026. 149 pages, 28 figures

点击查看摘要

Abstract:Artificial intelligence enables computational agents to represent political preferences and take part in collective decision-making. In this thesis, I investigate the opportunities and challenges of digital twins (DTs) based on Large Language Models (LLMs) as intermediaries in augmented democracy, focusing on individual preference representation, collective representation of political organizations, and the vulnerability of those representations to attackers. First, using data from an online experiment in Brazil, I examine whether personalized DTs can predict citizens' preferences for unseen policy proposals. Second, I extend the DT framework from individuals to political organizations. Using Swiss parliamentary data, I build topic-specific knowledge graphs from lawmakers' legislative records and connect them to LLM-based lawmaker agents, which are organized into party-level DTs representing collective positions. Agentic deliberation among these agents tests whether aggregated party representations capture a broader range of intra-party perspectives than official party communications. Finally, I study the vulnerability and robustness of LLM-mediated deliberation against prompt-injection attacks that amplify viewpoints, suppress opinions, or redirect consensus. Using data from a 2023 deliberative experiment in the United Kingdom, I analyze how attack effectiveness varies with the distribution of opinions and rhetorical strategies, and evaluate a pipeline combining injection detection, structured opinion representations, and reinforcement learning to improve resistance. These findings characterize the opportunities and challenges of LLM-based digital twins in augmented democracy, stressing accurate preference representation, faithful aggregation, and robustness to strategic interaction.

186. 【2610.04409】Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents

链接:https://arxiv.org/abs/2610.04409

作者:Peigui Qi,Kunsheng Tang,Yide Song,Weiming Zhang,Nenghai Yu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, Large language, invoke external tools, Hallucination Escape, increasingly serve

备注:

点击查看摘要

Abstract:Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial improvements, yet we identify a previously overlooked failure mode that we term Hallucination Escape. These methods reduce hallucination on the tool configuration they are tuned on but increase it on other configurations, canceling out the gain. We further investigate this phenomenon and find that hallucination rises sharply when a model's intrinsic tool-use tendencies conflict with the current tool configuration, and that existing methods reinforce rather than suppress these tendencies, which in turn contributes to hallucination escape. Building on these findings, we propose EscapeGuard, a training-free inference-time method that combines conflict-aware gating with configuration-derived attention enhancement to mitigate tool hallucination while preventing hallucination escape. Across six benchmarks on various models, EscapeGuard reduces tool-selection hallucination by 9.0 pp and suppresses hallucination escape, lowering the cross-configuration mean by 23.7 pp and achieving an 89.1% net improvement in paired-query evaluation. We hope this work can encourage evaluation beyond a single tool configuration and pave the way for more reliable tool-using LLM agents.

187. 【2610.04403】Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define

链接:https://arxiv.org/abs/2610.04403

作者:Saurabh Kumar Singh,Yogeshwar Singh Dadwhal,Malhar Vedak

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:GGUF format mixed-precision, format mixed-precision K-quants, reach consumer hardware, fine-grained lexical competence, Post-training quantization

备注: 41 pages, 15 figures, 1 table. Data, answer key, scored outputs and scorer code: [this https URL](https://github.com/NietzscheDostoevsky/snk-release)

点击查看摘要

Abstract:Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, evaluated down their published ladder to Q2_K (about 2.6 bits per weight), on 429 frequency-validated rare English words under two probes: surface inclusion of a prompt-supplied word and its one-sentence definition, scored by a tiered multi-synonym matcher, its error measured by a blind LLM-judge census of every definition, with human verification. Three regimes emerge at Q2: total collapse into unusable builds, severe semantic dissociation in sub-2B models, and mostly robust preservation above about 3B. In every sub-2B artifact, definitions fall 20-67% below the artifact's baseline, typically several times the inclusion loss. Two controls separate rarity from task difficulty: within the rare set, loss rises with rarity in six of seven sub-2B artifacts, and on a 100-word common-word set rare words lose more than common words in all eight, significantly in six. Tokenizer vocabulary size does not predict the damage (Spearman rho=0.12); parameter count dominates (rho=0.72), confirmed within five of six same-tokenizer families. Q4_K_M remains lexically clean at =1B. The damage is frequency-graded, provider-dependent, and not calibrated by WikiText-2 perplexity: across nine artifact-matched ladders, near-identical Q2 penalties (44.7%/47.6%) separate an artifact keeping its definitions (3.6%) from one losing them (43.6%). Aggressively quantized small models can keep generating fluent text while no longer knowing what it means, risking hardware-constrained deployments in domains where semantics carries consequences. Validation must be per artifact.

188. 【2610.04401】Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media

链接:https://arxiv.org/abs/2610.04401

作者:Safaruzzaman Shovo,Monowar Islam,Asif Hossain,Sameya Akhter,Md. Shamsul Islam

类目:Computation and Language (cs.CL)

关键词:Public vs. private, debatable issue, private universities, zero-shot Llama, Public

备注: Accepted for publication at the 2026 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), 13-14 August 2026, Bangladesh Army University of Engineering Technology (BAUET), Qadirabad, Natore, Bangladesh

点击查看摘要

Abstract:Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduces a manually annotated dataset of 4,060 Bengali comments labeled as Pro-Public, Pro-Private, or Neutral. We evaluated the quality of our annotations by Fleiss's Kappa agreement that was 0.89, corresponding to a high agreement among annotators. The classical ML (SVM, Random Forest, XGBoost), BiLSTM network, hybrid BanglaBERT+XGBoost models and the state-of-the-art zero-shot LLMs (Claude Sonnet 4, DeepSeek-V3.1, Llama 4 Maverick, Kimi K2 Thinking, Qwen3-235B Thinking) models are evaluated. The accuracy of BanglaBERT+XGBoost is 91.81% and macro F1 score is 91.70%, which is higher than all the supervised baselines. The zero-shot Llama 4 Maverick Thinking achieves a macro F1 of 0.931 (overall accuracy of 93.31%) without any fine-tuning. All machine learning (ML), deep machine learning (DL) and transformer models were outperformed by the zero-shot Llama 4 Maverick model. Polarization also is evident, in some ways more clearly in the Pro-Private comments, which emphasize modern facilities and timely graduation, versus the Pro-Public comments, which emphasize affordability and government jobs. Our findings open new directions for analyzing social media polarization in low-resource languages.

189. 【2610.04400】XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions

链接:https://arxiv.org/abs/2610.04400

作者:Zhanxun Liu,Yifan Duan,Hengtao Wu,Chen Yang,Qinyuan Cheng,Kun Wang,Xingyu Zeng,Xipeng Qiu,Chaochao Lu,Xie Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)

关键词:General turn-taking behavior, systems requires deciding, General turn-taking, real-time dialogue systems, dialogue systems requires

备注:

点击查看摘要

Abstract:General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them difficult to use as a unified causal controller with comprehensive context. We propose XTurnix, a compact text-based model that formulates turn control as two binary decisions conditioned on the AI's current listening or speaking state and predicts a single control token from the complete dialogue history. XTurnix is pretrained on 5.5 million causal action examples automatically derived from timestamped two-speaker transcripts, then fine-tuned on synthetic multi-turn examples with a flatter distribution across the four state-action labels. We evaluate XTurnix on four public benchmarks and a balanced self-curated benchmark. Across the public benchmarks, XTurnix achieves the best results on all SemanticVAD and LiveKit splits, ties the native Smart-Turn model on Smart-Turn Bench, and achieves the highest incomplete-turn accuracy on Easy-Turn. On the self-curated benchmark, it reaches 89.06% accuracy, more than 20 percentage points above the strongest third-party baseline at 68.75%, while maintaining F1 scores between 84.21% and 90.63% across all four categories. These results demonstrate unified listening- and speaking-state turn control in a single compact model. Code is available at this https URL, with an interactive demo at this https URL.

190. 【2610.04399】GlitchPatch: Repairing Glitch Tokens in Frozen Language Models via Local Retokenization

链接:https://arxiv.org/abs/2610.04399

作者:Kunsheng Tang,Peigui Qi,Yide Song,Peijun Huang,Weiming Zhang,Nenghai Yu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:anomalous vocabulary entries, produce outputs inconsistent, Glitch tokens, anomalous vocabulary, vocabulary entries

备注:

点击查看摘要

Abstract:Glitch tokens are anomalous vocabulary entries that can cause large language models (LLMs) to produce outputs inconsistent with their inputs. Existing repair methods require access to model internals, making them impractical for frozen checkpoints. We investigate whether glitch tokens can be repaired outside the model by optimizing the input tokenization. An empirical study on BPE merge-rule deletion reveals that (1)deleting a glitch token's merge rule can fix a substantial fraction of failures, yet disrupting normal tokens sharing intermediate merge nodes causes the overall glitch rate to rise, and (2)different decomposition granularities yield non-monotonic fix rates while collateral damage on normal tokens grows monotonically. Motivated by these findings, we propose GlitchPatch, a repair framework for frozen language models based on local retokenization, consisting of two stages: the offline stage uses Behavioral Path Optimization (BPO) to find the behaviorally optimal replacement token sequence for each glitch token and compiles validated replacements into a rule table; the online stage substitutes only the IDs of matched glitch tokens in the canonical token sequence, with no modification to model parameters or internal states. Experiments on ten models spanning six tokenizer families show that GlitchPatch achieves an 85.10% mean fix rate, outperforming the strongest baseline by 14.37 percentage points, and reduces the average glitch rate from 14.88% to 2.27%. GlitchPatch achieves a 0.00% RR in full-vocabulary evaluation and leaves rule-unmatched inputs unchanged by design. We further evaluate the practical impact of repair from the perspectives of time cost, language understanding, and capability, supporting its deployment feasibility. We hope this work provides a practical option for improving tokenizer reliability.

191. 【2610.04370】Boundaries Agree, Labels Do Not: Intra-Annotator Dynamics as a Kind of Training Data

链接:https://arxiv.org/abs/2610.04370

作者:Marharyta Shvets

类目:Computation and Language (cs.CL)

关键词:training language models, training language, training data, Data quality, Data

备注: 8 pages, 1 figure, 5 tables. Code and data: [this https URL](https://github.com/Malificenta883/intra-annotator-dynamics)

点击查看摘要

Abstract:Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combines annotators into a "ground truth" and measures how well they agree with each other; the other treats their disagreement as a signal. Both compare different people at one point in time. We measure something else: how well one reader reproduces their own reading of the same text over time. One expert human reader and three LLM families segmented three Sumerian myths and labelled the causal function of each segment with one of seven states. Across runs months apart, the human cut the text in much the same places but named the segments differently, in every myth. The models show no such consistent pattern: their gap between the two layers is positive in some myths and negative in others, and its size varies. The human's label changes are not random: the runs go through much the same functions but start them one step apart, while model runs start them at the same places. We argue that this pattern is a usable measure of data quality and a contamination check: a "human" annotation whose labels are as stable as its boundaries, and whose functions start in sync, looks like a model's.

192. 【2610.04347】Evidence and Intervention: A Coupled Active-Inference Extension of Rational Speech Act Models

链接:https://arxiv.org/abs/2610.04347

作者:Yonghyeon Gwon,Elliot Murphy,Chun Kee Chung

类目:Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:Identical utterance choices, identical interpretations, Rational Speech Act, Identical utterance, choices can arise

备注: 83 pages, 3 figures, 2 tables

点击查看摘要

Abstract:Identical utterance choices can arise from different communicative causes, and identical interpretations can leave different traces in what a listener learns. Rational Speech Act (RSA) models treat interpretation as inference over speaker meaning, but standard one-shot RSA does not intrinsically distinguish these causal update targets. We develop a coupled active-inference model of dialogue in which a listener's likelihood for an utterance is the policy distribution attributed to the speaker, placing the speaker's expected free energy within the listener's variational free energy. Each utterance is therefore both evidence about and an intervention on a partner. The model separates four updates that RSA approaches typically collapse: inference about the partner's current pragmatic state; learning of partner-specific parameters; prospective evaluation of clarification or repair; and a policy prior shaped by habit and a context-sensitive price of time. Under one-step, exact-inference restrictions, the model recovers the RSA speaker and listener, with RSA as the restricted single-turn limit of the coupled process. Outside these restrictions, the updates obey distinct rules and timescales, predicting which adaptations persist, remain partner-specific, or transfer. Worked examples show audience design reversing after clarification, self-confirming misunderstanding in which both interlocutors have low free energy while disagreeing about reference, and rational closing before uncertainty is resolved. Consistent with critiques of equating natural language with communication, the model treats communication as a downstream use of linguistic structure and recasts production and interpretation as coupled inference: speaking is both an intervention on a partner and an epistemic action that samples evidence for the speaker's model of that partner.

193. 【2610.04344】Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry

链接:https://arxiv.org/abs/2610.04344

作者:Qi Yu,Ruizhong Qiu,Zhichen Zeng,Xuying Ning,Yanjun Zhao,Dongqi Fu,Yinglong Xia,Hong Li,Hanghang Tong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:group-based RLVR methods, large language models, Reinforcement learning, RLVR methods, group-based RLVR

备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works refine credit assignment of GRPO based on local signals such as token locations or entropy, they often fail to capture the global semantic novelty of a reasoning behavior relative to the current policy. In this work, we propose a hierarchical credit assignment approach for group-based RLVR methods, called HarA, which identifies and encourages semantically novel reasoning behaviors during RLVR. HarA represents each sampled rollout as a distribution over the hidden states and locations of tokens, and computes the Fused Gromov-Wasserstein (FGW) barycenters of all rollouts with the same outcome, capturing the internal reasoning patterns in the latent space under the current policy. The semantic novelty of a reasoning element can then be measured by its contribution to the FGW distance between the current rollout and the barycenter. While solving the FGW formulation is expensive, we introduce an anchor-guided linearization that turns it into a Wasserstein formulation solvable via the Sinkhorn algorithm efficiently. By reweighing token-level advantage of group-based RLVR methods based on the novelty signals, HarA highlights novel reasoning behaviors at flexible granularities to encourage fine-grained LLM exploration. Extensive experiments across three group-based RLVR methods show that our plug-and-play method effectively enhances the exploration of LLMs, outperforming existing methods across diverse reasoning benchmarks.

194. 【2610.04339】ShadowMiner v1 - An Experience Report on Implementing and Measuring a Problem-and-Hypothesis Discovery Engine

链接:https://arxiv.org/abs/2610.04339

作者:Jinhyuk Choi

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:automatically discovers research, discovers research problems, system that automatically, automatically discovers, problems and generates

备注:

点击查看摘要

Abstract:ShadowMiner v1 is a system that automatically discovers research problems and generates hypotheses from AI papers. It is a nine-stage pipeline. It structures documents into a knowledge graph and finds graph gaps in it - structural blind spots in research. These graph gaps are included in the LLM generation prompt. Each generated hypothesis is then verified by checking whether it is already covered by existing research, scoring its quality, and checking that the facts it relies on are accurately drawn from its sources. This report does not propose a new generation or evaluation technique. It describes our experience of implementing and applying ideas from prior work, and measuring whether each one actually contributed.

195. 【2610.04329】Suppressing Pressure, Amplifying Evidence: Self-Guided Attention Steering to Mitigate Sycophancy and Stubbornness

链接:https://arxiv.org/abs/2610.04329

作者:Yinghao He,Mengyu Xu,Haixiang Sun,Donghan Li,Yibo Wang,Lixu Wang,Kezhen Chen,Chi Li,Chunwei Liu,Bharat Bhargava,Chongyang Gao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Reliable language models, Reliable language, unsupported user pressure, resist unsupported user, contextual information

备注:

点击查看摘要

Abstract:Reliable language models should resist unsupported user pressure while effectively using objective contextual information. However, models may exhibit sycophancy by yielding to unsupported user pressure or contextual stubbornness by failing to update their answers when relevant contextual information warrants revision. Evaluating interventions for these failures separately can obscure whether mitigating one failure exacerbates the other. To assess this trade-off, we introduce CoPE-Bench with six conditions per question: a neutral baseline, correct or incorrect user pressure, contextual information consistent with or conflicting with the neutral answer, and a joint condition combining incorrect claims with conflicting contextual information. To regulate the influence of user pressure and contextual information, we propose SPAE (Suppressing Pressure, Amplifying Evidence), a training-free framework that uses the model's own judgments to identify relevant tokens, suppressing user pressure and amplifying contextual information through token-level attention steering. On average across five backbones, SPAE reduces pressure following by 18.8 percentage points and increases joint-condition updating by 5.5 percentage points relative to the strongest baseline in the main comparison. In two-turn dialogue, it improves joint-condition updating by an average of 13.2 percentage points over the strongest prompting baseline. The source data and codes can be found at this https URL.

196. 【2610.04328】Bidirectional Preference Synthesis: Learning Prompt-Conditioned Preferences from Boundary Failures

链接:https://arxiv.org/abs/2610.04328

作者:Junbo Wang(1 and 2),Lidong Lu(2),Zhuoqun Li(1),Guiping Jiang(1),Xiangyu Wu(1),Tinghai Zhang(1),Tong Lu(2) ((1) Kuaishou Technology, (2) Nanjing University)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Correction-based offline preference, pipelines commonly treat, offline preference pipelines, preference pipelines commonly, Correction-based offline

备注: 18 pages, 3 figures

点击查看摘要

Abstract:Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby intent or constraint setting. We introduce Bidirectional Preference Synthesis (BPS), a data-construction method for standard Direct Preference Optimization (DPO) that makes this missing prompt dependence explicit. For each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, so the same response is rejected where it is wrong and chosen where it is right, without changing the DPO objective, training a reward model, or requiring online sampling. On Qwen3-4B-Instruct-2507, BPS preserves original-side pairwise ranking while raising achieved-side ranking accuracy from 6.8% to 62.3% on held-out crossed anchors, with a similar shift under a Kimi-K2.6 cross-teacher probe. A blind human audit supports the intended reverse preference direction, and downstream evaluations show the clearest separation from Forward-DPO in multilingual multi-turn instruction following, with consistent capability-retention patterns on agentic, tool-use, and code checks.

197. 【2610.04316】From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility

链接:https://arxiv.org/abs/2610.04316

作者:Mohammad Mosafer

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Output-only safety monitoring, Output-only safety, model input-output Jacobian, computes its answer, answer before emitting

备注: 26 pages

点击查看摘要

Abstract:Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the output vocabulary via the model's input-output Jacobian, have been proposed as a window on behaviorally accessible internal content, and a first safety protocol (JADR) showed that danger recognition is readable there. What remains unknown is how this accessible content relates to the latent-space safety geometry studied by representation engineering, and how safety training shapes it. We introduce two quantitative bridges: (i) a transport-amplification profile $A_\ell$, measuring how strongly a latent safety direction is carried toward output space by each layer's Jacobian, and (ii) a paired base-vs-tuned protocol that attributes accessibility to training provenance. Across four small model pairs (135M to 1.5B, three lens-fit seeds), tuned checkpoints concentrate recognition in upper-middle layers, and at 0.5B DPO installs refusal while J-space recognition drops to chance: training can widen the accessibility gap exactly where behavioral safety looks best. A deployment audit closes the loop: monitor-aware GCG-style suffixes suppress the prompt-side monitor at zero behavioral cost at every scale; only a learned monitor rung resists its own adaptive re-attack; the training-time defense fails its re-attack at every penalty weight; steering shows the amplification profile is descriptive, not causal; and continuous-prefix optimization fails where discrete search succeeds. Safety monitoring, training, and evaluation must operate on accessibility itself, not on outputs alone.

198. 【2610.04304】Evaluating Modeling Approaches for Experience-Level Classification in Job Description

链接:https://arxiv.org/abs/2610.04304

作者:Celia Liang,Eddie Wu,Shiqi Wang,Yonah You

类目:Computation and Language (cs.CL)

关键词:aiming to automatically, paper investigates, automatically identify, identify the qualifications, qualifications required

备注: 12 pages, 4 figures, 9 tables

点击查看摘要

Abstract:This paper investigates the task of predicting job experience levels in recruitment texts, aiming to automatically identify the qualifications required for positions. Unlike traditional text classification, recruitment texts typically possess explicit internal structures, with different paragraphs playing disproportionate roles in conveying experience clues. To address this, we propose a structure-aware Section-Aware BERT approach that segments and encodes key paragraphs (titles, responsibilities, requirements) for integrated modeling, building upon rule-based systems and classical baselines TF-IDF. Simultaneously, we evaluate large language models under both few-shot and fine-tuning settings on the same dataset to compare the capability boundaries of different modeling paradigms. Experimental results demonstrate that explicitly leveraging text structure significantly improves experience level prediction performance, particularly in scenarios with ambiguous job titles. Further error analysis reveals systemic challenges in this task, including confusion between Entry and Senior levels and the blurred boundaries of Mid-level positions. This research provides an effective modeling approach and analytical framework for understanding structured recruitment texts.

199. 【2610.04299】Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

链接:https://arxiv.org/abs/2610.04299

作者:Jinyuan Li,Chengsong Huang,Langlin Huang,Donghong Cai,Shiping Gao,Yuyi Yang,Jiaxin Huang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Self-evolving reasoning models, Self-evolving reasoning, Self-evolving, questions, invalid questions

备注:

点击查看摘要

Abstract:Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filtering further increases their proportion in training data. Second, existing question diversity controls based on lexical similarity can miss mathematically equivalent questions expressed in different ways, which leads to question diversity collapse in later training rounds. Building on these findings, we introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. We first train the solver to recognize and reject invalid questions, then use its judgments to guide questioner rewards and filter solver training data. To avoid question repetition, we use a frozen base model to compare sampled question pairs and provide novelty feedback. Empirically, our method consistently achieves the highest average performance on 12 benchmarks in mathematical reasoning, general-domain reasoning, and code generation across two model families. Additionally, R-Quest maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.

200. 【2610.04295】Language-Conditioned Token and Reasoning Efficiency in Large Language Models: A Paired Cross-Lingual Study Protocol

链接:https://arxiv.org/abs/2610.04295

作者:Genliang Zhu,Chu Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:incur language-dependent representation, Large language models, models incur language-dependent, Large language, conflate input language

备注: 32 pages, 2 figures, 3 tables. Prospective paired study protocol; no confirmatory model outcomes are reported

点击查看摘要

Abstract:Large language models incur language-dependent representation and inference costs, but existing comparisons often conflate input language, assigned observable-trace language, and answer realization. We specify a prospective paired study that separates these interfaces while holding the semantic item, checkpoint, and answer oracle fixed. The initial design instantiates 240 exactly scored items rendered from templates in English and seven non-English languages, three distinct-lineage open-weight checkpoints, three trace-token budgets, 22 input- and trace-language conditions, a fixed answer reserve, and a separately counted delimiter: 47,520 initial core runs before prospective sample-size selection. RQ1-RQ3 estimate input and trace effects by intention-to-treat with failure-inclusive terminal accounting and test answer realization by cloning a sealed prefix and runtime-native KV state into eight crossed branches. Pre-freeze independent language review, fixed-form ASCII selectors, and code/surface/solver agreement constrain the realization test. H1-H5 share one Holm family and a global simultaneous component band. A secondary randomized experiment compares one-long-attempt and complete K-short-attempt policies at equal trace allowance under frozen seeds and oracle-blind aggregation; it is a full-policy contrast because answer capacity differs. Outcomes include exact token spans, correctness, latency, runtime-exposed memory, and qualified same-host operating-system-reported energy over prespecified hardware rails. The protocol separates tokenizer expansion, observable-trace cost, and answer-realization cost without treating visible traces as internal cognition or operating-system estimates as physical cross-device energy. No confirmatory model outcome is reported; result fields remain disabled until the frozen evidence ledger passes independent verification.

201. 【2610.04283】First-Order Steering: Translating Weight Adaptation into Activation Steering

链接:https://arxiv.org/abs/2610.04283

作者:Sri Pranav Kunda,Alexander Kurz,Tomas Dominik,Uri Maoz

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:exploits interpretable directions, steering exploits interpretable, Activation steering, steering, Activation steering exploits

备注:

点击查看摘要

Abstract:Activation steering exploits interpretable directions in the residual stream to enable inference-time manipulation of model behavior. Composing steering vectors to apply multiple target behaviors simultaneously is important in various fields-including AI alignment and safety-but remains a challenge for existing activation steering methods. In contrast, prior work in model merging shows that target behaviors represented by learned weight adaptations can be combined with high accuracy. A method that translates weight adaptations into activation steering vectors could therefore extend prior work in model merging to generate composable steering vectors that better enable simultaneous inference-time behavioral control. For this, we introduce First-Order Steering, a formulation of activation steering as a first-order approximation of weight update matrices parameterized by a vector of steering strengths, and establish theoretical bounds on the approximation error of first-order steering. We then develop a novel model merging procedure, HeRD-Merging, which minimizes the first-order approximation error terms to enable higher first-order steering accuracy. Together, our method produces steering vectors that control both individual and composed behaviors more accurately than existing activation steering methods. Furthermore, HeRD-Merging matches the performance of conventional model-merging baselines, while producing weight adaptations that admit more accurate first-order steering vectors.

202. 【2610.04272】Rethinking Self-Distillation for Multi-Teacher Capability Merging

链接:https://arxiv.org/abs/2610.04272

作者:Roy Xie,Dan Friedman,Feng Nan,Yukun Huang,Zhichao Xu,Chengjiu Zhang,Jun Xu,Manaal Faruqui,Vivek Rathod,Bhuwan Dhingra

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:frontier language-model post-training, language-model post-training, Combining capabilities, base checkpoint, increasingly common

备注:

点击查看摘要

Abstract:Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training design choices and hyperparameter optimization, rather than the algorithm itself. We conduct a controlled self-distillation study across two multi-teacher settings, four models, and eleven benchmarks, comparing off-policy methods, namely supervised fine-tuning (SFT) and soft-label distillation, to hybrid teacher-prefix distillation and MOPD. We found that all four methods achieve \textit{nearly identical} accuracy. However, MOPD uses $14.8$--$23.1\times$ SFT's training GPU-hours. We also revisit four recently published comparisons between on-policy and off-policy distillation and find that the reported on-policy gains shrink substantially when SFT baselines are trained on rejection-sampled teacher trajectories and use independently tuned hyperparameters. As a training-free alternative, we also find that simple weight-merging methods can recover expert capabilities with minutes of CPU merging time, although their accuracy degrades as model size decreases and task interference increases. Overall, our results question recent gains reported due to MOPD and suggest careful tuning of more efficient off-policy baselines as a viable alternative.

203. 【2610.04267】AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

链接:https://arxiv.org/abs/2610.04267

作者:Steven Moore,Nicholas Diana

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:Generating multiple-choice questions, Generating multiple-choice, quality remains difficult, increasingly scalable, remains difficult

备注: 8 pages, 2 tables, Full paper accepted to the AIME Conference 2026

点击查看摘要

Abstract:Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.

204. 【2610.04263】Multimodal Dual-Encoder Retrieval for Automated ICD Coding

链接:https://arxiv.org/abs/2610.04263

作者:Abhinav Bohra,Anuj Bohra

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Accurate International Classification, Accurate International, International Classification, crucial for large-scale, ICD

备注: 6 pages, 2 figures, 1 table

点击查看摘要

Abstract:Accurate International Classification of Diseases (ICD) coding is crucial for large-scale clinical research, documentation, and billing. There are three primary problems with current ICD prediction methods: (1) They are unable to comprehend multimodal patient data because they rely on either structured EHR data or unstructured clinical notes. (2) They also struggle with scalability to a larger amount of ICD codes (9K+ codes in ICD-9), as traditional classifiers need dense output layers and often do not generalize well to long tail rare diseases. (3) They lack transparency for clinical use. To address these challenges, this research proposes a two-stage framework that first retrieves ICD codes using a multimodal dual-encoder retrieval model, where structured and unstructured patient data are integrated through gated fusion. The second stage refines the top-k retrieved candidates with an LLM-based re-ranker that provides ranked codes with clinically relevant explanations. Our experiments show that the proposed approach improves Micro-F1 and Precision over a multimodal dual-fusion classifier baseline. These improvements demonstrate that combining a gated multimodal retrieval system with LLM-based re-ranking is a practical alternative to dense multi-label classification for automated ICD coding.

205. 【2610.04261】Playing social deduction games with reinforcement fine-tuned large language models

链接:https://arxiv.org/abs/2610.04261

作者:Lingzhe Zhang,Yunpeng Zhai,Tong Jia,Kening Zheng,Chiming Duan,Minghua He,Zhaoyang Liu,Bolin Ding,Philip S. Yu,Ying Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)

关键词:large language models, language models, applications where large, large language, RFT

备注:

点击查看摘要

Abstract:Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state inference, social reading and vote steering. Our results show that LLM agents do not reliably acquire social-deduction ability by directly optimizing terminal win--loss outcomes, suggesting that final game results provide a sparse and noisy signal for socially interactive learning. However, RFT is particularly effective at improving social reading, including tasks that require agents to infer hidden roles from public discussion, update beliefs over time and predict other agents' future decisions. We further show that RFT can also improve social influence, including tasks that require agents to steer votes, team approvals and collective decisions, although these gains depend more strongly on behaviourally specific rewards and structured interaction settings. Finally, we show that LLMs' ability to play social deduction games can be further improved through multi-agent social-cognitive reinforcement fine-tuning, which combines social-reading and social-influence signals during same-side multi-agent training. These learned behaviours also receive more favourable human evaluations of strategic competence, persuasiveness and social usefulness. Together, these results enrich our understanding of how RFT changes LLMs' social behaviour and provide a step toward a behavioural learning theory for machine social intelligence.

206. 【2610.04239】Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

链接:https://arxiv.org/abs/2610.04239

作者:Jiayi Xin,Evan Qiang,Zihan Zhu,Xiang Li,Weijie J. Su,Qi Long

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:provide reliable measures, Uncertainty quantification, predictive confidence, large language models, aims to provide

备注: NeurIPS 2026 Poster

点击查看摘要

Abstract:Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework robust to semantic rewordings. Our approach trains a lightweight proxy model on LLM hidden states and aggregates its predictions across paraphrases to construct label-wise nonconformity scores. Under score exchangeability, conformal calibration retains marginal coverage. This guarantee can also hold under test-only rewording, provided that the paraphrase pipeline satisfies an additional distributional alignment condition. We evaluate three settings (normal, fully reworded, and semi-reworded) which apply rewording to neither dataset, both calibration and test datasets, or only the test dataset, respectively. Across seven multiple-choice QA benchmarks and multiple model families, our method produces compact prediction sets with empirical coverage generally near the nominal target, even in the semi-reworded setting. Ablation studies show that the learned proxy accounts for most of the reduction in set size, while paraphrase-augmented training and inference-time aggregation improve stability under rewording. Code is available at this https URL.

207. 【2610.04210】Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams

链接:https://arxiv.org/abs/2610.04210

作者:Sugam Panthi,Muhaiminul Yeamin,Rabab Abdelfattah

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, text, Large, message as plain

备注:

点击查看摘要

Abstract:Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the model treats a trailing user comment as part of the pasted text, returning it inside the edited text. This happens even though the user did not intend the comment to become part of that text. Existing instruction-data separation benchmarks tell the model which text is instruction and which is data, then test whether it obeys that separation. They do not test harmless user speech following an unmarked paste. We introduce SEAM, a controlled benchmark of 300 editing examples. Each example is tested under six matched conditions that vary how the boundary between pasted text and later user speech is expressed. Across 20 models, absorption at a bare newline ranges from 7.7% to 66.7%. Adding a blank line does not significantly reduce absorption in any model, while boundary markers reduce it in 19 of 20 models. Comments that fit the pasted text, such as a code comment typed after code, are absorbed significantly more often in 17 of 20 models. Models often fail to separate pasted material from later user speech, and explicit boundaries reduce but do not remove this failure.

208. 【2610.04204】Clean: Second-order LLM Training at Linear Memory Cost via Nyström Sketching

链接:https://arxiv.org/abs/2610.04204

作者:Beheshteh T. Rakhshan,Sahar Rajabi,Maziar Sargordi Shikai Fang,Guillaume Rabusseau,Sirisha Rambhatla

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Training large language, Adam discard cross-parameter, Training large, Adam discard, discard cross-parameter curvature

备注:

点击查看摘要

Abstract:Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer designed to resolve this bottleneck. Clean leverages the randomized Nystrom method to accurately approximate the left and right preconditioners in SOAP, and to reduce the optimizer's memory complexity from quadratic to linear in terms of model dimensions. We subsequently reintegrate the off-subspace components to capture curvature information beyond the low-rank approximation, preserving rich curvature at minimal memory cost. We further propose Q-Clean, a low-precision variant that aggressively compresses optimizer states. Q-Clean reduces optimizer memory consumption by \textbf{over 50\%} compared to Muon when pre-training a LLaMA-1.3B architecture, all while maintaining strong and competitive predictive performance. Notably, Clean operates with a smaller optimizer-state footprint than standard AdamW while reaching AdamW's final performance \textbf{26\% faster} in wall-clock time. Furthermore, our methods uniquely enable the pre-training of a 13B-parameter model on a single 80GB GPU, providing a scalable, efficient, and accessible approach to large-scale model optimization.

209. 【2610.04183】Language Model Activations Inhabit Privileged Error-Correcting Basins

链接:https://arxiv.org/abs/2610.04183

作者:Matthew Finlayson,Francisco Pernice,Eric Todd,Amir Zur,Daniel Wurgaft,Fenil R. Doshi,Vasudev Shyam,Matt Feiszli,Satchel Grant,Lucius Bushnaq,Tal Haklay,Usha Bhalla,Matthew Kowal,Thomas Fel,Jack Merullo,Atticus Geiger,Xiang Ren,Owen Lewis,Ekdeep Singh Lubana

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:exhibit remarkable robustness, produce coherent text, models exhibit remarkable, produce coherent, produce coherent outputs

备注:

点击查看摘要

Abstract:Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward "good" regions that produce coherent outputs. To investigate these hypothesized error-correcting mechanisms, we probe the geometry of language model activation space by observing the action of model layers on low-dimensional curves. In doing so, we discover that model activations occur within a cluster of distinct attracting basins, which differentiate natural activations geometrically from distributionally similar synthetic activations. Applying this lens to language model steering, we observe feature-specific basins along semantic steering directions, and find that steering moves activations between these basins. To demonstrate the active role of this geometry in neural computation, we show that adaptively modulating steering strength to transport activations across basins improves inter-language steering, significantly increasing the probability of sampling tokens from the target language compared to fixed-strength steering. Our findings establish analysis of activation space geometry as a promising approach to interpreting and controlling language models.

210. 【2610.04174】DimSteer: Steering LLM Authoring with Automatically Discovered Stylistic Controls

链接:https://arxiv.org/abs/2610.04174

作者:Ajit Mallavarapu,Ziwei Gu

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)

关键词:Large language model, language model writing, repeatedly articulating desired, Large language, model writing interfaces

备注: 24 pages, 5 figures

点击查看摘要

Abstract:Large language model writing interfaces often make users steer outputs by repeatedly articulating desired changes in natural language. Yet writers may recognize useful stylistic directions only after seeing alternatives, making revision recall-heavy. We present DimSteer, an authoring interface that samples prompt-local completions, discovers high-variance activation-space axes of variation, labels them, and exposes them as sliders with pole previews, diff comparison, and reset controls. Users can manipulate discovered dimensions, reducing the need to reformulate prompts for each stylistic adjustment. In a within-subjects study with 16 participants against a matched prompt-only baseline, DimSteer reduced mental demand, effort, and frustration while preserving comparable perceived success. Participants valued the surfaced dimensions, yet 15 of 16 disagreed that they would have thought to request the same changes in a prompt. Results suggest prompt-local controls can shift LLM authoring from recall-based prompting toward recognition-based exploration and direct manipulation, while preserving prompting for open-ended edits.

211. 【2610.04162】Principled Top-$k$ Selection for Language Models with Hybrid Gradients

链接:https://arxiv.org/abs/2610.04162

作者:Xuchen Gong,Junfei Sun,Tian Li

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:modern large language, Retrieval-Augmented Generation, large language model, language model systems, modern large

备注:

点击查看摘要

Abstract:Selecting the best $k$ items out of $m$ candidates is a critical component of modern large language model systems, such as document selection in Retrieval-Augmented Generation (RAG) and expert routing in Mixture-of-Experts (MoEs). However, training these selection modules remains challenging due to weak gradient signals and suboptimal exploration-exploitation tradeoffs. Furthermore, prior works often rely on heuristics, lacking principled objectives and approaches that explicitly model and solve the top-$k$ selection problem. In this work, we propose a principled objective for training selection modules, whose gradient naturally provides richer training signals in a hybrid form---containing a supervised-gradient component and a policy-gradient component. We show that the selection problem becomes harder as $m$ increases, and our algorithm converges at rate $O(1/\sqrt{T})$, with the optimal upper bound achieved by balancing between bias and variance. Practically, we apply our method to a set of tasks involving top-$k$ selection, including synthetic regression problems, RAG, and MoE systems, showing that our method outperforms the baselines in next-token prediction perplexity and QA accuracy.

212. 【2610.04156】rajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents

链接:https://arxiv.org/abs/2610.04156

作者:Mincheol Daniel Song,Joshua Ward,Jake Jung,Guang Cheng

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:applications reason autonomously, LLM agents, reason autonomously, autonomously over multiple, applications reason

备注: Accepted at the NeurIPS 2026 Workshop on Resource-Aware Agentic AI (RAAAI). 19 pages, 7 figures

点击查看摘要

Abstract:LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resource failures: an incorrect reasoning trajectory spends computation budget on outputs that must be discarded. We introduce Sentinel, a trajectory-derived, classifier-based confidence layer that analyzes an agent's reasoning, code and database outputs to decide at three points whether to stop: refusing unanswerable questions before the agent runs, halting doomed trajectories mid-run, and withholding untrustworthy answers at delivery. Here, utilizing Chow's rule, we optimize decisions under the EHRSQL shared task's Reliability Score, which penalizes incorrect answers given a utility weighting, and find on the benchmark EHRSQL that Sentinel raises this score from +0.08 to +0.24 when mistakes have a low utility weighting, well above the +0.03 earned by refusing every question, with delivered-answer accuracy rising from 54% to 69% as coverage falls from 84% to 55%. At higher stakes the agents we test rarely answer reliably enough to deliver, and Sentinel detects this on its own, abstaining to that same +0.03 where the unmonitored agent scores -3.02. The same stopping decisions cut computation: at low stakes, where the system still answers, Sentinel eliminates 13-28% of agent steps at little or no reliability cost.

213. 【2610.04140】ExpertMuon-Compass: Alignment-Guided Step Sizes for Mixture-of-Experts Training

链接:https://arxiv.org/abs/2610.04140

作者:Omatharv Bharat Vaidya,Ashwin Vinod,Pedram Akbarian,Aditya Sai Ellendula,Connor T. Jerzak,Nhat Ho

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Optimization and Control (math.OC)

关键词:language models send, expert, models send, Muon applies updates, Muon

备注: 37 pages, 9 figures

点击查看摘要

Abstract:Mixture-of-experts (MoE) language models send each token to a few experts, so each expert is trained on a different part of the data, and this part changes during training. With a shared learning rate, Muon applies updates of roughly the same size to expert matrices of the same shape, even when an expert's update is poorly aligned with its current gradient. We here propose ExpertMuon-Compass (Compass), which multiplies the Muon step of each expert by two factors. A family factor compares the cosine between the expert's orthogonalized update and its gradient with the same cosine for the other experts in its layer. A scalar radius aggregates the alignment between corresponding rows of the update and gradient into one step-length multiplier. Compass keeps the update direction and the momentum buffer of Muon. In pretraining on FineWeb-Edu, Compass with Nesterov momentum on all matrices performs as well as or better than Muon, NorMuon, and other optimizers, with weight decay matched to NorMuon in the longer runs. Adding its factors to NorMuon gives the same or a lower loss than NorMuon. Compass is the most effective when the data seen by each expert varies during training, for example, as when the languages of a multilingual corpus arrive in separate blocks. With Compass, the expert load stays balanced, and the router assigns tokens to experts more decisively. We also prove a perturbation bound for the two factors.

214. 【2610.04132】PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

链接:https://arxiv.org/abs/2610.04132

作者:Jingquan Wang,Jun Yin,Xu Han,Yongsheng Mei,Jie Hao,Bin Guo

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:requires Building LLMs, Building LLMs, locally helpful responses, producing locally helpful, LLMs that behave

备注:

点击查看摘要

Abstract:Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.

215. 【2610.04129】InvestigationWorlds: An Agentic Environment for Legal Investigation

链接:https://arxiv.org/abs/2610.04129

作者:Albert Yu Sun,Andrew Benard,Sil Hamilton,Anna Teresita A. Marcelo,Yong Jae Kim,Carl-Leander Henneking,Rundong Hu,Yuhong Wang,David Mimno,Bishan Yang,Igor Labutov

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:introduce InvestigationWorlds, legal investigation, Court Electronic Records, court-adopted hypothesis, Federal Court case

备注: Accepted to NeurIPS 2026 Evaluations Datasets

点击查看摘要

Abstract:We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a court-adopted hypothesis that is treated as ground truth for the purposes of deciding the motion. Each environment is built from a real U.S. Federal Court case retrieved from Public Access to Court Electronic Records (PACER) and augmented by an attorney-validated generation pipeline that synthesizes role-tagged documents around the original record. The resulting corpus admits multiple coherent factual readings, only one of which matches the court-adopted hypothesis. Evaluating on 100 cases, we find agents often commit to incorrect hypotheses despite retrieving relevant evidence, struggling to distinguish the court-adopted hypothesis from alternative hypotheses.

216. 【2610.04128】What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

链接:https://arxiv.org/abs/2610.04128

作者:Georgios Politis,Evangelos Pappas

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:server sends gradients, server, Split learning, client, gradients

备注:

点击查看摘要

Abstract:Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20% of tokens from the activations alone and 97.38% when it also sees the gradients, 3.17 percentage points more 95% interval [2.72, 3.64]. Counted by document, the difference is much larger. The attacker rebuilds 13.71% of 32-token documents exactly without the gradients and 37.77% with them, because a document only counts when every token is right. How we count also changes how good a defence looks. Secret mixup, which blends each outgoing vector with a decoy, stops the attacker from rebuilding almost any document exactly, yet the attacker still recovers 83-91% of tokens. In a second experiment on GPT-2 and Qwen3-0.6B, where the server trains only a run of consecutive layers, the layer at which the run starts changes both model quality and leakage, even when the run's length is fixed. We recommend reporting leakage both per token and per document, and treating what a split model sends as being as sensitive as the text itself.

217. 【2610.04125】Representational Control over Self-Report Behavior Coherence in LLM Risk-Taking

链接:https://arxiv.org/abs/2610.04125

作者:Rafal Kocielnik,Peiyang Song,Pengrui Han,Myrl G. Marmarelis,Ramit Debnath,Dean Mobbs,R. Michael Alvarez

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:appealing low-cost probe, recent work finds, appealing low-cost, low-cost probe, recent work

备注: 10 pages main content, 62 pages total, 21 figures, 19 tables

点击查看摘要

Abstract:Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model's own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model's own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each channel reachable by several independent constructions. (3) An intervention composing one behavior-moving and one self-report-moving direction moves both together; flipping one sign sets them in opposition, with reported and enacted risk pointing in opposite directions on 69-89% of flipped compositions in all models. These findings move the self-report-behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation.

218. 【2610.04119】Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?

链接:https://arxiv.org/abs/2610.04119

作者:Tejas Dahiya,Cole Blondin

类目:Computation and Language (cs.CL)

关键词:Indirect Object Identification, Object Identification task, Mechanistic interpretability, studies fully trained, interpretability usually studies

备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned once rather than the name mentioned twice. Pythia models pass through an early training window in which they prefer the repeated name, so accuracy in a choice between the two names falls below one half while language-model loss on a fixed text sample keeps decreasing across the same window. The window reflects a temporary imbalance between two computations. We identify one cause of the wrong preference by selecting a set of attention heads that write the repeated name, on prompts separate from those used for causal evaluation, keeping that selection fixed, and then replacing each head's final-token output with its average output on a separate set of non-repeated-name prompts. This improves the correct-minus-repeated logit difference in a separately trained 160M model and in the official 160M, 410M, and 1B models. At 160M, the head that lowers the repeated name in the mature model shows little of its mature behaviour at this point. It directs less than one percent of its attention to the repeated mention, and its output makes almost no direct contribution to lowering that name's logit. Both properties grow over the interval in which behaviour recovers. Across the 160M, 410M, and 1B models, transplanting the corresponding head's mature parameters into the early checkpoint recovers 35 to 68 percent of the total improvement in the correct-minus-repeated logit difference seen by the end of training. Related early-to-late reversals appear at further Pythia scales, in two independently trained GPT-2 models, and in OLMo. A mature circuit can therefore conceal a transient causal configuration that shaped behaviour earlier in training.

219. 【2610.04118】LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?

链接:https://arxiv.org/abs/2610.04118

作者:Xinyi Liu,Rinat Khaziev,Dilek Hakkani-Tür,Tarek F. Abdelzaher

类目:Computation and Language (cs.CL)

关键词:track parent-reply relations, ingest entire online, entire online discussion, turning points, scoped subtrees

备注: Accepted to EMNLP 2026 Main Conference

点击查看摘要

Abstract:Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social reasoning, we introduce LongSocialBench, a benchmark of 1,462 verified human-authored multiple-choice items drawn from 94 complete Hacker News, Stack Exchange, and Reddit r/ChangeMyView episodes, with a median length of approximately 73K tokens. Each item pairs a complete serialized discussion and reply structure with a four-option question, requiring models to recover structured social evidence. Released items are verified for answerability, option uniqueness, and evidence grounding. Across 18 models and 29 evaluation settings, current long-context workflows remain far below human performance. The best individual result comes from Claude-Opus-4.7, which reaches 63.0% when prompted to eliminate incorrect options before answering, compared with 72.4% for independent human readers. Averaged across all 18 models, the full-context Baseline scores 43.9%. Supplying the gold evidence scope raises this to 55.0%, showing that substantial errors remain even after the relevant thread region is identified. LongSocialBench shows that the missing capability is not context access or prompting alone, but social understanding over structured reply trees.

220. 【2610.04109】SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

链接:https://arxiv.org/abs/2610.04109

作者:Mingtian Tan,Palash Goyal,Mihir Parmar,Sarkar Snigdha Sarathi Das,Chun-Liang Li,Nanyun Peng,Thomas Hartvigsen,Jinsung Yoon,Tomas Pfister

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:numerical observations insufficient, historical numerical observations, Real-world time series, rendering conventional forecasting, conventional forecasting based

备注: 39 pages, 19 figures

点击查看摘要

Abstract:Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals, and an inability to reason causally about event impacts. We propose SEER (Self-Evolving Event Reasoning and Retrieval), a closed-loop framework that dynamically optimizes event conditioning for time series forecasting. SEER translates prediction errors into two decoupled feedback mechanisms: (i) a reflective retrieval memory that refines subsequent search queries and filters spurious noise, and (ii) a persistent causal knowledge base that distills transferable domain dynamics. SEER enforces strict chronological boundaries across both event retrieval and reflection, preventing look-ahead bias and data leakage. Across six volatile time-series benchmarks, SEER consistently outperforms state-of-the-art time series foundation models and language model baselines.

221. 【2610.04098】Representation-Aligned Auxiliary Supervision for Language Model Adaptation

链接:https://arxiv.org/abs/2610.04098

作者:Kyuyoung Kim,Peiyao Sheng,Ashwin Hebbar,Peiyang Xu,Yunfei Xie,Kevin Wang,Rui Xin,Chen Wei,Zhangyang Wang,Jinwoo Shin,Pramod Viswanath,Sewoong Oh

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Language models exhibit, strong reasoning capabilities, yield inconsistent outcomes, exhibit strong reasoning, domains remains challenging

备注: NeurIPS 2026 Workshop on LP4FM Oral

点击查看摘要

Abstract:Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the extent to which a model effectively processes a representation for a structured task, as a key factor in adaptation. We study this in chess, which provides a controlled testbed with precise semantics, computable optimal actions, and multiple state representations, including a symbolic encoding (FEN) and a spatial format (ASCII). We find that models often process semantically equivalent inputs substantially differently, affecting both learning and generalization. Building on this observation, we propose representation-aligned auxiliary supervision, which uses environment-derived tasks expressed in compatible representations to improve adaptation to structured domains. Across models and representations, auxiliary supervision consistently improves optimal-move prediction relative to target-only training under identical target data. Tasks that expose environment dynamics provide larger and most consistent gains than surface-level or static supervision, while remaining competitive with substantially increasing the amount of target-task data. Moreover, ASCII-trained models transfer more effectively to FEN than FEN-trained models do to ASCII, even surpassing the FEN target-only baseline on FEN evaluation. The gains also extend beyond optimal-move prediction to open-ended, factually grounded commentary generation. Overall, our results show that auxiliary supervision in model-compatible representations can enable effective adaptation in structured domains.

222. 【2610.04074】IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

链接:https://arxiv.org/abs/2610.04074

作者:Jiarui Liu,Renjie Tao,Yiwei Liao,Chuanyang Jin,Kai Sun,Xiao Yang,Xinyuan Zhang,Xilun Chen,Zhuangqun Huang,Lechen Zhang,Yongjin Yang,Yinghui He,Weihao Xuan,Rakesh Wanga,Anuj Kumar,Mona T. Diab,Wen-tau Yih,Xin Luna Dong

类目:Computation and Language (cs.CL)

关键词:automating scientific research, generating promising, grounded research solutions, research solutions remains, rapid progress

备注:

点击查看摘要

Abstract:Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Accordingly, we introduce IdeaScientist, which decomposes ideation into gap finding, innovation, and report writing, and trains each role with reinforcement learning. These roles identify limitations in related work, draw solution intuitions from analogous problem settings, and develop those intuitions into complete research proposals. To facilitate discovery of insights across domains, we construct the Svalbard Idea Vault, a corpus of 2.77M decomposed research ideas for retrieval, training, and temporally controlled evaluation. Our evaluation restricts access to literature available before a cutoff date and assesses how closely proposed directions align with those later explored in 15K papers authored by human researchers. On Qwen3.6-27B, IdeaScientist outperforms the strongest open-source autoresearch baseline by 14.0%, driven mainly by gains in novelty. On this 27B open backbone, IdeaScientist even outperforms Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4, by up to 5.9%.

223. 【2610.04047】Periscope: Extending Frozen Language Models Beyond Their Context Window

链接:https://arxiv.org/abs/2610.04047

作者:Mohamed Eltahir,Anas Obayd,Raed Rashid,Abdulrahman Alghamdi,Abdulrahman Mousa,Abdallah Ahmed,Tanveer Hussain,Naeemullah Khan

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:quadratic forward pass, quadratic forward, loses accuracy, accuracy with length, length before reaching

备注:

点击查看摘要

Abstract:A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.

224. 【2610.04017】Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models

链接:https://arxiv.org/abs/2610.04017

作者:Sankaran Vaidyanathan,Rafal Urbaniak,Emily Bunnapradist,Michelangelo Naim,Daniel Waxman

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:Localizing behavior, mechanistic interpretability, behavior to individual, central goal, goal of mechanistic

备注:

点击查看摘要

Abstract:Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactions via witnesses: variables that provide contextual information to resolve interaction terms. However, estimation with witnesses typically requires combinatorial enumeration and is infeasible in practice. We introduce the witness-integrated set effect (WISE), a family of causal estimands that build on the witness mechanism while taking expectations over sets of causes and witnesses to remain computationally feasible. Building on this approach, we introduce JuntaLearner, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying-sized sets of components and witnesses. Alongside faithfulness metrics, we introduce measures of necessity and task specificity, and the circuit recognition score (CRS) to summarize each metric across circuit sizes while emphasizing effects achieved by small circuits. Across tasks and models of increasing size, JuntaLearner achieves higher mean CRS compared to attribution baselines on all metrics. Since its cost does not grow with the number of candidate components, JuntaLearner scales to large models while accounting for set-level interactions and avoiding first-order approximations.

225. 【2610.04002】Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

链接:https://arxiv.org/abs/2610.04002

作者:A. Bochkov

类目:Computation and Language (cs.CL)

关键词:embedding table assigns, assigns each vocabulary, vocabulary item, input embedding table, input

备注:

点击查看摘要

Abstract:A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.

226. 【2610.03998】Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas

链接:https://arxiv.org/abs/2610.03998

作者:Khashayar Pourtaheri,Ahmad Zareei

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:Large language models, Large language, language models, representing survey respondents, behavioral history

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add progressively richer information: no personal information, demographics, personality traits, cognitive scores, and finally the respondent's earlier survey choices as behavioral history. We use a two-wave panel of 845 US adults who completed measures of 14 behavioral biases (spanning risk, time preferences, overconfidence, and reasoning), so each respondent's earlier answers provide a human test-retest benchmark; in the behavioral-history condition, all items that score the target bias are withheld. At the population level, the average number of biases per respondent in every condition is close to the human average (7.1-8.1 biases, against 7.1 for humans). This aggregate similarity masks differences in variance: persona descriptions recover only 53-67% of human between-person variation, whereas adding behavioral history restores it to approximately the human level. At the individual level, description-based personas achieve only 7-12% of the informedness observed in human test-retest responses, while adding behavioral history raises this to 28%. The condition including behavioral history has the highest estimated informedness in all 17 demographic groups, whereas description-based conditions provide little or no information for some groups. Synthetic responses also exhibit stronger education- and income-related differences than human responses. For LLM synthetic personas, a respondent's past answers add more to individual-level prediction than a description of who they are.

227. 【2610.03962】BAIBAICHUCHU at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text?

链接:https://arxiv.org/abs/2610.03962

作者:Zong-Han Bai,Po-Yen Chu

类目:Computation and Language (cs.CL)

关键词:BAIBAICHUCHU team participated, Maximum Possible Profit, Chinese investor posts, ranking Chinese investor, BAIBAICHUCHU team

备注: 15 pages, 1 figure, 6 tables. Both authors contributed equally. Selected for oral presentation at NTCIR-19 (FinArg-3 Social Media Subtask)

点击查看摘要

Abstract:The BAIBAICHUCHU team participated in the Social Media Subtask of NTCIR-19 FinArg-3, ranking Chinese investor posts by Maximum Possible Profit (MPP). A three-track ensemble of lexical features, a FinArg-2-pre-finetuned MacBERT ranker, and an LLM judge reaches 0.734 in post-grouped development evaluation, but our best official run scores 0.517. All twelve submitted runs lie between 0.4598 and 0.5402, and our 26 unanimous three-track pairs score 0.500. A post-hoc audit finds that the submitted judge applied a long-only rule to bearish posts although MPP is stance-aware. Correcting it changes 28 of 87 official predictions without changing accuracy, yet lowers development accuracy from 0.680 to 0.622: a regime-specific semantic shortcut improved validation fit. We decompose ranking into directional text, volatility and horizon, and pairwise margin. A running-extremum model predicts $\sigma\sqrt{T}$ scaling, observed ex post in 502 price-aligned posts from a separate July 2026 collection. Pre-posting volatility scaled by the post-specific observation horizon, $\sigma_{\mathrm{pre},20}\sqrt{N_i}$, is associated with a later truncated-horizon MPP proxy (Spearman $\rho=0.320$; ticker-cluster 95% CI [0.133,0.466]) and correctly orders 61.4% of unequal-outcome pairs. Transferred text is nearly uncorrelated with the July outcome ($\rho=0.055$) and adds little conditional on this score and stance. Development reliability rises from about 0.60 to 0.93 as the labeled MPP gap widens. An earlier ERAI result of 0.6207 rules out universal unpredictability as a simple explanation. Evaluation should jointly consider text, market regime, historical volatility, horizon, and pair composition.

Comments:
15 pages, 1 figure, 6 tables. Both authors contributed equally. Selected for oral presentation at NTCIR-19 (FinArg-3 Social Media Subtask)

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2610.03962 [cs.CL]

(or
arXiv:2610.03962v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.03962

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
228. 【2610.03940】A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

链接:https://arxiv.org/abs/2610.03940

作者:Valeria Ruscio,Seth Nabarro,Keiran Thompson

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:language model, model, answers, answer mass, current gradient

备注:

点击查看摘要

Abstract:During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimiser memory contributes to forgetting by separating old-task loss into confusion among its answers and leakage of probability outside the answer set. Across three language-model families, answer mass consistently declines while discrimination among old answers usually improves: the model becomes less likely to produce answers that it can still distinguish correctly. Decomposing Adam updates reveals opposing contributions to this loss of answer mass. Over training, accumulated history favours leakage, while the current gradient opposes it. Resolving history by age shows that the harmful contributions come mainly from older gradients of the new task, whereas recent gradients tend to protect the old answers. Changes in history's effect are dominated by its orientation relative to the old-task gradient. Interventions that reset momentum while matching the initial update norm establish that stored history affects retention, with state-dependent immediate effects and lower final old-task loss over longer Adam continuations, mainly through recovered answer mass. Finally, integration along finite updates shows that most sampled large loss increases are captured by local projections, while curvature along history amplifies some events. Together, these findings reveal how an optimiser's memory can erode learned behaviour even as its current gradient acts to preserve it.

229. 【2610.03935】General Decision Models: Benchmarking and Insights Beyond Jev

链接:https://arxiv.org/abs/2610.03935

作者:Feiyu Duan,Jiayu Lin,Jia Wang,Jun Xiang,Jialiang Wu,Xinnong Zhang,Hanqi Yan,Siyuan Wang,Zhongyu Wei

类目:Computation and Language (cs.CL)

关键词:General decision models, judgment and selection, decision models, recently emerged, emerged as efficient

备注:

点击查看摘要

Abstract:General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVal, a bilingual benchmark comprising 11,257 instances from 36 datasets across 10 application domains, and evaluate 25 model configurations spanning general decision models and generative LLMs. Our results show that (1) general decision models are most competitive when decisions can be resolved from available evidence, but weaken when they require specialist knowledge or faithful uncertainty estimation: they can often identify the most likely outcome while substantially overstating its probability. (2) In more dynamic and realistic systems involving long-horizon, multi-step interactions, the advantages of fast local decision making are offset by reliability failures at the system level. on $\tau$-bench, faster local decisions reduce median episode time but lower task success as decision errors accumulate over long trajectories. (3) In large-scale social simulation, decision models approach strong generative LLMs on individual response prediction at substantially lower inference cost, yet remain weaker in user profiling and exhibit larger aggregate estimation errors and systematic bias. Finally, we propose InnerJev-4B and InnerJev-27B, which internalize an open-weight LLM's own reasoning into a single-pass first-token decision through Reasoning-to-Readout Self-Distillation, with InnerJev-27B performing on par with Jev on JEVal while answering a typical query in about 0.1 s.

230. 【2610.03902】When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents

链接:https://arxiv.org/abs/2610.03902

作者:Wenhui Chu(University at Albany, State University of New York)

类目:Computation and Language (cs.CL)

关键词:agent derived facts, re-read current evidence, revoked or replaced, derived facts, facts are revoked

备注: 43 pages, 5 figures, 28 tables

点击查看摘要

Abstract:When documents supporting an agent's derived facts are revoked or replaced, should it repair memory or re-read current evidence? We introduce an evidence-revision evaluation on medication- and problem-list tasks from public ICU records. Under revocation, replacement and control events, we compare full and source-filtered re-reading with caching, rebuilding and graph-local repair across two 7B models. Memory is supplied in full without retrieval, and costs include ingest, revision and every use. On short records, local repair uses 5-10$\times$ fewer revision tokens than rebuilding, yet every memory pipeline costs at least twice full re-reading in held-out conditions. In a small pre-specified development sweep, adding task-ineligible documents extended records to about 10,000 tokens; at that length, memory's mean cumulative cost fell below full re-reading's after 2-14 uses, partly through truncated extraction, while source-filtered re-reading remained cheapest. In the replacement study, none of the four primary confirmatory tests reached statistical significance. These results show why the cost of agent memory after evidence revision must be assessed against source-filtered re-reading over the full pipeline.

231. 【2610.03894】Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities

链接:https://arxiv.org/abs/2610.03894

作者:Yikai Zhao,Saurabh Pandey,Pradeep Kumar Misra

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:deployed LLM agent, LLM agent emits, deployed LLM, LLM agent, agent emits tool

备注: 20 pages, 6 figures, 12 tables

点击查看摘要

Abstract:A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, reproducing the same call across samples. We recover the missing signal from a low-cost open-weight surrogate run in parallel. It reads the same context, schema, and proposed action as the agent, then scores the call from its own log-probabilities through a family of complementary readouts: teacher forcing and request-PMI weigh the likelihood of each argument value, a discriminative verdict judges the call as a whole, and tool-choice competition tests the function against its siblings. One principle says which to trust: a generative likelihood localizes wrong argument values, while the verdict catches holistically wrong calls. When the error type is unknown, an ensemble is the low-regret default. The readout is training-free, needs no access to the agent's internals, and costs one prefill pass alongside the tool call. On difficult coding tasks it reaches AUROC 0.825 where the actor's stated confidence is near chance (0.598), and the generative readouts beat it by +0.07 to +0.28 across three further actors. Against self-consistency it gains +0.14 to +0.19 on near-deterministic actors, at 1/K the cost. The signal drives two deployment modes: a real-time gate escalating the least-trustworthy calls for review (+0.05 to +0.30 accepted-action accuracy at 50% coverage), and confidence feedback, returning the tool result with the score so the agent adapts its next step -- lifting task success on live-execution benchmarks (+0.119 and +0.137, p = 1e-4) and beating a random-value control where step errors are silent (+0.078, p = 0.003).

Comments:
20 pages, 6 figures, 12 tables

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2610.03894 [cs.AI]

(or
arXiv:2610.03894v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2610.03894

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
232. 【2610.03872】raining Numerical Intelligence via Auto-Diagnosis and Skill Discovery

链接:https://arxiv.org/abs/2610.03872

作者:Peter Chen,Wotao Yin

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:generating scientific code, generating code, generating scientific, scientific code, increasingly capable

备注: 20 pages

点击查看摘要

Abstract:AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework that links numerical diagnosis to reusable solver self-improvement. ADSD follows a diagnosis-first paradigm that first explains why a solver performs poorly, then uses this diagnosis to guide the discovery of appropriate numerical methods. The resulting knowledge is packaged into reusable solver skills, turning solver improvement from trial-and-error editing into a structured process of diagnosis, discovery, and implementation. Across four challenging numerical domains--power flow equation, AC optimal power flow control, stiff ordinary differential equations, and heterogeneous diffusion PDEs--ADSD consistently improves solver accuracy, robustness, and efficiency. On GOC-500 power flow, for example, ADSD reduces mean solver error by nearly $71\times$, with improvements further transferring to unseen grid topologies and operating regimes.

233. 【2610.03839】SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning

链接:https://arxiv.org/abs/2610.03839

作者:Yifeng Zhao,Hongjun Yu,Shibo Wang,Yunjiao Zhou,Zixiao Zhu,Zhipeng Ning,Kezhi Mao,Junlang Qian

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:traces impose substantial, substantial output-token costs, impose substantial output-token, traces impose, output-token costs

备注: 5 pages, 3 figures, 2 tables

点击查看摘要

Abstract:Long chain-of-thought (CoT) traces impose substantial output-token costs. Under constrained budgets, compression must preserve answer-critical information, making boundary placement central. Token-level and fixed-length boundaries can fragment coherent spans such as phrases, formulas, and local derivations, whereas step-level boundaries can bind content requiring different compression actions. We introduce SynLat, a text-latent CoT framework that aligns compression boundaries with syntactic structure through non-overlapping Syntax-Aligned Units (SAUs). An answer-conditioned Teacher constructs progressive KEEP/LATENT targets for a single compression-conditioned Student, which generates mixed reasoning from only the question and requested compression level at inference. Across two Qwen3 Student scales, Standard-CoT and Long-CoT groups, and three compression levels, SynLat matches or exceeds the strongest evaluated baseline in all 12 task-group aggregates and strictly leads in 11 under the reported achieved-CR selection protocol. Overall gains reach 3.6/2.6 points at MEDIUM and 7.0/5.5 points at HIGH for Qwen3-8B/14B, with larger advantages under stronger compression, particularly on Long-CoT groups.

234. 【2610.03829】OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes

链接:https://arxiv.org/abs/2610.03829

作者:Wuraola Oyewusi,Eliana Vasquez Osorio,Goran Nenadic,Gareth Price

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:tumour staging expressions, Real-world outpatient oncology, institution-specific de-identification markers, Real-world outpatient, outpatient oncology notes

备注: Accepted at the AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models Workshop at NeurIPS 2026

点击查看摘要

Abstract:Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.

235. 【2610.03827】he Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer

链接:https://arxiv.org/abs/2610.03827

作者:Saman Rahbar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)

关键词:Researchers often support, model shares structure, support the claim, shares structure, languages

备注: 12 pages, 2 figures, 1 table. Accepted as a poster at the NeurIPS 2026 - Linguistic Principles for Foundation Models (LP4FM). Code: [this https URL](https://github.com/saman-rahbar/score-is-not-the-structure)

点击查看摘要

Abstract:Researchers often support the claim that a model shares structure with the brain, or across languages, by reporting a similarity score. We ask what such a score reads when the shared structure is absent, or when the tool that measures it does not work. We check two settings, and in both the score is not what it appears. First, a probe trained to tell grammatical from ungrammatical sentences in one language transfers worse to more distant languages, the usual evidence for shared structure. But the probe itself gets worse along the same axis: in four of seventeen languages it performs at chance, so 64 of 272 language pairs are scored with a tool that does not work. Dropping those languages halves the strength of the relationship, but they are also the most distant, and this design cannot separate the two effects. Counting matters too: the same data give p = 0.0006 when the 272 pairs are treated as independent and p = 0.155 when the seventeen languages are, which is the correct unit. Second, training a language model to match human brain responses raises its similarity score from 0.10 to 0.34, against a ceiling of 0.54. A model trained on a target whose correspondence to the brain was destroyed still scores 0.31, so only 0.028 to 0.068 of the rise is specific to the brain. With no model at all, a destroyed target already sits at 0.204 from the real one, a floor that tracks the target's rank divided by the number of sentences. Finally, steering a language along its own direction works (+6.2 over a random direction in sixteen of seventeen languages) yet shows no effect that varies with language distance. Before asking whether a correspondence helps, ask how much of the score would survive without it.

236. 【2610.03825】Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score

链接:https://arxiv.org/abs/2610.03825

作者:Parth Kulshreshtha,Shivali Dalmia,Abhishek Mukherji

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:methodological attention falls, benchmark score compares, methodological attention, attention falls, score compares

备注: 13 pages (8 main), 1 figure, 8 tables. Accepted as a poster at TAE (Trust-AI-Eval): Can We Trust AI Evaluation?, a NeurIPS 2026 workshop (non-archival). The workshop name comes from the acceptance email in [this http URL](http://reviews.md)

点击查看摘要

Abstract:A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shipped as gold, and a reviewer gold from independent expert re-annotation of a sample. Using it, we hold the scored output fixed and exchange the reference. The score moves by 4.95 F1 points for one output and 2.00 for the other (95% CIs [3.23, 6.27] and [0.26, 3.37]), and by 7.55 in the worst language. The comparison between two outputs moves as well: the paired interaction between reference and system is +2.95 points (CI [+1.97, +3.87]), survives correction for multiple testing, and changes one language's margin outright. The cause is an undocumented aggregation default that usually kept one annotator when adjudication did not fire, making the shipped reference a partial copy of an output being scored. We report the resulting reference-sensitivity band, show how to compute one from any retained annotation record, and argue that the quantity belongs beside the score.

237. 【2610.03802】Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models

链接:https://arxiv.org/abs/2610.03802

作者:Quim Motger,Carlota Catot,Marc Oriol

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Software Engineering (cs.SE)

关键词:polarity-based opinion mining, including emotionally informed, feature-oriented feedback analysis, emotionally informed issue, informed issue prioritisation

备注:

点击查看摘要

Abstract:Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied. Objectives: Building on a previously published annotation framework and human-labelled ground truth adapted from Plutchik's taxonomy, this paper investigates how large language models can be leveraged for automatic multi-label emotion classification under severe class imbalance. Methods: We compare encoder-only fine-tuning under multi-label and binary-ensemble formulations, decoder-only zero- and few-shot prompting across open-source and proprietary models, and a catalogue of imbalance mitigation strategies (loss reweighting, resampling, generative data augmentation), with the synthetic-review generator and prompting strategy selected via an intrinsic augmentation-utility ranking. Results: Fine-tuned encoders trail the best decoder-only few-shot prompting (macro-F1 0.642) by a wide margin at baseline (multi-label: 0.387; binary ensemble: 0.450); pairing the best multi-label encoder with generative data augmentation and positive-weighted loss closes most of this gap (+0.204) at up to three orders of magnitude lower inference latency than the decoders, with the largest gains on the rarest emotions, from undetected to gains of up to +0.501 F1. Conclusion: Large language models make fine-grained, multi-label emotion classification of app reviews feasible for requirements engineering pipelines, with modest macro-F1, and the best formulation and mitigation strategy are backbone- and formulation-dependent. We release the experimental pipeline, synthetic corpora, and fine-tuned checkpoints for replication and reuse.

238. 【2610.03765】Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

链接:https://arxiv.org/abs/2610.03765

作者:Varsha Suresh,Divij Jain,Jia Liu,M. Hamza Mughal,Vera Demberg

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:co-speech gesture generation, atomic units, gesture generation, motion tokenizers encode, tokenizers encode motion

备注: Accepted at MINT EMNLP 2026

点击查看摘要

Abstract:Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descriptors spanning from raw motion to abstract communicative function. Results show that geometry and handedness are readily decodable from token embeddings, while motion category is only weakly decoded despite showing systematic differences in discrete code usage. This gap between reconstruction quality and descriptor decodability suggests that reconstruction objectives alone do not guarantee that gesture semantics are captured, and that evaluating codebooks on such properties can guide the design of more semantic motion tokenizers.

239. 【2407.21783】he Llama 3 Herd of Models

链接:https://arxiv.org/abs/2407.21783

作者:Aaron Grattafiori,Abhimanyu Dubey,Abhinav Jauhri,Abhinav Pandey,Abhishek Kadian,Ahmad Al-Dahle,Aiesha Letman,Akhil Mathur,Alan Schelten,Alex Vaughan,Amy Yang,Angela Fan,Anirudh Goyal,Anthony Hartshorn,Aobo Yang,Archi Mitra,Archie Sravankumar,Artem Korenev,Arthur Hinsvark,Arun Rao,Aston Zhang,Aurelien Rodriguez,Austen Gregerson,Ava Spataru,Baptiste Roziere,Bethany Biron,Binh Tang,Bobbie Chern,Charlotte Caucheteux,Chaya Nayak,Chloe Bi,Chris Marra,Chris McConnell,Christian Keller,Christophe Touret,Chunyang Wu,Corinne Wong,Cristian Canton Ferrer,Cyrus Nikolaidis,Damien Allonsius,Daniel Song,Danielle Pintz,Danny Livshits,Danny Wyatt,David Esiobu,Dhruv Choudhary,Dhruv Mahajan,Diego Garcia-Olano,Diego Perino,Dieuwke Hupkes,Egor Lakomkin,Ehab AlBadawy,Elina Lobanova,Emily Dinan,Eric Michael Smith,Filip Radenovic,Francisco Guzmán,Frank Zhang,Gabriel Synnaeve,Gabrielle Lee,Georgia Lewis Anderson,Govind Thattai,Graeme Nail,Gregoire Mialon,Guan Pang,Guillem Cucurell,Hailey Nguyen,Hannah Korevaar,Hu Xu,Hugo Touvron,Iliyan Zarov,Imanol Arrieta Ibarra,Isabel Kloumann,Ishan Misra,Ivan Evtimov,Jack Zhang,Jade Copet,Jaewon Lee,Jan Geffert,Jana Vranes,Jason Park,Jay Mahadeokar,Jeet Shah,Jelmer van der Linde,Jennifer Billock,Jenny Hong,Jenya Lee,Jeremy Fu,Jianfeng Chi,Jianyu Huang,Jiawen Liu,Jie Wang,Jiecao Yu,Joanna Bitton,Joe Spisak,Jongsoo Park,Joseph Rocca,Joshua Johnstun,Joshua Saxe,Junteng Jia

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:Modern artificial intelligence, Modern artificial, artificial intelligence, systems are powered, Llama

备注:

点击查看摘要

Abstract:Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.

240. 【2610.06615】COMPASS 2.0: psychometric representational similarity analysis distinguishes symptom structure from personal signal

链接:https://arxiv.org/abs/2610.06615

作者:Baihan Lin

类目:Neurons and Cognition (q-bio.NC); Computation and Language (cs.CL)

关键词:Language models, score psychiatric questionnaires, psychiatric questionnaires, Language, questionnaires from speech

备注: 28 pages, including Extended Data and Supplementary Information. Code for reproducibility: [this https URL](https://github.com/linlab/xpsych)

点击查看摘要

Abstract:Language models can score psychiatric questionnaires from speech, but agreement with self-report may reflect the questionnaire rather than the person. We introduce psychometric representational similarity analysis, a framework for comparing the structure of speech-derived scores, self-report, item wording and theory, and implement it alongside person-level construct scoring in COMPASS 2.0. We show how similarly worded items induce covariance without psychological signal. In pre-registered discovery and confirmation analyses of clinical interviews from 275 participants, language-derived symptom geometry resembled wording more than self-report, with no structure beyond wording detected by the registered tests. Geometric agreement with self-report survived assigning participants someone else's answers, whereas person-paired scores captured distress more than specific symptoms. Complementary analyses examined counselling quality and wording structure across 34 instruments and the Research Domain Criteria (RDoC) framework. These findings distinguish agreement about psychological structure from evidence that language-derived assessments track individual people.

241. 【2610.06562】Synthetic Cultural Agents from Aggregate Anchors

链接:https://arxiv.org/abs/2610.06562

作者:Augusto Gonzalez-Bonorino(1 and 2),Kseniia Biriukova(2 and 3),Monica Capra(2 and 4) ((1) Department of Economics, Arizona State University, (2) EconLLM Lab, (3) Department of Information Systems, Arizona State University, (4) Department of Economics, Claremont Graduate University)

类目:General Economics (econ.GN); Computation and Language (cs.CL)

关键词:combine information supplied, generate synthetic survey, Global Preferences Survey, Direct Preference Optimization, encoded during pretraining

备注: Working paper, September 2026. 20 pages

点击查看摘要

Abstract:Population prompts are widely used to generate synthetic survey responses, but they combine information supplied at inference with associations already encoded during pretraining. We introduce an alternative construction that maps declared aggregate preference anchors into group-indexed choice policies. For each population, the signs of six Global Preferences Survey (GPS) coordinates deterministically label a shared bank of paired synthetic responses, and Direct Preference Optimization fits a parameter-efficient adapter to those comparisons. We evaluate the adapters on candidate World Values Survey (WVS) items using prompts that omit country names and distinguish four questions: recovery of the imposed labels, transfer of the anchor signal to new text, coherence between the GPS anchors and human WVS responses, and agreement between adapter and human scores. The adapters recover the imposed pairwise labels. On a purposively selected sixteen-country development panel, adapter trust scores completely separate the two GPS-sign groups and have a rank correlation of (0.74) with continuous GPS trust scores. Human-GPS and adapter-human associations remain unresolved on the same panel, and results for the other preference dimensions are heterogeneous. These findings show that an anchored policy can retain a declared aggregate signal without thereby reproducing human response patterns. The contribution is therefore both an inspectable construction and an evaluation framework that separates anchor transfer from human criterion agreement.

242. 【2610.06176】What Does It Cost to Simulate a Quantum Sentence Classifier? An Energy and Compute Perspective on Near-Term QNLP

链接:https://arxiv.org/abs/2610.06176

作者:Kishlay Kashyap,Sandipan Ganguly

类目:Quantum Physics (quant-ph); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Near-term quantum natural, natural language processing, practical compute burden, field practical compute, quantum natural language

备注: 11 pages, 4 figures

点击查看摘要

Abstract:Near-term quantum natural language processing (QNLP) experiments often run on classical simulators, so simulator cost is part of the field's practical compute burden, yet accuracy tables do not show it. We measure that cost for a variational quantum classifier (VQC) on binary SST-2 sentiment classification, using PennyLane's state-vector simulator over a controlled grid of 27 configurations: three balanced training-set sizes (N = 200, 500, 1000), three qubit counts (4, 6, 8), and three circuit depths. Each VQC is compared with logistic regression on the same PCA-reduced input; full TF-IDF logistic regression gives an uncompressed reference. The VQC beats its matched baseline in 6 of 27 single-seed comparisons. After reruns at two further seeds, only 1 of these 6 keeps a positive mean advantage larger than its paired seed-to-seed variability, and paired tests on the fixed validation set do not establish it. VQC training is 886-21,127 times slower in measured wall-clock time than the matched classical fit (median 3,158 times); going from 4 to 8 qubits roughly doubles simulator time, and within the tested range per-step cost is well approximated by a linear function of the parameter count. CodeCarbon energy and CO2 estimates are secondary: they imply an almost constant power of about 41 W, so they add little beyond runtime, and we do not build an energy ratio from them. The study is narrow (one dataset, representation, ansatz, simulator, and CPU environment) and is a reproducible feasibility measurement, not a general verdict on QNLP.

243. 【2610.04690】SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

链接:https://arxiv.org/abs/2610.04690

作者:Séverin Baroudi,Hervé Bredin,Ricard Marxer

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Self-supervised learning, speech representation learning, representation learning, single-speaker audio, limiting their usefulness

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction with a pseudo-source-separation objective over frozen random-projection codebooks. By adopting a novel mask-free, multiresolution approach, SepRQ achieves state-of-the-art performance in Speaker Diarization and Speech Separation on the SUPERB benchmark, surpassing WavLM and other cocktail-party derived SSLs at both Base and Large scales, while requiring only 85.68M inference parameters. SepRQ also demonstrates strong performance across target-speaker tasks requiring enrollment (such as Target-Speaker Automatic Speech Recognition), and on the challenging multi-domain DIHARD 3 diarization dataset. Notably, we report strong separation capabilities on three-speaker mixtures (WSJ0-3Mix), where current SSL literature struggles. While cocktail-party SSLs remain scarce and closed-source, limited to C-HuBERT and the enrollment-based SA-WavLM, we open-source SepRQ to the community.

244. 【2610.04333】Factorized Delayed Streams Modeling for LLM-based Streaming ASR

链接:https://arxiv.org/abs/2610.04333

作者:Tatsunari Takagi,Kai Washizaki,Atsushi Kojima,Lianbo Liu,Koki Nikaido,Yui Sudo

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

关键词:Delayed Streams Modeling, enables LLM-based streaming, Delayed Streams, Streams Modeling, LLM-based streaming automatic

备注: Submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token p and the word-start token w to the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show that w can be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability for p from the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.

信息检索

1. 【2610.06703】Reading the Mood: Emotion-Guided Book-to-Music Recommendation via CGANs and LLMs

链接:https://arxiv.org/abs/2610.06703

作者:Manousos Linardakis,Georgios Alexandridis

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:motivating recommender systems, make readers feel, Generative Adversarial Network, Aware Generative Adversarial, Conditional Generative Adversarial

备注: 9 pages, 5 figures, 5 tables. Accepted at SENTIRE 2026 (ICDM 2026 Workshops)

点击查看摘要

Abstract:Background music that matches the mood of a text has been shown to make readers feel more immersed and improve their reading experience, motivating recommender systems that pair books with mood-matched music. In this direction, we present Sentiment Aware Generative Adversarial Network for Cross Domain Recommendation (SAGA-CDR), a two-phase cross-domain recommendation framework that personalizes music suggestions and emotionally aligns them with the book being read. In the first phase, transformer-based sentiment embeddings are constructed from user reviews and mapped across domains via a Conditional Generative Adversarial Network, whose mask-conditioned generator handles missing sentiment components and injects stochasticity for richer preference transfer. A compact rating neural network then fuses sentiment-specific interaction scores with a collaborative filtering prior to predict music ratings. In the second phase, large language models classify each book into a valence-arousal emotional quadrant, and candidate tracks are filtered to match that quadrant. Experiments on both the English Amazon and Chinese Douban datasets show that SAGA-CDR achieves the best rating prediction accuracy on Amazon (RMSE 0.98) and the lowest RMSE on Douban (0.91), with ranking performance competitive with the strongest sentiment-aware baseline, even in cross-lingual settings.

2. 【2610.06590】SPRIG: Semantic-ID-enhanced Paths for Knowledge Graph-based Generative Recommendation

链接:https://arxiv.org/abs/2610.06590

作者:Justin Hangoebl,Marta Moscati,Alessandro B. Melchiorre,Shah Nawaz,Markus Schedl

类目:Information Retrieval (cs.IR)

关键词:Recommender systems leveraging, item identifiers directly, systems leveraging generative, ranking catalog items, identifiers directly

备注: Accepted as a short paper at CIKM 2026. 5 pages, 1 figure, 2 tables

点击查看摘要

Abstract:Recommender systems leveraging generative models often generate item identifiers directly, rather than ranking catalog items by a recommendation score. Recent work extends beyond pure sequential interaction signals by incorporating item content and structured relationships among items, with two distinct directions emerging. Semantic IDs (SIDs) enrich item representations by replacing opaque, randomly initialized embeddings with hierarchically quantized discrete codes derived from item content. Knowledge-graph (KG) path reasoning instead generates entity-relation paths that ground recommendations in structured relationships between items, attributes, and external entities, thereby enriching the relational context. These two lines have complementary limitations: SID-based models lack relational grounding, while KG-based generative recommenders still represent items as arbitrary, opaque tokens tied to large embedding tables, limiting parameter sharing and generalization. We propose SPRIG, a generative recommender that integrates content-derived SIDs into KG path reasoning. SPRIG is trained on information-rich KG paths that terminate in items represented as discrete, content-derived tokens, combining the advantages of both approaches. We evaluate SPRIG on movie and music recommendation datasets against baselines spanning sequential language models, KG-augmented methods, and SID-based approaches. Our results show that SPRIG achieves competitive performance over prior generative models while using fewer parameters and a lower compute cost. Code: this https URL

3. 【2610.06582】Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

链接:https://arxiv.org/abs/2610.06582

作者:Shengtao Wen,Xiang Chen,Yu Tian,Lingbing Guo,Lina Gong,Sheng-Jun Huang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:World-model controllers rely, real control systems, execute commands asynchronously, commands asynchronously due, prediction and planning

备注:

点击查看摘要

Abstract:World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model controllers, rather than treating it only as an external control disturbance. Through controlled interventions, we identify two architecture-dependent failure modes: planning-based controllers such as TD-MPC2 suffer from a future-action timeline mismatch between imagined and executed action sequences, while recurrent world models such as DreamerV3 can attribute observed transitions to commands that were not actually applied. Our analysis shows that TD-MPC2 requires the correct future action sequence during latent dynamics rollout, whereas DreamerV3 requires timely attribution of each transition to the action that generated it. Based on these findings, we introduce two lightweight execution-consistent interfaces, Future-Sequence for TD-MPC2 and Applied-Action Feedback for DreamerV3, that correct these mismatches without modifying the pretrained world models. Experiments across delays, packet loss, reordering, multiple control domains, measured network traces, and a process-separated asynchronous stack consistently support both diagnoses and the corresponding architecture-specific corrections.

4. 【2610.06205】Commercial Intent in Human-AI Conversations: A Corpus Audit and Architecture for Website Sales Agents

链接:https://arxiv.org/abs/2610.06205

作者:Benjamin Tannenbaum

类目:Information Retrieval (cs.IR)

关键词:Conversational sales agents, preserve explicit requirements, current business information, purchase commitments, Conversational sales

备注: 8 pages, 7 figures, 2 tables, 21 references. Technical white paper with an aggregate corpus audit and reference architecture; text-free aggregate counts and reproducibility code included as ancillary files

点击查看摘要

Abstract:Conversational sales agents must distinguish questions about products from purchase commitments, preserve explicit requirements, and ground the next action in current business information. We report an aggregate census of 725,219 records in an accessible conversation table provided by Aiso and develop a reference architecture for this setting. All records have distinct non-null conversation hashes. Existing metadata labels identify 41,800 commercial records (5.76%) and 2,387 transactional records (0.33%); their union contains 44,187 records (6.09%). Within the commercial category, 54.31% are labeled English, 41.59% have recorded depth of at least two, and 13.51% have depth of at least four. Commercial-label prevalence varies from 4.87% to 6.35% across three source batches. These measurements motivate explicit separation of corpus inventory, commercial relevance, training eligibility, and observed business outcomes. The proposed architecture combines business-grounded knowledge, provenance-bearing conversation state, and a constrained next-action policy. A quality specification addresses source rights, privacy, label validation, deduplication, and training-test separation. The study is a metadata audit and technical design, not a validation of label accuracy, model training volume, or sales conversion. No raw conversation text or personal identifiers are released.

5. 【2610.06060】Beyond States: Investigating the Effects of Context on User Modeling with Feature-Conditioned Markov Models

链接:https://arxiv.org/abs/2610.06060

作者:Jana Isabelle Friese,Andreas Konstantin Kruff,Timo Breuer,Philipp Schaer,Norbert Fuhr

类目:Information Retrieval (cs.IR)

关键词:classical state-based approaches, information retrieval systems, Markov models, evaluate interactive information, interactive information retrieval

备注: This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), [this http URL](http://dx.doi.org/10.1145/3799682.3840602)

点击查看摘要

Abstract:User behavior simulation is widely used to evaluate interactive information retrieval systems, but classical state-based approaches (e.g., Markov models) have limited ability to incorporate contextual information relevant for decision-making. We address this limitation by introducing a feature-conditioned Markov-style user model, in which transition probabilities are modeled as functions of positional, content-based, and interaction-derived features, enabling context-aware decision making while preserving the structural simplicity and computational efficiency of state-based models. Applying a multi-level framework that assesses predictive fit and behavioral fidelity, we analyze how different sources of contextual information contribute to realistic user simulation across multiple datasets, search settings, and feature configurations. Our results show that incorporating contextual features improves the models' ability to reproduce key aspects of real user interactions, but that their effectiveness hinges on search scenario and modeling objective. Instead of a one-size-fits-all solution, effective simulation requires task- and setting-specific feature selection. Our framework provides a practical and interpretable basis for making these choices.

6. 【2610.06050】MATE: Adaptive Long- and Short-Term User Memory for LLM-Based Recommendation

链接:https://arxiv.org/abs/2610.06050

作者:Yu Hou

类目:Information Retrieval (cs.IR)

关键词:enhanced recommender systems, Large language model, recommender systems leverage, systems leverage rich, leverage rich item

备注:

点击查看摘要

Abstract:Large language model (LLM)-enhanced recommender systems leverage rich item semantics to support personalized recommendation. However, semantic representations alone do not determine which historical behaviors reflect persistent preferences and which mainly indicate recent interests, leaving an important aspect of user understanding unresolved. Recent advances in LLM inference show that newly available information can be used to refine the internal state during inference, thereby improving subsequent predictions. Inspired by this principle, we propose MATE (Memory Adaptation with Temporal Evidence), an adaptive user modeling framework for LLM-enhanced sequential recommendation. MATE first evaluates each newly observed interaction from two temporal perspectives: whether it is repeatedly supported by historical behaviors and whether it is consistent with recent interactions. The resulting temporal evidence controls the updates of two user-specific memories, where the long-term memory conservatively preserves persistent preferences while the short-term memory rapidly adapts to recent interests. For each recommendation, a recent-context representation dynamically determines how strongly the two memories contribute to the current user representation. During offline training, next-item prediction is jointly optimized with temporal supervision, while during online adaptation, the shared model remains fixed and only the two user memories are updated from newly observed interactions. Experiments on MovieLens-10M, Amazon Luxury Beauty, and KuaiRec show that MATE improves mean NDCG@10 over the strongest external baseline by 7.0--13.2%. Further analyses support its ability to adapt to recent interests while retaining useful information about recurring earlier preferences.

7. 【2610.05945】OntoInk: Interactive Ontology Visualization, Validation, and Reasoning

链接:https://arxiv.org/abs/2610.05945

作者:Ebrahim Norouzi,Jörg Waitelonis,Harald Sack

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:separate tools, Abstract, run OWL reasoning, Ontology, OWL

备注: Demo paper at ISWC 2026 Companion Volume, October 25 to 29, 2026, Bari, Italy

点击查看摘要

Abstract:Ontology documentation, visualization, and validation are usually carried out with separate tools. This split workflow slows down development and makes knowledge transfer harder. We present OntoInk, an open-source MkDocs plugin that brings these activities together. Within a single documentation-as-code pipeline, OntoInk renders interactive ontology diagrams, validates instance data against SHACL shapes, runs OWL\,DL reasoning, and supports inline Turtle editing. General-purpose diagram plugins for MkDocs cannot parse RDF, dereference IRIs, overlay SHACL constraints, or run OWL reasoning. Compared with standalone ontology visualization tools, OntoInk embeds interactive and editable diagrams directly into documentation pages. A live demo and source code are available at \url{this https URL}.

8. 【2610.05807】Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

链接:https://arxiv.org/abs/2610.05807

作者:Hang Xiao,Janet Sung,Zhaoyi Li,Gangzhen Qian,Chuhong Xu

类目:Databases (cs.DB); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Distributed File System, Hadoop Distributed File, log anomaly detection, change the conclusions, conclusions drawn

备注: Accepted at DASC 2026. 8 pages, 1 figure, 8 tables. Reproduction support artifact: [this https URL](https://doi.org/10.5281/zenodo.23151002)

点击查看摘要

Abstract:Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.

9. 【2610.05787】Constraint-Aware Conversational Job Recommendation in Code-Mixed Low-Resource Settings

链接:https://arxiv.org/abs/2610.05787

作者:Md Arman Hossain,Mubashir Jawad,Fariha Khandaker Moon,Sonia Binte Siraj,Masfiqur Rahaman,Raihan ul Islam,Ahmed Wasif Reza,Nafis Sadeq

类目:Information Retrieval (cs.IR)

关键词:real-world career discussions, requires jointly modeling, recommendation requires jointly, Conversational job recommendation, jointly modeling semantic

备注: Submitted to WSDM 2027

点击查看摘要

Abstract:Conversational job recommendation requires jointly modeling semantic relevance, user preferences, eligibility requirements, and the noisy language used in real-world career discussions. These challenges are especially pronounced in low-resource, code-mixed settings, where strict constraint matching can incorrectly eliminate otherwise suitable jobs. We introduce JobCCC, a conversational job recommendation benchmark for Bangladesh comprising 22,410 structured job postings and 988 multi-turn career-advice dialogues derived from regional Reddit communities. Each dialogue is annotated with evolving seeker preferences and linked to a ground-truth job, and is evaluated in semantically equivalent English and Romanized Bangla--English variants. We compare sparse BM25 retrieval, multilingual dense retrieval, and their hard-constraint-filtered counterparts against Weighted Soft-Constraint-Aware Ranking (W-SCAR), our multi-criteria ranking framework that combines lexical relevance, semantic relevance, and graded utilities for experience, location, education, and salary using the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS). Experiments reveal that strict filtering consistently degrades retrieval because incomplete extraction and brittle attribute matching irreversibly remove relevant jobs. W-SCAR avoids destructive pruning and achieves more balanced performance across the two language conditions, obtaining 37.37% and 38.43% Hit@10 on English and Banglish, respectively. The code and dataset are publicly available at \href{this https URL}{GitHub} and \href{this https URL}{Hugging Face}, respectively.

10. 【2610.05750】Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

链接:https://arxiv.org/abs/2610.05750

作者:Reza Esfandiarpoor,Radek Osmulski,Yauhen Babakhin,Gabriel de Souza P. Moreira,Oliver Holworthy,Jie He,Ronay Ak,Jiarui Cai,Ryan Chesler,Bo Liu,Even Oldridge

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:explore large amounts, Large Language Models, retrieval, agentic retrieval, amounts of unstructured

备注: Code: [this https URL](https://github.com/NVIDIA/NeMo-Retriever/tree/main/retrieval-bench)

点击查看摘要

Abstract:Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.

11. 【2610.05732】PACMI: Provenance-Aware Cascading Memory Invalidation for Long-Term LLM Agents

链接:https://arxiv.org/abs/2610.05732

作者:Yiqi Wang,Jiaqi Liu,Jiaqi Zhang,Zhangkai Wu,Yiqun Duan,Mingkai Zheng,Taotao Cai

类目:Machine Learning (cs.LG); Information Retrieval (cs.IR)

关键词:LLM agents rely, LLM agents, long horizons, agents rely, rely on long-term

备注:

点击查看摘要

Abstract:LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dependent records, and retain value as historical evidence. This calls for two capabilities: dependency tracking to identify downstream effects and historical preservation to retain useful past records. We propose Provenance-Aware Cascading Memory Invalidation (PACMI), a framework that represents memories and new evidence in a provenance graph with typed dependency edges. PACMI assigns records to a four-state validity lattice, propagates validity changes to dependent memories, and uses the resulting states for retrieval and stale-premise detection. We also introduce a diagnostic benchmark with 100 cases and 300 queries across five domains. The evaluation separates node, context-, and answer-level performance. PACMI achieves the highest final-answer accuracy on this benchmark, and its paired difference from the strongest baseline is significant under an exact McNemar test. The premise checker achieves perfect precision, recall, and F 1 on the controlled query distribution. Cascading propagation primarily improves memorystate correctness: removing it increases final-answer errors from 3 to 11, but the paired difference does not reach the 0.05 significance threshold. Code and data will be made publicly available.

12. 【2610.05690】Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts

链接:https://arxiv.org/abs/2610.05690

作者:Yanjun Chen(1,2),Yongfeng Zhang(3),Lanjing Zhang(1, 3, 4, 5, 6) (1 Department of Chemical Biology, Ernest Mario School of Pharmacy, Rutgers University, Piscataway, NJ. 2 East Brunswick High School, East Brunswick, NJ. 3 Department of Computer Science, Rutgers University, Piscataway, NJ. 4 Department of Pathology, Princeton Medical Center, Plainsboro, NJ. 5 Department of Pharmacology, Physiology, and Neuroscience, New Jersey Medical School, Rutgers University, Newark, NJ. 6 Rutgers Cancer Institute, New Brunswick, NJ.)

类目:Digital Libraries (cs.DL); Information Retrieval (cs.IR); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)

关键词:Large language models, environmental science, Large language, Nature Climate Change, Lancet Planetary Health

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.

13. 【2610.05670】Generate What You Can Trust: Content Credibility in Generative Recommenders

链接:https://arxiv.org/abs/2610.05670

作者:Zhuo Cai,Guanghao Wu,Shoujin Wang,Peilin Zhou,Victor W. Chu

类目:Information Retrieval (cs.IR)

关键词:Generative recommendation, generates target item, semantic IDs, discrete token sequences, generates target

备注:

点击查看摘要

Abstract:Generative recommendation (GR) represents items with semantic IDs (i.e., discrete token sequences) and generates target item tokens as recommendations. Despite its promising results, existing methods predominantly optimize for accuracy while neglecting the credibility of the recommendations they generate. This oversight inevitably exposes users to uncredible content (e.g., fake news) with serious societal consequences, including user distrust, reputation harm to platforms, and broader social instability. To address this critical yet underexplored challenge, we propose CreGR, the first credible GR model that jointly tackles content credibility across the two core stages of GR: tokenization and generation. In the tokenization stage, we design a new credibility-aware tokenizer that explicitly encourages the model to learn discriminative tokens respectively for credible and uncredible items, thereby disentangling credibility signals at the token level. Building on this, in the generation stage, we propose a novel accuracy-preserving and credibility-oriented generator grounded in discrete diffusion. Specifically, we introduce an asymmetric masking probability reduction strategy that selectively diminishes the contribution of tokens associated with uncredible content to the generation process, while leaving tokens encoding user preference signals unaffected so as to preserve recommendation accuracy. Experiments on three real-world datasets demonstrate the effectiveness of CreGR.

14. 【2610.05619】SCOUT: Supply-Aware Cold-Start Proactive Query Suggestion for Travel Search

链接:https://arxiv.org/abs/2610.05619

作者:Hao Li,Shashank Reddy,Kedar Bellare,Ashish Jain,Stephanie Moyerman

类目:Information Retrieval (cs.IR)

关键词:Large Language Models, Language Models, Large Language, guide intent formulation, powered by Large

备注: Accepted at the CIKM 2026 Workshop on Generative, Retrieval-augmented, and Agentic Intelligence for Personalization

点击查看摘要

Abstract:Generative query suggestion, powered by Large Language Models (LLMs), has become increasingly popular in search and conversational systems to reduce user friction and guide intent formulation. Existing approaches align suggestions with user preferences (e.g., clicks or conversions). This works for open-ended applications like chatbots and personal assistants, where the result space is unconstrained or historical user free-text queries are abundant. However, applying these methods to travel search presents two limitations. First, travel search is fundamentally constrained by physical inventory; a query (e.g., "romantic beachfront villa") may yield abundant results in Bali but few in Tokyo, so aligning with user preferences is not by itself grounded in what can be offered. Second, travel platforms traditionally rely on faceted search interfaces with no free-text queries. This creates a cold-start problem: without historical query logs there is no demand-side data for alignment, and without a seed query at request time, suggestions must be generated proactively from structured context alone. To address these challenges, we propose SCOUT, a bootstrapping framework for supply-aware proactive query suggestion. SCOUT overcomes the data gap by substituting missing demand-side user feedback with supply-side system feedback. It treats the search engine as a reinforcement learning environment, deriving a dense reward from the production reranker's query-listing match scores, and optimizes the policy with Group Relative Policy Optimization (GRPO). SCOUT improves inventory match rate (IMR@18) by 12.3% while preserving diversity, matching a compute-intensive best-of-8 policy at zero marginal inference cost and making supply-aware suggestion deployable on a real-time travel search path.

Comments:
Accepted at the CIKM 2026 Workshop on Generative, Retrieval-augmented, and Agentic Intelligence for Personalization

Subjects:

Information Retrieval (cs.IR)

ACMclasses:
H.3.3; I.2.6

Cite as:
arXiv:2610.05619 [cs.IR]

(or
arXiv:2610.05619v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.05619

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
15. 【2610.05559】Cut Binary Cross Entropy: Efficient Large-Vocabulary Loss and Gradient Kernels for Sequential Recommendation

链接:https://arxiv.org/abs/2610.05559

作者:Yaoyiran Li,Haowen Ning,Mohamed Hammad

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Information Retrieval (cs.IR)

关键词:Industrial sequential recommender, sequential recommender systems, recommender systems operate, Industrial sequential, High Bandwidth Memory

备注:

点击查看摘要

Abstract:Industrial sequential recommender systems operate over massive item catalogs (e.g., 10^5--10^7 items). Multi-label recommendation models are trained with Binary Cross-Entropy (BCE) loss over the full vocabulary, but standard BCE materializes a dense [B, N, V] logits tensor in High Bandwidth Memory (HBM), incurring prohibitive $O(BNV)$ memory and fatal Out-Of-Memory (OOM) errors. While chunked loss optimizations exist for Softmax Cross-Entropy in LLMs, large-scale multi-label BCE optimization remains unexplored across deep learning ecosystems. We propose CutBCE, an exact, hardware-accelerated BCE loss and gradient operator implemented in JAX and Pallas for large-vocabulary workloads. CutBCE introduces (1) an exact fused reformulation evaluating dense background loss and sparse target corrections; (2) a custom Vector-Jacobian Product (VJP) with a dedicated Pallas TPU backward kernel computing logit tiles on-chip in both passes so logits and their gradients never reside in HBM; (3) dynamic VMEM budgeting and sharding-aware collective hoisting for distributed meshes; and (4) count-based zero-overhead training metrics. On single-chip TPU v5e/v6e mini-benchmarks, CutBCE eliminates OOM errors with up to 91.9% speedup. On 8-chip TPU slice training for multi-label SASRec with 876k items (Yambda-50M), CutBCE reduces peak HBM by 65.7% (14 GiB saved per chip) and increases training speed by 225.9% with comparable accuracy. CutBCE is open-sourced at this https URL.

Subjects:

Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Information Retrieval (cs.IR)

Cite as:
arXiv:2610.05559 [cs.LG]

(or
arXiv:2610.05559v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2610.05559

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
16. 【2610.05432】OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation

链接:https://arxiv.org/abs/2610.05432

作者:Yueqi Wang,Zitian Guo,Yupeng Hou,Yifei Wang,Kibum Kim,Zhenrui Yue,Shuo Xing,Haodong Li,Heming Xia,Renrui Zhang,Zhengzhong Tu,Julian McAuley

类目:Information Retrieval (cs.IR)

关键词:Recent advances, substantially improved multimodal, retrieval and reasoning, modeling have substantially, substantially improved

备注:

点击查看摘要

Abstract:Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching. Mainstream approaches primarily adopt independent encoding of vision and language modality followed by rigid late fusion such as concatenation, inherently omitting native vision-language interactions and introducing cross-modal semantic distortion. To address this challenge, we propose OpticalRec, the first visual-space unified encoding paradigm for multimodal collaborative filtering, a fundamental recommendation setting. Instead of isolated modality-specific encoding, OpticalRec renders item textual metadata as visual glyphs, enabling native image-text interaction within the visual encoder - the perceptual encoding level. The resulting representations are further processed by the language decoder - the semantic encoding level, allowing OpticalRec to exploit the dual-attention mechanism of modern vision-language models that previous encoding methods omitted. OpticalRec's efficacy is theoretically supported by mutual information analysis and empirically demonstrated through superior performance across strong baselines and benchmarks. As a plug-and-play module, OpticalRec (1) introduces minimal cost, (2) is robust against rendered text font, color and layout, etc., and (3) integrates seamlessly into existing multimodal collaborative filtering models.

17. 【2610.05348】Search Engines Never Say No: How Frozen Agents React When the Retrieval Tool Refuses

链接:https://arxiv.org/abs/2610.05348

作者:Ramraj Chandradevan,Sayontan Ghosh,Vinoth Selvendran

类目:Information Retrieval (cs.IR)

关键词:miss signal, returns its top-k, irrelevant text, frozen search agents, top-k passages

备注:

点击查看摘要

Abstract:A search tool never says no: it returns its top-k passages even when the index holds no answer, so the agent sees irrelevant text instead of a miss signal. We ask what frozen search agents do when the tool refuses instead. On an index-hole testbed (257 NQ and 300 HotpotQA questions run with and without their gold passages in a 21M-passage BM25 index), seven agents receive one of five refusal wordings. An un-announced one-sentence refusal raises abstention on unanswerable questions from 23% to 97% on average for Qwen3-8B/32B and from 28% to 57% for Claude Haiku 4.5, cutting wrong answers almost one-for-one and beating a system-prompt instruction by 51 points on average. Search-R1 ignores the refusal and fabricates retrievals; Claude Sonnet 5.5 and Opus 5.5 answer from memory (abstention +2 points) and obey a system-prompt directive instead (+16). We also found that wording matters: an explanation beats a bare token; a directive inside the observation is decisive for Haiku; a soft warning is useless. Realistic triggers, from a lightweight score-based predictor to an LLM grounding judge, fall well short of the oracle, and all land on a benefit-versus-signal-quality curve that prices any trigger by its recall at a fixed false-refusal budget: for compliant agents the bottleneck is the detector inside the tool, not the agent, and the curve tells future detector work what each point of recall is worth.

18. 【2610.05343】MemStrata: 95% and 90.91% Source-Aware Accuracy on LongMemEval-500 and LoCoMo-1540 with a Local Qwen 3.8 27B Q4_K_M Reader

链接:https://arxiv.org/abs/2610.05343

作者:Neeraj Yadav(Called It Inc.)

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:incomplete benchmark reference, adequate conversational answer, adequate conversational, short or incomplete, incomplete benchmark

备注: 29 pages, 23 tables, 1 figure. Ancillary files contain per-question grades and reproducible analyses, plus explicitly labelled exports from audited follow-up reports

点击查看摘要

Abstract:An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing system-blinded answers; original reference-only grading is reported alongside. With a local Qwen 3.8 27B Q4_K_M reader and a 24,000-token evidence ceiling, MemStrata CL1 scores 475/500 (95.0%) on LongMemEval-S and 1,400/1,540 (90.91%) on LoCoMo categories 1-4 under source-aware GPT-5.5 adjudication, against 463/500 (92.6%) and 1,205/1,540 (78.25%) under reference-only grading of the same answers. It preserves a retrieval backbone and adds nonduplicated, dated, speaker-attributed source spans. A same-reader full-history control with about 4.7 times the evidence scores 464/500 reference-only and 470/500 (94.0%) source-aware; neither difference is decisive. Keyword-only selection at the same budget scores 425, and a matched-reader Letta arm 438. On LongMemEval-M, where the packet holds about 1.6% of each history, MemStrata CL1 scores 427/500, with losses concentrated in multi-session and temporal questions. On 300 BEAM-1M questions it outscores dense retrieval, 0.738 to 0.706 (Wilcoxon p = 0.011). A same-seed replay of unchanged requests changed 1.5-2.3% of labels. On identical packets GLM 5.3 flash is non-inferior within 3 points (462 versus 463); Muse Spark 1.3 did not show non-inferiority on 269 questions. None of four pre-registered interventions met all of its registered advancement or feasibility criteria. Signed read-side artifacts support inspection but do not regenerate the private retrieval pipeline. The superiority of source-aware grading to human adjudication is not established, and development exposure, automated-judge dependence and the absence of held-out data preclude an independent-replication or leaderboard claim.

19. 【2610.05285】Quality-Aware Cross-Model Computation Reuse

链接:https://arxiv.org/abs/2610.05285

作者:Jin Cheng,Xiangxiang Dai,Maoli Liu,Ziyi Han,Zhuohua Li,John C.S. Lui

类目:Information Retrieval (cs.IR)

关键词:intermediate result computed, scheduling, quality, reuse, quality uncertainty

备注:

点击查看摘要

Abstract:An intermediate result computed by one model can be reused by other models to perform their tasks. Existing work mainly focuses on practical execution, leaving a theoretical gap in optimizing reuse decisions. This optimization faces two challenges: quality uncertainty, because the effect of reuse on task quality is uncertain across models, and coupled scheduling, because tasks need to share the cost of preparing reusable results. These challenges compound each other: quality must be learned online, but the shared preparation structure makes scheduling NP-hard even with known quality, breaking the key assumption in existing methods. We formulate cross-model computation reuse as an online decision problem and develop the Quality-Aware Reuse Scheduling (QARS) algorithm to address it. For quality uncertainty, QARS learns task-dependent reuse quality from selected, possibly delayed feedback and uses optimistic estimates to guide decisions. For coupled scheduling, it jointly chooses which results to prepare and which tasks should use them, adapting scheduling accuracy to the remaining quality uncertainty. For the considered problem, our analysis separates learning and optimization error in regret and quantifies the tradeoff between scheduling accuracy and computation. Completing the quality-aware stopping rule yields $\widetilde O(\sqrt{T})$ regret while preserving feasibility. Experiments demonstrate the effectiveness of QARS in optimizing cross-model reuse, reducing the combined cost of computation and quality loss by up to 18.0%, and mean regret by 63.9% over the strongest scheduling baseline.

20. 【2610.05157】Cross-Modal Contrastive Learning for the Retrieval of Immunotherapy-Associated Molecular Signatures from Histopathology

链接:https://arxiv.org/abs/2610.05157

作者:Sigrid Vila-Bagaria,Mar Teixidó,Miquel Piñol,Felip Vilardell,Robert Montal,Veronica Vilaplana

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Gastric Adenocarcinoma, Multiple Instance Learning, cancer mortality, Contrastive Multiple Instance, Cross-modal Contrastive Multiple

备注: Accepted to MICCAI 2026 CaPTion Workshop

点击查看摘要

Abstract:Gastric Adenocarcinoma is a leading cause of cancer mortality. Although "Inflamed/Non-Inflamed" subtypes have been proposed to predict immunotherapy response, their identification relies on a costly 10-gene RNA signature. We propose a Cross-modal Contrastive Multiple Instance Learning (CCMIL) framework for cross-modal retrieval, imputing these molecular signatures directly from standard Hematoxylin Eosin (HE) slides. By leveraging a supervised contrastive objective, CCMIL aligns visual morphological patterns with molecular phenotypes into a shared latent space. This establishes an interpretable search-by-case retrieval engine, enabling pathologists to query a whole slide image to surface transcriptomically coherent neighbors and approximate RNA signatures without genomic sequencing at inference. Our results demonstrate that this retrieval-first approach captures the continuous phenotypic spectrum of tumor inflammation and yields clinically interpretable attention heatmaps. Furthermore, the learned representation also supports competitive downstream classification, providing a practical molecular pre-screening strategy.

21. 【2610.05107】SearchJev: A Fast and Calibrated System-1 Model for Search Agents

链接:https://arxiv.org/abs/2610.05107

作者:Congfeng Cao,Lipeng Zuo,Konstantinos Papakostas,Qiwei Xu,Songwei Xu,Lun Zhou,Zhaochun Ren,Yougang Lyu,Xiaohui Yan

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:repeatedly make short, agents repeatedly make, evidence sufficiency, repeatedly make, Search

备注:

点击查看摘要

Abstract:Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.

22. 【2610.05034】Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads

链接:https://arxiv.org/abs/2610.05034

作者:Jingjie Ning,Xueqi Li,Yibo Kong

类目:Information Retrieval (cs.IR)

关键词:retrieval-augmented generation span, agentic retrieval-augmented generation, generation span questions, retrieval-augmented generation, generation span

备注:

点击查看摘要

Abstract:Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers standard error by 33\% versus five reads and 12.6\% versus three trajectories. Archived nested and Q-only forecasts predict these allocations within 4.0\% and 3.5\%, respectively. Depth subsets establish no clear forecasting advantage beyond the two-trajectory audit. One-read variance penalties relative to the fitted optimum at the same token budget are 0--9.9\%, with substantial Pro uncertainty. Under recorded model fees, more questions beat more trajectories at search prices of \$0--1 per 1,000 requests; question-versus-read fee rankings remain unresolved. Temperature zero cuts answer disagreement from 14.3\% to 3.4\% while comparison precision stays similar. \par\medskip\noindent\textbf{Keywords:} Agentic RAG; Evaluation budget; Generalizability theory; Repeated sampling.

23. 【2610.05028】Do We Still Need Gazetteers in the Era of LLMs? Chaining Retrieval with a Spatial Neuro-Symbolic Index

链接:https://arxiv.org/abs/2610.05028

作者:Horde-Vo Alexis,Duckham Matt,He Estrid

类目:Information Retrieval (cs.IR)

关键词:tasks require systems, interpret ambiguous toponyms, tasks require, downstream applications, require systems

备注: Accepted to ACM SIGSPATIAL '26

点击查看摘要

Abstract:Geographic information retrieval (GeoIR) tasks require systems to interpret ambiguous toponyms for downstream applications. Traditionally, toponym resolution relies on gazetteers to provide an explicit index of place entities and spatial relationships. Recently, gazetteer-free approaches seek to reduce dependence on handcrafted searches: dense retrieval utilizes text encoders to capture rich context, moving beyond the limitations of lexical search. However, text encoders implicitly assume that learned representations can function as reliable spatial-semantic indexes. In this paper, we evaluate this assumption through a spatial-semantic indexing setup: given a contextualized toponym mention, we retrieve the corresponding gazetteer entity represented by text derived from a gazetteer knowledge graph. We benchmark five frozen text encoders under two retrieval strategies: brute-force nearest-neighbor retrieval over entity representations, and a neuro-symbolic hierarchical beam search that constrains retrieval (i.e. chaining the search with gazetteer hierarchy). Experimental results reveal a distinct coarse-versus-fine trade-off. Unconstrained dense retrieval frequently incurs catastrophic spatial errors. Conversely, hierarchical constraints improve coarse geographic grounding, but still yield limited benefit for fine-grained localization metrics: vanilla text encoders fail to capture the fine-scale spatial fidelity encoded in gazetteers. Our code is publicly available at: this https URL

24. 【2610.04904】ModelLakeFishing: Efficient Retrieval over Million-Scale Model Lakes

链接:https://arxiv.org/abs/2610.04904

作者:Xiaoyang Liu,Zhengyuan Dong,Renée J. Miller

类目:Information Retrieval (cs.IR); Databases (cs.DB)

关键词:Open model lakes, identify suitable models, Navigable Small World, Hierarchical Navigable Small, Open model

备注:

点击查看摘要

Abstract:Open model lakes may contain millions of reusable models, making it costly to identify suitable models for a new dataset. We present ModelLakeFishing, a model-retrieval framework for queries specifying a target dataset, prediction task, and evaluation metric. It consolidates metadata and historical evaluations into a model-dataset-task evidence graph, learns model and query embeddings with a relation-aware graph encoder, and indexes model embeddings using Hierarchical Navigable Small World (HNSW) search. At query time, HNSW retrieves 1,000 candidates without scoring every model, after which a training-side task prior reranks candidates for the requested metric and returns the top 10. We evaluate on a lake of 3,016,439 models and 247,803 observed model-dataset performance pairs using three root-aware splits that hold test performance edges out of representation learning and retrieval. ModelLakeFishing achieves a mean eligible-query gold@10 of 0.2968, recovering the observed-best model in the top 10 for 29.68% of eligible queries and retaining 93.47% of the gold@10 of an exhaustive baseline using the same scoring and reranking procedure. Given precomputed query embeddings, retrieval and reranking take 0.747 ms median and 1.102 ms at the 95th percentile. These results demonstrate efficient retrieval over million-model lakes from sparse relational evidence.

25. 【2610.04352】Query Generation with Direct Preference Optimization for Document Expansion in E-commerce Search

链接:https://arxiv.org/abs/2610.04352

作者:Kaihao Li,Feng Liu,Juexin Lin,Xunfan Cai,Zhen Yang,Tony Lee,Ciya Liao

类目:Information Retrieval (cs.IR)

关键词:document expansion technique, popular document expansion, generate relevant queries, vocabulary mismatch, expansion technique

备注:

点击查看摘要

Abstract:Doc2Query, a popular document expansion technique, leverages sequence-to-sequence models to generate relevant queries, effectively addressing the "vocabulary mismatch" problem in information retrieval. However, these models often suffer from generating either hallucinations unrelated to the document or repetitive content already present in the document. Training sequence-to-sequence models to produce high-quality, novel, and relevant tokens remains a significant challenge. To address these issues, we introduce a novel approach, QGDPO, that employs direct preference optimization (DPO) to guide the generation process. We first fine-tune a base sequence-to-sequence model and subsequently utilize a relevance model to score its predictions. Based on these scores, we construct pairs of winning and losing predictions as relevance preferences for the DPO training. Furthermore, we enhance our pipeline by using the relevance model to filter out poor predictions, retaining only the most relevant generated content for indexing. QGDPO effectively eliminates 50% of irrelevant predictions comparing against Doc2Query baselines, while the relevance filter removes an additional 14.61%. This feature has been successfully deployed to production on this http URL for full traffic, with a substantial improvement in relevance and user engagement.

26. 【2610.04302】From Valid to Useful: Post-Verification Acquisition for Recursive Self-Improving Recommendation

链接:https://arxiv.org/abs/2610.04302

作者:Tonmoy Hasan,Taylor Foust,Shao Tang,Leonardo Neves,Aman Gupta,Hiroto Udagawa,Helder Dias,Daniel Silva,Rohan Ramanath

类目:Information Retrieval (cs.IR)

关键词:generate synthetic interaction, verified sequences, Sequential recommenders, synthetic interaction sequences, sequences

备注: 15 pages

点击查看摘要

Abstract:Sequential recommenders can generate synthetic interaction sequences and retrain on the augmented corpus in a recursive self-improvement loop. To limit error accumulation, current methods verify each generated sequence remains predictive of the user's real interactions and discard those that drift away from it. Verification does not, however, determine which verified sequences should train the next model. With every verified sequence used for training, source sequences yielding more verified sequences or longer continuations have more influence, although neither quantity indicates how much those sequences will help the next model. We formulate the decision of which verified sequences are used to train the next model as \emph{post-verification acquisition} and introduce {\bf Disagreement-Aware Recursive Self-Improving Recommendation (DA-RSIR)}. DA-RSIR caps each source sequence's contribution and ranks its verified sequences by how much the model's predictions disagree over their augmented interactions. It uses a score derived from Bayesian Active Learning by Disagreement (BALD) and estimated with Monte Carlo (MC) dropout. DA-RSIR requires no extra labels, teacher model, or quality scorer. Across four datasets, three recommender models, and two metrics, it improves on the retain-all approach in all $24$ comparisons and attains the highest mean in $23$ of $24$ overall; the aggregate improvement is statistically significant on both metrics. A single DA-RSIR round exceeds the retain-all approach's best gain over five recursive rounds. These findings establish post-verification acquisition as a separate control point in recursive self-improvement, separating which sequences pass verification from which verified sequences are used to train the next model.

Comments:
15 pages

Subjects:

Information Retrieval (cs.IR)

Cite as:
arXiv:2610.04302 [cs.IR]

(or
arXiv:2610.04302v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.04302

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
27. 【2610.04263】Multimodal Dual-Encoder Retrieval for Automated ICD Coding

链接:https://arxiv.org/abs/2610.04263

作者:Abhinav Bohra,Anuj Bohra

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Accurate International Classification, Accurate International, International Classification, crucial for large-scale, ICD

备注: 6 pages, 2 figures, 1 table

点击查看摘要

Abstract:Accurate International Classification of Diseases (ICD) coding is crucial for large-scale clinical research, documentation, and billing. There are three primary problems with current ICD prediction methods: (1) They are unable to comprehend multimodal patient data because they rely on either structured EHR data or unstructured clinical notes. (2) They also struggle with scalability to a larger amount of ICD codes (9K+ codes in ICD-9), as traditional classifiers need dense output layers and often do not generalize well to long tail rare diseases. (3) They lack transparency for clinical use. To address these challenges, this research proposes a two-stage framework that first retrieves ICD codes using a multimodal dual-encoder retrieval model, where structured and unstructured patient data are integrated through gated fusion. The second stage refines the top-k retrieved candidates with an LLM-based re-ranker that provides ranked codes with clinically relevant explanations. Our experiments show that the proposed approach improves Micro-F1 and Precision over a multimodal dual-fusion classifier baseline. These improvements demonstrate that combining a gated multimodal retrieval system with LLM-based re-ranking is a practical alternative to dense multi-label classification for automated ICD coding.

28. 【2610.04047】Periscope: Extending Frozen Language Models Beyond Their Context Window

链接:https://arxiv.org/abs/2610.04047

作者:Mohamed Eltahir,Anas Obayd,Raed Rashid,Abdulrahman Alghamdi,Abdulrahman Mousa,Abdallah Ahmed,Tanveer Hussain,Naeemullah Khan

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:quadratic forward pass, quadratic forward, loses accuracy, accuracy with length, length before reaching

备注:

点击查看摘要

Abstract:A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.

29. 【2610.04028】Learning Subject-Specific Anatomical Representations via Manifold Expansion: Application to Accelerated Multi-Contrast MRI

链接:https://arxiv.org/abs/2610.04028

作者:Ruimin Feng,Wanyu Bian,Albert Jang,Zachary Stewart,Fang Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:Clinical MRI routinely, MRI routinely acquires, complementary tissue characterization, routinely acquires multiple, acquires multiple contrast-weighted

备注:

点击查看摘要

Abstract:Clinical MRI routinely acquires multiple contrast-weighted images of the same anatomy for complementary tissue characterization. However, current accelerated MRI methods typically reconstruct each contrast independently, without fully exploiting shared anatomical information. This work aims to learn anatomical representations invariant to contrast-dependent appearance for reconstruction of accelerated multi-contrast MRI. We propose MAX (MAnifold eXpansion), a subject-specific framework that learns anatomical representations from a single fully sampled reference contrast. To address the under-constrained separation of shared anatomy and contrast-dependent components from a single image, MAX expands the multi-contrast manifold using anatomy-preserving intensity augmentations. A disentangled implicit neural representation models augmented samples using shared spatial coordinates for anatomy and spatially invariant coordinates for contrast appearance. The learned anatomical representation is then fixed, with the contrast representation adapted to the undersampled target data, followed by unrolled refinement. Theoretical analyses further provide insight into the disentangled representation learning and explain how the learned anatomical representation improves the target contrast reconstruction. At R = 8 for brain MRI and R = 6 for knee MRI, MAX achieves the highest mean PSNR and SSIM across all tasks, improving PSNR by more than 1 dB over the strongest baseline for both brain contrasts. MAX more faithfully recovers subtle anatomical and pathological structures and remains robust to inter-contrast motion, structural heterogeneity between reference and target contrasts, and measurement noise. Therefore, MAX provides a general strategy for leveraging high-quality reference scans in accelerated MRI and has the potential to be extended to other reference-assisted MRI inverse problems.

30. 【2610.03958】FICO: Find-Then-Compute for Corpus-Level Spreadsheet Question Answering

链接:https://arxiv.org/abs/2610.03958

作者:Sandarsita Guntupalli,Lu He,Kang Li

类目:Information Retrieval (cs.IR)

关键词:spreadsheet collections requires, collections requires finding, Structured Query Language, constrained Structured Query, answering over spreadsheet

备注:

点击查看摘要

Abstract:Question answering over spreadsheet collections requires finding the correct workbook and computing over complete tables. We introduce Find-then-Compute (FiCo), which retrieves document summaries, disambiguates similar workbooks, and executes constrained Structured Query Language (SQL) over the selected full table. On DataBench (80 datasets, 1,810 questions), FiCo reaches 76.2% accuracy: 9.9 points above a strong TableRAG-style baseline on the same frozen workbook choices (66.3%) under the tracks' prespecified evaluators, and 63.4 points above prefix RAG (12.8%). On 508 MiMoTable questions, FiCo reaches 79.7%, versus 22.2% for prefix RAG. Giving the strong baseline the gold workbook raises it from 66.3% to 76.3%, exposing a 10.0-point source-selection cost under fixed compute. Despite 95.1% document recall and 98.5% executable SQL, only 81.3% of questions execute on the gold workbook. FiCo's advantage comes from integrating semantic source selection with exact, schema-grounded computation.

31. 【2610.03923】Learning Robust Personalized Prompts for LLM-Driven Sequential Recommendation

链接:https://arxiv.org/abs/2610.03923

作者:Xiaolin Zheng,Qiyong Zhong,Jiajie Su,Xiang Chen

类目:Information Retrieval (cs.IR)

关键词:LLM-driven sequential recommendation, sequential recommendation formulates, recommendation formulates next-item, formulates next-item prediction, autoregressive generation conditioned

备注:

点击查看摘要

Abstract:LLM-driven sequential recommendation formulates next-item prediction as autoregressive generation conditioned on natural-language prompts. However, minor wording changes in semantically equivalent prompts can cause substantial performance fluctuations, undermining robustness and requiring costly manual prompt engineering. Continuous prompt learning reduces template dependence but faces two interacting challenges: shared task-level instructions lack user-specific reasoning guidance, while gradient updates can push continuous prompts outside the LLM's effective semantic space. Injecting personalized signals can further amplify this semantic drift. To address these challenges, we propose LRPRec, a learnable prompting framework that initializes continuous instruction prompts from discrete templates and introduces two complementary mechanisms. Personalized prompt injection encodes user behavior into a preference embedding and additively injects it into shared prompts, enabling parameter-efficient user-level adaptation. A semantic drift constraint regularizes the shared prompts within a trust region around their initialization anchors to preserve semantic validity during optimization. By constraining the shared component while allowing additive personalization, LRPRec decouples stability from expressiveness. Extensive experiments on three benchmark datasets demonstrate consistent improvements over strong baselines while eliminating the need for manual tuning of background and task inference templates.

32. 【2610.03802】Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models

链接:https://arxiv.org/abs/2610.03802

作者:Quim Motger,Carlota Catot,Marc Oriol

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Software Engineering (cs.SE)

关键词:polarity-based opinion mining, including emotionally informed, feature-oriented feedback analysis, emotionally informed issue, informed issue prioritisation

备注:

点击查看摘要

Abstract:Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied. Objectives: Building on a previously published annotation framework and human-labelled ground truth adapted from Plutchik's taxonomy, this paper investigates how large language models can be leveraged for automatic multi-label emotion classification under severe class imbalance. Methods: We compare encoder-only fine-tuning under multi-label and binary-ensemble formulations, decoder-only zero- and few-shot prompting across open-source and proprietary models, and a catalogue of imbalance mitigation strategies (loss reweighting, resampling, generative data augmentation), with the synthetic-review generator and prompting strategy selected via an intrinsic augmentation-utility ranking. Results: Fine-tuned encoders trail the best decoder-only few-shot prompting (macro-F1 0.642) by a wide margin at baseline (multi-label: 0.387; binary ensemble: 0.450); pairing the best multi-label encoder with generative data augmentation and positive-weighted loss closes most of this gap (+0.204) at up to three orders of magnitude lower inference latency than the decoders, with the largest gains on the rarest emotions, from undetected to gains of up to +0.501 F1. Conclusion: Large language models make fine-grained, multi-label emotion classification of app reviews feasible for requirements engineering pipelines, with modest macro-F1, and the best formulation and mitigation strategy are backbone- and formulation-dependent. We release the experimental pipeline, synthetic corpora, and fine-tuned checkpoints for replication and reuse.

33. 【2609.19456】Beyond Private Training: The New Landscape of AI Privacy

链接:https://arxiv.org/abs/2609.19456

作者:Sean Culatana,Kang Li

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Retrieval-augmented systems increasingly, systems increasingly rely, Retrieval-augmented systems, retain deleted items, systems increasingly

备注:

点击查看摘要

Abstract:Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embeddings during graph traversal. We formalize this distinction as output safety versus traversal safety, and introduce TSD-AUDIT, a framework for auditing and enforcing traversal-safe deletion in graph-based approximate nearest-neighbor retrieval. On Faiss IndexHNSWFlat, native filtering leaves the number of distance computations unchanged relative to unfiltered search; at a 70% deletion rate, trace-faithful replay detects deleted-vector scoring in all 100 audited queries. Code inspection of hnswlib's mark_deleted path reveals the same scoring-before-liveness pattern. TSD-AUDIT enforces an alive-before-scoring invariant, repairs connectivity using only live candidates, and emits per-query scored-trace certificates that an independent verifier can check against the deletion snapshot. Under region-targeted deletion, TSD-AUDIT improves Recall@10 over native filtering by 4.3--42.2 percentage points across deletion fractions from 0.5 to 0.9, while remaining comparable under random deletion. These results show that output-only deletion audits can miss process-level exposure: auditing deletion in vector retrieval requires accounting for the vectors scored during search, not only the identifiers returned.

34. 【2601.18267】Orchestrating Specialized Agents for Trustworthy Enterprise RAG

链接:https://arxiv.org/abs/2601.18267

作者:Xincheng You,Qi Sun,Neha Bora,Huayi Li,Shubham Goel,Kang Li,Sean Culatana

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:high-stakes decision settings, enterprise knowledge work, structured Memory Bank, Adaptive Deep Orchestration, require deep synthesis

备注:

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) shows promise for enterprise knowledge work, yet it often underperforms in high-stakes decision settings that require deep synthesis, strict traceability, and recovery from underspecified prompts. One-pass retrieval-and-write pipelines frequently yield shallow summaries, inconsistent grounding, and weak mechanisms for completeness verification. We introduce ADORE (Adaptive Deep Orchestration for Research in Enterprise), an agentic framework that replaces linear retrieval with iterative, user-steered investigation coordinated by a central orchestrator and a set of specialized agents. ADORE's key insight is that a structured Memory Bank (a curated evidence store with explicit claim-evidence linkage and section-level admissible evidence) enables traceable report generation and systematic checks for evidence completeness. Our contributions are threefold: (1) Memory-locked synthesis - report generation is constrained to a structured Memory Bank (Claim-Evidence Graph) with section-level admissible evidence, enabling traceable claims and grounded citations; (2) Evidence-coverage-guided execution - a retrieval-reflection loop audits section-level evidence coverage to trigger targeted follow-up retrieval and terminates via an evidence-driven stopping criterion; (3) Section-packed long-context grounding - section-level packing, pruning, and citation-preserving compression make long-form synthesis feasible under context limits. Across our evaluation suite, ADORE ranks first on DeepResearch Bench (52.65) and achieves the highest head-to-head preference win rate on DeepConsult (77.2%) against commercial systems.

35. 【2610.06702】Optimal compression with quantum retrieval

链接:https://arxiv.org/abs/2610.06702

作者:Shyam Dhamapurkar,Mohit Garg,Manaswi Paraashar,Jaikumar Radhakrishnan

类目:Quantum Physics (quant-ph); Data Structures and Algorithms (cs.DS); Information Retrieval (cs.IR); Information Theory (cs.IT)

关键词:data compression problem, compression problem, data compression, Abstract, logarithmic factor

备注: 13 pages, 3 figures

点击查看摘要

Abstract:We consider the following data compression problem. Given a string $x \in \{0,1\}^m$ of Hamming weight at most $n$, compress it into a shorter string $y \in \{0,1\}^s$ so that any bit $x_i$ of $x$ can be retrieved without any error using at most $t$ quantum queries to the standard oracle encoding of $y$. If queries are allowed to be adaptive we show how optimal compression up to a logarithmic factor can be achieved. If the queries are required to be made non-adaptively, we show schemes whose space is optimal in its dependence on $m$ except for a logarithmic factor, and is at most quadratically worse when compared to the optimum in its dependence on $n$.

计算机视觉

1. 【2610.06852】One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

链接:https://arxiv.org/abs/2610.06852

作者:Shih-Chen Tseng,Chih-Hsuan Chen,Ryan Yang,Hsi-An Chen,Chun-Wei Tuan Mu,Yu-Lun Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:including paper columns, 9:16 phone previews, including paper, paper columns, 16:9 slides

备注: Project page: [this https URL](https://onefigureeverycanvas.vercel.app/)

点击查看摘要

Abstract:Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulate aspect-ratio-adaptive flowchart relayout as a distinct task: given a raster flowchart and a target ratio, produce a structurally faithful, hallucination-free, editable layout. Existing methods fail characteristically: image-to-image models stretch blocks and reject extreme ratios, text-to-image agentic systems hallucinate content, and parse-then-render systems mis-route edges. We propose an agentic pipeline factored into Parse, Style, and Layout stages, each pairing a main agent with a critic that combines deterministic constraint checks with VLM visual feedback so connectivity is explicitly checked and prevented from being silently broken. Outputs are this http URL-editable mxGraph XML. On a curated benchmark of 100 flowcharts at five aspect ratios, evaluated by Gemini 3.1 Pro and validated against human judgments, our method reaches 68.6% Content Fidelity versus 11.2-41.4% for prior work. Project page: this https URL

2. 【2610.06850】InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation

链接:https://arxiv.org/abs/2610.06850

作者:Yucheng Zhang,Sirui Xu,Jinhong Li,Liuyu Bian,Anatulya Nandi,Derek Zhang,Xiangchen Liu,Xueting Li,Umar Iqbal,Yu-Xiong Wang,Liang-Yan Gui

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Captured human-object interactions, Captured human-object, provide rich supervision, rich supervision, humanoid robot

备注: Project Page: [this https URL](https://sirui-xu.github.io/InterMimicGen)

点击查看摘要

Abstract:Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.

3. 【2610.06847】S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

链接:https://arxiv.org/abs/2610.06847

作者:Jeffrey Hu,Daniel Olmeda Reino,Ayush Tewari

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:effectively unlimited in-distribution, unlimited in-distribution data, simple symbolic rules, violate physical laws, procedural generators

备注: Project Page: [this https URL](https://jefequien.github.io/S2PD/)

点击查看摘要

Abstract:Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.

4. 【2610.06844】Learning to Read the Contextual Tokens in Diffusion Transformers

链接:https://arxiv.org/abs/2610.06844

作者:Omer Dahary,Etai Sella,Hadar Averbuch-Elor,Daniel Cohen-Or,Or Patashnik

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Machine Learning (cs.LG)

关键词:Multimodal Diffusion Transformers, Diffusion Transformers, Multimodal Diffusion, jointly process visual, contextual tokens

备注: Project page: [this https URL](https://omer11a.github.io/learning_to_read/)

点击查看摘要

Abstract:Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate contextual tokens into the input space of a frozen Large Language Model (LLM), allowing the LLM to answer questions about the emerging image directly from these hidden representations. Our reader reveals that contextual tokens encode a rich, global representation of the emerging scene: generation-specific semantics, including attributes left underspecified by the prompt, are accessible surprisingly early in denoising, while increasingly fine-grained details become readable over time. Remarkably, this information remains decodable even when the MM-DiT receives an empty prompt, showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself. We further find that generations with more readable contextual representations tend to receive higher human-preference scores. Building on these observations, we introduce Contextual Alignment, a training technique that explicitly reinforces the visual-semantic information encoded in the contextual tokens, improving generation quality and distributional coverage. Together, our results establish contextual tokens as both an interpretable view into the internal dynamics of MM-DiTs and an effective target for improving generative models.

5. 【2610.06837】Anatomy-aware Fine-grained Multimodal Fusion for Laryngopharyngeal Cancer T-Staging Prediction Using CT and Radiology Report

链接:https://arxiv.org/abs/2610.06837

作者:Xingyue Zhao,Yanzhou Su,Fang Zhang,Zhanghexuan Ji,Yirui Wang,Dazhou Guo,Sibo Ju,Yuehua Cheng,Yuzhen Chen,Ming Feng,Le Lu,Tsung-Ying Ho,Jian Wang,Dakai Jin,Na Shen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:guiding personalized treatment, personalized treatment strategies, Accurate T-staging, crucial for guiding, guiding personalized

备注: Accepted by IEEE Transactions on Medical Imaging (IEEE TMI)

点击查看摘要

Abstract:Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.

6. 【2610.06831】UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

链接:https://arxiv.org/abs/2610.06831

作者:David Serrano-Lozano,Duygu Ceylan,Yannick Hold-Geoffroy,Iliyan Georgiev,Javier Vazquez-Corral,Anna Frühstück

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:provide an intuitive, intuitive interface, strength, slider, Sliders provide

备注: Project page: [this https URL](https://color.cvc.uab.cat/unislider)

点击查看摘要

Abstract:Sliders provide an intuitive interface for continuous image editing. In current generative approaches, however, the slider is simply a rescaling of the method's strength parameter, such as an adapter coefficient, a prompt weight, or an interpolation factor. This strength relates poorly to perceptual change. The image can partially revert as the slider moves, long stretches of the range produce no visible difference, and short intervals transform the image abruptly. Remapping the strength could fix this uneven pace, but only if the trajectory is monotone, which current methods do not enforce. We therefore distinguish the slider from the strength, and require perceptual distance from the input to grow linearly with the slider value. We introduce UniSlider, a lightweight LoRA trained on a few-step editing backbone so that its strength approximates this ideal slider. Few-step sampling lets us impose this objective in pixel space without intermediate ground truth, and the backbone's output is preserved at full strength. However, a low-rank adapter cannot make the strength fully uniform. Our slider is thus an inference-time remapping of the strength, obtained by adaptive sampling. Since training optmizes to make the trajectory monotone, this remapping closes the remaining gap without extra training or parameters. On a new benchmark of 300 continuous edits evaluating uniformity, monotonicity, edit fidelity, and identity preservation, UniSlider outperforms all prior methods and is preferred in a user study.

7. 【2610.06825】PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

链接:https://arxiv.org/abs/2610.06825

作者:Yaohui Zhang,Binxu Li,Haoyi Duan,Jiacheng Miao,Yixin Wang,Xinran Du,Chenyue Li,Shilong Liu,Kevin Wu,James Zou

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:reusing published findings, making accurate plot, real scientific figures, Scientific figures, encode quantitative results

备注:

点击查看摘要

Abstract:Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing benchmarks rely largely on synthetic charts or cover only a limited range of chart types. We introduce PlotGround, an automated pipeline for building plot digitization benchmarks from real scientific figures and their author-released source data. PlotGround maps figures to source tables, identifies reconstructable panels, and generates quantitative questions with source-grounded reference values. We use PlotGround to construct PlotGround-1k, a human-verified benchmark of 1,119 questions from 1,066 bioRxiv preprints. Across sixteen multimodal models, the best reaches 87.5% accuracy at a $\pm 5\%$ relative-error tolerance. Tightening the tolerance to $\pm 2\%$ lowers every model's accuracy by 11-24 percentage points, revealing a gap between approximate visual reading and precise quantitative recovery. PlotGround's paired figure-source structure lets us compare how accurately the same values are recovered from figures and from source tables. Providing source tables instead of figures raises a coding agent's accuracy from 90.0% to 97.4% while cutting cost by 72%.

8. 【2610.06814】APDreamer: Transferable Adversarial Patches for World Action Models

链接:https://arxiv.org/abs/2610.06814

作者:Xuanyu Lu,Fengqing Jiang,Kaiyuan Zheng,Yichen Feng,Yaorui Ding,Yuetai Li,Zhen Xiang,Bhaskar Ramasubramanian,Basel Alomair,Luyao Niu,Radha Poovendran

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:general-purpose robotic control, world action models, World models learn, environment will evolve, robotic control

备注: Project Page: [this https URL](https://tapdreamer.github.io)

点击查看摘要

Abstract:World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.8% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 2.1% and 0.8% on two DreamWAM configurations and to 10.0% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.

9. 【2610.06813】Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

链接:https://arxiv.org/abs/2610.06813

作者:Zhimin Shao,Xijun Liu,Zhaoliang Zhang,Yutao Tang,Abhay Yadav,Rama Chellappa,Cheng Peng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent progress, enabled rapid, leveraging learned, priors from vast, spatial data

备注:

点击查看摘要

Abstract:Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.

10. 【2610.06801】MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

链接:https://arxiv.org/abs/2610.06801

作者:Jiarui Chen,Zeqiang Lai,Jiangshan Wang,Ziheng Ouyang,Ye Huang,Xiangyu Yue,Cewu Lu,Chunchao Guo

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:long-sequence generation tasks, primary approach, approach to reducing, reducing the latency, latency of diffusion

备注: 11 pages, 8 figures

点击查看摘要

Abstract:Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a $1.80\times$ denoising speedup on Minimax-H3-Base and a $2.32\times$ speedup on 3D asset generation, both with negligible quality loss.

11. 【2610.06704】Extending Dynamic World Surface Water Mapping to Sentinel-1 with AlphaEarth Embeddings

链接:https://arxiv.org/abs/2610.06704

作者:Rohit Mukherjee,Frederick Policelli,Beth Tellman,TC Chakraborty,Jonathan Giezendanner,Jonathan A. Sullivan,Ning Sun

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Dynamic World, land cover globally, land cover, maps land, cloud-free observations

备注: 8 pages, 3 figures, 5 tables; includes 3 pages of supplementary material

点击查看摘要

Abstract:Dynamic World (DW) maps land use and land cover globally at 10 m from Sentinel-2 (S2) imagery, but only for cloud-free observations, which limits where and when surface water can be mapped. We use the DW water class as weak supervision for a Sentinel-1 (S1) synthetic aperture radar (SAR) model so that DW-like water maps can be produced for every S1 acquisition. Google's AlphaEarth Foundations (AEF) annual embedding supplies spatial context, while S1 backscatter supplies the acquisition-time observation. On 53 globally distributed scenes with independent annotations of 3 m PlanetScope imagery acquired within 48 h of the S1 overpass, the S1-only model already reaches a pooled water intersection over union (IoU) of 0.77, comparable to 0.75 for the operational OPERA DSWx-S1 product, and adding AEF raises it to 0.85. The fused model improves on the S1-only model on 44 of 53 scenes and exceeds OPERA on 48, and on the independent S1S2-Water benchmark it reaches 0.94, compared with 0.87 for OPERA. Optical land-cover products can thus provide scalable training labels for SAR surface water mapping.

12. 【2610.06688】GS-Pool: Object-Level Change Detection in 3D Gaussian Splatting

链接:https://arxiv.org/abs/2610.06688

作者:Boaz Keren-Gil,James Gain,Patrick Marais

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:museums and surveyors, Gaussian Splatting, Factories, Gaussians, space months

备注:

点击查看摘要

Abstract:Factories, museums and surveyors photograph the same space months apart and need to know which objects changed. When each visit is reconstructed with 3D Gaussian Splatting (3DGS), a direct comparison of the two reconstructions does not answer this. Training is stochastic, so two reconstructions of an unchanged space never coincide, and the second visit is often a quick re-scan with far fewer photographs. We propose GS-Pool, which takes two independently reconstructed Gaussian fields of the same space and returns the changed objects in each, together with their masks. SAM2 masks of each visit's photographs are lifted onto the Gaussians that render them and merged into an object pool, so every decision is taken once per object in 3D. We introduce a photographic carrier, the 3DGS training loss of each input reconstruction against the other visit's photographs, backpropagated to the Gaussians that rendered each pixel. We combine it with GS-Diff's geometry and colour terms and our distilled DINOv3 features. This evidence is compared with that of the objects present in both visits, which sets a change threshold for each scene. On PASLCD, GS-Pool reaches mIoU/F1 scores of 0.751/0.846 against 0.644/0.758 for GS-Diff, the strongest prior method, a gain of 17%/12%. Its mIoU is also 36%, 40% and 57% above that of O-SCD, PlenoCI and MV-3DCD, and it reaches 0.855 mIoU on CL-Splats, 33% above MV-3DCD. Each changed object is returned as a set of Gaussians with the evidence behind its decision, which an inspector can review in 3D.

13. 【2610.06687】ChronoWorld: Camera-Controlled Consistent 4D World Generation via Spatiotemporal Cues and Geometric Reflections

链接:https://arxiv.org/abs/2610.06687

作者:Xiaoyu Zhou,Dingwei Xian,Zhenyu Wang,Yajiao Xiong,Yongtao Wang,Ming-Hsuan Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:visually compelling sequences, coherence remains challenging, produce visually compelling, spatiotemporal coherence remains, existing camera-controllable video

备注:

点击查看摘要

Abstract:While existing camera-controllable video generation models can produce visually compelling sequences, preserving intrinsic 4D spatiotemporal coherence remains challenging. To address this limitation, we propose ChronoWorld, an "Observation--State--Reflection" framework that leverages spatiotemporal causal cues and reconstruction priors to generate globally consistent, free-view 4D scenes. Given a context video, we introduce a Spatiotemporal Epipolar Causal Attention mechanism that enforces multi-view epipolar constraints and temporal causality throughout the generation process. In addition, we develop a reconstruction-driven geometric reflection pipeline with a 4D retrieval strategy to enable dynamic self-assessment and correction of generated outputs, improving consistency and accuracy. Extensive experiments show that ChronoWorld achieves state-of-the-art performance in spatiotemporally consistent, cinematic-quality 4D scene generation, with strong generalization and high-fidelity geometry across diverse scenarios.

14. 【2610.06674】Detecting Nighttime Anomalies from NASA Black Marble Using a Generalized Spatio-Temporally Robust Framework of Machine Leaning Ensembles

链接:https://arxiv.org/abs/2610.06674

作者:Srija Chakraborty

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:NASA Black Marble, events including fires, product suite capture, anomalous events including, Marble product suite

备注: 8 pages, 5 figures, 2 tables

点击查看摘要

Abstract:Nighttime lights from NASA's Black Marble product suite capture thermal and light emission signals from anomalous events including fires, volcanic eruptions, and gas flaring. Existing detection approaches rely primarily on thermal bands, limiting sensitivity to weaker signals. We propose a novel machine learning framework that jointly models Black Marble M-band and Day/Night Band (DNB) signals to derive a generalized, spatio-temporally robust ensemble of anomaly detectors. The framework iteratively builds detectors that scale across regions, seasons, anomaly classes, and extends over land and ocean. Detection sets at varying confidence levels are derived based on relevant bands and detector agreement. The approach improves true detection rate while reducing spurious detections and results demonstrate strong generalizability with applications in natural hazard monitoring and energy extraction.

15. 【2610.06672】VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding

链接:https://arxiv.org/abs/2610.06672

作者:Yucheng Liu,Yufei Yin,Mingxiao Feng,Jiajun Deng,Wengang Zhou,Houqiang Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Long-video understanding places, understanding places substantial, places substantial demands, requires retrieving information, retrieving information distributed

备注:

点击查看摘要

Abstract:Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensitive to localization errors, and query-independent memory construction, which may omit question-specific details. We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. The preconstructed memory organizes video content into three levels, capturing global narrative context, event-level temporal structure, and fine-grained relational evidence, respectively. To support coarse-to-fine localization and observation, we assign a specialized agent to each level, keeping retrieval and refinement within a scale-specific context. Guided by the query, these agents revisit relevant video regions and enrich layer-wise memories with targeted multimodal observations. Their refinements are assembled according to the original hierarchy into a composite query-adaptive memory, preserving global context in a compact form while retaining fine-grained evidence along query-relevant branches for final reasoning. Compared with direct GPT-5.5 inference, VideoTapestry achieves absolute accuracy gains of 17.2%, 14.9%, 9.8%, and 7.0% on LVBench, LongVideoBench (Long), Video-MME (Long), and EgoSchema, respectively, achieving the state-of-the-art results among all competitors.

16. 【2610.06663】Cross-dataset harmonization for robust endoscopic image analysis

链接:https://arxiv.org/abs/2610.06663

作者:Romil Imtiaz,Dimitris K. Iakovidis

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:endoscopic image analysis, machine learning, training set, purpose usually underperform, underperform when applied

备注:

点击查看摘要

Abstract:A significant problem in endoscopic image analysis is that the machine learning (ML) models used for this purpose usually underperform when applied on images acquired from endoscopes that are different from those used to acquire the images of their training set. The main difference of the images originating from different endoscopes is their color distributions, which depend both on the image sensors and the light sources used. Although previous studies have highlighted this challenge, to the best of our knowledge it has not been previously explicitly tackled. This study focuses on this problem and proposes very simple but impactful method. It implements a reference-based image harmonization that reduces global appearance differences between endoscopic datasets. Specifically, it extracts global color statistics from a chosen reference dataset in the CIE-Lab color space and applies a statistical channel-wise transformation to map each target image toward the appearance of the images of the reference dataset. The method is evaluated in the context of polyp detection in both flexible colonoscopy and capsule endoscopy datasets using a dataset-level cross validation protocol. The results indicate that the proposed harmonization consistently improves cross-dataset performance up to 30.7%, outperforming relevant baseline and state-of-the-art methods. The results indicate that a substantial part of the generalization gap is driven by low-level appearance variation that can be mitigated without retraining.

17. 【2610.06658】alk Like You: Imitating How You Speak in Real-Time Talking Head Generation

链接:https://arxiv.org/abs/2610.06658

作者:Baiqin Wang,Zhixing Ding,Jijie Li,Jiankuo Zhao,Zhen Lei,Xiangyu Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:consistent lip-shape variations, exhibits unique speaking, person exhibits unique, daily life, exhibits unique

备注: 18 pages,10 figures. Project Page: [this https URL](https://bq-wang0511.github.io/TalkLikeYou/)

点击查看摘要

Abstract:In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: this https URL

18. 【2610.06643】AffordCraft: Scalable Construction of Task-Ready Simulation Assets from Single Images

链接:https://arxiv.org/abs/2610.06643

作者:Haoyun Yang,Xueyang Zhou,Ziyi Xie,Yongchao Chen

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Robot learning, simulator offers, learning in simulation, simulation depends, Robot

备注: 33 pages, 14 figures, 17 tables. Project page: [this https URL](https://affordcraft.github.io) Code: [this https URL](https://github.com/AffordCraft/AffordCraft)

点击查看摘要

Abstract:Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew for every image: generative models predict parts and joints that mostly fail to settle or move in simulation, and general-purpose agents need a long session of model calls for each photograph. AffordCraft builds such an asset from a single RGB image and a task instruction by retrieval instead of generation: it locates the object and the part to operate, selects a matching entry from a library of articulated assets, and fits it to the image while keeping its parts and joints intact. Without any box or mask marking the object, AffordCraft produces a physically valid asset for 1,703 of 2,000 photographs from 31 categories. Five generative methods pass on at most 45% of the same photographs and, at the median, need 10 to 78 times our GPU time per valid asset. On 50 cluttered images, 162 of 237 annotated objects pass the same physical test after automatic detection. Growing the library from 141 to 11,372 entries needs no change to the method and raises category coverage from 46% to 100% and the share of selections with the requested label from 18% to 51%. We also build manipulation tasks from the constructed assets, both with single objects and in composed scenes; policies trained on scripted demonstrations complete both kinds of tasks from initial states unseen in training.

19. 【2610.06617】RealtimeWAM: One-Step Asynchronous World Action Models

链接:https://arxiv.org/abs/2610.06617

作者:Chengtao Lv,Jinyang Du,Shuyi Feng,Yang Yong,Shiqiao Gu,Shunzi Yang,Ruihao Gong,Shen Ren,Tianwei Zhang,Wenya Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)

关键词:incorporate visual representations, World Action Models, guide action prediction, World Action, incorporate visual

备注: The code and checkpoints are available at $\href{ [this https URL](https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam) }{\text{this https URL}}$

点击查看摘要

Abstract:World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, $1\%$ drop) across these benchmarks while delivering significant end-to-end speedup (\eg, $\sim25\times$ on H100). Our code and checkpoints are available via this \href{this https URL}{link}.

20. 【2610.06616】Video Encoders Built on Image Representations

链接:https://arxiv.org/abs/2610.06616

作者:Jusheng Zhang,Wenhao Wang,Longqi Cai,Liangzhe Yuan,Yuxiao Wang,Ming-Hsuan Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:evidence remains accessible, video encoder determines, language model, begin to interact, remains accessible

备注:

点击查看摘要

Abstract:The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve independently computed frame representations but incur a much larger visual-token cost when all image tokens are forwarded. We ask a basic question: whether a compact video encoder can instead be built on image representations. To answer this question, we separate three operations that are often coupled: per-frame representation, cross-frame token allocation, and temporal interaction. A frozen image encoder first produces frame-specific candidates. A question-aware selector then allocates a fixed token budget across frames using relevance, diversity, and cross-frame correspondence, after which a lightweight learned refiner reads neighboring-frame context and writes residual updates only to the retained anchors. This preserves source positions and keeps the visual output at the fixed budget. Across 13 benchmarks and three vision-language backbones, the resulting pathway matches full-image aggregate performance while using only about 28%-35% of its visual tokens. Specifically, on Qwen3-VL-8B, it achieves a 13-benchmark macro-average of 62.75 with 1,535 visual tokens, compared with 62.58 for the full Image pathway at 4,424 tokens and 59.49 for native Conv3D at 2,212 tokens. On Qwen3-VL-32B, it reaches a 13-benchmark macro-average of 66.28, compared with 66.09 for Image, while providing a 2.16x end-to-end speedup. These results show that compact video encoding does not require early temporal mixing: frame-specific evidence can be preserved first, allocated jointly, and temporally contextualized after selection.

21. 【2610.06611】Lens3D: Target-Conditioned Visual Foveation for Fine-Grained 3D Understanding

链接:https://arxiv.org/abs/2610.06611

作者:Junming Huang,Zini Chen,Shuaiying Hou,Chi Wang,Qiang Dai,Weiwei Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, visually salient objects, overlook fine-grained attributes, large language, scene videos

备注:

点击查看摘要

Abstract:Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.

22. 【2610.06602】Multitask Conditional Generative Adversarial Network Enables Automatic Whole Knee Cartilage and Menisci Segmentation and Reliable T1\r{ho} and T2 Quantification Without High-Resolution Morphological Images

链接:https://arxiv.org/abs/2610.06602

作者:Ahmed Tahseen Minhaz,Richard Lartey,Zhiyuan Zhang,Jeehun Kim,Kunio Nakamura,Mingrui Yang,Jiasen Zhang,Weihong Guo,Naveen Subhas,Carl S. Winalski,Xiaojuan Li

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:Early osteoarthritis detection, high-resolution Double Echo, Double Echo Steady-State, traditionally necessitating time-consuming, high-resolution Double

备注:

点击查看摘要

Abstract:Early osteoarthritis detection through quantitative MRI (qMRI) requires accurate cartilage and meniscus segmentation, traditionally necessitating time-consuming, costly 3D high-resolution Double Echo Steady-State (DESS) MRI scans. This study developed a multi-task conditional generative adversarial network (MT-cGAN) to simultaneously synthesize DESS-like images and segment tissues directly from qMRI echo images. This retrospective study evaluated 508 knee MRI volumes from 361 subjects (mean age: $40.4 \pm 12.2$ years; 179 female) across three cohorts. Ground truth segmentation masks were generated from DESS images using a pretrained model with manual correction, and $T_{1\rho}$ and $T_2$ maps were computed from magnetization-prepared angle-modulated partitioned $k$-space spoiled gradient echo snapshots (MAPSS) echo images. MT-cGAN was trained to jointly synthesize DESS-like images and segment cartilage and meniscus directly from echo images. Model performance was evaluated using Dice score for segmentation accuracy and coefficient of variation (CV) for $T_{1\rho}$ and $T_2$ quantification. MT-cGAN achieved the highest segmentation performance, mean Dice score 0.84 (range: 0.80--0.86) across all cartilage and meniscus compartments and significantly outperformed the state-of-the-art conditional GAN model with transfer learning (mean Dice, 0.82; $p 0.001$, Wilcoxon signed-rank test). For relaxometry quantification, MT-cGAN demonstrated the highest consistency with the reference DESS protocol, yielding the lowest CV ($T_{1\rho}$: 1.84%, $T_2$: 1.81%). The proposed MT-cGAN accurately segmented cartilage and menisci while providing reliable $T_{1\rho}$ and $T_2$ quantification directly from echo images. By eliminating the need for separate morphological DESS scans, this workflow reduces required scan times to facilitate the clinical translation of qMRI.

23. 【2610.06598】SimForcing: Distilling Simulation Motion Priors into Real-Domain Robot World Models

链接:https://arxiv.org/abs/2610.06598

作者:Xiaodong Wang,Tianle Li,Chuanxin Song,Junliang Xie,Zhanmi Zhong,Suiying Wu,Peixi Peng

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:realistic visual dynamics, preserving realistic visual, videos remains challenging, Action-conditioned robot world, Action-conditioned robot

备注: Code: [this https URL](https://github.com/Wang-Xiaodong1899/SimForcing)

点击查看摘要

Abstract:Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. \url{this https URL}

24. 【2610.06596】Analysis of SWIR Imaging Detection Performance Under Adverse Environmental Conditions for Autonomous Driving Systems

链接:https://arxiv.org/abs/2610.06596

作者:Rohan Mehra,Alexandre Riffard,Yannis Loumouamou,Mathieu Labussière

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:remain poorly characterized, Short-wave infrared, RGB remain poorly, imaging has emerged, autonomous driving

备注: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:Short-wave infrared (SWIR) imaging has emerged as a promising modality for autonomous driving, yet its practical benefits over RGB remain poorly characterized across diverse conditions. This paper presents a systematic comparative study of paired RGB and SWIR object detection on the RASMD dataset, covering four weather conditions and two real-time detection architectures, with various fine-tunings evaluated against a unified ground truth. Overall, RGB demonstrates comparable or superior performance in most scenarios, while RF-DETR exhibits greater robustness across varying conditions. Beyond aggregate metrics, we propose a sensor-dominance mining framework that combines multi-model agreement with targeted manual inspection to identify scenarios where one sensing modality provides more reliable detections using largely unannotated paired data. This analysis reveals that SWIR offers clear advantages in four safety-critical situations, including windshield glare, water droplets on the windshield, low-contrast object visibility, and long-range vehicle detection. The findings suggest that SWIR should be viewed as a complementary modality that enhances perception in rare but challenging conditions. The datasets will be available upon request, and all code and trained model weights are publicly released at this https URL.

25. 【2610.06594】VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges

链接:https://arxiv.org/abs/2610.06594

作者:Sungjae Choi,Hanna Bae,Sunghyun Baek,Junmo Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Feed-forward visual geometry, visual geometry transformers, Feed-forward visual, VGGT reconstruct dense, single forward pass

备注: Accepted to ACCV 2026

点击查看摘要

Abstract:Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.

26. 【2610.06588】Keepsake: Selective Spatial Memory for Long-Horizon Video Generation

链接:https://arxiv.org/abs/2610.06588

作者:Abdul Mohaimen Al Radi,Kunyang Li,Yuzhang Shang,Mubarak Shah,Yu Tian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Long-horizon camera-controlled video, camera-controlled video generation, video generation relies, maintain scene consistency, Long-horizon camera-controlled

备注:

点击查看摘要

Abstract:Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduce redundancy, but make one-time retention decisions that are never revisited, even as an observation's value changes with the evolving memory bank. Both strategies leave a shared question unresolved: as the generated history evolves, which stored observations should still remain in memory? Our key insight is that the value of a stored observation is not fixed, but relational: it depends on the alternatives currently available in the memory bank. A view supported by many geometrically and visually similar substitutes can be relinquished with little loss of coverage, whereas an observation with few viable alternatives should remain regardless of age. We introduce Keepsake, an online, training-free controller for fixed-capacity spatial memory. At each update, Keepsake constructs a pose-appearance graph over retained and newly generated observations, combining camera-pose proximity with visual similarity. A retention priority jointly captures the number of strong substitutes and the similarity of the closest alternative, allowing Keepsake to continually reassess memory value, preserve observations with little alternative support, and evict highly replaceable ones under a fixed budget. The controller modifies only the persistent-memory update; the host generator, denoising schedule, and retrieval rule remain unchanged. Across MemCam and WorldMem, Keepsake improves FVD and LPIPS under a fixed memory budget. On 180-second MemCam trajectories, it retains only 32 of 5,397 frames while reducing FVD by 35.1%.

27. 【2610.06575】FrontVeg V2: A Training-Free Software Framework for Foreground-Aware Zero-Shot Plant Trait Segmentation in High-Resolution Images of Trellised Crops

链接:https://arxiv.org/abs/2610.06575

作者:Abdoul Djalil Ousseini Hamza,Herearii Metuarea,Corentin Lothod{é}(IRHS-IMHORPHEN,DPT SPE),Morgane Roth(GAFL),Jacem Ben Hamden,Eric Duch{ê}ne(SVQV,BAP),Lionel Ley,David Alletru(UE ARBO),David Rousseau

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:training-free software framework, foregroundaware zero-shot segmentation, Graph-Based Mask Assembly, training-free software, trellised crops

备注:

点击查看摘要

Abstract:FrontVeg V2 is an open-source, training-free software framework for foregroundaware zero-shot segmentation of plant traits in high-resolution images of trellised crops. The pipeline combines monocular depth estimation, automatic foreground extraction using Valley-Aware Depth Thresholding, tiled zero-shot segmentation, Graph-Based Mask Assembly, and geometry-aware fusion. This design enables plant organs and disease symptoms to be segmented while reducing detections arising from neighboring vegetation rows. The current implementation integrates Depth Anything V2 (DAV2) and SAM3 and can be used through both command-line batch processing and a Napari graphical interface. FrontVeg V2 provides a reusable framework for multi-crop, multi-trait digital phenotyping without task-specific model retraining.

28. 【2610.06571】BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI Clinical Reasoning

链接:https://arxiv.org/abs/2610.06571

作者:Qizhen Lan,Mengchen Fan,Hang Zhang,Jingwei Duan,Moule Lin,Jialin Chen,Baocheng Geng,Xiaoqian Jiang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:radiologists compare serial, Brain MRI, Brain MRI interpretation, compare serial studies, MRI

备注: 35 pages. Accepted to NeurIPS 2026

点击查看摘要

Abstract:Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into report-grounded assessments. Existing medical VQA and 3D imaging benchmarks capture important parts of this workflow, but often evaluate brain MRI through isolated images, static volumes, or ungrounded report-style answers, thereby obscuring failures in the evidence chain that support clinical validity. We introduce BrainTRACE, a report-grounded benchmark for evaluating whether vision-language models can trace the evidence structure required for longitudinal brain MRI interpretation. BrainTRACE contains 7,273 scored VQA instances derived from 1,778 longitudinal patients, 7,299 MRI studies, and approximately 29k co-registered 3D MRI sequence volumes. The benchmark is organized by five levels of clinical reasoning, from acquisition recognition to case-level synthesis, and by evidence demands covering longitudinal comparison, report-grounded references, multi-sequence integration, and volumetric spatial evidence. BrainTRACE supports rendered inputs compatible with standard VLM interfaces, a 3D-evidence condition, and a decomposed case-reasoning track that audits six steps in a longitudinal evidence chain. Evaluation of 20 VLM configurations shows that current systems can identify isolated visual cues but rarely compose them into grounded longitudinal interpretations. We release the benchmark specification, evaluation lists, scoring implementation, scoring rubrics, and audit-record format to support reproducible progress in brain MRI VLM evaluation.

29. 【2610.06508】A General Pipeline for Dense Illuminant Estimation via Physically Based Synthetic Data

链接:https://arxiv.org/abs/2610.06508

作者:Luca Cogo,Gianmarco Corti,Simone Bianco,Raimondo Schettini

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:color shifts induced, varying lighting conditions, computational photography, fundamental problem, problem in computational

备注: Accepted at the 34th Color and Imaging Conference (CIC 2026), hosted by the Society for Imaging Science and Technology (IST)

点击查看摘要

Abstract:Illuminant estimation is a fundamental problem in computational photography, as it enables the correction of color shifts induced by varying lighting conditions. While learning-based methods have demonstrated strong performance, their progress is hindered by the limited availability of large-scale datasets with accurate illuminant ground-truth. In this work, we propose a general and reusable pipeline to derive dense illuminant chromaticity maps from physically based 3D-rendered scenes. By repurposing an existing 3D scene collection, our approach enables the systematic generation of pixel-wise illuminant annotations under controlled lighting conditions, effectively lowering the barrier to data acquisition for learning-based illuminant estimation. Using this pipeline, we generate a large-scale synthetic set of 74,321 images, which we employ for pre-training both single- and multi-illuminant estimation models. Extensive experiments with state-of-the-art architectures show that synthetic pre-training consistently improves performance, with gains of up to 28% for single-illuminant estimation and up to 57% for multi-illuminant estimation, particularly in data-scarce regimes. These findings demonstrate that synthetic data generation pipelines offer an effective and scalable solution for the pre-training of illuminant estimation methods.

30. 【2610.06505】Improving Proactive AI Assistance with Hierarchical Procedural Understanding

链接:https://arxiv.org/abs/2610.06505

作者:Jin-Seop Lee,TaeYeon Won,SeongJun Jung,JungHoon Kim,Boyang Albert Li,Jin-Young Park,Jaehong Yoon,Jee-Hyong Lee

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:assistants continuously observe, guidance, task progress, remain silent, assistants continuously

备注: 30 pages

点击查看摘要

Abstract:Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should adapt to task progress and user needs. However, existing datasets either focus on detection-based proactive understanding or provide procedural guidance at a fixed granularity. Fixed-granularity guidance provides limited information about fine-grained progress and broader procedural context, making it difficult to determine completion and adapt guidance granularity. To address these limitations, we introduce the ProactiveCoach suite, comprising ProactiveCoach-Instruct for training, ProactiveCoachBench for evaluation, and fine-tuned VLMs with an adaptive guidance system. ProactiveCoach-Instruct provides hierarchically structured guidance at the phase, step, and action levels for learning task progress and procedural context. ProactiveCoachBench evaluates whether models provide appropriate guidance at the right time across different guidance levels and adapt when the requested level changes. We fine-tune pretrained VLMs on ProactiveCoach-Instruct and demonstrate its effectiveness across backbones. Compared with fixed-granularity supervision, hierarchical supervision improves overall performance across backbones by up to 9.6%p. We further build an adaptive guidance system by combining our fine-tuned model with a lightweight guidance router. Without additional fine-tuning, our system outperforms the in-context adaptation baseline by 57.1%p across four guidance-level transitions. Our project page is available at this https URL.

31. 【2610.06503】Harmful Content Generation in Text-to-Image Models: Capabilities and Moderation Limitations

链接:https://arxiv.org/abs/2610.06503

作者:Paschalis Giakoumoglou,Manos Schinas,Symeon Papadopoulos

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:raise concerns, highly realistic imagery, produce highly realistic, content, harmful

备注: Accepted for publication in ACM Transactions on Intelligent Systems and Technology (TIST)

点击查看摘要

Abstract:Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain limited. We present a systematic evaluation of harmful content generation across five open text-to-image models using an automated pipeline that transforms legitimate news captions into unsafe prompts targeting sexually explicit content, violence/gore, harmful stereotypes, self-harm, and hate speech. We evaluate both standard models with built-in safety mechanisms and community fine-tuned variants that bypass content restrictions. A human evaluation of 1,500 generated images shows high harmful-content generation rates: 89.2% for gore-related prompts, 47.6% for sexually explicit content, 43.6% for harmful stereotypes, 46.0% for hate speech, and 34.5% for self-harm, predominantly through graphic violence. Models show substantial capability for generating violent and stereotypical content, while community fine-tuned variants are particularly vulnerable to sexually explicit prompts. Generation quality is largely preserved under harmful prompting, producing imagery of sufficient fidelity to pose risks for disinformation and abuse; FLUX.1-dev produces clearly realistic harmful images in 30.9% of cases. We further evaluate automated moderation systems and find substantial detection gaps that allow unsafe images to evade filtering. Finally, we assess synthetic image detectors and show that models trained only on benign datasets perform worse on explicit content, while more diverse training data improves detection, highlighting semantic distribution gaps in current approaches. These findings expose limitations in current generation safeguards, moderation systems, and synthetic image detection, highlighting the need for stronger defenses against misuse at scale.

32. 【2610.06502】NeuroCBIR: A Fast and Accurate Image Retrieval System for Whole-Brain and Region-Specific MRI

链接:https://arxiv.org/abs/2610.06502

作者:Felix Nieto-del-Amor,Jingru Fu,J.-Sebastian Muehlboeck,Eric Westman,Daniel Ferreira,Rodrigo Moreno

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)

关键词:Content-based image retrieval, image retrieval, Toggle, Accurate Image Retrieval, Image Retrieval System

备注: Neuroimaging, Content-Based Image Retrieval, MRI, Zero-Shot Learning

点击查看摘要

Abstract:Content-based image retrieval (CBIR) in neuroimaging enables the identification of structurally similar brain scans, supporting diagnosis, prognosis, and treatment planning; however, existing methods are often limited to small datasets, single brain regions, or coarse class labels, thereby restricting their clinical utility and generalizability. Here, we present NeuroCBIR, a framework for fast and flexible retrieval of both whole-brain and region-specific 3D T1w MRI scans. A total of 103 cortical and subcortical regions are extracted to enable both whole-brain and region-level queries. NeuroCBIR leverages latent representations learned by a variational autoencoder (VAE) combined with contrastive learning, producing scan-specific embeddings that capture anatomical patterns. These embeddings were evaluated for subject re-identification, zero-shot age prediction, and zero-shot multi-class pathology stratification. Re-identification performance was high across both whole-brain and brain-region levels (mean average precision across the top-5 retrieved images (mAP@5) = 98.4%), with robust generalization across datasets and acquisition conditions. While NeuroCBIR is not trained for age prediction or pathology stratification, zero-shot evaluations for these two tasks demonstrate that the embeddings encode meaningful information for downstream tasks. Embedding extraction on a 4-core CPU required approximately 18.7 s per scan, whereas similarity search was effectively instantaneous (less than 0.01 s). NeuroCBIR is publicly available for brain MRI with more than 26,000 precomputed T1w MRI embeddings. It supports reproducible research, region-specific flexibility, and clinically meaningful personalized diagnostic support. The software is available at this https URL.

Comments:
Neuroimaging, Content-Based Image Retrieval, MRI, Zero-Shot Learning

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)

Cite as:
arXiv:2610.06502 [cs.CV]

(or
arXiv:2610.06502v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.06502

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Felix Nieto-Del-Amor [view email] [v1]
Mon, 5 Oct 2026 15:22:09 UTC (6,440 KB)

Full-text links:
Access Paper:

View a PDF of the paper titled NeuroCBIR: A Fast and Accurate Image Retrieval System for Whole-Brain and Region-Specific MRI, by Felix Nieto-del-Amor and 5 other authorsView PDFHTML (experimental)TeX Source

view license

Current browse context:
cs.CV

prev

|
next

new
|
recent
| 2026-10

Change to browse by:

cs
cs.LG
q-bio
q-bio.NC

References Citations

NASA ADSGoogle Scholar
Semantic Scholar

export BibTeX citation
Loading…

BibTeX formatted citation

loading…

Data provided by:

Bookmark

checked="checked"class=“labs-tab-input”>
Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author
Venue
Institution
Topic

    About arXivLabs

arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.

Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)

mathjaxToggle();

    We gratefully acknowledge support from
    our major funders,
    member institutions, ,
    and all contributors.

About

Help

Contact

Subscribe

Copyright

Privacy

Accessibility

Operational Status (opens in new tab)

Major funding support from

33. 【2610.06494】opology-Informed Prompt-Conditioned Universal Segmentation of Uterine Structures from Ultrasound and MRI

链接:https://arxiv.org/abs/2610.06494

作者:Yongheng Sun,Yuexi Gu,Jingwen Sun,Maureen Kohi,Mingxia Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:complementary clinical information, provide complementary clinical, MRI provide complementary, Multi-structure segmentation, computer-assisted screening

备注: 4 pages, 2 figures, 3 tables. Code: [this https URL](https://github.com/YonghengSun1997/TPUS)

点击查看摘要

Abstract:Multi-structure segmentation of the uterus is important for computer-assisted screening, diagnosis, and treatment planning of uterine diseases, where ultrasound and MRI provide complementary clinical information. However, developing a unified model across these modalities is challenging due to their substantially different image appearances, anatomical contexts, spatial resolutions, and label spaces. Moreover, existing datasets often define different segmentation targets, making joint learning challenging and potentially leading to negative transfer across heterogeneous tasks. To this end, we propose a Topology-informed Prompt-conditioned Universal Segmentation (TPUS) framework for segmenting multiple uterine structures across ultrasound and MRI. TPUS introduces a graph-based multi-dataset backbone comprising modality-specific stems and a modality-shared graph-based encoder-decoder to support modality-sensitive input adaptation, structural feature reasoning, and joint representation learning across heterogeneous uterine segmentation tasks. In addition, TPUS uses task-aware class prompts to condition the segmentation process for different datasets and label spaces, a dynamic convolutional adaptation module to generate task-specific output responses, and a topology-informed loss to encourage anatomically consistent predictions. Experiments on a uterine ultrasound dataset and a T2-weighted uterine myoma MRI dataset demonstrate that TPUS achieves Dice scores of 0.898 and 0.693 on the two held-out test sets, respectively, outperforming several generic and universal segmentation baselines. Source code can be accessed at this https URL.

34. 【2610.06472】MaRO-GS: Mask-Robust Object-Centric Gaussian Splatting from Inconsistent Multi-view Masks

链接:https://arxiv.org/abs/2610.06472

作者:Eunji Kim,Gahyeon Kim,Gianella Cravioto,Dong-hun Lee,Chaewon Moon,Chae-yeong Song,Sang-hyo Park

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, challenge of accurate, address the challenge, target object, object

备注: Accepted to ACCV 2026. Project page: [this https URL](https://eunjikim02.github.io/marogs/)

点击查看摘要

Abstract:We address the challenge of accurate 3D object reconstruction from multi-view images in Gaussian Splatting. Existing object-level 3DGS methods reconstruct the entire scene rather than directly optimizing the target object, even when only the target object is needed, which incurs substantial computational overhead. They also rely on 2D segmentation masks to associate Gaussians with objects, but these masks are often inconsistent across views. Such inconsistencies corrupt Gaussian optimization and produce incorrectly supervised Gaussians that degrade object reconstruction fidelity. To overcome these limitations, we propose MaRO-GS, a 3DGS framework that directly optimizes target-object Gaussians from object-masked multi-view images and remains robust to inconsistent supervision. For reliable supervision, mask-reliability view filtering excludes unreliable views. Object-supported Gaussian density control suppresses Gaussians irrelevant to the target object and prevents background densification, while Silhouette-aligned Object Loss maintains object-focused optimization. Extensive experiments across diverse datasets demonstrate that MaRO-GS improves PSNR, segmentation accuracy, and computational efficiency, with the largest PSNR gain of 2.05 dB on the small-object LERF-Mask dataset.

35. 【2610.06423】oward Reliable Infant Pose Estimation: A Training-Dynamics Approach to Noisy Annotation Detection

链接:https://arxiv.org/abs/2610.06423

作者:Emanuele Cardinale,Sara Moccia,Alessandro Cacciatore,Lucia Migliorelli

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Spontaneous movement analysis, clinically relevant motion, relevant motion biomarkers, motion biomarkers directly, infants relies increasingly

备注:

点击查看摘要

Abstract:Spontaneous movement analysis in preterm infants relies increasingly on markerless pose estimation (PE) to derive clinically relevant motion biomarkers directly from video recordings. Training accurate infant PE models requires large sets of manually annotated keypoints, and human annotation is inherently prone to error. Noisy keypoints (i.e., keypoints mislocalized with respect to their true anatomical position) are especially problematic in this clinical setting, since they can propagate as artificial artifacts into the reconstructed joint trajectories. Building on the small-loss hypothesis and training-dynamics-based sample selection established in the noisy-label learning literature, we propose a novel framework for detecting noisy keypoint annotations. A hybrid convolutional-attention model is trained to predict the anatomical category of each keypoint from its spatial coordinates and local visual features; the resulting cross-entropy training dynamics are then used to derive per-keypoint descriptors, which are partitioned into clean and noisy subsets via unsupervised clustering. We validate the approach on NeoPose, a newly collected dataset of 65 hospitalized preterm infants, under two realistic noise scenarios (random positional perturbation and left-right swapping) across multiple noise levels. Results show that the proposed approach achieves an F1-score of up to 91.9% in noisy-keypoint detection. The framework further generalizes to the heterogeneous COCO benchmark, where filtering CE-detected noisy keypoints from the training set also yields measurable improvements (up to 7.4 AP points) in downstream pose estimation accuracy at moderate-to-high noise levels.

36. 【2610.06413】SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs

链接:https://arxiv.org/abs/2610.06413

作者:Rafael Teixeira Sousa,Vinícius Paulo Lopes de Oliveira,Elisa Ayumi Masasi de Oliveira,Luiza Martins de Freitas Cintra,Fernanda Bufon Färber,Igor Dias Aguiar,Julia Yasmim de Almeida Nobre,Arlindo Rodrigues Galvão Filho

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:reflects faithful spatial, prediction reflects faithful, faithful spatial reasoning, report ever-higher accuracy, correct prediction reflects

备注: Accepted at the 2nd Workshop on Embodied Spatial Reasoning (ESR), NeurIPS 2026. 29 pages (8 main), 9 figures, 18 tables. Code and data: [this https URL](https://github.com/spatialchain/SpatialChainBenchmark)

点击查看摘要

Abstract:Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve $\geq$79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking ($\rho$ = 0.88). Data, generation scripts, and evaluation code are released at this https URL.

37. 【2610.06394】Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection

链接:https://arxiv.org/abs/2610.06394

作者:Zhaolin Cai,Huiyu Duan,Liu Yang,Yanjun Qin,Bo Ai,Wei Chen,Xiongkuo Min,Guangtao Zhai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:localize human-object pairs, localize human-object, human-object pairs, aims to localize, pairs and recognize

备注:

点击查看摘要

Abstract:Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad visual-semantic knowledge and versatile perceptual and reasoning capabilities. However, existing approaches largely invoke these capabilities through loosely coordinated inference stages. This fragmented execution restricts the role of interaction hypotheses in guiding visual exploration, leaving key participants overlooked and local ambiguities unresolved. Furthermore, propagating early semantic assumptions through subsequent visual grounding and relation prediction induces self-reinforcing semantic circularity. To resolve these challenges, we propose HarnessHOI, a training-free framework that transforms passive MLLM inference into an active interaction-centric harness. Specifically, we introduce an interaction-guided perception mechanism that projects emerging interaction hypotheses back into the visual space to discover missing participants and refine ambiguous evidence through targeted observation. Furthermore, a relation-agnostic geometric adjudication module reconciles multi-source evidence to establish a unified spatial basis for grounded interaction reasoning across multiple actions and semantic roles. Extensive experiments on HICO-DET and V-COCO demonstrate that HarnessHOI achieves state-of-the-art performance among training-free methods, confirming the effectiveness of the proposed harness for complex interaction understanding. Code will be released upon publication.

38. 【2610.06389】Multi-Task Partially Supervised Learning for Super-Resolution and Semantic Segmentation on Earth Observation data

链接:https://arxiv.org/abs/2610.06389

作者:Hoàng-Ân Lê,Minh-Tan Pham,Solange Lemai-Chenevier,Daniel Greslou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Earth observation context, Earth observation, observation context, Earth, semantic segmentation

备注:

点击查看摘要

Abstract:Super-resolution and semantic segmentation are known to benefit one another, especially in the Earth observation context. However, learning both tasks in a joint model often requires both task annotations, which is impractical and expensive. In this paper, we study the multi-task partially supervised learning paradigm for both tasks, where each example is assumed to have only a single-task annotation. To that end, we examine two multi-task architectural variations, the sequential and shared variants, and then propose a hybrid variant and a re-projection loss to benefit from the shared representation and enforce image quality of super-resolution when training with semantic segmentation. Experiments show favorable results compared to the SOTA sequential variant. Source code will be published at this https URL.

39. 【2610.06378】MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

链接:https://arxiv.org/abs/2610.06378

作者:Hang Wang,Chao Shen,Lei Zhang,Zhi-Qi Cheng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:detection increasingly challenging, making generalizable AI-generated, video detection increasingly, making generalizable, increasingly challenging

备注: 18 pages, 4 figures, 19 tables

点击查看摘要

Abstract:The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at this https URL.

40. 【2610.06369】Environmental sensor readings in two crop disease image datasets identify the session in which each image was taken

链接:https://arxiv.org/abs/2610.06369

作者:Sungwoo Kang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Integrating environmental sensor, Integrating environmental, disease classification accuracy, environmental sensor data, classification accuracy

备注:

点击查看摘要

Abstract:Integrating environmental sensor data with leaf imagery is widely reported to boost crop disease classification accuracy. In this work, we reveal that these reported gains are often artifacts of dataset construction: because a single sensor reading is shared across many images collected in a single session (one farm on one date), multimodal networks can predict disease simply by memorizing session identities. Analyzing two widely used Korean datasets, the Crop Disease Diagnosis (CDD) benchmark and an AI Hub pest/disease dataset, we demonstrate that nearly all images share sensor values, with 91.9% of CDD test images having exact sensor duplicates in the training set. Remarkably, an image-free classifier given only timestamps matches or exceeds sensor-driven predictions across all seven evaluated crops, and matches the published macro-F1 of a state-of-the-art CDD fusion model. These results indicate that performance gains on standard random splits cannot be disentangled from session leakage. We propose that multimodal crop studies must evaluate on session-held-out splits and report performance against sensor-free date-time baselines to ensure genuine generalization.

41. 【2610.06349】KineWorld: Action-Induced Transport Fields for Embodied World Modeling

链接:https://arxiv.org/abs/2610.06349

作者:Ziying Song,Yuchen Liu,Zhuoran Xu,Ziyang Liu,Jian Jin,Jiangtao Su,Haibao Yu,Lei Yang,Yuanpei Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:world models predict, actions before execution, Embodied world models, consequences of candidate, candidate actions

备注: 36 pages. Project page and code: [this https URL](https://modaxiansheng.github.io/KineWorld/)

点击查看摘要

Abstract:Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Even with explicit motion conditioning, these objectives can underemphasize spatially sparse changes that are critical to interaction. We propose KineWorld, a transport-aware world-modeling framework that extends robot kinematics from motion conditioning to the spatial allocation of generative supervision. Kinematic Transport Lifting (KTL) constructs renderer-derived, camera-aligned transport fields from commanded robot motion. Transport-Aware World Diffusion (TAWD) calibrates their motion support on the video-latent grid and reweights future-RGB flow matching through a normalized mixture of uniform and transport-focused distributions. We train KineWorld using ALOHA-AgileX bimanual manipulation data from RoboTwin 2.0. KineWorld achieves an EWMScore-P of 68.95 in single-view evaluation and a TWB-Score of 54.82 in multi-view evaluation. These results support a shift from appearance fitting toward action-consequence modeling for embodied decision-making.

42. 【2610.06342】MeSD: Multi-Evidence Self-Distillation for VideoLLM

链接:https://arxiv.org/abs/2610.06342

作者:Weijie Zhu,Han Fang,Hanyu Fu,Yuzhe Zhang,Xin Wei,Zhaoyan Pan,Feiran Liu,Xunjie Jin,Hongbo Sun,Zhiyu Lin,Tianyi Gao,Tianyi Ding,Ye Yuan,Zhongjiang He,Hao Sun,Zhiheng Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:sequence-level rewards offer, rewards offer limited, reliable outcome supervision, offer limited token-level, sequence-level rewards

备注:

点击查看摘要

Abstract:While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.

43. 【2610.06339】BabelFake: A Multilingual Audio-Visual DeepFake Benchmark

链接:https://arxiv.org/abs/2610.06339

作者:Carlotta Segna,Joel Tschesche,Anna Rohrbach

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reflect diverse linguistic, diverse linguistic contexts, Reliable and practical, reflect diverse, diverse linguistic

备注: 25 pages (8 main paper + ack), 25 pages total, 8 figures, under submission

点击查看摘要

Abstract:Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.

44. 【2610.06334】SPIN: Image Immunization Against Diffusion Editing via Single-Step Projection in Stochastic Neighborhoods

链接:https://arxiv.org/abs/2610.06334

作者:Fengming Gu,Jie Zhang,Zhongqi Wang,Qiankun Li,Shiguang Shan,Xilin Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:greatly advanced instruction-guided, Diffusion models, unauthorized image manipulation, advanced instruction-guided image, models have greatly

备注:

点击查看摘要

Abstract:Diffusion models have greatly advanced instruction-guided image editing, while also raising concerns about unauthorized image manipulation. Image immunization addresses this risk by adding imperceptible perturbations to an input image to disrupt subsequent edits. Since editing requests are unknown at image release, protection should remain effective beyond the instruction used to construct the perturbation. Existing immunization methods either require costly full-trajectory backpropagation or use intermediate objectives whose effects may be weakened by subsequent denoising. Meanwhile, a single inference path provides limited feedback about alternative denoising continuations. To address these challenges, we propose \textsc{SPIN}, a framework for image immunization via one-step projection over local stochastic trajectory neighborhoods. Starting from an early denoising state, \textsc{SPIN} generates stochastic neighboring states under the same instruction and predicts their clean latents through one-step projection without full unrolling. We then optimize a bounded input perturbation to maximize the average deviation of these predictions from a clean-edit reference, encouraging the perturbation to disrupt multiple possible editing outcomes. Experiments on two image editors demonstrate substantial gains in protection performance, with \textsc{SPIN} outperforming compared methods across all six metrics under seen instructions and in the more challenging unseen instruction setting.

45. 【2610.06327】Dual Variational Autoencoders for Efficient Sim-to-Real Transfer in Low-Cost Robotic Navigation

链接:https://arxiv.org/abs/2610.06327

作者:Álvaro Díez(Department of Computer Science and Artificial Intelligence, University of Alicante),Fidel Aznar(Department of Computer Science and Artificial Intelligence, University of Alicante)

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Vision-based autonomous navigation, Vision-based autonomous, real-world operational conditions, simulated training environments, fundamental challenge

备注: 30 pages, 13 figures. Published in Image and Vision Computing under a CC BY 4.0 license

点击查看摘要

Abstract:Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffective, while training exclusively on real data is impractical. We propose a hybrid transfer learning framework that effectively bridges the sim-to-real gap by combining domain randomization with feature-level domain adaptation. Our method employs a dual convolutional variational autoencoder architecture with a shared decoder, trained on an extensive set of 45225 simulated images and a minimal set of only 4556 real-world samples. This architecture learns a compact, common latent representation space that aligns the distributions of both domains. The adaptation process is further enhanced by two complementary data augmentation techniques designed to expand the limited real-world data. Experimental evaluation demonstrates that our method achieves an average success rate of almost 91% on image classification tasks for real-world indoor navigation, significantly outperforming both simulation-only and real-world-only training. We validate these findings through a direct, real-world deployment, where the proposed policy successfully guides a low-cost robot in a reactive exploration task. Furthermore, we validate the model's efficiency through a rigorous computational estimation, confirming its suitability for resource-constrained embedded platforms such as the Raspberry Pi 4 and NVIDIA Jetson Nano. This work presents a practical solution for developing effective and efficient navigation policies for low-cost robotic systems.

46. 【2610.06324】Readout Blindness: VLM Scores Miss the Spatial Direction Their Frozen Encoders Retain

链接:https://arxiv.org/abs/2610.06324

作者:Guangyuan Li,Tianming Du,Yan Jiang,Bihan Wen,Jiancheng Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:CLIP-like vision-language models, directed spatial relations, vision-language models remain, CLIP-like vision-language, multimodal systems

备注:

点击查看摘要

Abstract:CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.

47. 【2610.06318】Wiring Matters: Injection Topology and Initialization of Affordance Heads in Vision-Language-Action Policies

链接:https://arxiv.org/abs/2610.06318

作者:Zijian An,Linhan Wang,Jiayan Wang,Shijie Geng,Ran Yang,Yiming Feng,Lifeng Zhou

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Dense affordance supervision, appealing auxiliary signal, severely damage instruction, Dense affordance, appealing auxiliary

备注: 8 pages, 4 figures, 2 tables

点击查看摘要

Abstract:Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.

48. 【2610.06293】VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

链接:https://arxiv.org/abs/2610.06293

作者:Qiutong Chen,Yuchan Guo,Zhenlong Yuan,Haobo Yang,Fangfang Lin,Xinyi Long,Yin Wang,Zijian Song,Rui Lan,Shi Qiu,Boyuan Pan,Yang Luo,Yuyin Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, demonstrated remarkable potential

备注:

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.

49. 【2610.06260】CentriQ: Calibration-Free Quantization of Diffusion Transformers via Exact Mean Centering

链接:https://arxiv.org/abs/2610.06260

作者:Nataša Jovanović,Mathieu Salzmann,Saqib Javed

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:cost limits deployment, Diffusion transformers, sampling cost limits, limits deployment, sampling cost

备注: Code and project page will be released soon

点击查看摘要

Abstract:Diffusion transformers (DiTs) achieve state-of-the-art image generation, but their sampling cost limits deployment. Quantizing both weights and activations to 4 bits reduces this cost, yet existing methods fall short in one of two ways. Calibration-based methods are tied to a specific checkpoint and prompt distribution, whereas data-free Hadamard rotation, effective for LLMs, loses quality on DiTs. We show that this loss has a structural cause. Adaptive layer-norm conditioning adds a per-token mean to the activations, and at the widths of the evaluated DiTs, the Hadamard rotations used by data-free methods cannot spread this mean uniformly across coordinates. A single dominant direction therefore survives the rotation and sets the quantization range. We introduce CentriQ, a calibration-free quantizer that centers each token before rotation and restores the mean exactly through a rank-1 full-precision branch, so that per-token scales follow in closed form without data. Weights are fitted under a robust $\ell_p$ objective that tracks the dense mode of each group and discounts heavy tails. Across three DiTs, CentriQ matches the quality of calibrated SVDQuant at 4 bits, whereas calibration-free weight quantizers with plain per-token activation quantization collapse or degrade substantially. CentriQ outperforms the strongest calibration-free method reported to date at 2-bit weights. It is also the first calibration-free method to retain usable image quality at 2-bit activations.

50. 【2610.06234】Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning

链接:https://arxiv.org/abs/2610.06234

作者:Lingyu Shen,Wei Tang,Fakhri Karray,Min-Ling Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Multi-instance partial-label learning, Multi-instance partial-label, addresses inexact supervision, addresses inexact, Multi-instance

备注:

点击查看摘要

Abstract:Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.

51. 【2610.06226】LeAVJEPA: A Minimalist Architecture for Audio-Visual Self-Supervised Learning

链接:https://arxiv.org/abs/2610.06226

作者:Benjamin Robson,Santeri Mentu,Wenshuai Zhao,Arno Solin

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:EMA target encoders, Prior audio-visual self-supervised, self-supervised learning methods, learning methods rely, EMA target

备注:

点击查看摘要

Abstract:Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.

52. 【2610.06221】Frequency-Decoupled Diffusion Guidance for Non-Blind Image Deblurring

链接:https://arxiv.org/abs/2610.06221

作者:Sihan Wang,Jinshu Huang,Haibin Su,Yunhua Xue

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Pretrained diffusion models, Pretrained diffusion, powerful image priors, provide powerful image, diffusion models provide

备注: 31 pages, 12 figures. Project page: [this https URL](https://github.com/Sea-serpents/frequency-decoupled-diffusion-guidance)

点击查看摘要

Abstract:Pretrained diffusion models provide powerful image priors for training-free posterior sampling in image restoration. To guide this sampling process, frequency-aware methods progressively incorporate measurement information across frequency bands, facilitating coarse-to-fine reconstruction. However, existing methods typically do not explicitly separate frequency activation from degradation-induced attenuation, leaving attenuation differences among inactive frequencies insufficiently modeled. In this work, we propose frequency-decoupled posterior guidance to separate frequency activation from attenuation-aware spectral regularization. Specifically, a progressive low-to-high frequency schedule determines the active measurement band, while a kernel-derived attenuation map defines a selective spectral prior over inactive components. To stabilize the sampling process, we also introduce a local trajectory regularizer that suppresses spatially irregular state-to-clean deviations. For a fixed endpoint energy, we provide a KL-regularized path-space interpretation. In practice, we construct time-dependent guidance through local energy corrections using a Tweedie plug-in approximation. Experiments on natural-image benchmarks demonstrate strong PSNR and SSIM performance across challenging non-blind deblurring settings, even at higher measurement noise levels.

53. 【2610.06210】MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting

链接:https://arxiv.org/abs/2610.06210

作者:Yiming Xu,Hao Cheng,Monika Sester

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Autoregressive generation, local coordinate frames, natural for language, trajectory forecasting lacks, clean token

备注: Accepted at NeurIPS 2026. Camera-ready version

点击查看摘要

Abstract:Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that evolve with the predicted trajectory. We present MoCAR (Motion-code Coordinate-aware AutoRegression), a decoder-only framework that casts trajectory forecasting as next-code prediction in a coordinate-aware continuous latent space. MoCAR learns a continuous motion-code space from endpoint-normalized trajectory segments, where each code jointly captures local trajectory geometry and the reference-frame transition induced by that segment. Historical motion codes are used as a teacher-forced prefix, future codes are generated autoregressively under temporal, map, agent, and mode interactions, and predicted codes persist in latent memory while decoded endpoints update the local scene context. This enables rollout without trajectory-space re-tokenization, trajectory queries, goal candidates, or proposal-and-refinement pipelines. On Argoverse (AV) benchmarks, MoCAR achieves top-tier performance with a simple single-stage architecture, transfers strongly from AV2 to AV1 in zero-shot evaluation, and improves on turn-heavy scenarios. Ablations confirm that the learned continuous motion-code space, latent alignment, weak KL regularization, and joint tokenizer-predictor optimization are essential for stable latent autoregression.

54. 【2610.06196】EORestore-Agent: Fidelity-Guided Agentic Restoration of Remote Sensing Images with Composite Degradations

链接:https://arxiv.org/abs/2610.06196

作者:Heli Qi,Zeqi Zhou,Jingjun Yi,Kunyi Liu,Ziyang Lihe,Junjue Wang,Osamu Yoshie,Naoto Yokoya

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:low resolution coexist, low light, low resolution, resolution coexist, Remote sensing

备注: 20 pages, 6 figures, 11 tables, including appendices

点击查看摘要

Abstract:Remote sensing images often carry composite degradations, in which haze, cloud, noise, blur, low light, and low resolution coexist. Restoring them requires deciding which tool to apply, in what order, and when to stop, yet no clean reference is available at inference time to verify these decisions. All-in-one models trained on single degradations converge to a narrow PSNR band as degradations accumulate. To formulate real-world remote sensing restoration as a traceable trajectory, we present EORestore-Agent, which replaces this unmeasurable objective with reference-free, verifiable per-step decisions. A fine-tuned vision-language model reports all residual degradation types, whose tool pools are scored together, so the restoration order emerges from step-wise selection. A relative quality scorer, trained with full-reference supervision on synthetic degradation chains, predicts the changes in PSNR, SSIM, and LPIPS from the current image to each candidate. A step is accepted only when no predicted change is negative and the predicted PSNR gain is positive. Otherwise, the agent keeps the current image. On a synthetic Landsat-8 benchmark with six degradation types, EORestore-Agent improves PSNR by 2.3 to 3.2 dB over the strongest retrained all-in-one baseline on composites of two to six degradations, whereas zero-shot natural-image agents fall below the degraded input in PSNR in 17 of 18 settings. Replacing the learned scorer with no-reference quality differences costs 1.1 to 4.6 dB. The remaining harmful steps are small and cluster near the acceptance threshold. Sentinel-2 examples illustrate transfer to real atmospheric degradation without retraining.

55. 【2610.06195】Loss-Invariant Projections as Passive Probes of Learned Representations

链接:https://arxiv.org/abs/2610.06195

作者:Akshay Chandrasekhar,Pavlo Melnyk

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:final task output, neural networks, passive probes, task output, Learned feature representations

备注:

点击查看摘要

Abstract:Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study this structure using $\textit{passive probes}$ that apply fixed, untrained, property-independent projections to representations as they evolve during training. We motivate this approach through the task of prediction on $S^2$ where equivalent vector and Hermitian parameterizations reveal an additional loss-invariant trace coordinate. This motivates a general construction in which fixed random projections serve as observers of learned features. Because the observer is loss-invariant and independent of the property being studied, changes in accessibility reflect changes in the representation relative to the fixed observer rather than adaptation of the observer itself. We show that ensembles of passive probes can directly reflect task-relevant information such as target alignment. Under our constructions, the accessibility of eventual difficulty evolves differently across tasks. It increases during training in the regression tasks of surface-normal estimation and image inpainting but remains near its initial level in image classification. Comparisons with learned linear probes further show that recoverability and passive accessibility can evolve differently during training. Together, these results show how passive probes can separately characterize changes in representation geometry and the accessibility of eventual task difficulty.

56. 【2610.06167】Bayesian Optimization in Sequence-to-Architecture Latent Space for Zero-Shot NAS

链接:https://arxiv.org/abs/2610.06167

作者:Ondrej Tybl,Lukas Neumann

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Zero-shot Neural Architecture, Neural Architecture Search, Bayesian Optimization objective, Bayesian Optimization, single Bayesian Optimization

备注:

点击查看摘要

Abstract:Zero-shot Neural Architecture Search removes the prohibitive cost of traditional NAS, but its search process is typically based on the evolutionary algorithm (EA); lacking an explicit model of the objective, it often resorts to a near-random search through mutation. Bayesian Optimization offers a principled alternative by modeling the objective and aggregating information across iterations, but scales poorly to the high-dimensional, discrete, graph-structured spaces of modern NAS, restricting its use to only small networks. In this paper, we bring Bayesian Optimization to zero-shot NAS for large-scale architectures by learning a latent space via a Variational Autoencoder trained to reconstruct a novel prefix encoding of architectures and propose a proxy scalarization that combines several zero-shot proxies into a single Bayesian Optimization objective. After only 10,000 iterations of the proposed search algorithm (8 hours on a single GPU), our method found a network architecture which under the given model parameter count constraints achieves state-of-the-art results on three separate tasks -- image classification, object detection and semantic segmentation.

57. 【2610.06147】Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts

链接:https://arxiv.org/abs/2610.06147

作者:Seungmin Oh,Seunghun Kang,Jongbin Ryu

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Vision-language models, achieve strong zero-shot, strong zero-shot transferability, achieve strong, inference time

备注: Accepted for publication in Advances in Neural Information Processing Systems (NeurIPS) 2026

点击查看摘要

Abstract:Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose Test-Time Correction (TTC), a hypothesis-based correction framework guided by a simple principle: hypothesize, reconstruct, correct. Given a test feature and its top-k candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting divergence shift. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This training-free candidate-verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to 2x speedup, over 3x lower CPU memory usage, and up to 1.4x lower GPU memory usage than the lowest-memory training-free baseline.

58. 【2610.06146】Impact of Data Augmentation on Confidence Calibration in Melanoma Classification

链接:https://arxiv.org/abs/2610.06146

作者:Morgan May,Simon Caton,Pierpaolo Dondio

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Logic in Computer Science (cs.LO)

关键词:Accurately quantifying, medical image classification, accurate uncertainty quantification, patient care, quantifying the predictive

备注: Presented at the 30th Conference on Medical Image Understanding and Analysis (MIUA 2026)

点击查看摘要

Abstract:Accurately quantifying the predictive uncertainty or improving model calibration plays an important role in medical image classification, in particular in melanoma diagnosis, where accurate uncertainty quantification can have significant implications for patient care. One of the methods for calibration improvement is data augmentation. In addition, data augmentation as a method for synthetically increasing the size of the dataset has been proven to improve the performance of models trained on imbalanced datasets. However, the impact of data augmentation, as a transformation of a part of the original data, on calibration of models trained on imbalanced datasets, in particular in melanoma classification is under-explored. We train neural networks on SIIM-ISIC 2020 melanoma classification dataset under two conditions: with and without data augmentation, and compare the differences in AUC and expected calibration error (ECE) in both scenarios. Our results shows improvements in uncertainty calibration using different augmentation methods.

59. 【2610.06139】On Impact of Loss Function on the Performance of Neural Networks in Melanoma Diagnosis

链接:https://arxiv.org/abs/2610.06139

作者:Morgan May,Pierpaolo Dondio,Simon Caton

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:skin cancer, patients' survival, deadliest type, type of skin, crucial for patients'

备注: Presented at the 30th Conference on Medical Image Understanding and Analysis (MIUA 2026)

点击查看摘要

Abstract:Melanoma is the deadliest type of skin cancer, whose early diagnosis is crucial for patients' survival. Image classification using deep learning models has shown promising results for melanoma diagnosis. However, the performance of these models on the melanoma datasets such as SIIM-ISIC melanoma classification dataset is a challenge due to the class imbalance. One of the methods to deal with this challenge is using loss function modifications. In this work, we have investigated the effect of different loss functions on the performance of deep neural networks. We trained these networks using focal loss, logit-adjusted softmax cross-entropy (CE) loss, and weighted softmax CE loss, and we report different metrics for evaluating performance and uncertainty calibration. Our results suggest that focal loss delivers a good combination of performance in terms of AUC and uncertainty calibration in terms of expected calibration error (ECE) simultaneously.

60. 【2610.06135】AnchorGen: Anchored Optimization for Customizable Generative 3D Design

链接:https://arxiv.org/abs/2610.06135

作者:Hantao Zhang,Oliver Heinimann,Jieke Wu,Yingxuan You,Emilien Seiler,Pascal Fua

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:learned generative priors, Engineering design, shape optimization relies, generative priors, geometry valid

备注: 29 pages

点击查看摘要

Abstract:Engineering design often starts from a 2D sketch that fixes style and proportions, yet the subsequent 3D shape optimization relies on learned generative priors to keep the geometry valid. However, these priors are agnostic to the sketch: while they admit a valid design by correcting a drifted proposal back to its training distribution, they often correct it towards the high-density region, ignoring the specified design. We introduce \emph{AnchorGen}, a rectified-flow framework trained unconditionally on the concatenated shape and sketch latents of paired data. The learned manifold represents the joint distribution of shape-sketch pairs, so constraining the sketch component restricts the iterate to the sub-manifold of shapes consistent with a target style. Since training employs no conditioning signal, the constraint is imposed at inference: gradient descent optimizes the shape latent to minimize a differentiable drag surrogate, while constraining the sketch latent to remain close to the target sketch via a token-wise cosine penalty. A single model thereby supports design-preserving optimization, dimensionally explicit design edits, and sketch-only synthesis.

61. 【2610.06102】Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation

链接:https://arxiv.org/abs/2610.06102

作者:Fernando Alonso-Fernandez,Kevin Hernandez-Diaz,Jose Maria Buades,Josef Bigun

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:investigate CLIP, CLIP, Abstract, CLIP backbones, periocular

备注: Accepted for publication at 25th International Conference of the Biometrics Special Interest Group, BIOSIG 2026

点击查看摘要

Abstract:We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without task-specific training. For periocular, zero-shot predictions are strongly biased towards males, primarily due to a misaligned decision boundary. Threshold alignment substantially reduces this bias, reaching 85.29% accuracy. Linear SVMs trained on CLIP features provide only marginal gains, with a best periocular accuracy of 86.17%, approximately 2.8% above previous Adience results in the literature. Nevertheless, the gap with full-face performance confirms the greater difficulty of periocular gender estimation

62. 【2610.06094】Anatomy-preserving unpaired cone-beam CT refinement for image-guided radiotherapy using pseudo-label guided diffusion

链接:https://arxiv.org/abs/2610.06094

作者:Qi Lai,Yutong He

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Cone-beam computed tomography, Cone-beam computed, limit image quality, CBCT, beam hardening

备注:

点击查看摘要

Abstract:Cone-beam computed tomography (CBCT) is widely used in image-guided radiotherapy, but scatter, beam hardening, noise, truncation, and other artifacts limit image quality and CT number accuracy. Paired CBCT and CT data are difficult to obtain clinically because of motion, anatomical changes, and acquisition mismatch. We present RefineCBCT, an unpaired CBCT refinement framework that uses pseudo-label guidance and short-step diffusion to reduce artifacts while preserving patient-specific anatomy. RefineCBCT was trained and evaluated on unpaired CBCT and planning CT data from public LUNG TCIA and PELVIC TCIA datasets and compared with representative GAN and diffusion based methods. On LUNG TCIA, it achieved the best results across all metrics, with MAE 19.411, RMSE 62.758, PSNR 30.845 dB, and SSIM 0.931. On PELVIC TCIA, it achieved the best MAE, PSNR, and SSIM, with values of 14.905, 36.671 dB, and 0.876. The refined images showed fewer streaking and shading artifacts, clearer anatomical boundaries, and improved soft tissue uniformity, with line profile and ROI analyses showing closer agreement with planning CT. These results suggest that RefineCBCT provides efficient and effective CBCT refinement under clinically realistic unpaired training conditions and may support more reliable CBCT use in image-guided radiotherapy workflows. Code is publicly available on GitHub, and the evaluated datasets are available from The Cancer Imaging Archive.

63. 【2610.06063】Vision Transformer Ensembles for Panoramic Street Segmentation

链接:https://arxiv.org/abs/2610.06063

作者:Yunus Serhat Bıçakçı

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:support detailed descriptions, unequal training costs, training costs make, model selection difficult, costs make model

备注: 18 pages, 4 figures. Code available at [this https URL](https://github.com/yunusserhat/palmcity_challenge) . Trained models available at [this https URL](https://huggingface.co/yunusserhat/palmcity-eomt-dinov3-large)

点击查看摘要

Abstract:Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.

64. 【2610.06056】ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs

链接:https://arxiv.org/abs/2610.06056

作者:Yijing Du,Xiangcheng Zhan,Shuo Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large Vision-Language Models, Large Vision-Language, frequently suffer, suffer from object, Large

备注: Accepted in EMNLP 2026 Oral

点击查看摘要

Abstract:Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.

65. 【2610.06052】Local2Mesh: Spatially Localized Contour-to-Mesh for Left Ventricular Reconstruction from Sparse 2D Cardiac MRI

链接:https://arxiv.org/abs/2610.06052

作者:Haoyu Wu,Ling Lin,Pascal Lefèvre,Ruizhe Li,Xiaowu Sun

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:imaging remains challenging, remains challenging due, cardiac magnetic resonance, left ventricular, insufficient local spatial

备注: submit ICASSP 2027

点击查看摘要

Abstract:Three-dimensional (3D) left ventricular (LV) reconstruction from sparse cardiac magnetic resonance (CMR) imaging remains challenging due to inter-slice misalignment and insufficient local spatial information between slices. Global aggregation of contour features may obscure local contour-to-surface relationships. We propose Local2Mesh, a spatially localized contour-to-mesh framework that deforms a template mesh to reconstruct 3D LV geometry from sparse 2D contours without 3D mesh annotations. The framework introduces geometry-aware alignment to correct inter-slice misalignment and a plane-aware Local Router that routes contour features to template vertices using vertex-to-plane distances. Local and global contour features then jointly guide graph-based template deformation for 3D LV reconstruction. Experiments on two public datasets, M\Ms-2 and ACDC, demonstrate superior geometric reconstruction and functional estimation over existing methods. Zero-shot transfer from M\Ms-2 to ACDC demonstrates strong cross-dataset generalization. Reconstructed meshes also improve disease classification over sparse contours, supporting their utility for downstream cardiac analysis. These results demonstrate that combining geometry-aware alignment with local contour-to-vertex modeling improves LV reconstruction from sparse 2D contours and supports downstream cardiac analysis. The code is available at \url{this https URL}.

66. 【2610.06041】Representation Disentanglement for Fair Chest X-Ray Diagnosis

链接:https://arxiv.org/abs/2610.06041

作者:Yujie Sun,Ruizhe Li,Xiaowu Sun

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:advanced chest X-ray, chest X-ray, Deep learning, advanced chest, biases in learned

备注: submit to ICASSP 2027

点击查看摘要

Abstract:Deep learning has advanced chest X-ray (CXR) diagnosis, yet demographic biases in learned representations may contribute to performance disparities across intersectional groups. We propose a single-encoder framework combining dual-level decorrelation with prototype-guided cross-group contrastive learning to reduce demographic dependence while accounting for within-class variation. We further propose Demographic Representation Alignment Reduction (DRAR), a new metric that quantifies the reduction in demographic structure within disease representations. The framework is evaluated on four classification tasks using 34,809 CheXpert test images across eight intersectional groups, defined by age, sex and ethnicity. Compared with empirical risk minimization (ERM), our method reduces the mean equalized-odds gap from 15.41\% to 10.86\% and the AUC gap from 5.95\% to 5.01\%. Our method achieves a DRAR of 59.04\% relative to ERM, with only a slight decrease in mean AUC. These results demonstrate that representation disentanglement can reduce demographic bias and improve intersectional fairness. Code is available at \url{this https URL}.

67. 【2610.06035】Casual Flash Lighting for Gaussian Splat Inverse Rendering

链接:https://arxiv.org/abs/2610.06035

作者:Jiamin Xu,Dongheng Wei,Jiarong Zhao,Qi Wang,James Tompkin,Weiwei Xu,Gang Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recovering geometry, photographs is highly, highly ambiguous, Recovering, flash

备注:

点击查看摘要

Abstract:Recovering geometry, materials, and lighting from photographs is highly ambiguous when only static illumination is available. Active-lighting setups reduce the ambiguity but require dark rooms or specialized hardware. Instead, we synergize both static and flash lighting from casual indoor capture, with the flash on or off, each from independent viewpoints. The flash residual constrains albedo and the BRDF, while static lighting captures grazing-angle specular highlights that flash misses. With a 2DGS reconstruction framing, our key contribution is a GS-anchored diffuse field: a hash-encoded MLP is queried at the rasterized 2DGS depth. As it depends only on world position, it is view consistent in 3D and allows the flash residual to drive material decomposition instead of being absorbed by alpha-blending drift across views. At the same time, we render static lighting with deferred shading such that it can also supervise material decomposition. On five synthetic and three real indoor scenes, our method outperforms six recent baselines on diffuse color, albedo and roughness material parameters, and in relighting where PSNR improves by 4.17 dB over the next-best baseline.

68. 【2610.06021】Scalable Minimal-Change Learning for Controllable Image Editing

链接:https://arxiv.org/abs/2610.06021

作者:Shuo Chen,Fengming Huang,Yu Yao,Mingming Gong,Tongliang Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:current methods, methods often make, make unintended changes, make unintended, Image editing

备注: 29 pages, 2 figures. Accepted at NeurIPS 2026. Revised version with additional off-target, reward and SFT control, human evaluation, and OmniGen2 transfer results; clarified related work and experimental scope

点击查看摘要

Abstract:Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: this https URL

69. 【2610.06020】Patch-based Querying Identifies Structures of Interest in Electron Microscopy

链接:https://arxiv.org/abs/2610.06020

作者:Niels Vyncke,Nicolas Nadisic,Yvan Saeys,Aleksandra Pižurica

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:essential sensing technique, Volume electron microscopy, electron microscopy, biomedical research, nanometer-scale resolution

备注: 41 pages, 20 figures, 5 tables. Accepted for publication in Computers in Biology and Medicine

点击查看摘要

Abstract:Volume electron microscopy (vEM) has emerged as an essential sensing technique in biomedical research, allowing the three-dimensional imaging of biological cells and tissues at nanometer-scale resolution. The ability to generate extensive datasets has reached the limitations of downstream analysis processes, which depend significantly on the intervention of human experts for preprocessing and annotation. We propose an efficient and reliable patch-based retrieval framework based on self-supervised learning of local image descriptors to locate self-similar structures in vEM datasets. Given a few manual annotations of a given cellular structure, our method can retrieve similar structures across the EM volume. Our framework is interactive, allowing the human expert to refine the search queries and retrieve relevant image patches quickly and using little labeled data. Experiments on real-world vEM images of biological tissues demonstrate that our framework can reliably identify relevant cellular structures, generalize across different organelles and acquisition modalities, and substantially reduce the search space for downstream analysis.

70. 【2610.06018】Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

链接:https://arxiv.org/abs/2610.06018

作者:Eryk Kołodziejczyk,Alberto Presta,Karol Szurkowski,Michal Byra

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:natural language queries, Spatio-temporal video grounding, STVG models, aims to localize, space and time

备注: Accepted on EMNLP 2026 Findings

点击查看摘要

Abstract:Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.

71. 【2610.06008】Ultrasound Operator Guidance Using World Modeling and Retrieval Based Action Planning

链接:https://arxiv.org/abs/2610.06008

作者:Noortje I.P. Schueler,Hans van Gorp,Ruud J.G. van Sloun

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:knowledge and expertise, quality is heavily, heavily dependent, acquisition quality, operator knowledge

备注: 11 pages, 8 figures, 5 tables

点击查看摘要

Abstract:Ultrasound is widely used, but acquisition quality is heavily dependent on the operator's knowledge and expertise. With demand for examinations outpacing the supply of trained sonographers, operator-guidance systems aim to close this gap by instructing a less trained user how to move the probe toward a target view. In this paper, we propose a retrieval-induced latent transition model for ultrasound acquisition dynamics, formulating ultrasound operator guidance as multi-step planning and retrieval in a world model. Using a V-JEPA 2.1 backbone, observations are first encoded into a latent space where anatomically related views lie close together. We then retrieve similar views from a reference database containing encoded latent states and corresponding probe positions and orientations. Rather than learning a parametric transition function, we directly use physically executed transitions from the database to establish our nonparametric, retrieval-induced transition model that supports receding-horizon planning. At deployment, guidance is generated from the live ultrasound image feed alone, without any probe tracking hardware. Applied to carotid ultrasound, the proposed planner reaches the target view in 86% of retrospective closed-loop episodes, versus 52% and 43% for representative baselines, outperforming both on every target view, including the challenging longitudinal internal and external carotid artery views. A prospective feasibility study on unseen volunteers, run in real time on a CPU using distillation, reaches 83% target-view reachability. Because planning is driven by proximity to any encodable goal latent, the same world model can navigate back to any previously acquired, patient-specific frame, supporting reproducible longitudinal imaging for e.g. perioperative or follow-up monitoring.

72. 【2610.05993】From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval

链接:https://arxiv.org/abs/2610.05993

作者:Yihe Zhao,Songhe Feng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:modification text, desired target image, aims to retrieve, modification text specifies, retrieve a desired

备注: 24 pages, 8 figures, including appendices

点击查看摘要

Abstract:Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the reference state, whereas retrieval candidates depict completed target states. This creates a representation mismatch for zero-shot methods that query pretrained vision-language spaces directly with transformation-oriented language. We study this mismatch and reformulate zero-shot composed image retrieval as target-state reconstruction followed by retrieval. We instantiate this formulation with ASAP-CIR, a training-free framework that reconstructs a static target representation using a frozen multimodal large language model (MLLM). The representation combines multiple holistic descriptions with a variable set of importance-weighted atomic semantics, thereby preserving both overall target identity and fine-grained visual constraints. Retrieval then integrates holistic state alignment, atomic constraint grounding, and calibrated target-state evidence aggregation. A controlled text-only diagnostic shows that target-side static query formulations achieve more reliable retrieval than dynamic composed query formulations, particularly when source-state semantics must be suppressed or transformed. Experiments on FashionIQ, CIRR, and CIRCO further characterize the effectiveness and limitations of this representation principle, with the clearest gains on the multi-target CIRCO benchmark. These results show that how composed intent is represented before retrieval is a consequential design choice, distinct from the choice of retrieval backbone itself.

73. 【2610.05987】Label-Free Coreset Selection with Foundation Models for Efficient Annotation in Computational Pathology

链接:https://arxiv.org/abs/2610.05987

作者:Tuo Yin,Jennifer Dhont

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:improve clinical outcomes, Computational pathology, prognostic accuracy, potential to improve, demonstrated increase

备注: 32 pages, 7 figures

点击查看摘要

Abstract:Computational pathology has the potential to improve clinical outcomes through a demonstrated increase in diagnostic and prognostic accuracy. However, the development and validation of deep learning algorithms still require annotated data, a costly procedure involving expert pathologists who already face critical workforce shortages. Existing coreset selection methods to optimize annotation efforts currently all rely on hyperparameters tuned on natural-image benchmarks that do not transfer to histopathology and are cumbersome to use in clinical practice. In this study, we present GCcore, a novel label-free coreset selection method that embeds every image of a dataset with any pathology foundation model and greedily selects the samples that collectively maximize the global coverage of the embedding space. The proposed method provides a lower-bound guarantee on the global coverage of the returned coreset for any coreset size, while being completely hyperparameter-free and deterministic. We demonstrate GCcore's superior performance over 14 baselines including state-of-the-art methods across 10 tasks and datasets spanning whole slide image classification, tile classification, and tissue segmentation, where it ranks first on six and within the top three on nine, while also demonstrating how existing methods can shift by up to five rank positions depending on their hyperparameter settings. Code is publicly available at this https URL.

74. 【2610.05967】JLD: Perceptual Distance Through A Jacobian Lens

链接:https://arxiv.org/abs/2610.05967

作者:Shreshth Saini,Balu Adsumilli,Alan C. Bovik

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:generation all require, Jacobian Lens Distance, JLD, Jacobian Lens, resolution

备注:

点击查看摘要

Abstract:Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution is doubled, the correlation of DISTS with human scores on TID2013 drops from 0.815 to 0.717. We introduce the Jacobian Lens Distance (JLD), which derives its perceptual geometry from a frozen vision encoder rather than from human labels. JLD combines the locality of early patch features with the perceptual sensitivity captured by later encoder representations. Specifically, we use the encoder Jacobian to identify directions in the early feature space that most strongly affect the encoder output, producing a fixed metric tensor, $E[J^\top J]$, which we call the Jacobian lens. The lens is fitted only once from 100 unlabeled images, taking about 35 seconds. Locally, this construction defines a pullback metric in pixel space, giving JLD a clear geometric interpretation that can be directly analyzed on real images. Across four standard perceptual databases, JLD achieves state-of-the-art performance and consistently outperforms LPIPS, DISTS, PieAPP, and DreamSim. JLD is also robust to changes in image resolution, on TID2013, its lens-term correlation remains nearly unchanged when the resolution is doubled, decreasing only from 0.850 to 0.845. We further introduce JLD-fast, which is $4\times$ faster than LPIPS-VGG while achieving a mean correlation of 0.911. Finally, JLD naturally extends to video, reaching a correlation of 0.786 on Waterloo IVC 4K compared with 0.611 for VMAF.

75. 【2610.05954】MEND: RL For Flow Models via Proximal Velocity Matching

链接:https://arxiv.org/abs/2610.05954

作者:Shreshth Saini,Neil Birkbeck,Yilin Wang,Balu Adsumilli,Alan C. Bovik

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:worth its size, MEND, Reward, Reward post-training, frozen reference

备注:

点击查看摘要

Abstract:Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.

76. 【2610.05940】ReMem: Streaming Video Understanding With Long Context Retention

链接:https://arxiv.org/abs/2610.05940

作者:Li Yiheng,He Xu,Wang Shaobo,Shao Ling,Lu Shijian

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Vision Language Models, Language Models, current Vision Language, video understanding tasks, low latency response

备注:

点击查看摘要

Abstract:Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify that most existing works tend to progressively lose long context information as length of input stream increases. To address this, we propose ReMem, a novel training-free adaptation technique that enables VLMs to process streaming videos of arbitrary lengths while improving their long context information retention capability. ReMem exploits memory from two perspectives, implemented as two core components. The Streaming Context Memory (SCM) continuously compresses historical context with query-independent attention. The Retrieved Vision Memory (RVM) then retrieves the most salient, query-relevant context from memory to augment the VLM's input. Comprehensive experiments demonstrate that the proposed ReMem achieves state-of-the-art (SOTA) performance across a variety of widely used benchmarks, spanning both streaming video and general long video understanding tasks.

77. 【2610.05932】UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance

链接:https://arxiv.org/abs/2610.05932

作者:Gaoxiang Cong,Liang Li,Jianwei Wen,Zhedong Zhang,Zheng-Jun Zha,Qingming Huang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)

关键词:Visual voice cloning, cloning requires intelligible, speaker-consistent speech synchronized, voice cloning requires, Visual voice

备注:

点击查看摘要

Abstract:Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.

78. 【2610.05921】Beyond Transport Cost: Routing Differences between Flow Matching and Optimal Transport

链接:https://arxiv.org/abs/2610.05921

作者:Eungyeol Han,Jong-Seok Lee

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:improve Flow Matching, Flow Matching, improve Flow, Optimal Transport, reducing noise-data coupling

备注:

点击查看摘要

Abstract:In generative models, Optimal Transport (OT) is used to improve Flow Matching (FM) by reducing noise-data coupling cost. However, different noise-to-output assignments can yield nearly equal costs, raising a key question. Is cost alone sufficient to guide coupling design? We address this question by separating transport cost from routing, i.e., the destination reached by each noise sample. We show numerically how FM and OT can differ in routing while remaining close in cost. We examine its consequences in learned neural FM. Using the exact FM routing as an oracle, we further construct a routing-aware training coupling and find that it yields a directionally consistent improvement in generation over a cost-matched, cost-only counterpart. Our findings highlight what cost minimization can overlook and motivate using both cost and routing to evaluate the design of OT-based FM couplings. Code will be released upon acceptance.

79. 【2610.05918】Prompt and Refinement: Asymmetric Mutual Learning for Infrared Small Target Detection with Noisy Labels

链接:https://arxiv.org/abs/2610.05918

作者:Yimin Fu,Songbo Wang,Lizhuo Liu,Baicheng Pan,Zhunga Liu,Michael K. Ng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Existing data-driven infrared, methods typically require, typically require large-scale, data-driven infrared small, small target detection

备注: The code will be released at [this https URL](https://github.com/fuyimin96/PAR) upon acceptance

点击查看摘要

Abstract:Existing data-driven infrared small target detection (ISTD) methods typically require large-scale datasets with accurate pixel-level annotations for model training. However, such labor-intensive requirements are difficult to satisfy in real-world applications due to the heavy reliance on expert knowledge and the inherently weak distinctiveness of infrared small targets. Consequently, the presence of noisy labels during model training is inevitable, which can severely mislead the learning of target perception toward spurious patterns. To address this challenge, we propose Prompt and Refinement (PAR), a label-noise-robust asymmetric mutual learning paradigm for ISTD. Specifically, PAR comprises a pretrained Segment Anything Model (SAM) and an ISTD-specific detector trained from scratch, which learn collaboratively through a peer-teaching scheme. Coupled with local contrast regularity, the predictions of the two asymmetric peer models are mutually exploited as rectification cues for the supervisory masks of their counterparts. The interaction between complementary inductive biases effectively prevents the label correction process from degenerating into the self-confirmation loop of a single model, enabling progressive refinement of the annotations toward intrinsic target characteristics. In addition, the detector predictions are utilized as corrective mask prompts to facilitate task-specific adaptation of the vision foundation model. Moreover, an evidential uncertainty estimation strategy is introduced into the optimization process to further alleviate the adverse effects of noisy labels. Extensive experiments under diverse noisy label scenarios on three ISTD datasets demonstrate that PAR consistently achieves state-of-the-art performance.

80. 【2610.05911】Every View Counts: View-Consistent Panoptic Quality for Multi-view Panoptic Segmentation

链接:https://arxiv.org/abs/2610.05911

作者:Youngmin Lee,Byungha Ko,Guhnoo Yun,Dong Hwan Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:single forward pass, models predict, forward pass, assigns a semantic, semantic class

备注: 23 pages, 8 figures. Under review. Youngmin Lee and Byungha Ko contributed equally

点击查看摘要

Abstract:Multi-view panoptic segmentation assigns a semantic class and a scene-level instance ID to every pixel of an unordered set of images, and recent feed-forward 3D models predict these labels for the input views in a single forward pass. Their predictions, however, have been evaluated with the scene-level PQ (PQ^scene) borrowed from per-scene optimization methods, typically on rendered held-out views. PQ^scene tiles all views of a scene into a single image, so that a missed appearance or a change of ID lowers the score of the matched pair only in proportion to its area. We propose View-Consistent Panoptic Quality (VC-PQ), which extends PQ from a single image to a set of input views, counts equally every view in which an instance is visible, and penalizes a prediction that is not visible in the same views as its ground truth. A decomposition of VC-PQ attributes the score a method loses to mask accuracy, view consistency, and the matching threshold. A single additional parameter recovers the area weighting of tiling for comparison. Under a fixed evaluation protocol on ScanNet++ and ScanNetv2, recent feed-forward methods are evaluated with VC-PQ and PQ^scene, and the decomposition shows where each of them loses its score. Controlled perturbations of the ground truth show that VC-PQ responds to the number of views in which an instance is missed or changes ID, whereas PQ^scene responds to their area. The aim of this work is to make view consistency part of the evaluation of multi-view panoptic segmentation, with VC-PQ reported alongside PQ^scene.

81. 【2610.05910】AstraSR: Real-World Thermal Super-Resolution with GPT-6 Astra

链接:https://arxiv.org/abs/2610.05910

作者:Mengyuan Li,Changhong Fu,Jun Zhang,Ziyu Lu,Yuhang Zhang,Haobo Zuo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:limited sensor resolution, obtaining corresponding high-resolution, constrained by limited, limited sensor, sensor resolution

备注:

点击查看摘要

Abstract:Real-world thermal super-resolution (SR) is constrained by limited sensor resolution and the difficulty of obtaining corresponding high-resolution (HR) observations for direct model supervision. Conventional SR methods typically construct training pairs by treating captured thermal images with real-world degradations as HR references and applying predefined degradation to generate synthetic low-resolution (LR) inputs. Such a construction not only introduces a domain gap between synthetic and captured LR observations but also retains acquisition degradations in the supervision. To address this issue, we propose AstraSR, a real-world thermal SR method guided by GPT-6 Astra, a frontier multimodal generative model endowed with emergent and transformative visual capabilities. Specifically, we construct a dataset of image pairs by using captured LR thermal images to condition GPT-based HR reference. We develop a direct generative supervision strategy that learns from captured thermal inputs paired with GPT-generated HR references. Pixel, gradient, and perceptual losses jointly supervise the transfer of intensity patterns, structural boundaries, and visual details from the generated references. Qualitative comparisons with seven existing state-of-the-art real-world SR methods show continuous object contours, distinct structural boundaries, and smooth intensity transitions in the thermal scenes. These results demonstrate that AstraSR outperforms existing real-world SR methods in both thermal clarity and structural coherence.

82. 【2610.05908】Safe Image Generation via Reinforcement Learning

链接:https://arxiv.org/abs/2610.05908

作者:Eungyeol Han,Jong-Seok Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models achieve remarkable, achieve remarkable visual, remarkable visual image, models achieve, including violent

备注:

点击查看摘要

Abstract:Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.

83. 【2610.05902】On Hyperparameter Tuning on the Test Set

链接:https://arxiv.org/abs/2610.05902

作者:Matteo Fregonara,Tom Viering,Jan van Gemert

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:machine learning textbooks, test set, test, learning textbooks, stated in machine

备注:

点击查看摘要

Abstract:"Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fraud. Yet evidence suggests that test set hyperparameter tuning does occur in practice, making it all the more important to understand its actual consequences. So how bad is it, really? In this work we question this dogma and put it to an empirical test. We systematically study the magnitude of the performance inflation caused by tuning the hyperparameters on the test set for MNIST-1D, CIFAR-10, and three tasks from the GLUE benchmark. Our experiments show that while the effect is real and significant, it is frequently small relative to other sources of noise. In many cases, we find that tuning on the test set recovers exactly the same model as when tuning on the validation set. Most importantly, we find that the rankings of models remain essentially preserved after tuning on the test set and therefore that consistent test-set tuning may not invalidate benchmarks or model selection. Our results call for a more nuanced view of tuning hyperparameters on the test set, stimulating researchers to openly report test tuning.

84. 【2610.05900】End-to-End Autonomous Recursive Arborescence Deformable Flow and Non-Linear Hemodynamics for Patient-Specific Coronary Centerline Extraction

链接:https://arxiv.org/abs/2610.05900

作者:Zeyu Jia,Xin Ming

类目:Computer Vision and Pattern Recognition (cs.CV); Tissues and Organs (q-bio.TO)

关键词:Extracting patient-specific vascular, volumetric medical images, Extracting patient-specific, Euclidean Minimum Spanning, patient-specific vascular trees

备注: 10 pages, 4 figures

点击查看摘要

Abstract:Extracting patient-specific vascular trees from volumetric medical images is fundamental to computational angiography and non-invasive hemodynamic assessment. Conventional voxel segmentation models often sever delicate bifurcations, while heuristic Euclidean Minimum Spanning Trees introduce non-anatomical shortcuts. Moreover, linear Poiseuille flow neglects quadratic kinetic dissipation across arterial narrowings, underestimating ischemia. We formulate an end-to-end framework decoupling continuous geometric arborescence generation from non-linear hemodynamics. First, an autonomous 3D Ostium Landmark Localization Head with dual-sinus query channels and spherical-gated refinement eliminates centerline seeding dependency, achieving cohort mean localization error of 7.63 mm (7.43 mm LCA, 7.83 mm RCA; 71.4% = 8.0 mm) from raw contrast context. Second, a Spatially-Grounded Deformable Step Flow Architecture queries continuous 3D feature pyramids via trilinear sampling, sequentially generating trajectories with anchor boundary enforcement (X(0) = P_start). Third, a Top-Down Recursive Arborescence State Machine detects bifurcation peaks via Tree-NMS and parameterizes predecessor parent pointers (p_k k), guaranteeing single connected acyclic tree topology (beta_0 = 1, beta_1 = 0) with differentiable step termination. Fourth, an iterative Picard non-linear Kirchhoff solver with Young-Tsai / Gould quadratic dissipation enforces machine-precision mass conservation (residual 5.82e-11 mL/s). Across 14 development patients under standardized in-silico stenosis stress testing (Q_0 = 4.0 mL/s), linear Poiseuille flow misclassifies 75% diameter lesions as non-ischemic (FFR 0.80) in 14/14 cases, whereas our non-linear solver captures functional ischemia (FFR = 0.5864, lesion disparity 32.89 mmHg, p = 6.10e-5) with 3.66x collateral shunting. Test set firewall isolation was maintained.

85. 【2610.05899】LoDEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery

链接:https://arxiv.org/abs/2610.05899

作者:Kai Li,Zigan Zhou,Zhenyang Li,Hui Shan,Zhe Chen,Yupeng Deng,Zhihao Xi,Yu Meng,Yifan Peng,Xiangyu Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:off-nadir imagery, five-dimensional offset token, central to extracting, offset token, Instance-level

备注: 13 pages, 2 figures, 5 tables, including appendices

点击查看摘要

Abstract:Instance-level roof-to-footprint offset (RFO) prediction is central to extracting building footprints from off-nadir imagery. Query-based pipelines commonly use high-dimensional instance tokens to predict signed two-dimensional RFOs. We investigate whether RFO prediction can instead use a compact offset token. Under local pinhole projection and vertical-extrusion assumptions, the idealized RFO map admits a five-parameter sufficient descriptor comprising intrinsic shape, composite amplitude, and relative geometry. This factorization provides a structural prior for a five-dimensional offset token, whose channels learn task-relevant latent representations through end-to-end training. Based on this design, we propose LoDEOT, which retains high-dimensional instance tokens for detection and segmentation but maps instance-token, concentration-gated roof, and box-mask evidence to a five-dimensional offset token followed by an independent two-dimensional readout. Known denoising-query target indices further align each supervised decoder-layer estimate with the same clean instance RFO, organizing successive predictions as target-aligned recovery under perturbed query conditions. Experiments on five real-world building datasets demonstrate the effectiveness of LoDEOT for building footprint extraction. Experiments on real-world building datasets demonstrate that a five-dimensional offset token can support accurate RFO prediction. On BONAI, LoDEOT achieves the best roof-detection bAP and bAP50 and leads all five offset-corrected footprint metrics among the evaluated end-to-end methods, with FAP50 of 54.58 and mEPE of 5.23 pixels. Its FAP50 exceeds those of the evaluated end-to-end baselines by 7.56-16.85 percentage points.

86. 【2610.05897】Fitting Vision Adapters at Frontier Scales

链接:https://arxiv.org/abs/2610.05897

作者:Jaehoon Lee,Harry Partridge,Mudith Jayasekara,Charles O'Neill,Max Kirkby,Michael Psenka

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:frozen vision encoder, vision capabilities, language model, vision, established approach

备注: NeurIPS 2026 Workshop: Grounded and Faithful Vision-Language Models for Real-World Deployment

点击查看摘要

Abstract:Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.

87. 【2610.05896】asteRoute: Personalized Routing for Video Generation

链接:https://arxiv.org/abs/2610.05896

作者:Zhi Rui Tam,Chao-Chung Wu,Sin-Han Yang,Peyton Ku,Brendan Kuang,Tzu-Ting Hsieh,Min-Fang Hsu,Fang-Ling Tsai,Yun-Nung Chen,Wei-Chiu Ma,Chieh-Yen Lin

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Rapid progress, differ substantially, substantially in capability, Rapid, Abstract

备注:

点击查看摘要

Abstract:Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees with each annotator's own favorite only 34-55% of the time. Motivated by this observation, we introduce TasteRoute, a personalized video-generation router that selects a generator jointly based on the input request, user preferences, and available generation budget. Across text-to-video and image-to-video settings, TasteRoute is competitive with strong simple baselines on preference routing while reducing average generation cost. The cost saving increases under higher budget caps. Finally, we release TasteRoute-3k, a human-annotated dataset containing multi-model video comparisons, quality judgments, preference rankings, and user-profile signals to facilitate future research on personalized and cost-aware video routing.

88. 【2610.05891】Spatial Supervision Without Attribution Optimization: Improving Post-Hoc Class Activation Maps via Box-Guided Evidence Routing

链接:https://arxiv.org/abs/2610.05891

作者:Wenhao Liang,Liangwei Nathan Zheng,Lin Yue,Wei Emma Zhang,Mingyu Guo,Weitong Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image classifier predictions, class activation maps, ordinary training encourages, Post-hoc class activation, class activation

备注: 24 pages, 12 figures. Appendix included in the main PDF (pages 10-24)

点击查看摘要

Abstract:Post-hoc class activation maps (CAMs) are a standard tool for inspecting the evidence behind an image classifier's predictions, yet nothing in ordinary training encourages these maps to be spatially appropriate. We study whether inexpensive spatial supervision can improve a classifier's own predicted-class Grad-CAM without ever optimizing an attribution map. Box-Guided Evidence Routing (BGER) trains a lightweight gate on the final feature map under box or mask supervision and routes classification through the gated features, while Grad-CAM is computed separately at the pre-gate representation, so the evaluated map never enters the training objective. With a BCE routing loss, BGER raises MaxBoxAccV2 from $0.584$ to $0.715$ on CUB-200-2011 and from $0.757$ to $0.832$ on Stanford Dogs at comparable accuracy. Matched controls attribute most of the ResNet-50 gain to the spatial supervision reshaping the backbone rather than to routing itself: when classification bypasses the gate, most of the improvement remains, and detaching gradients through the gate leaves the ResNet-50 result nearly unchanged. The same detachment preserves most of the gain in two DenseNet-121 chest X-ray settings but removes the apparent gain on Swin-T, and directly supervising the CAM reaches stronger localization at a larger accuracy cost. Overall, spatial supervision can improve separately evaluated post-hoc CAMs, but both the mechanism and the size of the benefit depend on the architecture and the evaluation setting.

89. 【2610.05880】fMRI-TAMCL: Text-Anchored Supervised Multimodal Contrastive Learning for fMRI-Based Brain Disorder Classification

链接:https://arxiv.org/abs/2610.05880

作者:Juliana Mantebea Danso,Enoch Opanin Gyamfi,Mylene C.Q. Farias

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:strong multisite heterogeneity, exhibits strong multisite, Resting-state fMRI, brain disorders, multisite heterogeneity

备注: 10 pages, 6 figures

点击查看摘要

Abstract:Resting-state fMRI is important in the classification of brain disorders, but highly multimodal and exhibits strong multisite heterogeneity. Existing methods fuse images, BOLD-based functional connectivity, and phenotypic data modalities. Unlike other medical imaging datasets, rs-fMRI datasets rarely include a text modality, so they are generated from phenotypic data or BOLD activations. These text generation methods rely on fixed assumptions for subjects, sites, devices, and protocols, leading to poor generalization across datasets. We propose fMRI-TAMCL, a text-anchored multimodal contrastive learning framework that integrates fMRI images, sparse FC, and generated subject-specific text. Its Subject-Adaptive Threshold Derivation module generates BOLD activation text, while Feature-Value Serialization module generates phenotypic text. All three modalities are encoded as clustered graphs, projected onto a shared unit hypersphere space, aligned using pairwise, text-anchored supervised contrastive learning, and fused with attention. fMRI-TAMCL proves its generalization capability across five datasets outperforming 29 baselines with 78.6%-86.4% accuracy in downstream classification.

90. 【2610.05870】Certification of Real Images through Calibrated Content Authentication

链接:https://arxiv.org/abs/2610.05870

作者:Sarim Hashmi,Abdelrahman Elsayed,Mohammed Talha Alam,Samuele Poppi,Nils Lukas

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:synthesize high-quality inauthentic, high-quality inauthentic multimedia, Generative models, inauthentic multimedia content, misused at scale

备注:

点击查看摘要

Abstract:Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further reduce every baseline detector to below 2% accuracy, effectively inverting the detector's assigned label. We argue that this unreliability reflects a fundamental ambiguity: generators can reproduce authentic content exactly (e.g., through memorization), so content alone cannot reveal the true provenance this http URL this reason, content produced by a generator must admit a faithful reconstruction by that same generator, and finding such a reconstruction makes synthetic provenance plausible and authenticity plausibly this http URL therefore propose and evaluate a detection paradigm that outputs a calibrated prediction of whether authenticity is plausibly deniable: a faithful reconstruction by any known generator establishes plausible deniability, while calibration bounds how often content from known generators fails to be reproduced. Our evaluation shows that (i) our detector can be calibrated so that at most 1% of generated content is wrongly certified, an operating point at which most baseline detectors reach near-zero recall, including the strongest with 93% accuracy; (ii) calibrating a stricter security threshold on attacked samples preserves this bound against adaptive adversaries within the evaluated bounded-perturbation attack space, whose perturbations break every baseline, but does not cover arbitrary adversarial transformations; and (iii) post-hoc verifiability is eroding, as 1,116 of 3,000 Reddit images resist reproduction by a 2022 generator, but only 55 to 79 resist reproduction by 2024 generators.

91. 【2610.05865】Weave Mamba Fusion: Global Cross-Scale Interaction for Lightweight Face Detection

链接:https://arxiv.org/abs/2610.05865

作者:Dohun Kim,Jinmyung Jung

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:achieved strong performance, FPN to BiFPN, fusing multi-scale features, achieved strong, strong performance

备注:

点击查看摘要

Abstract:Feature pyramid methods, from FPN to BiFPN, have achieved strong performance in face detection by fusing multi-scale features. However, detecting faces under unconstrained conditions, such as small scale, occlusion, and extreme pose, remains difficult, as it requires global cross-scale dependencies that local fusion cannot model. State space models such as Mamba provide global context with linear complexity by scanning features as a sequence, and therefore offer a promising direction for this problem. Nevertheless, such a scan needs the two pyramid scales combined into a single feature map, and the way they are combined determines whether cross-scale structure is preserved. Summation collapses the two scales before the scan, so the scan has no cross-scale structure to exploit, while concatenation keeps both scales but at far higher cost. To address this, we propose \textbf{Weave Mamba Fusion (WMF)}, which interleaves two adjacent pyramid scales column by column so that each step of a horizontal bidirectional SS2D scan moves from one scale to the other. With partial-channel processing and parameter-free de-weaving, WMF enables efficient cross-scale interaction while preserving feature structure. Integrating WMF into every fusion node yields \textbf{WeaveBiFPN}, the neck of our \textbf{WeaveFace} detector. On WIDER FACE, WeaveFace achieves 91.41\% mean AP with only 0.34M parameters and 1.16 GFLOPs, outperforming prior detectors under 0.5M parameters. Its largest gains are on the Hard subset, where it reaches 87.14\% AP. The code is publicly available at \url{this https URL}.

92. 【2610.05861】Imagine to Act: High-Fidelity Data Synthesis via Image Editing World Model for Scalable GUI Agent Training

链接:https://arxiv.org/abs/2610.05861

作者:Yongxin Ning,Runliang Niu,Qianli Xing,Zhiyi Duan,Qingzu He,Pan Wang,Qi Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Graphical User Interface, Graphical User, User Interface, automating complex digital, complex digital workflows

备注:

点击查看摘要

Abstract:Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-action trajectories, which are notoriously difficult to acquire. While human demonstrations are unscalable, existing GUI world models rely on text descriptions or HTML rendering, discarding crucial pixel-level visual details like icons and layout styles. To address this issue, we introduce Infinite-Dreamer, a simulation-free data synthesis method powered by a pixel-level Image Editing World Model. By conceptualizing GUI transitions as image editing tasks, we leverage Vision-Language Models (VLMs) to describe action-induced UI changes as structured delta-text. We then fine-tune an image editing backbone to controllably synthesize realistic screenshot transitions. We utilize this model to generate both single-frame visual robustness data and multi-step imaginary trajectories. To validate the effectiveness of our approach, we fine-tune the Qwen3-VL baseline solely on the synthesized data to obtain Infinite-Actor, and evaluate it on AndroidWorld, MobileWorld, and AndroidControl-Curated benchmarks. Infinite-Actor consistently outperforms the Qwen3-VL baselines across scales: Infinite-Actor-8B improves AndroidWorld Pass@1 by +4.45 and nearly doubles the MobileWorld Pass@3 success rate, while Infinite-Actor-2B improves Pass@1 by +9.05. Code is available at this https URL.

93. 【2610.05839】Dual-Rate Force-Image Control with Model-Based Orientation Limits for Robotic Ultrasound

链接:https://arxiv.org/abs/2610.05839

作者:Tyler Foster,Qiang Zhang,A B M Tahidul Haque,Anh Thu Nguyen

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)

关键词:Robotic ultrasound couples, image-guided ultrasound probe, ultrasound probe rotation, high-rate contact-force loop, Robotic ultrasound

备注: 8 pages, 4 gigures, conference

点击查看摘要

Abstract:Robotic ultrasound couples a high-rate contact-force loop with slower, delayed image feedback, so image-guided ultrasound probe rotation can perturb contact force before the resulting image response is observed. We derive a closed-form orientation-rate limit that bounds the modeled rotation-induced estimated-force excursion over a finite horizon while accounting for disturbance rejection by the fast force loop. The limit depends on local contact stiffness, force-loop gains, a conservative rotation-to-force gain bound, the excursion budget, and the prediction horizon. We implement this model in a dual-rate controller with timestamp-based delay reconstruction and joint-torque-based force estimation, and evaluate it on a curved gelatin phantom using paired controller comparisons and component ablations. Relative to unconstrained image guidance, the proposed rate-limited controller reduced first-second root-mean-square (RMS) estimated-force error by 0.40 N while increasing cue-convergence time by 0.94 s. A fixed rate cap near the analytically predicted ceiling produced no resolvable difference in force error and converged 0.32 s faster, indicating that the principal practical value of the model is the rate-design rule rather than online prediction. Delay reconstruction had no resolvable effect at the tested latency. A single-subject popliteal scan demonstrated feasibility, although the image cue was noise-limited on heterogeneous tissue.

94. 【2610.05816】Level-of-Token Diffusion

链接:https://arxiv.org/abs/2610.05816

作者:Kiyohiro Nakayama,Brian Chao,Jan Ackermann,Hansheng Chen,Federico Tombari,Leonidas Guibas,Lior Yariv,Gordon Wetzstein

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:models allocate equal, intended scene calls, allocate equal computation, diffusion models allocate, models allocate

备注:

点击查看摘要

Abstract:Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at this https URL.

95. 【2610.05801】Gauss-Map Variation for Image Denoising: Geometric Analysis and an Anderson--Accelerated Majorization--Minimization Method

链接:https://arxiv.org/abs/2610.05801

作者:Haibin Su

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gauss-map variation, scaled image graph, propose a Gauss-map, spatial variation, denoising that measures

备注:

点击查看摘要

Abstract:We propose a Gauss-map variation (GMV) model for image denoising that measures the spatial variation of the tangent-plane projectors of the scaled image graph. We establish an equivalent representation of the regularizer in terms of the corresponding Gauss map and, using differential geometric tools including tubular coordinates and the Frenet frame, analyze its behavior across general $C^2$ and piecewise $C^2$ boundaries. The resulting estimates provide edge- and corner-contrast preservation properties. To solve the proposed model, we introduce a bilinear decomposition involving a unit normal field and a scalar magnitude field and develop an Anderson-accelerated majorization--minimization algorithm. The normal field subproblem admits an explicit pointwise majorization--minimization update, which is combined with an Anderson acceleration. For both $L^1$ and $L^2$ data fidelity terms, we establish sufficient decrease and boundedness of the iterates and prove that the generated sequence converges to a critical point of the penalized model. Numerical experiments on synthetic and natural images demonstrate the boundary preserving capability of the proposed model and its competitive performance in removing Gaussian and impulsive noise.

96. 【2610.05790】FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models

链接:https://arxiv.org/abs/2610.05790

作者:Md Aminur Hossain,Omkumar Vaghasiya,Rajeev Ranjan Dwivedi,Vinod Kurmi,Biplab Banerjee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:hide systematic performance, commonly evaluated, hide systematic, systematic performance disparities, ecological regions

备注: 13

点击查看摘要

Abstract:Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: this https URL.

97. 【2610.05779】A Spatiotemporal Semantic Importance-Guided Unified Compression and Editing Framework for AI-Generated Videos

链接:https://arxiv.org/abs/2610.05779

作者:Xihua Sheng,Dong Liu,Chang Wen Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:creating growing demands, AI-generated videos, increasing in volume, creating growing, storage and transmission

备注:

点击查看摘要

Abstract:AI-generated videos are rapidly increasing in volume, duration, and resolution, creating growing demands for efficient storage and transmission. Unlike natural videos captured from the physical world, AI-generated videos are samples from a learned generative distribution, where semantic structures are critical to content consistency, while many local textures and stochastic details can be plausibly regenerated. This distinction suggests that compression should preserve semantically important spatiotemporal information rather than reconstruct every pixel of a particular generative sample. Beyond reconstruction, AI-generated videos also create a practical need for prompt-based editing, where users expect to modify generated content while preserving its original spatiotemporal semantics. Motivated by these observations, we propose a unified compression and editing framework for AI-generated videos that incorporates a frozen video generator as a reusable generative prior. Within this framework, we design three spatiotemporal semantic importance-guided techniques that respectively address what to transmit, how much to transmit, and how to use the transmitted side information. First, an innovation selection method projects the latent discrepancy using spatiotemporal semantic importance, so that the selected innovations prioritize semantic invariants over replaceable generative variations. Second, a frame-adaptive bit allocation method estimates the nonuniform semantic demands of latent frames and allocates more innovations to frames requiring stronger semantic preservation. Third, a unified reconstruction and editing method continuously adjusts the influence of the transmitted side information, enabling the same compressed representation to provide strong guidance for faithful reconstruction or serve as a flexible semantic anchor for structure-preserving prompt-driven editing.

98. 【2610.05775】InteractionBench: A Real-Time Interaction Benchmark for Streaming Video Systems

链接:https://arxiv.org/abs/2610.05775

作者:Enxin Song,Suhao Yu,Yifei Xu,Barbara Su,Weili Xu,Wenhao Chai,Yao Tang,Jie Deng,Haiyang Xu,Jiatao Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:assistant must speak, instruction warrants, stay silent, https URL, URL

备注: Project page: [this https URL](https://www.enxinsong.com/projects/interactionbench/) Code: [this https URL](https://github.com/Espere-1119-Song/InteractionBench) Data: [this https URL](https://huggingface.co/datasets/InteractionBench/InteractionBench)

点击查看摘要

Abstract:A video assistant must speak when its instruction warrants a response and stay silent otherwise. We introduce a benchmark that evaluates this decision for the complete system of model, memory, and response controller. InteractionBench covers query responses, event triggers, and ongoing updates in 1,060 interactions over 812 videos, with 69 negative streams and 53 suites that pair counted events with look-alike near misses. It scores content accuracy, timing accuracy, and silence compliance on the video clock. Timely speech costs silence across systems. Polled Qwen3-VL-8B reaches 77.8 timing accuracy but 10.9 silence compliance. A native real-time interaction system reaches 29.2 silence compliance at 66.8 timing accuracy, yet emits on 89.9% of negative streams. No open-weight system clears a third of the near-miss suites. Fewer replies help only when chosen, as random deletion merely trades timing for silence. Offline scores miss these failures and mispredict online behavior. Adding restraint is costly, as the native system's controller adds little by itself and agentic systems add it only at about 30 s per this http URL page: this https URL Code: this https URL Data: this https URL

99. 【2610.05771】Controllable Road Marking Generation

链接:https://arxiv.org/abs/2610.05771

作者:Zhiyu(Joey)Cai,Yufan Zhang,Ruichen Tan,Zengxiang Lei,Satish Ukkusuri

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:limits quantitative analysis, markings provide critical, road markings provide, Road Marking Generation, Controllable Road Marking

备注:

点击查看摘要

Abstract:Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet authoring them at scale remains a manual workflow that limits quantitative analysis and scenario testing. We introduce Controllable Road Marking Generation, which synthesizes a missing center-region marking layout from a drivable-area mask, optional outer-ring markings, and a textual description. Our benchmark uses deterministic, metadata-derived prompts and three output channels: lane dividers, road dividers, and pedestrian crossings. We develop a conditional bird's-eye-view (BEV) pipeline that combines (i) a text-conditioned latent rectified-flow DiT trained with a topology-aware auxiliary loss, (ii) Gaussian-blurred training targets that stabilize learning of thin, sparse markings, and (iii) Structured Gaussian Render (SGR), a training-free post-process that recovers crisp divider geometry by extracting polylines, fitting cubic Bézier curves, and re-rendering them as anisotropic super-Gaussian primitives. On 4,597 Argoverse~2 test tiles, our system achieves Buffered F1 of 80.8 and clDice of 50.2, compared with 38.8 and 24.6 for an adapted state-of-the-art mask-refinement baseline. On Waymo dataset, it yields 88.0 Buffered F1 and 66.2 clDice. Component ablations show complementary connectivity gains from topology-aware supervision and SGR. Text-editing experiments reveal that stronger guidance improves edit success but also increases changes to non-target structures. We see this framework as a step toward simulation-ready road-marking variation, automated map completion, and early-stage infrastructure design exploration.

100. 【2610.05758】A Three-Dimensional Reverse-Projection Method for Sparse Point Cloud Completion and Its Application to High-Speed Train Nose Reconstruction

链接:https://arxiv.org/abs/2610.05758

作者:Xiaozhen Ma,Zhao Tang,Hanbin Lai,Ruiqi Chen,Jin Jin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remain an unresolved, unresolved problem, problem in three-dimensional, point cloud, LiDAR scans

备注: 16 pages, 6 figures

点击查看摘要

Abstract:Holes in LiDAR scans of environments with glass windows remain an unresolved problem in three-dimensional reconstruction. This study presents a pipeline for point cloud acquisition, filtering, completion, and surface reconstruction to address sparse sampling and missing window regions in scans of a high-speed train nose. FAST-LIVO2 provides the initial point cloud through multisensor odometry and mapping, and moving least squares (MLS) smooths the observations. We then introduce three-axis projection-based subdivision and interpolation with reverse hole boundary identification, referred to as three-axis reverse completion. The method interpolates missing regions from observations around each hole. Greedy projection triangulation, Poisson surface reconstruction, and a Marching Cubes-based pipeline generate meshes from the completed point cloud. Experiments on a proportionally scaled display model of a high-speed train nose show that the proposed method fills missing point cloud regions around the glass windows. Under the evaluation setting used in this study, greedy projection triangulation yields lower geometric distance errors than the other two reconstruction pipelines. The pipeline supports non-contact digital modeling of train nose geometry and provides a practical approach to reconstructing objects with glass windows.

101. 【2610.05756】Vision-enabled detection of safety helmet compliance in construction zones

链接:https://arxiv.org/abs/2610.05756

作者:Tri Nhut Do*,Ba Loc Pham

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:rapidly evolving field, top priority, rapidly evolving, evolving field, remains a top

备注:

点击查看摘要

Abstract:In the rapidly evolving field of construction management, worker safety remains a top priority. This paper introduces an innovative vision-based system for real-time detection of helmet compliance, specifically designed for construction sites, utilizing advanced computer vision techniques and machine learning algorithms within the YOLO (you only look once) framework. Our system leverages high-resolution video feeds from strategically positioned cameras to monitor adherence to safety regulations regarding helmet usage. By employing deep learning methodologies, the system effectively identifies individuals not wearing helmets, thereby significantly mitigating the risk of head injuries among workers. Our training and validation results revealed an impressive precision exceeding 97% at mAP@0.5 for both helmeted and non-helmeted individuals. Furthermore, our experiments demonstrate exceptional detection accuracy, demonstrating the system's resilience under varying lighting conditions and diverse worker movements. The consistent decrease in loss and improvement in metrics throughout training validates the effectiveness of the YOLOv8 model in enhancing recognition performance. The implications of this research extend beyond mere regulatory compliance, opening avenues for innovative applications in occupational safety management. This study highlights the critical role of technology in protecting lives and lays the groundwork for future advancements in smart construction environments.

102. 【2610.05749】A new design of a fall detection system integrating landmark identification and deep learning techniques

链接:https://arxiv.org/abs/2610.05749

作者:Tri Nhut Do,Thi Thuy Le

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:integrates landmark identification, article introduces, introduces an innovative, enhance fall detection, Media Pipe

备注:

点击查看摘要

Abstract:This article introduces an innovative system that integrates landmark identification with deep learning to enhance fall detection accuracy and reliability. By utilizing advanced computer vision techniques, such as Media Pipe for spatial recognition, the system effectively differentiates between routine movements and actual falls. The integration of landmarks with a deep learning prediction algorithm minimizes false alarms, ensuring timely responses to genuine falls. Comprehensive experimentation underscores the system's versatility across various scenarios, emphasizing its potential to improve safety and independence for older adults. The training process demonstrates a steady increase in accuracy, stabilizing by the 40th cycle, while error rates decline significantly during the initial cycles. Real-time experiments, involving both male and female participants aged 8 to 50, recorded a remarkable 95% detection rate of falls, demonstrating the system's effectiveness and promising future applications in elder care and smart health monitoring environments.

103. 【2610.05743】Robust Local Optimization Done Right

链接:https://arxiv.org/abs/2610.05743

作者:James Pritts,Kevin Köser

类目:Computer Vision and Pattern Recognition (cs.CV); Mathematical Software (cs.MS)

关键词:inlier scale, impose different robustness, robustness requirements, motivating the separation, separation of hypothesis

备注:

点击查看摘要

Abstract:RANSAC scoring and local optimization (LO) impose different robustness requirements, motivating the separation of hypothesis selection from refinement. We systematically isolate the effects of robust-loss shape, incorrectly specified inlier scales, and optimization strategy on essential matrix, fundamental matrix, and homography estimation. A profile-marginal score marginalizes the nuisance inlier scale and selects an inlier partition, from which we estimate the scale that sets the LO loss width; this makes LO robust to an inlier scale specified too large, whereas one specified too small degrades selection itself. Refinement needs gradient from correspondences the seed currently rejects: optimizers that reweight from current residuals stay pinned to their seed, whereas methods with broad basins recover strongly perturbed seeds yet degrade accurate score-selected hypotheses, so basin size alone is insufficient to assess RANSAC LO. Joint half-quadratic optimization balances the two and is the most consistent strategy across model classes. An optimizer matched to the profile-marginal score, which never decreases it, does not reach the best accuracy, challenging the prescription that scoring and refinement objectives should match. Composed from these findings, our RANSAC reduces the median essential-matrix pose error of a state-of-the-art RANSAC on PhotoTourism from 2.23 degrees to 1.58 degrees with a correctly specified inlier scale and from 38 degrees to 6.2 degrees when it is grossly misspecified (128x too large).

104. 【2610.05739】HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

链接:https://arxiv.org/abs/2610.05739

作者:Zhuokun Chen,Feng Chen,Xi Lin,Xiyu Wu,Jiahao He,Jianfei Cai,Bohan Zhuang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Long-horizon video world, video world models, world models require, models require persistent, Long-horizon video

备注:

点击查看摘要

Abstract:Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the $60$-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a $0.74$ dB PSNR gain and a $28.5\%$ reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over $547$ samples. At a $60$-second context, HLA-WM reduces historical-state memory by $12\times$ relative to full KV caching while incurring at most a $1.6\%$ reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: this https URL

105. 【2610.05731】-JEPA: A Temporal Joint-Embedding Predictive Architecture for Learning Better Remote Sensing Representations

链接:https://arxiv.org/abs/2610.05731

作者:Bowen Peng,Li Liu,Yongxiang Liu,Weijie Li,Jie Zhou,Zhen Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing remote sensing, remote sensing foundation, sensing foundation models, imposing predefined pairwise, predefined pairwise relations

备注:

点击查看摘要

Abstract:Earth observation (EO) data provide rich temporal supervision, yet existing remote sensing foundation models mainly exploit sequential observations through imposing predefined pairwise relations or aggregating holistic reconstruction context. We seek to further exploit the sparse and nonuniform temporal sampling inherent in EO sequences as supervisory signals. To this end, we propose T-JEPA, a temporal joint-embedding predictive architecture that learns time-gap-conditioned latent transitions. A shared single-frame encoder processes each observation, while a temporal predictor estimates the complete target latent field from a masked source latent representation and the actual elapsed time. Across multiple temporal intervals, these predictive constraints organize observed states into structured latent trajectories. Asymmetric metadata injection mitigates shortcut learning, and direct supervision across multiple temporal scales proves more effective than recursively rolling out intermediate states. In parallel, masked pixel reconstruction provides complementary supervision for preserving spatial details. Under matched pre-training data and throughput, T-JEPA achieves leading transfer performance on both static and temporal tasks. Analyses further reveal that T-JEPA learns representations with time-gap-dependent transition predictability and coherent latent dynamics, while maintaining strong cross-period consistency, representation diversity, and semantic discriminability.

106. 【2610.05715】Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs

链接:https://arxiv.org/abs/2610.05715

作者:Zhaochen Wang,Yujun Cai,Huangbo Zou,Hower Yang,Naipeng Dong,Miao Xu,Haibin Ling

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Vision-language models, rotated across views, object has rotated, Vision-language, circ

备注: 37 pages, 8 figures

点击查看摘要

Abstract:Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12 VLMs, the gap is stark: the strongest models approach 100% accuracy on detection, yet even coarse magnitude estimation is near chance. When asked for exact angles, models place 91.8--100% of their predictions on just $0^\circ$, $90^\circ$, and $180^\circ$, a failure we term canonical-angle collapse. This collapse persists even without visual input. Representation probing shows that missing information is only part of the explanation. Although rotation information becomes less recoverable at finer granularity, substantial coarse-grained information remains, and a simple linear probe outperforms the models' generated answers. This suggests that VLMs underuse rotation information they already encode. We therefore propose RotationCue, a lightweight decoder that recovers coarse rotation information from the VLM's own frozen representations and feeds it back to the model as intermediate textual context. Across three VLMs, RotationCue improves every model--task combination on OR-Bench, raising macro-average accuracy by 7.9--12.6 points while preserving general capabilities.

107. 【2610.05711】From Pixels, Without Pre-training: Joint Generative and Self-Supervised Representation Learning in One Model

链接:https://arxiv.org/abs/2610.05711

作者:Vicente Balmaseda,Ching-Long Lin,Tianbao Yang

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:Strong image generation, Strong image, separately trained autoencoders, frozen pretrained encoders, aligned to frozen

备注:

点击查看摘要

Abstract:Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders or autoencoders pretrained for the target domain. We study joint generative and self-supervised representation learning in a single model, enabling self-conditioned generation without labels or pretrained models. This is challenging because the objectives are mismatched: contrastive learning consumes clean augmented views and favors coarse, invariant semantics, while flow matching consumes noisy images and must preserve the fine detail and spatial layout that contrastive learning discards. We propose SCION (Self-conditioned Generation on Self-supervised representation), whose core is a single pixel-space encoder conditioned on the flow timestep and an embedding. For representation learning, this conditioning embedding is a learned global vector shared across images, with the encoder's [CLS] token yielding the semantic representation trained by the contrastive loss. For generative training, the conditioning embedding is the image's own [CLS] representation, while patch tokens pass through a decoder to predict the image. To sample without a reference image at inference, we jointly learn a prior over the embedding. Gradient-norm balancing and stop-gradient mechanisms enable joint optimization in one run. SCION is self-supervised and self-contained, with no labels or pretrained models. On ImageNet 256x256, with the JiT-B recipe and no representation guidance, SCION reaches 8.92 FID, surpassing class-unconditional iREPA, which aligns to pretrained DINOv2 (46.44), and RCG, which conditions on it (14.27). With JiT-L, SCION achieves 5.89 FID without guidance and 3.47 with representation guidance, outperforming RCG with the ADM recipe (6.24).

108. 【2610.05707】Difference Feature Map Distillation: Transferring Inter-Sample Relational Knowledge Towards Efficient Transformer-Based Tracking

链接:https://arxiv.org/abs/2610.05707

作者:Zhicheng Ding,Xinyu Chu,Qing Tian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous driving perception, satisfy stringent latency, driving perception, object tracking systems, autonomous driving

备注: Published in the 2026 IEEE Intelligent Vehicles Symposium (IV)

点击查看摘要

Abstract:In autonomous driving perception, visual object tracking systems must satisfy stringent latency and power constraints while remaining robust in complex and dynamic environments. Although transformer-based trackers achieve state-of-the-art accuracy, their substantial computational and memory overheads hinder deployment on real-time, resource-constrained platforms. To move toward this goal, we propose Difference Feature Map Knowledge Distillation (DFM-KD), a novel relational distillation framework tailored for transformer-based visual object tracking. Unlike conventional feature distillation methods that minimize point-wise discrepancies (e.g., mean squared error) between teacher and student feature representations, DFM-KD transfers knowledge through inter-sample feature differences, explicitly aligning the relational structure of the feature space. By distilling how the teacher models appearance variation and consistency across samples, rather than enforcing similarity in absolute activations, DFM-KD enables the student to better capture the structural dynamics of visual changes within a batch. As a result, the distilled model exhibits enhanced feature robustness and improved tracking performance. Extensive experiments demonstrate that DFM-KD consistently outperforms conventional feature-level distillation methods in both tracking precision and success rates.

109. 【2610.05674】Bayesian Data Augmentation for DNN Retraining with Binomial Outcomes in Vision-Based UAV Landing

链接:https://arxiv.org/abs/2610.05674

作者:Ashik E Rasul,Hyung-Jin Yoon

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:reliable UAV missions, cluttered urban environments, UAV missions, reliable UAV, urban environments

备注:

点击查看摘要

Abstract:In GPS-denied or cluttered urban environments, vision-based landing is essential for reliable UAV missions. Real-world landing sites are often unstructured and highly variable, requiring strong generalization by the perception system. Deep Neural Networks (DNNs) trained with synthetic data augmentation offer a scalable solution for learning landing-site features across diverse vehicle and environmental states. However, computationally expensive DNN retraining, along with challenging performance validation via test flights, limits exhaustive model fine-tuning and necessitates an optimized retraining pipeline. In this work, we deploy a Bayesian data augmentation framework integrated with a photorealistic simulator featuring high-fidelity vehicle dynamics to iteratively retrain the helipad detector DNN, maximizing landing performance as the objective function. We validate our framework with experiments in a photorealistic simulator under different environmental conditions and vehicle states, demonstrating improved landing performance and tighter confidence intervals on predicted landing outcomes.

Subjects:

Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.05674 [cs.RO]

(or
arXiv:2610.05674v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2610.05674

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
110. 【2610.05664】StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation

链接:https://arxiv.org/abs/2610.05664

作者:Anh Dao,Quan-Dung Pham,Le Danh Vinh, TheAnh Nguyen,Nguyen Viet Tri Pham,Yiyu Chen,Pham Tuyen Le,Van-Truong Nguyen,Quan Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:policies increasingly benefit, strong semantic priors, semantic priors provided, large vision-language models, policies increasingly

备注:

点击查看摘要

Abstract:Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3\% SR and 51.4\% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3\% SR without additional navigation training data or a geometry encoder at inference.

111. 【2610.05637】Visual Grounding Safety in Vision-Language Models

链接:https://arxiv.org/abs/2610.05637

作者:Erfan Shayegani,Kundan Krishna,Yue Dong,Nael Abu-Ghazaleh,Leon Gatys,Shruti Palaskar

类目:Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:generate structured outputs, Vision-language models, downstream interfaces, systematically analyzed, increasingly trained

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety benchmarks spanning direct harm (VLSU), social bias (BBQ-V), and situational safety (Asimov-2.0) into 15,401 matched pairs of harmful requests that differ only in the requested output: a free-text answer (VQA) or a grounding (point or bounding box). Across five VLMs, models that refuse a harmful request posed as a question often comply when the same request asks for a grounding: averaged over models, grounding refusal trails VQA refusal by 31-59 percentage points, depending on the domain, and safety system prompts do not close this gap. We propose a fine-tuning approach that combines grounding-form refusals with capability grounding data and self-distilled benign data to counter over-refusal. For Qwen3-VL-8B and VisionReasoner-7B, it improves grounding refusal by 77-95 percentage points on VLSU and BBQ-V and by 64-85 points on the held-out Asimov-2.0 domain, while also improving VQA refusal, preserving grounding capability, and keeping over-refusal limited. Representation analysis shows that fine-tuning moves harmful requests toward each model's refusal direction, most strongly for grounding, while leaving benign requests near the harmless reference.

112. 【2610.05630】Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection

链接:https://arxiv.org/abs/2610.05630

作者:Nallathambi Vethiappan,Derya Soydaner,Gijs Wijnholds

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:image supports, Visual entailment, leaves undecided, undecided a textual, Atomic Visual Entailment

备注: 15 pages, 16 figures

点击查看摘要

Abstract:Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.

113. 【2610.05615】DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate

链接:https://arxiv.org/abs/2610.05615

作者:Khanh-Binh Nguyen,Van Dai Do,Tien Anh Nguyen,Svetha Venkatesh,Hung Le

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, Dynamic Resolution Assignment, existing multimodal MAD, language models, multimodal MAD frameworks

备注:

点击查看摘要

Abstract:Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variation in the visual scale needed across samples and agents. In addition, these frameworks frequently suffer from groupthink, a phenomenon where agents prematurely abandon correct deductions to conform with confident but hallucinated peer responses. To address these bottlenecks, we introduce DREAM (Dynamic Resolution Assignment For Multimodal Multi-Agent Debate), which operates via two core components: (1) Dynamic Resolution Assignment, a zero-shot probe round where agents test multiple resolutions, quantify uncertainty using Average Normalized Log-Likelihood (ANLL), and use an adaptive threshold to assign each agent to its empirically optimal resolution; (2) Uncertainty-Guided Rollback Aggregation counters groupthink by tracking each agent's uncertainty over rounds and restoring early low-uncertainty answers overridden by group pressure. On six multimodal datasets, DREAM improves the accuracy-token trade-off over multi-agent debate baselines by 1.5-3.2% accuracy without dataset-specific tuning.

114. 【2610.05608】Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

链接:https://arxiv.org/abs/2610.05608

作者:Team Kandinsky,Julia Agafonova,Bulat Akhmatov,Mikhail Aksyutin,Grigorii Alekseenko,Anastasia Aliaskina,Olga Androsova,Vladimir Arkhipkin,Anna Averchenkova,Alexander Belykh,Serafima Bocharova,Sofiya Bogakovskaya,Anton Bukashkin,Mark Bulygin,Kirill Buzygin,Irina Cheremnykh,Kirill Chernyshev,Mikhail Chernyshov,Vladimir Chernyy,David Chikovani,Georgy Daniltsev,Denis Dimitrov,Anna Dmitrienko,Vladimir Dokholyan,Sergey Emelyanov,Dmitry Ermilov,Georgii Fedorov,Polina Gavrilova,Nikolai Gerasimenko,Aleksandr Gordeev,Andrey Inozemtsev,Andrei Ivaniuta,Alexander Ivanov,Mikhail Karaev,Anastasiia Kargapoltseva,Ivan Kirillov,Nikita Kiselev,Valeria Kobenko,Yury Kolabushin,Denis Koposov,Anatoly Korobov,Vladimir Korviakov,Kirill Kozlov,Denis Krzhivokolskiy,Konstantin Kuklev,Alexander Kunitsyn,Sergey Kuzin,Vladislav Lakhtionov,Alexey Letunovskiy,Maxim Litvinov,Alexander Lyulkov,Georgy Makarov,Kirill Malakhov,Egor Malykh,Mikhail Mamaev,Dmitrii Mikhailov,Polina Mikhailova,Ivan Mikheev,Elizaveta Muromtseva,Nikolai Nazarkin,Tatiana Nikulina,Lev Novitskiy,Stanislav Onuchin,Nikita Osterov,Denis Parkhomenko,Anatoliy Parpara,Vladimir Polovnikov,Konstantin Reznikov,Azat Saginbaev,Nikita Samsonov,Alexander Sentsov,Nikita Shaimov,Artem Sherstyuk,Andrey Shutkin,Egor Silvestrov,Bulat Suleimanov,Matvey Suprunov,Sergey Taranov,Irina Tolstykh,Tatiana Trofimuk,Ilya Trushkin,Aleksandra Tsybina,Olga Varlashina,Viacheslav Vasilev,Ilya Vasiliev,Eugeny Vilisov,Sergey Yakubson,Konstantin Zakharov

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)

关键词:Video Lite, Video Pro, Kandinsky, foundation diffusion models, Video

备注: Technical report on the open-source T2AV model. GitHub: [this https URL](https://github.com/kandinskylab/kandinsky-6)

点击查看摘要

Abstract:We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.

115. 【2610.05603】EchoDino: A pediatric foundation model for transferable echocardiographic analysis across the lifespan

链接:https://arxiv.org/abs/2610.05603

作者:Sheng Cheng,Donnchadh M. O'Sullivan,Daniel J. Penny,Craig G. Rusin,Minh B. Nguyen,Devika Subramanian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:interpretation demands integrating, demands integrating visual, integrating visual evidence, cardiac imaging modality, dynamic cardiac motion

备注: 33 pages, 5 figures, including Supplementary Information

点击查看摘要

Abstract:Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.

116. 【2610.05587】Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife

链接:https://arxiv.org/abs/2610.05587

作者:Yuzhuo Li,Di Zhao,Xinyu Zhang,Daniel Wilson,Yun Sing Koh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:data scarcity, diverse poses, identification often suffers, suffers from data, Individual-level wildlife identification

备注: 29 pages, 12 figures, 9 tables

点击查看摘要

Abstract:Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a single reference image. However, existing I2V models mainly emphasize global layout, semantics, and motion, and therefore often fail to preserve fine-grained local appearance cues that distinguish one wildlife individual from another, such as fur texture, stripe boundaries, spot configurations, and contour transitions. We observe that these identity-critical cues are closely related to high-frequency information. To address this challenge, we propose WildIcon, a high-frequency-guided I2V framework for wildlife individual consistency. Specifically, WildIcon introduces a frequency-aware identity encoding branch that extracts individual-specific high-frequency cues from the reference image. Combined with isolated foreground information, the resulting identity tokens are then injected into cross-attention blocks as identity conditioning. Building on a frozen backbone with lightweight identity adaptation, WildIcon preserves fine-grained identity cues visible in the reference image while retaining the motion controllability and semantic fidelity of the base I2V model. In addition, to support the training and evaluation of wildlife individual-consistent I2V, we construct WildlifeVid, a wildlife-centric video dataset with high-quality, temporally coherent clips and individual-level identity labels. Experiments on I2V generation and downstream animal re-identification (ReID) show that WildIcon achieves stronger individual consistency than existing baselines, and that its filtered outputs can serve as useful candidate training augmentations for downstream ReID.

117. 【2610.05578】Rethinking Streaming-Perception Evaluation on Heterogeneous Edge Platforms

链接:https://arxiv.org/abs/2610.05578

作者:Misun Yu,Jinyoung Moon,Jemin Lee

类目:Distributed, Parallel, and Cluster Computing (cs.DC); Computer Vision and Pattern Recognition (cs.CV)

关键词:Multi-camera streaming perception, streaming average precision, Multi-camera streaming, heterogeneous edge platforms, edge platforms shared

备注: accepted in ACCV 2026

点击查看摘要

Abstract:Multi-camera streaming perception is increasingly deployed on heterogeneous edge platforms shared with co-resident workloads, yet accelerator placement is often evaluated using isolated single-stream experiments and mean streaming average precision (sAP). Using two end-to-end pipelines on a single GPU--NPU platform, we show that isolated evaluation can mis-rank deployment-time placement. Although the GPU pipeline is preferred in isolation, GPU-localized contention introduces deadline misses that make detections stale and can reverse the preferred placement before full GPU saturation. The NPU pipeline is less accurate than the GPU pipeline on small and medium objects in isolation, but nearly matches it on large objects. The largest absolute sAP losses in our latency and contention experiments occur for large objects. In our four-stream experiments, the preferred placement depends on which path becomes stale, and increasing GPU-side contention shifts the best placement from All-GPU to All-NPU. Under a GPU-saturating vision--language co-tenant, All-NPU achieves $5.2\times$ the worst-stream sAP of All-GPU. Because mean sAP can hide severe single-stream degradation, evaluation should report contention sweeps, deadline-miss rates on both paths, and worst-stream sAP alongside mean sAP.

118. 【2610.05576】SteadySplats: Resampling of Low-Variance Gaussians for High-Fidelity Stochastic Rendering

链接:https://arxiv.org/abs/2610.05576

作者:Felix Windisch,Thomas Köhler,Lukas Radl,Chris Wyman,Georgios Kopanas,Bernhard Kerb,Markus Steinberger

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Gaussian Splatting models, Gaussian Splatting, order-independent transparency enables, transparency enables efficient, primitive-based radiance fields

备注:

点击查看摘要

Abstract:Stochastic order-independent transparency enables efficient and elegant rendering of primitive-based radiance fields like 3D Gaussian Splatting models, but remains impractical due to the inherent visible noise in the output. We propose a principled approach to minimize high-frequency noise, addressing its sources at the representation and image synthesis level. During stochastic rendering, our history-based spatial resampling scheme drastically accelerates image convergence, while temporal importance resampling ensures coherence under camera movement. During training, a color regularizer implicitly reduces the variance along view rays in the 3DGS models. With these properties, our optimized, Vulkan-based renderer effectively mitigates output noise at low and high sample counts, achieving a substantial 13~dB PSNR increase in quality over previous stochastic methods at 1 sample per pixel and quickly converging to sorted 3DGS with an average L1 error of less than $10^{-4}$.

119. 【2610.05552】Monocular markerless biomechanics for clinically interpretable gait assessment in spinal cord injury

链接:https://arxiv.org/abs/2610.05552

作者:Shreyasvi Natraj,Mathieu Ruepp,Yanke Li,Robert Riener,Inge Eriks-Hoogland,Diego Paez-Granados

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Three-dimensional gait analysis, spinal cord injury, Three-dimensional gait, analysis guides rehabilitation, marker-based motion capture

备注:

点击查看摘要

Abstract:Three-dimensional gait analysis guides rehabilitation after spinal cord injury but depends on marker-based motion capture and force plates, which few clinics have. Monocular markerless pipelines have been established in fewer healthy adult cohorts but not in neurological cohorts. We present the SCAI SCI Gait dataset, comprising 239 adult individuals with spinal cord injury with synchronized video, motion capture, and force-plate measurements, we fitted a parametric body mesh to a single sagittal-view video, driving an anthropometrically scaled OpenSim model via virtual markers. Markerless lower-body kinematics showed state-of-the-art agreement with motion-capture measurements (r = 0.68-0.90, p 0.001, and RMSE = 4.18-6.49 degrees), and accurate kinematics-based predicted ground-reaction forces closely matched those measured by force plates (r = 0.85-0.87, p 0.001, and RMSE = 2.13-2.19 Newton per kg). Furthermore, conditional-dependence graph analysis with Markov blankets revealed that waveform components were conditionally associated with functional independence, and speed-stratified clustering revealed distinct mechanical strategies among individuals walking at similar speeds. These findings establish the use of monocular video as a scalable approach for clinically meaningful biomechanical assessment and data-driven phenotyping in patients with spinal cord injury. Github: this https URL

120. 【2610.05531】Deep Prior Learning for Embodied Perception

链接:https://arxiv.org/abs/2610.05531

作者:Yimou Wu,Jiaxin Guo,Yun-hui Liu,Zheng Li

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:emph, Geometry Grounded Transformer, exploits available observations, observations beyond images, Embodied systems

备注:

点击查看摘要

Abstract:Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.

121. 【2610.05529】DynaMesh: Dynamic 3D Texture Generation

链接:https://arxiv.org/abs/2610.05529

作者:Raj Hansini,Guan Chen,Rana Hanocka,Itai Lang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:video, dynamic texture generation, texture generation method, object geometry, texture generation

备注: Project page: [this https URL](https://threedle.github.io/dynamesh/)

点击查看摘要

Abstract:We present DynaMesh, a dynamic texture generation method for 3D meshes. Given a textureless shape and a text prompt describing an effect, our method produces an appearance that evolves while the object's geometry remains unchanged. Previous works on dynamic 3D content generation have focused on motion, where an object's geometry and position change while keeping its appearance the same. Methods on texture generation sit on the other side of the problem, painting appearance onto a shape as a fixed surface property and not as an evolving process. Neither addresses a visual effect that propagates on a 3D object. A natural route consists of two generators: a video model that shows the effect from a single view, and an image-to-3D generator that lifts each frame to 3D. However, the latter has no notion of time, so running it per video frame produces a sequence that flickers, loses effect details, and yields a different mesh at every video frame. Our method addresses these failures by conditioning a video model on a render of the mesh and the prompt to obtain a reference video, then running a frozen image-to-3D generator on the video with two changes. The conditioning of each frame is blended over a temporal window, and low-rank adapters are fit per shape to restore the lost details. The mesh is encoded once for the whole sequence, so geometry is constant by construction, and the output is a single mesh with a texture per frame. Applied to various objects and effects, DynaMesh substantially improves over recent video-to-4D and texturing methods, and can generalize its temporal effect to different shapes never seen during training. Our project page is at this https URL.

122. 【2610.05505】Robust 2D Traversability Mapping for Construction AMRs via Failure-Mode-Aware Fusion of LiDAR Geometry and Monocular Semantics

链接:https://arxiv.org/abs/2610.05505

作者:Manoj Karnekar,Om Mandhane,Gautham Ramkumar

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Autonomous Mobile Robots, severe navigational challenges, misses visually hazardous, produces phantom obstacles, Autonomous Mobile

备注: Extended version of a paper presented at the 5th Workshop on Future of Construction, IROS 2026

点击查看摘要

Abstract:Autonomous Mobile Robots (AMRs) on active construction sites face severe navigational challenges: geometry-based traversability mapping (e.g., LiDAR) misses visually hazardous but geometrically flat surfaces like wet mud and ponding concrete, while abrupt geometry on drivable speed-breakers and inclines produces phantom obstacles. We propose a real-time, failure-mode-aware multimodal traversability pipeline on an NVIDIA Jetson AGX Orin, where LiDAR is the primary geometric safety estimate and monocular semantics act as a selective, class- and confidence-gated corrective signal. The representation retains distinct traversable classes, namely flat road, terrain, and rocky terrain, while flagging construction hazards. We also release a multimodal construction-site dataset from a custom AMR: four closed-loop ROS 2 sequences from two active sites (RGB, depth, LiDAR, IMU, GPS-RTK, odometry) plus 506 annotated frames across 28 semantic classes. By projecting LiDAR onto dense semantic masks, resolving sparsity via morphological in-painting, and applying failure-mode-aware fusion with Patchwork++, the system corrects complementary geometric failure modes for a local AMR costmap.

123. 【2610.05491】Robust Surgical Robotic Instrument Tracking via Sequential Multi-Cue Fusion and Sim-to-Real Self-Training

链接:https://arxiv.org/abs/2610.05491

作者:Hanyang Hu,Zekai Liang,Florian Richter,Michael C. Yip

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:minimally invasive surgery, robot-assisted minimally invasive, remains challenging due, surgical robotic instruments, tracking of surgical

备注:

点击查看摘要

Abstract:Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based approaches are efficient, but their performance depends on reliable feature detection. Improving these detectors with real-world supervision is difficult because accurate real-world annotations are costly to obtain at scale. To address this limitation, we introduce a tracker-guided self-training framework that adapts a model pretrained on synthetic images to unlabeled real-world videos. Given measured robot joint states, an uncertainty-aware EKF recursively corrects the instrument pose and the observable joint angles by comparing projected model features with detected keypoints, shaft boundaries, and mask-derived cues. An RTS smoother subsequently refines the resulting trajectory, which is projected into pseudo-labels for fine-tuning the feature detector without laborious pose annotations. Experiments on real-world videos demonstrate consistent improvements from self-training across all evaluated keypoint metrics, and the resulting model outperforms prior approaches in both accuracy and runtime. The code and data will be released upon publication.

124. 【2610.05484】Universal Test-Time Training

链接:https://arxiv.org/abs/2610.05484

作者:Zefan Cai,Qinzhe Hu,Ziqiao Ma,Hao Tan,Junjie Hu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent Test-Time Training, architectures compress context, Universal Test-Time Training, Recent Test-Time, Test-Time Training

备注: 37 pages. Project page: [this https URL](https://zefan-cai.github.io/uTTT.github.io/) ; code: [this https URL](https://github.com/Zefan-Cai/uTTT)

点击查看摘要

Abstract:Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.

125. 【2610.05463】Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs

链接:https://arxiv.org/abs/2610.05463

作者:Renchi Zhang,Joost C. F. de Winter,Dimitra Dodou,Harleigh C. Seyffert,Yke Bauke Eisma

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:fundamental cognitive ability, Large Language Models, Multimodal Large Language, cognitive ability, Visual search

备注:

点击查看摘要

Abstract:Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates ($\rho = 0.82$), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.

126. 【2610.05453】he Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models

链接:https://arxiv.org/abs/2610.05453

作者:Tobias Braun,Jonas Henry Grebe,Emil Sivic,Patrick Mohr Gordillo,Hossein Shakibania,Marcus Rohrbach,Anna Rohrbach

类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:shared conversational context, increasingly shifting, architectures that understand, generate text, Multimodal models

备注: Code: [this https URL](https://github.com/multimodal-ai-lab/PLW)

点击查看摘要

Abstract:Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.

127. 【2610.05425】Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

链接:https://arxiv.org/abs/2610.05425

作者:Ali Vosoughi,Akhil Kasturi,Chenliang Xu,Axel Wismueller

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Automated checks, radiology reports, reports may rely, rely on AI-generated, leave findings unmentioned

备注: 40 pages, 7 figures, 17 tables (main text and references pp. 1-19; Supplementary Information as an appendix, pp. 20-40). Submitted to npj Digital Medicine. Code: [this https URL](https://github.com/ali-vosoughi/VerifyGRPO-Rad)

点击查看摘要

Abstract:Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.

128. 【2610.05418】EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation

链接:https://arxiv.org/abs/2610.05418

作者:Yuheng Na,Zhide Zhong,Junjie He,Junfeng Li,Haodong Yan,Jiaan Wang,Jiaguan Zhu,Yangyang Zheng,Tianyu Huang,Haoang Li

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:models rely, limiting performance, lose task-relevant evidence, rely on current, current observations

备注:

点击查看摘要

Abstract:Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7\% on RMBench, 82.0\% on RoboMME and 83.8\% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.

129. 【2610.05417】Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs

链接:https://arxiv.org/abs/2610.05417

作者:Yuqun Wu,Yao Xiao,Chuhang Zou,Shenlong Wang,Derek Hoiem

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent works augment, works augment Vision-Language, Recent works, augment Vision-Language Models, boost spatial reasoning

备注:

点击查看摘要

Abstract:Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: this https URL.

130. 【2610.05416】Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

链接:https://arxiv.org/abs/2610.05416

作者:Shuyuan Tu,Qi Tian,Yinming Huang,Yue Wu,Xintong Han,Kaihang Pan,Weijie Kong,Jiangfeng Xiong,Jian-Wei Zhang,Zuxuan Wu,Yu-Gang Jiang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:higher resolutions empowers, sharper motion dynamics, Natively training joint, learn richer visual, richer visual details

备注:

点击查看摘要

Abstract:Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informative content and disrupting pretrained priors. Existing sparse attention methods either target training-free acceleration or overlook the unique structure of joint video-audio data, where cross-modal interactions are inherently concentrated around sound-producing regions. To address this, we propose Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K. In particular, Prism organizes the token sequence into spatiotemporal macro-zones, enabling the attention structure to adapt to local content. For each zone, it estimates local information structure via video feature variance along the channel and feature norms from the audio-to-video cross-attention, jointly capturing how visual content varies directionally and how strongly audio influences each visual region. Based on these signals, Prism dynamically assigns a tailored block shape to each zone, applying finer partitioning along axes of rapid visual content variation and strong audio-visual coupling. This encourages tokens within each block to remain semantically coherent, allowing block-level features to capture both visual content and joint video-audio interaction patterns. Prism further adopts a hybrid block selection strategy to dynamically determine per-query sparsity. Experiments show that Prism achieves 2.5$\times$ training speedup compared to full attention, while surpassing it in generation quality.

131. 【2610.05413】A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models

链接:https://arxiv.org/abs/2610.05413

作者:Yilin Yang,Jun-Tao Tang,Kengyi Wang,Siyuan Su,Gaoyong Luo,Mingda Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:large language models, multimodal large language, language models, reliably predict, predict their downstream

备注: Code is available at [this https URL](https://github.com/JuntaoTang/MLLM-VisionEncoder-Eval)

点击查看摘要

Abstract:Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.

132. 【2610.05411】CleanMDM: Clean Motion Diffusion Model for Multimodal Motion Cleanup

链接:https://arxiv.org/abs/2610.05411

作者:Zhe Li,Shicheng Wang,Bowen Cai,Huan Fu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:rarely directly usable, typically exhibit missing, exhibit missing segments, Motion capture data, directly usable

备注:

点击查看摘要

Abstract:Motion capture data is rarely directly usable, as they typically exhibit missing segments, jitter, drift and contact artifacts. Traditionally, corrupted motions are cleaned by animators through the manual identification of keyframes from noisy motion, subsequent keyframe correction, and interpolation between corrected keyframes to reconstruct coherent motion. While the rise of generative motion models has made automatic cleanup feasible, most approaches operate as black box denoisers with limited controllability, making it difficult to preserve reliable segments or enforce specific user intents. Inspired by animation workflows, we present CleanMDM, a unified multimodal motion cleanup framework that formulates cleanup as masked conditional generation with plug-and-play conditions. This single model supports arbitrary combinations of noisy 3D motion, sparse 2D keyframes, sparse 3D keyframes, and text. This design enables both automatic cleanup without additional user annotation and controllable cleanup under multimodal guidance. To further improve motion realism, we incorporate the Latent Motion Quality Discriminator (LMQD) to better match kinematic distributions and reduce skating, jitter, and interpenetration artifacts, and we apply Mesh-Aware Contact Projection as a test-time optimization step to enhance contact and physical consistency. Experiments across multiple datasets demonstrate that CleanMDM consistently outperforms prior cleanup and generation baselines, and that low cost conditions (text and 2D keyframes) provide reliable controllability gains in multimodal cleanup scenarios.

133. 【2610.05400】CASE: Cost-Aware Stopping for Efficient Long-Video Agents

链接:https://arxiv.org/abs/2610.05400

作者:Yiming Du,Chenghao Liu,Zhiyuan Liu,Fangxing Zheng,Zhao Wang,Junnan Nie,Songfang Huang

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:actively gather question-relevant, central decision implicit, gather question-relevant evidence, actively gather, gather question-relevant

备注: 40 pages, including references and appendices

点击查看摘要

Abstract:Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.

134. 【2610.05353】FACET: Factorized Asymmetric Conditioning for Efficient Transport in High-Fidelity Fluorescence Microscopy Synthesis

链接:https://arxiv.org/abs/2610.05353

作者:Sazan Mahbub,Caleb N. Ellington,Eric P. Xing

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Fluorescence microscopy reveals, Fluorescence microscopy, cell morphological context, morphological context enables, microscopy reveals

备注:

点击查看摘要

Abstract:Fluorescence microscopy reveals where proteins localize, but only a limited number of proteins can be imaged in the same cell; generating these images from amino-acid sequence and the cell's morphological context enables in silico localization of unimaged proteins. The two conditions, however, play asymmetric roles: morphological context is spatially aligned with the target, whereas sequence is non-spatial and must specify protein-dependent localization within it, with recurring coarse patterns shared across proteins and finer protein-specific variation. Existing generators condition on both jointly, without separating what each explains. We introduce FACET (Factorized Asymmetric Conditioning for Efficient Transport), a probabilistic generative framework that encodes this structure as an explicit inductive bias: sequence semantics are learned from what context leaves unexplained, coarse localization regularities are shared across proteins through a semantic memory, and protein-specific variation is a bounded residual around them. A variance-preserving state projection further lets FACET perform continuous stochastic transport through a pretrained diffusion predictor with minimal parameter overhead. On held-out proteins, FACET improves spatial overlap by 34.3% on the Human Protein Atlas and 14.0% on OpenCell over a backbone-matched baseline, and reduces FID by 27.2% and 46.5%, respectively, with 75% fewer network evaluations. It also substantially improves protein-association structure recovery and yields better-calibrated predictions, while detailed ablations show complementary contributions from its design choices. These results identify factorized asymmetric conditioning, rather than generator capacity alone, as a key lever for high-fidelity, efficient, and biologically meaningful cellular image synthesis.

135. 【2610.05349】Learning Conditional Source Distribution via Flow Reversal for Temporal Flow Matching

链接:https://arxiv.org/abs/2610.05349

作者:Kuan-Hsun Tu,Hsuan-Chi Liu,Jia-Wei Liao,Chien-Sheng Chiang,Tsung-Wei Ke

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:flow matching framework, learns conditional source, conditional flow matching, flow reversal, flow matching

备注:

点击查看摘要

Abstract:We introduce CNP-Flow, a flow matching framework for temporal generation that learns conditional source distributions through flow reversal. Whereas standard conditional flow matching (FM) incorporates conditioning through the vector field and draws source samples from a standard Gaussian, CNP-Flow uses a conditional noise predictor (CNP) to produce an isotropic Gaussian source for each temporal condition. The CNP is supervised by source samples obtained through flow reversal, which maps observed targets backward through a pretrained FM model. A three-stage pipeline pretrains the FM model, trains the CNP, and fine-tunes the FM model using the learned source distribution, while preserving the FM backbone architecture. Across video prediction, video interpolation, and 7-DoF Franka robot motion planning, CNP-Flow consistently improves generation quality. It also matches baseline performance with fewer function evaluations. Project page: this https URL

136. 【2610.05342】IRSTD-Agent: Agentic Infrared Small Target Detection via Zoom-Guided Interaction Learning

链接:https://arxiv.org/abs/2610.05342

作者:Jiawen Xi,Yu Zhang,Tianyi Zhao,Zhu Liu,Maoxun Yuan,Xingxing Wei

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:small-target detection plays, Infrared small-target detection, aerial surveillance, plays an important, important role

备注:

点击查看摘要

Abstract:Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisely localize infrared small targets. In this paper, we propose IRSTD-Agent, an agentic framework for infrared small target detection through dynamic visual search. The framework enables an MLLM to adaptively determine where and at what scale to inspect an image and progressively gather fine-grained visual evidence for precise target localization. Five complementary visual tools (PROPOSAL, ZOOM, DETECT, DROP and REFINE) support object candidate discovery, adaptive observation, target localization, hypothesis rejection, and target extent refinement, together enabling a coordinated search process over original-resolution images. To teach the MLLMs to conduct this search, we introduce Zoom-guided Interaction Learning, which uses annotation-derived interaction trajectories to supervise tool selection and the corresponding arguments. Through extensive experiments on WideIRSTD-Full and IRSTD-1k datasets, we demonstrate that IRSTD-Agent outperforms the evaluated vision-language models and enhances the precise localization capabilities of MLLMs in IRSTD tasks.

137. 【2610.05341】WILLIE: A Unified Framework and Benchmark for Wound Classification, Segmentation, and Localization

链接:https://arxiv.org/abs/2610.05341

作者:Gopi Trinadh Maddikunta,Shannan Hamlin,Hsin-Mei Chen,Kimaya Barnes,Peizhu Qian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:United States, Chronic wound management, States and imposes, imposes substantial clinical, wound management affects

备注: 21 pages, 4 figures, 9 tables. Published in Proceedings of the 11th Machine Learning for Healthcare Conference (MLHC 2026), PMLR 340:1243-1263. Code: [this https URL](https://github.com/Qian-Group-HRI/Willie)

点击查看摘要

Abstract:Chronic wound management affects over 8.2 million patients in the United States and imposes substantial clinical and economic burden. Clinical wound assessment commonly involves three coupled tasks: identifying wound type, delineating wound boundaries, and localizing the wound region for measurement and monitoring. Despite this clinical coupling, existing machine learning approaches typically address wound classification, segmentation, and localization using separate models. We present WILLIE, a unified framework and benchmark for wound classification, segmentation and localization that enables systematic evaluation of multi-task wound analysis under a common protocol. WILLIE harmonizes three public wound datasets into a shared benchmark and compares unified models across three scaling configurations against 10 single-task baselines. The best model achieves 91.88% classification accuracy, 91.41% Dice, and 96.23% AP@0.5 while producing all three outputs in a single forward pass. Beyond aggregate performance, our results show that segmentation-derived localization outperforms dedicated detection baselines in this benchmark, suggesting that box-based localization may be unnecessary for spatially coherent wound targets. Our findings highlight that effective multi-task learning in healthcare imaging depends not only on shared representations, but also on task formulation, compatibility, and benchmark design.

138. 【2610.05326】Riemannian Shape Analysis of the Corpus Callosum in Kendall Space: Aging and Alzheimer's Disease

链接:https://arxiv.org/abs/2610.05326

作者:Olakunle S. Abawonse,Fatou Fall

类目:Computer Vision and Pattern Recognition (cs.CV); Differential Geometry (math.DG); Optimization and Control (math.OC)

关键词:major white-matter structure, corpus callosum, boundary geometry, major white-matter, white-matter structure

备注: 19 pages, 3 figures

点击查看摘要

Abstract:The corpus callosum (CC) is a major white-matter structure and a well-established marker of brain aging, but most studies quantify it using scalar summaries that discard its boundary geometry. We present a Riemannian shape-space framework for analyzing age-related morphological change in the midsagittal CC, applied to the OASIS-1 cohort. Each contour is represented by $128$ landmarks and embedded into Kendall shape space, where translation, rotation, and scale are removed. We derive a multivariate geodesic regression with exact Riemannian gradients and use the fitted age-velocity field to localize age-related deformation to five anatomical sub-regions. In the cognitively normal cohort ($n = 252$), geodesic regression outperforms the Euclidean linear benchmark ($R^2 = 0.1355$ vs.\ $0.1216$). Regional energy is posterior-dominant: the Splenium carries $41.4\%$ and the Isthmus $23.8\%$ of total age-related shape change, together accounting for $\sim 65\%$ despite comprising only $\sim 35\%$ of landmarks. Signed projections confirm the ordering (Splenium $r = 0.570$; Isthmus $r = 0.473$). In contrast, age explains less than $0.5\%$ of shape variance in Alzheimer's disease ($n = 88$), indicating that the disease disrupts the healthy aging trajectory. A tangent-space classifier achieves an age-group AUC of $0.791$ from the 2D contour alone, exceeding a recent volumetric benchmark ($0.67$).

139. 【2610.05325】Shadow Feature Refinement Network: Progressive Feature Refinement based on Knowledge Distillation for Effective Shadow Removal

链接:https://arxiv.org/abs/2610.05325

作者:Donghyun Han,Byoung-Dai Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:persistent challenge owing, significant advancements, shadow removal remains, field of deep, remains a persistent

备注:

点击查看摘要

Abstract:In the field of deep learning, has seen significant advancements; however, shadow removal remains a persistent challenge owing to the variable sizes and colors of shadows influenced by lighting conditions. This study proposes a novel shadow feature refinement network (SFR-Net), which leverages supervised learning, feature refinement loss, and knowledge distillation to enhance shadow removal performance. A dedicated post-processing algorithm is further introduced to restore natural color consistency in the generated shadow-free images. We evaluated our method on two public datasets: the adjusted image shadow triplet dataset (ISTD+) and the shadow removal dataset (SRD), which demonstrate strong generalization capabilities under diverse conditions. On ISTD+, our model achieved a root mean square error (RMSE) of 3.4627 and structural similarity index measure (SSIM) of 0.9382 across the entire image. On SRD, it recorded an RMSE of 4.3781 and an SSIM of 0.9341. These comprehensive results show that our approach performs competitively across both shadow and non-shadow regions while setting a promising direction for robust and perceptually natural shadow removal. Code is available at this https URL.

140. 【2610.05306】Kinematics-Centric Continuous Sign Language Retrieval with Gloss-Guided Boundary-Aware Alignment

链接:https://arxiv.org/abs/2610.05306

作者:Chang Liu,Ke Han,Davide Talon,Elisa Ricci,Nicu Sebe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Sign language-text alignment, language-text alignment remains, sign language understanding, remains a fundamental, fundamental challenge

备注: Accepted by ACM Multimedia 2026 (ACM MM 2026)

点击查看摘要

Abstract:Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sign language-text alignment in a structured kinematic space and propose a kinematics-centric framework that adopts 3D SMPL-X motion as the primary representation. By explicitly modeling the kinematic dynamics of signing in a unified motion space, our approach reduces reliance on appearance signals and yields more semantically consistent representations. To capture the compositional nature of sign language, we introduce a gloss-guided local alignment mechanism that leverages gloss temporal spans as weak supervision to decompose continuous motion into coherent segments and establish fine-grained motion-text correspondences, thereby reducing ambiguity in localizing word-level semantics in continuous signing. Furthermore, we develop a visual distillation strategy, where RGB signals serve as privileged supervision during training to provide complementary contextual cues, while being completely removed at inference time. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art bidirectional retrieval performance on CSL-Daily and competitive results on PHOENIX-2014T. These results highlight the effectiveness of kinematic representations and explicit local grounding for sign language-text alignment.

141. 【2610.05298】Revisiting Ground-Truth Synthesis from High-Speed Video: Exact Validity Conditions and an Audited Consumer Capture Corpus

链接:https://arxiv.org/abs/2610.05298

作者:Abdullah Al Shafi,Sumaiya Rahim Suma

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Motion-deblurring datasets, synthesised by averaging, commonly synthesised, labelling the result, Motion-deblurring

备注:

点击查看摘要

Abstract:Motion-deblurring datasets are commonly synthesised by averaging $N$ consecutive high-frame-rate frames and labelling the result with the central frame. We show that this label is unbiased for every capture timing only when the window is odd and the frames' sample durations are equal. An even window shifts every label by a fixed fraction of the blur length, even under perfect timing. Unequal durations are subtler: their mean misalignment is zero, so dataset statistics cannot reveal them, yet when the offset cannot be read from the blur they convolve the supervision rather than adding noise to it. Auditing 51 smartphone clips recorded at a nominal 240 fps, we find 19 captured near 176 fps, a behaviour recorded only in the container's timing tables. In a controlled test the parity choice costs a fitted linear deblurring filter far more than these timing irregularities do, and interpolation labels, unlike blur labels, can be repaired with the true frame times. We provide a tool that checks both conditions without decoding, and will release the clips and their timing tables.

142. 【2610.05293】BossouChimpanzee: Long-term Chimpanzee Video Dataset

链接:https://arxiv.org/abs/2610.05293

作者:Daniel Schofield,Susana Carvalho,Vladimir Iashin,Andrew Zisserman,Max Bain,Arsha Nagrani,David Ng,Claudia Sousa,Boniface Zogbila,Jules Doré,Dora Biro,Misato Hayashi,Tetsuro Matsuzawa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unique long-term visual, long-term visual record, continuous video recordings, video recordings collected, BossouChimpanzee video dataset

备注: 12 pages, 3 figures, 5 tables. Dataset: [this https URL](https://robots.ox.ac.uk/~vgg/data/BossouChimpanzee/)

点击查看摘要

Abstract:We describe the BossouChimpanzee video dataset, a unique long-term visual record of wild chimpanzees at an outdoor laboratory for field experiments in Bossou, Guinea, spanning three decades (1988-2018) and comprising over 1,200 hours of continuous video recordings collected through collaborative fieldwork and research. In this paper, we outline the history and scientific contributions of the experimental paradigm and video archive, provide key statistics and details on the structure of the main video dataset, and release an initial ~74h snapshot, BossouChimpanzee70h, covering 23 identified individuals focused on chimpanzee individual and action recognition, ahead of the full video resource. This dataset represents a valuable resource for cognitive and behavioural research in ethology and a rich benchmark for training and evaluating machine learning models on audiovisual data from the wild.

143. 【2610.05289】Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting

链接:https://arxiv.org/abs/2610.05289

作者:Xiaobiao Du,Beixi Hao,Zhen Fang,Tianqing Zhu,Richard Hartley,Xin Yu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:costly per-frame computation, achieved remarkable performance, remains challenging due, Recent advances, devices remains challenging

备注: Code has been released: [this https URL](https://xiaobiaodu.github.io/mobile-4dgs-project/)

点击查看摘要

Abstract:Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable performance in novel view synthesis, yet deploying both static and dynamic Gaussian representations on resource-constrained mobile devices remains challenging due to heavy storage, redundant primitives, and costly per-frame computation. We present Mobile-4DGS, a unified lightweight framework for high-fidelity real-time static and dynamic Gaussian rendering on mobile platforms. For compact appearance modeling, we introduce a Monte Carlo Specular Energy Aggregator that compresses high-order radiance residuals into the first-order Spherical Harmonics (SH), together with an Attribute-Conditioned SH Enhancement module whose predicted offsets are pre-baked before inference. We further propose a Multi-View Alpha-Based Densification and Pruning strategy to suppress redundant primitives while maintaining multi-view consistency. For dynamic scenes, we develop a compact explicit 4D representation by constructing second-order Gaussian motion, learnable temporal support, and a binary static-dynamic partition, enabling continuous-time modeling without runtime deformation networks. Based on this partition, a Depth-Order Certificate selectively reuses previously committed depth orders to reduce re-projection, sorting, merging, and index-buffer updates during playback. Extensive experiments on static and dynamic scenes demonstrate that Mobile-4DGS substantially reduces storage and rendering overhead while maintaining competitive visual quality, enabling real-time 3D and 4D Gaussian Splatting on mobile devices. \textcolor{magenta}{\href{this https URL}{Code has been released: this https URL}}.

144. 【2610.05274】Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes

链接:https://arxiv.org/abs/2610.05274

作者:Runwei Guan,Rongsheng Hu,Shangshu Chen,Ningwei Ouyang,Shaofeng Liang,Heyi Lin,Jinjing Zhu,Yang Shi,Dongming Wu,Daizong Liu,Henghui Ding,Hui Xiong

类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:Roadside traffic reasoning, traffic reasoning requires, visual evidence, reasoning requires, backed by visual

备注: 14 pages, 7 figures

点击查看摘要

Abstract:Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and boxes separately leave this failure unpenalized. We trace the mismatch to the conventional answer-then-ground factorization, which commits to a numerical claim before any object is enumerated. To measure it, we build RoadSceneVQA-G, a benchmark of 34.7K question-answer pairs in which every free-form answer is linked to the set of boxes that witnesses it, and we propose the Answer-Grounding Consistency (AGC) evaluation suite. To address it, we introduce Enumerate-then-Answer (EtA), which reverses the generation order so that answer-evidence agreement becomes a property of the output structure, and Enumeration-Consistent Policy Optimization (ECPO), a reinforcement learning stage that uses the union of multiple rollouts as a recall teacher without ground-truth boxes. EtA raises say-point consistency from 26.6\% to 93.7\% and grounding F1 from 52.2\% to 73.0\%, and ECPO further increases F1 to 75.6\% without per-box supervision. On gRefCOCO, the same framework outperforms the strongest compared method, indicating that it transfers beyond traffic scenes. The project is available at \url{this https URL}.

145. 【2610.05273】When and What to Prune? Stage-Aware Visual Token Pruning for Efficient VLA

链接:https://arxiv.org/abs/2610.05273

作者:Tianjun Shi,Haotian Xiong,Ziyu Gong,Qi Lu,Li Li

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:predicting robot actions, accelerate vision-language models, pruning, accelerate vision-language, processed before predicting

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.

146. 【2610.05254】Hybrid-Basis Feature Forecasting for Diffusion Sampling Acceleration

链接:https://arxiv.org/abs/2610.05254

作者:Kai-Liang Cheng,Yuan-Yuan Cheng,Yu-fan Jin,Xiao-Ming Fu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Hybrid-Basis Feature Forecasting, accelerating diffusion sampling, propose Hybrid-Basis Feature, Feature Forecasting, framework for accelerating

备注:

点击查看摘要

Abstract:We propose Hybrid-Basis Feature Forecasting (HybridFF), a training-free, plug-and-play framework for accelerating diffusion sampling. To capture local smoothness, long-range trends, and complex non-monotonic variations when modeling feature evolution, HybridFF first estimates coefficients using moving least squares (MLS) for each of multiple complementary basis families and then combines the corresponding predictors using fusion weights. In addition to the choice of basis functions, the fusion weights also play a critical role. We introduce two strategies to balance quality and speedup. HybridFF (Fixed) prioritizes efficiency with model-specific fusion weights calibrated on a small set and held constant during inference. HybridFF (Adaptive) updates the fusion weights online using branch reliability scores computed from an exponential moving average of full-step prediction errors, improving prediction fidelity and generation quality under aggressive caching while retaining substantial acceleration. Experiments across DiT-XL/2, FLUX.1-dev, SD3.5-Large, and HunyuanVideo demonstrate a favorable speedup--quality trade-off over representative single-basis forecasters and caching baselines.

147. 【2610.05252】MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation

链接:https://arxiv.org/abs/2610.05252

作者:Yunyi Chen,Chenru Wang,Xinyi Ye,Zexin Zheng,Chi Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Diffusion-based dataset distillation, decision boundaries required, fundamental objective mismatch, discriminative decision boundaries, likelihood-driven diffusion models

备注: 32 pages

点击查看摘要

Abstract:Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.

148. 【2610.05249】ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image

链接:https://arxiv.org/abs/2610.05249

作者:Kai Lv,Yibo Yin,Lijun Guo,Heng Fan,Kaihao Zhang,Xingping Dong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Embodied agents benefit, Embodied agents, agents benefit, scene layout, scene layout consistent

备注: 25 pages

点击查看摘要

Abstract:Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.

149. 【2610.05233】PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment

链接:https://arxiv.org/abs/2610.05233

作者:Gavriel Habib,Dvir Samuel,Or Shimshi,Rami Ben-Ari

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:long-term identity stability, head-avatar reenactment aims, aims to animate, robust motion transfer, requiring robust motion

备注:

点击查看摘要

Abstract:Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.

150. 【2610.05207】F$^2$ SLAM: Turning Feed-Forward Geometry into Persistent Factors for SLAM

链接:https://arxiv.org/abs/2610.05207

作者:Zhisong Xu,Fan Zhu,Jiawei Qian,Ziyu Chen,Zhenjun Zhao,Javier Civera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:line simultaneous localization, multi-view geometric priors, models provide strong, provide strong multi-view, strong multi-view geometric

备注: 21pages

点击查看摘要

Abstract:Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically treat feed-forward predictions as an external geomet- ric state that is aligned or fused with the online estimate after the fact, which keeps broader multi-view evidence outside the optimizer that refines the SLAM state. We present F2SLAM, which instead converts feed-forward geometry di- rectly into optimization-native target-weight measurements attached to a persis- tent dense factor graph. A high-frequency stream maintains local tracking con- straints and graph connectivity, while a low-frequency stream uses wider multi- view context to selectively refresh existing measurements after a state-consistency check. Both streams constrain the same poses, inverse depths, and optional cam- era intrinsics through a single dense bundle adjustment. Experiments on multiple benchmarks demonstrate consistently strong trajectory estimation and improved dense reconstruction in both calibrated and uncalibrated settings. Notably, the uncalibrated configuration reduces the average ATE RMSE from 0.030 m for the strongest feed-forward baseline to 0.002 m on the Replica dataset.

151. 【2610.05195】Long-MDR: Long-Context Reinforcement Learning for Multimodal Deep-Research Agents

链接:https://arxiv.org/abs/2610.05195

作者:On Tai Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:short model completions, long-lived research histories, multimodal research agents, model completions, reason over long-lived

备注:

点击查看摘要

Abstract:The next generation of multimodal research agents must reason over long-lived research histories rather than short model completions. During a single task, an agent may repeatedly search the web, inspect visual evidence, revisit earlier hypotheses, and accumulate tens of thousands of tokens of multimodal context. Despite this trend, online RL for multimodal research agents remains largely confined to shorter contexts and interaction horizons. We push online RL training to 128k context and 75+ tool-interaction turns. To our knowledge, this is the first online multimodal deep-research RL study trained at 128k context, and the first trained with a 75 tool-turn horizon. Scaling to this regime exposes several practical limitations of conventional RL training. Early in training, weak policies make poor use of large interaction budgets, causing expensive rollouts with little reward improvement. Later, policy entropy can collapse before performance has saturated, prematurely ending useful learning. We introduce Long-MDR, a three-component training recipe designed specifically for this setting: On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue. Together, these techniques improve both the learning efficiency and stability of long-horizon RL, enabling continued gains in a regime where direct training is slow and costly. At a 50-turn evaluation budget, our RL-trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B-9B agents.

152. 【2610.05191】Order Matters: Competition-Guided Query Ordering for RNN-Based Object Detection

链接:https://arxiv.org/abs/2610.05191

作者:Shengjian Wu,Li Sun,Yu Shangguan,Qingli Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:bipartite matching, non-maximum suppression, assign object queries, NMS, assign object

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:DETR-style detectors use one-to-one bipartite matching during training to assign object queries to ground-truth objects, enabling end-to-end set prediction without non-maximum suppression (NMS). However, without an explicit de-duplication procedure, multiple queries can still produce highly similar hypotheses for the same object, making training unstable and predictions less decisive. Inspired by the sequential ordering of NMS, we propose DETRNN, a plug-and-play module that turns unordered object queries into a competition-aware sequence for recurrent refinement. DETRNN builds an explicit confidence-and-similarity based order from prior predictions, then refines queries with an RNN along this order to model competition inside the decoder. This ordered recurrent refinement reduces redundant predictions, stabilizes optimization, and improves final detection accuracy. Experiments on multiple DETR-style detectors show consistent gains with comparable efficiency.

153. 【2610.05185】Recurrent Latent Visual Search for GUI Grounding

链接:https://arxiv.org/abs/2610.05185

作者:Kaiyu Wu,Beichen Zheng,Weiyao Huang,Keze Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:GUI agents powered, execute user instructions, GUI agents, GUI grounding, capability for GUI

备注:

点击查看摘要

Abstract:GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motivating multi-step visual search. However, existing approaches commonly rely on textual reasoning misaligned with visual space or costly multi-round interactions with external visual tools. To make multi-step visual search an explicit spatial process within the model, we propose ReLaViS, which performs Recurrent Latent Visual Search in a single interaction round. At each step, a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that explicitly represents the search focus. This distribution then aggregates the visual tokens into latent visual evidence, which is recurrently fed back as the next input embedding to condition subsequent search. We further introduce a GUI-aware coarse-to-fine inductive bias through trajectories constructed from flat element annotations, supervising search from the global interface through intermediate element groups to the target. Built on Qwen2.5-VL-7B, ReLaViS improves ScreenSpot-Pro accuracy by 3.1 percentage points to 56.3% with only a 3.5% increase in inference FLOPs and outperforms the matched single-step baseline on all five benchmarks.

154. 【2610.05154】Construction and Evaluation of Machine Learning Models for Near-Real-Time Fire Detection from MTG FCI Imagery

链接:https://arxiv.org/abs/2610.05154

作者:Asaf Vanunu,Boaz Nadler,Arnon Karnieli

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Geostationary satellite observations, Geostationary satellite, MTG FCI, satellite observations, observations are important

备注:

点击查看摘要

Abstract:Geostationary satellite observations are important for wildfire detection and monitoring. The current study evaluates machine learning models for MTG FCI near-real-time fire detection in 1- and 2-km spatial configurations and compares them with threshold-based algorithms. The models were trained and evaluated using VIIRS fire reference data across diverse ecological regions in Europe, Africa, and the Middle East. The key results are that 1-km models significantly outperform both their 2-km variants and operational threshold products. The constructed 1-km models achieved F1 scores higher by up to 0.36 compared to baseline products. Importantly, the 1-km models detected small fires with higher probability compared to competing models. Finally, the models robustly detected fires up to 260 min earlier than baseline products. To support opensource applications, our trained models are publicly available.

155. 【2610.05141】SemCam: Semantic Camera Motion Control for Video Generation

链接:https://arxiv.org/abs/2610.05141

作者:Janna Bruner,Omer Talmi,Ianir Ideses,Lior Fritz,Lior Wolf,Sagie Benaim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:frontal view require, subject changing position, desired trajectory difficult, position and orientation, frontal view

备注:

点击查看摘要

Abstract:Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.

156. 【2610.05135】How Does Geometry Enter Generated Motion?

链接:https://arxiv.org/abs/2610.05135

作者:Weihan Li,Junhao Wu,Yuhan Song,Xiaofeng Lin,Xinlei Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:scene determines, fixed physical law, visible geometry, physical law, motion

备注: 43 pages, 13 figures, including supplementary material. Project page: [this https URL](https://liweihan1107.github.io/samelaw/)

点击查看摘要

Abstract:Under a fixed physical law, the visible geometry of a scene determines how motion must change. We ask how video generators realize this relationship. We fix the law and the initial state and change only the geometry drawn in the first frame, within matched families of tracks and deflectors, and compare each generated trajectory with the simulator prediction for that geometry. Paired interventions change one thing at a time: a local bump, the height of a barrier, the words of the prompt, the length of the clip. Across nine image-to-video models, geometry is preserved and shapes the motion: the speed of the ball follows the drawn undulation of a track. A physical state would carry this response forward, and here the generated motion parts from the law. The mean slope barely accelerates the ball, successive contacts fail to compose through a consistent state, an edit ahead of the ball alters its motion before it arrives, and the ball climbs over barriers higher than its release point. Two global conditions organize the global trajectory: text strongly controls the destination, while clip length strongly controls timing in the open-weight models tested. The pattern persists with photographed first frames. Current video generation thus behaves as geometry-conditioned motion synthesis whose evolution of state differs systematically from that of a fixed physical law.

157. 【2610.05131】RoMod: Temporal Routing Modulation via Mixture-of-Experts for Video Anomaly Detection

链接:https://arxiv.org/abs/2610.05131

作者:Chao Huang,Pengfei Wei,Benfeng Wang,Chengliang Liu,Wei Wang,Li Shen,Wenqi Ren,Xiaochun Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:power remains unclear, multimodal large language, shown strong potential, discriminative power remains, large language models

备注:

点击查看摘要

Abstract:Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, whose explicit expert structure and sparse activation make their internal computation easier to inspect. With a fully frozen backbone and no additional training, we find that anomaly-related evidence is concentrated in a small set of experts. These experts recur across layers, spontaneously specialize in different anomaly types, and together form a dynamic routing subnetwork. We further show that the output channels most strongly influenced by these experts are also the hidden dimensions that contain the most anomaly-relevant information. Routing statistics can therefore serve as an internal anomaly cue that complements semantic this http URL on these findings, we propose RoMod, an efficient VAD framework trained with only \(5\%\) of weakly labeled videos. RoMod includes a Routing-Modulated Fusion module, RoMF, and a Routing-aware Temporal Network, RoTN. RoMF uses routing signals to adaptively recalibrate hidden semantic channels. Its design also prevents the routing branch from bypassing semantic features and making predictions on its own. RoTN captures the temporal evolution of anomalies from onset to persistence and termination. Experiments on three benchmarks show that RoMod achieves state-of-the-art performance while running substantially faster than dense backbones of comparable size.

158. 【2610.05129】Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection

链接:https://arxiv.org/abs/2610.05129

作者:Chao Huang,Pengfei Wei,Kaige Li,Chengliang Liu,Wei Wang,Wenqi Ren,Xiaochun Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, Language Models

备注:

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012\% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.

159. 【2610.05115】PCLM: Small-target localization with frozen CLIP via prototype contrast and local magnification

链接:https://arxiv.org/abs/2610.05115

作者:Zhipeng Ye,Feng Jiang,Qiufeng Wang,Hao Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:mix object appearance, frozen CLIP encoder, surrounding content, mix object, object appearance

备注:

点击查看摘要

Abstract:Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.

160. 【2610.05109】LoopMoEVR: Loop-Based Degradation-Aware Mixture-of-Experts for Unified UHD Video Restoration

链接:https://arxiv.org/abs/2610.05109

作者:Yucheng Xin,Runci Bai,Yongcong Wang,Guangwei Gao,Jiao Liu,Dianjie Lu,Linwei Fan,Zhuoran Zheng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unified high-definition image, attracted considerable attention, existing models tend, high-definition image restoration, unified high-definition

备注:

点击查看摘要

Abstract:Recently, unified high-definition image restoration has attracted considerable attention; however, existing models tend to excessively increase their depth in pursuit of improved generalization, which often yields only limited gains. Meanwhile, loop-based learning paradigms have drawn widespread attention due to their low parameter counts and strong regression capability, as exemplified by GPT-6 and looped Transformers. In this paper, we introduce the loop learning paradigm to address restoration tasks that require cross-domain learning. Specifically, we propose LoopMoEVR, a loop-based mixture-of-experts model capable of handling degraded ultra-high-definition (UHD) inputs. First, a degradation-conditioned low-rank loop embedding is designed to construct input-dependent stage conditions. Second, a spatio-temporal iterative adaptive normalization module, termed IterAda3DN, is developed to fuse local features with global loop context, thereby performing position-wise affine modulation. Finally, the expert branches further integrate the attention-updated local and global video states with the loop conditions to generate dedicated modulation parameters, while an input-conditioned depth predictor adaptively configures the number of loop iterations. With only approximately 0.884M trainable parameters, the proposed model uniformly handles UHD video dehazing, deraining, denoising, and low-light enhancement tasks, achieving state-of-the-art restoration performance on both public benchmarks and real-world scenarios.

161. 【2610.05097】ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory

链接:https://arxiv.org/abs/2610.05097

作者:Hao Jiang,Zhanyu Guo,Chenwei Wu,Yichen Guo,Qizhe Zhang,Junchi Yao,Jixian Wu,Jinhao You,Kai Tang,Jiajun Cao,Tinghao Wang,Mengyu Wang,Leo Anthony Celi,Shanghang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:multimodal large language, weakening visual grounding, visual input diminishes, initial visual input, large language models

备注: 32 pages. Code coming soon

点击查看摘要

Abstract:As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that local evidence benefits from global context, compact latent representations balance accuracy and visual-context cost, and the utility of memory access depends on the reasoning state. Guided by these findings, we propose ReMAP (Reasoning-Time Memory-Augmented Perception), which couples two complementary latent memories: a static, question-conditioned Global memory that preserves scene and cross-image context, and a dynamic Local memory that uses this context as an anchor while selecting and re-encoding region-level evidence according to the current reasoning state. Both memories return compact latent tokens inserted into the reasoning sequence, and a reinforcement-learning access policy trained with branched rollouts decides when to continue reasoning or invoke Global or Local memory. On ten benchmark families, ReMAP outperforms prior visual-memory methods on all four multi-image benchmarks, exceeding the strongest prior results on MuirBench and MIMIC by 8.38 and 14.84 percentage points. Across four backbone families, enabling memory access improves over the same trained model with memory disabled, and on shared V*Bench, CV-Bench-2D, and MuirBench questions ReMAP reduces the visual tokens entering the reasoning sequence by 51.0-76.8% relative to the native-resolution backbone. Further analyses show that Global and Local memory form distinct yet complementary latent representations. Together, these components restore the perceptual cycle by letting the reasoning state trigger targeted visual retrieval, with the retrieved evidence guiding subsequent reasoning.

162. 【2610.05078】CGDD-Net: Context-Guided Dynamic Detail Modeling for Retinal Vessel Segmentation

链接:https://arxiv.org/abs/2610.05078

作者:Xincheng Li,Xiaoqi Sheng,Xinyu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate retinal vessel, capture vascular geometry, fine-scale reconstruction, Accurate retinal, geometry while preserving

备注:

点击查看摘要

Abstract:Accurate retinal vessel segmentation requires features that capture vascular geometry while preserving information for fine-scale reconstruction. We propose CGDD-Net, a context-guided dynamic detail modeling network that connects adaptive feature extraction to a shared decoder pathway. Context-Guided Scale-Adaptive Deformable Encoding (CSDE) combines fixed-grid convolution with deformable local attention to capture vascular patterns at multiple spatial extents. Spatially Adaptive Multi-Kernel Gating (SAMG) selects receptive-field responses at each location. Dynamic Cross-Scale Detail Fusion (DCDF) aligns the gated intermediate features and compresses them into an eight-channel representation, which is reused at three decoder resolutions together with selected encoder skips. This design consolidates intermediate information before decoding instead of transferring each middle-stage feature through a separate direct skip. On DRIVE, CHASE\_DB1, STARE, and HRF, CGDD-Net achieves AUC values of 0.9824, 0.9938, 0.9895, and 0.9874, with F1 scores of 0.8323, 0.8102, 0.8510, and 0.8157, respectively. The complete model contains 1.96 million trainable parameters. In cumulative ablations, the full model improves F1 over the internal baseline by 2.48, 0.61, and 3.49 percentage points on DRIVE, CHASE\_DB1, and STARE. Twelve directed cross-dataset experiments further characterize transfer without target-domain adaptation. The results support shared intermediate detail delivery as an effective, parameter-compact architecture for retinal vessel segmentation. Code is available at \url{this https URL}.

163. 【2610.05066】Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models

链接:https://arxiv.org/abs/2610.05066

作者:Shengyin Sun,Yiming Li,Yingzhao Lian,Xing Li,Xingzhi Zhou,Anxin Tian,Zhili Wang,Haoyang Li,Ziqiang Cui,Chen Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:post-training incurs additional, incurs additional computational, additional computational costs, Adapting step-distilled, post-training incurs

备注: 29 pages

点击查看摘要

Abstract:Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit stylistic capabilities through language. Pursuing this direction requires textual guidance that captures how visual attributes jointly define a style and remain applicable as the depicted content changes. To explore this approach, we introduce StyleForge, a fully automatic, training-free framework that expresses reference styles as reusable rendering instructions. By integrating overall rendering characteristics with local color and lighting behavior, StyleForge organizes visual evidence from reference images into a coherent specification of how the target style should be expressed. The specification is then compiled into textual guidance that can be reused across content prompts, enabling frozen step-distilled T2I models to render different subjects and scenes in the reference style while retaining native few-step generation. Extensive experiments show relative gains of up to 29.47\% in generation quality scores over the strongest baseline, while Pareto analysis indicates that improved stylization is accompanied by strong adherence to the requested content.

164. 【2610.05053】CoDG-Net: Structure-Guided Style Diffusion and Collaborative Learning to Mitigate Catastrophic Forgetting in Medical Image Domain Generalization

链接:https://arxiv.org/abs/2610.05053

作者:Yucheng Song,Jincan Wang,Haokang Ding,Zhiqiang Tian,Kangxu Fan,Zhifang Liao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:critically important, highly challenging, challenging and critically, Catastrophic Forgetting, Domain Generalization

备注: Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence Main Track. Pages 1622-1630. [this https URL](https://doi.org/10.24963/ijcai.2026/181)

点击查看摘要

Abstract:Domain Generalization (DG) for medical image segmentation is both highly challenging and critically important. However, existing medical DG methods largely overlook the issue of Catastrophic Forgetting (CF): \textbf{Models often sacrifice their ability to retain source-domain knowledge while pursuing cross-domain robustness.} This can directly threaten diagnostic safety in already-deployed clinical scenarios. To address this, we investigate data augmentation strategies and catastrophic forgetting for medical image DG segmentation. First, we propose a structure-guided style diffusion augmentation method. Constrained by anatomical structure consistency in the frequency domain, this method performs cross-domain diffusion on the amplitude spectrum, generating samples with more diverse and broader style coverage to better support domain generalization. Then, we design a collaborative learning network with a dual-branch interactive architecture (CoDG-Net), together with a novel learning bias-guided strategy that adaptively regulates knowledge transfer at both the layer level and the task level, thereby effectively mitigating catastrophic forgetting on the source domain. Experiments and ablation studies on single-source and multi-source medical DG benchmark datasets demonstrate that CoDG-Net not only outperforms existing state-of-the-art methods in target-domain segmentation performance, but also achieves a lower forgetting rate on the source-domain data. The code is available at: this https URL.

165. 【2610.05033】Code2Games: Enabling Coding Agents for Gaming World Generation

链接:https://arxiv.org/abs/2610.05033

作者:Wei Wu,Ziyang Xu,Zeyu Zhang,Yang Zhao,Hao Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires joint reasoning, Generating a high-quality, intent requires joint, executable gameplay logic, spatial layout

备注: Code: [this https URL](https://github.com/AIGeeksGroup/Code2Games) , Website: [this https URL](https://aigeeksgroup.github.io/Code2Games)

点击查看摘要

Abstract:Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual assets, scenes, or scripts, but often struggle to maintain consistency across these components. We propose Code2Games, an agentic framework that builds a structured gaming world upon a base Blender world generated from the same game intent. Code2Games coordinates scene analysis, gameplay planning, constrained gaming-world generation, and gaming-engine customization through a shared scene-gameplay representation with persistent element correspondence. After world generation, Code2Games adapts the generated world to Unreal Engine 5 and employs an execution-guided reconstruction process that uses compilation diagnostics, runtime feedback, and gameplay test results to resolve inconsistencies arising during engine adaptation. To systematically evaluate gaming-world generation, we introduce the GameCode4D benchmark, which comprises ten fixed game prompts spanning different levels of scene and gameplay complexity. We evaluate the generated results across four dimensions: visual quality, interactive fidelity, multimodal artifact quality, and playable-game quality. Experiments demonstrate that, compared with direct gaming-world generation by coding agents and existing baseline methods, Code2Games consistently improves the visual quality and interactive fidelity of generated gaming worlds, as well as the quality of the resulting games after engine adaptation.

166. 【2610.05029】SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation

链接:https://arxiv.org/abs/2610.05029

作者:Hyun Song,Taewan Cho,Kangmin Kim,Andrew Jaeyong Choi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision-language foundation models, provide strong semantic, CLIP provide strong, Vision-language foundation, strong semantic representations

备注:

点击查看摘要

Abstract:Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.

167. 【2610.05026】GeoBridge-VLA: Geometry-Aware Residual Adaptation for Vision-Language-Action Models

链接:https://arxiv.org/abs/2610.05026

作者:Hyun Song,Kangmin Kim,Loren Jinsoo Um,Minhui Han,Jaehyeok Park,Taewan Cho,Andrew Jaeyong Choi

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:models encode semantic, precise spatial reasoning, encode semantic information, requires precise spatial, models encode

备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Stage I trains a feature bridge and geometry decoder with depth supervision. Stage II freezes these modules and trains a gated residual interface together with the action-side projections and action expert. The residual augments the existing visual tokens without adding a second image encoder or increasing the token count. Deployment requires RGB, robot state, and language, but no depth observations. Under matched evaluation conditions, GeoBridge-VLA achieves 70.9% success on LIBERO, compared with 60.0% for SmolVLA. Disabling the residual in the same trained checkpoint reduces success from 70.90% to 69.85%, with mixed effects across suites. On a physical ROBOTIS OMY robot, GeoBridge-VLA succeeds in 148 of 200 trials (74.0%) across four tasks, compared with 108 of 200 (54.0%) for SmolVLA.

168. 【2610.05025】riggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA

链接:https://arxiv.org/abs/2610.05025

作者:Hyemin Yang,Wooseong Jeong,Giwon Lee,Kuk-Jin Yoon

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:improve real-time robotic, specialist action expert, models improve real-time, fast specialist action, real-time robotic control

备注: Accepted by NeurIPS 2026

点击查看摘要

Abstract:Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.

169. 【2610.05024】LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation

链接:https://arxiv.org/abs/2610.05024

作者:Yiming Zhao,Tianshun Li,Jingle He,Ruonan Chai,Xinhu Zheng

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:complex three-dimensional environments, execute long-horizon natural-language, long-horizon natural-language instructions, enables unmanned aerial, unmanned aerial vehicles

备注: 8 pages, 5 figures

点击查看摘要

Abstract:Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale vision-language backbones and dense visual histories, imposing substantial computation and memory costs that hinder onboard deployment. We propose LightVLN, a lightweight history-aware aerial VLN framework that combines a compact 0.5B language backbone with compact representations of both historical and current observations. LightVLN compresses each historical frame into a single token using visual features already computed by the policy. It further introduces history- and instruction-conditioned local aggregation to reduce the current observation from 256 to 32 visual tokens while preserving navigation-relevant spatial information. With up to 16 historical frames, the policy uses at most 48 observation-derived tokens. On the public OpenFly dataset, LightVLN achieves 50.93% Test-Seen and 36.14% Test-Unseen success rates (SR), outperforming the evaluated 7B language-backbone baselines on most reported metrics. It also achieves 25.83% SR on AerialVLN-S Val-Seen. In a reconstructed unseen campus, we deploy LightVLN on a DJI M350 RTK with an external Jetson Orin NX 16 GB for closed-loop onboard-compute real-to-sim hardware-in-the-loop (HIL) evaluation, achieving 14.61 Hz model inference and 11.13 Hz end-to-end decision updates. These results demonstrate the effectiveness and efficiency of LightVLN for aerial navigation.

170. 【2610.05023】Look Where You Say You're Looking: Self-Grounded Attention for Visual Reasoning

链接:https://arxiv.org/abs/2610.05023

作者:Uri Berger,Gal Chechik,Gal Dalal

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:training Vision-Language Models, visual attention, training Vision-Language, attention, model visual attention

备注:

点击查看摘要

Abstract:We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.

171. 【2610.05013】EMBER-Bench: Benchmarking Cross-Event Causal Memory in Long-Horizon Embodied Tasks

链接:https://arxiv.org/abs/2610.05013

作者:Aoyang Cai,Boning Zhao,Shaoxuan Xie,Dahui Gao,Huan Yang,Zhongyuan Wang,Zhiwei Yu,Guocai Yao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Lifelong physical agents, Lifelong physical, disappear from view, reason over extended, extended interactions

备注: 26 pages, 4 figures, 14 tables. The first two authors contributed equally. Project page: [this https URL](https://zhaoalexgoat.github.io/EMBER-Bench/)

点击查看摘要

Abstract:Lifelong physical agents must reason over extended interactions where past events continue to shape the world long after they disappear from view. Beyond recalling what happened, agents must infer how history changes the current state and constrains future actions. Yet existing embodied and video-memory benchmarks largely focus on historical retrieval and summary, leaving such history-dependent causal reasoning underexplored. We introduce EMBER-Bench, an egocentric benchmark for cross-event causal reasoning in long-horizon embodied tasks, for which we newly created the task design, video recording, and data annotation. It contains 189 household tasks and 699 QA pairs, spanning task progress, failure recovery, external interventions, and compound long-horizon tasks with distant dependencies and prerequisites, with fine-grained event and causal-chain annotations. EMBER-Bench evaluates reasoning in both directions: next-action prediction selects the next action from history, and causal traceback, given that action, identifies the historical event that makes it necessary. Input ablations that add action logs or privileged cause-and-consequence annotations to the video indicate which kind of historical information models fail to use. Among the 16 evaluated models, the highest overall accuracy is 61.2%, compared with a mean of 98.3% across two human evaluators. At paired decision points, correct traceback is not associated with correct next-action prediction. Adding action logs yields a gain of 1.6 points, whereas cause-and-consequence annotations yield an additional gain of 13.0 points on top of that. These results suggest that extracting causal information from past events and converting it into constraints on current actions remains a key difficulty for long-horizon embodied agents. Project Page: this https URL

172. 【2610.05011】FreeLoc: Online Floorplan Localization via Diffusion-Aided Pose Refinement

链接:https://arxiv.org/abs/2610.05011

作者:Haocheng Peng,Boyang Zhou,Jiarui Hu,Xiyue Guo,Ziyang Zhang,Boming Zhao,Yifan Gao,Xiao Li,Hujun Bao,Zhaopeng Cui

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:existing high-performing floorplan-based, high-performing floorplan-based methods, Floorplans provide compact, dense scene-specific offline, discretized pose space

备注: Accepted at the Conference on Robot Learning (CoRL) 2026

点击查看摘要

Abstract:Floorplans provide compact and widely available geometric maps for indoor localization, but existing high-performing floorplan-based methods still convert them into dense scene-specific offline databases, tying accuracy, storage, and runtime to the sampling resolution of the discretized pose space. We present FreeLoc, an online RGB-based floorplan localization framework that treats the floorplan as a directly queryable geometric map. FreeLoc introduces an efficient online geometric querying and diffusion-aided refinement scheme, which retrieves plausible pose anchors through on-the-fly floorplan ray querying and refines them into accurate continuous pose estimates. For sequential localization, FreeLoc develops an online likelihood construction strategy that bridges single-frame localization and probabilistic temporal fusion by constructing likelihoods from coarse-sampled candidates and refined pose hypotheses, enabling histogram-filter-based temporal fusion without offline databases. Experiments demonstrate real-time online inference and state-of-the-art performance in both single-frame and sequential localization, while real-world results validate practical deployability in indoor robotic localization scenarios.

173. 【2610.05010】PortraitAes: Intent-Conditioned Structured Portrait Aesthetics Assessment

链接:https://arxiv.org/abs/2610.05010

作者:Junzhou Xie,Haozhong Xiong,Xunyun Tian,Kaile Du,Tianchen Yu,Qiang Li,Wei Liu,Jiaming Liu,Ruihua Huang,Yang Shi,Guangcan Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:effectively human-centered images, human-centered images fulfill, assigns comparable scores, assessment assigns comparable, assigns comparable

备注:

点击查看摘要

Abstract:Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Existing methods typically predict a single aesthetic score or use general-purpose MLLMs without conditioning on photographic intent. This omission matters because the same blur, pose, lighting, or framing choice may serve one photographic intent but undermine another. These models thus learn context-agnostic aesthetic priors and yield inconsistent, inaccurate, misleading judgments for portraits with distinct photographic objectives. We introduce PortraitAes-Bench, an 11K-scale benchmark that decomposes this task into intent-conditioned subjudgments. Expert-authored rubrics define nine photographic intents, six first-level dimensions, and 22 secondary criteria. They support a structured pipeline for intent routing, specialist assessment, verification, and score fusion. Following this structure, we train PortraitAes with multi-task supervision. We then improve score comparability through Gaussian score calibration and within-dimension cross-image ranking. On the standard benchmark, PortraitAes achieves a Pearson correlation of 0.924 and a Spearman rank correlation of 0.934. On the hard-case set, its Pearson correlation is 0.829 and its Spearman rank correlation is 0.795. Across both sets, PortraitAes outperforms the evaluated general-purpose MLLMs and specialized aesthetic baselines.

174. 【2610.05000】VisualErase: Dual-Branch Visual Trajectory Redirection for Robust Concept Erasure in Text-to-Image Diffusion Models

链接:https://arxiv.org/abs/2610.05000

作者:Qianlong Xiang,Miao Zhang,Kun Wang,Yupeng Hu,Junhui Hou,Liqiang Nie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unconstrained large-scale data, privacy-sensitive content learned, reproduce harmful, large-scale data, safe deployment

备注: The project page is [this https URL](https://qianlong0502.github.io/VisualErase-Homepage)

点击查看摘要

Abstract:Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting target-related text-to-image mappings. However, recent studies show that erased models may still retain visual generative trajectories of target concepts, leaving them vulnerable to adversarial recovery attacks and revealing a fundamental gap between redirecting text-to-image mappings and truly removing visual knowledge. To bridge this gap, we propose VisualErase, a new paradigm that redirects concept-bearing visual generative trajectories toward explicitly defined concept-removed outcomes. To enable this redirection, we use structure-preserving image editing to construct content-aligned, concept-removed counterparts for source images, providing explicit visual endpoints that retain non-target content. We then derive a denoising target from each source-to-counterpart pair and use a dual-branch redirection loss to align both text-conditioned and unconditional predictions with this target, since conditional supervision alone does not explicitly constrain generation without textual guidance. To mitigate the adverse effects of concept erasure on non-target generation, we jointly optimize the redirection loss with a counterpart retention loss that matches denoising predictions from the frozen pretrained model. Across style, celebrity, and nudity erasure, VisualErase limits the maximum attack success rate over seven attacks to 0%, 8%, and 0.1%, respectively, while retaining general generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.

175. 【2610.04980】Geometry-Aware Preference Optimization for Text-to-Image Diffusion Models

链接:https://arxiv.org/abs/2610.04980

作者:Lei Wang,Zhen Wang,Yuexiang Xie,Yaliang Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:standard practice, Preference Optimization, Direct Preference Optimization, Preference, Optimization

备注:

点击查看摘要

Abstract:Preference alignment has become a standard practice for text-to-image diffusion models. Direct Preference Optimization (DPO) simplifies this process by eliminating explicit reward modeling. Its diffusion variant, Diffusion-DPO, has become a widely adopted baseline. Diffusion-DPO essentially encourages the likelihood of preferred samples while suppressing dispreferred ones. In this paper, we revisit DPO-style alignment methods for diffusion models from the perspective of the manifold hypothesis. Under this view, natural images concentrate near a low-dimensional manifold embedded in the high-dimensional ambient space, whereas DPO directly optimizes preference distributions in the full space without accounting for this geometric structure. This creates a mismatch in the optimization dynamics: it suppresses geometry-preserving tangential updates, while insufficiently restricting hazardous normal-direction updates. This mismatch gradually degrades image quality and diversity. To address this issue, we propose Anisotropic Geometry-Aware Preference Optimization (APO), which replaces the uniform Euclidean treatment of prediction errors with a geometry-aware anisotropic metric derived from the reference model. Concretely, APO adaptively strengthens regularization in directions where the reference denoising function is highly sensitive, while relaxing constraints in directions that permit safe semantic adjustment. This recalibrates preference optimization according to the local manifold geometry, and maintains the original manifold structure. Experiments show that APO achieves strong performance and an average win rate exceeding 60\% against various existing alignment methods across diverse benchmarks. It requires significantly fewer training steps than prior methods, and preserves generation diversity throughout training.

176. 【2610.04979】Robust Tensor Completion via Reflective Convolution Nuclear Norm Minimization

链接:https://arxiv.org/abs/2610.04979

作者:Weiguo Zhou,Feng Zhang,Wenjin Qin,Jianwen Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:completion recovers multidimensional, partial observations corrupted, Robust tensor completion, recovers multidimensional data, tensor completion recovers

备注: 24 pages, 7 figures, 6 tables

点击查看摘要

Abstract:Robust tensor completion recovers multidimensional data from partial observations corrupted by sparse gross errors. Existing convolutional low-rank models typically construct translated copies using circular continuation, which introduces artificial wrap-around neighborhoods for finite nonperiodic data. We propose reflective convolution nuclear norm minimization (RCNNM), which replaces circular shifts with endpoint-nonrepeating reflection. The resulting lifting has nonuniform entry multiplicities and satisfies a weighted Gram identity that supports both the recovery analysis and the optimization method. Under random sampling and sparse corruption, we establish high-probability exact recovery of the underlying tensor and sparse errors, together with stability under bounded dense perturbations. We further develop a two-block ADMM with a closed-form entrywise tensor update, while singular-value thresholding is implemented through the smaller right Gram matrix. Experiments on synthetic tensors, BSDS color images, and CAVE multispectral images show that RCNNM consistently improves over its circular-lifting counterpart, with the clearest gains near image boundaries. In particular, average boundary-PSNR improvements reach 3.26 dB while global reconstruction quality remains competitive.

177. 【2610.04922】RACE: Time-Adaptive Residual Attention Control with Content-Style Decomposition for Training-Free Diffusion Style Transfer

链接:https://arxiv.org/abs/2610.04922

作者:Duc Khoan Le,Kim Ngoc Tran,Minh Nhat Le,Thanh An Tran,Viet Toan Nguyen,Khanh An Lay,Tran Thai Son,Hoang Pham Minh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reference-guided style transfer, Reference-guided style, style, transferring the visual, visual appearance

备注: Accepted to ACCV 2026

点击查看摘要

Abstract:Reference-guided style transfer aims to preserve the semantic structure of a content image while transferring the visual appearance of a style reference. Recent diffusion-based methods achieve impressive stylization quality by exploiting strong pretrained generative priors. However, training-free approaches still face a difficult trade-off among style fidelity, content preservation, and content leakage. Direct style injection may unintentionally transfer semantic content from the style image, while fixed guidance schedules often ignore the time- and state-dependent nature of diffusion sampling. To address these limitations, we propose TRACE, a training-free diffusion style transfer framework with Time-adaptive Residual Attention Control and Content-Style Decomposition. TRACE first performs offline CLIP-based subspace analysis to separate content and style directions from paired data. During inference, it removes content-related components from the style reference and style-related components from the content reference to reduce leakage. It then injects style information through residual cross-attention and applies uncertainty-aware guidance to adapt the guidance signal at each denoising step. Experiments show that TRACE achieves a favorable trade-off between stylization and preservation. Compared with optimal-control-based baselines, TRACE substantially improves style fidelity (+17.28 CSD and +34.10 SRA). While, compared with stylization methods, it better preserves content structure (+12.80 DINO, +5.52 CLIP-I, and -8.19 LPIPS) and reduces directional semantic leakage by 29.5% in DCL. Our code is publicly available at this https URL.

178. 【2610.04920】PWM: Personalized World Models with Online Reinforcement Learning

链接:https://arxiv.org/abs/2610.04920

作者:Zhexin Lou,Guancheng Lu,Zeyu Zhang,Yi Zhang,Yang Zhao,Hao Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generate diverse environments, Personalized World Models, world models, Pretrained world models, generate diverse

备注: Code: [this https URL](https://github.com/AIGeeksGroup/PWM) , Website: [this https URL](https://aigeeksgroup.github.io/PWM)

点击查看摘要

Abstract:Pretrained world models can generate diverse environments, yet users often want to explore a particular scene specified by their own video. This requires learning the scene's visual identity while retaining the quality of action-conditioned generation. We introduce Personalized World Models (PWM), a framework for customizing interactive world models from short scene videos through online reinforcement learning. In PWM, the support trajectory and its associated controls provide reward feedback on continuations sampled from the current policy. In the GRPO instantiation, group-relative optimization updates a compact LoRA adapter using a unified reward for scene appearance, visual continuity, and motion, while base-policy anchoring regularizes changes to the pretrained generation prior of a frozen Yume-5B backbone. The same adaptation procedure is applied across real and rendered environments. We also instantiate PWM with DiffusionNFT as an alternative reward-guided optimization method for learning the scene-specific adapter. We also introduce PWM-Bench, comprising 150 customization tasks across Indoor, Outdoor, and Gaming, with paired evaluation on held-out continuations. The GRPO and DiffusionNFT instantiations of PWM improve customization over native Yume in 71.3% and 65.3% of the evaluated scenes, respectively, with positive mean gains across all three domains. For the GRPO instantiation, matched SFT comparisons further demonstrate higher mean customization gains and better mean image-quality scores in every domain, while retaining frame-level visual quality close to the pretrained model.

179. 【2610.04911】VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research

链接:https://arxiv.org/abs/2610.04911

作者:Yuhang Zhou,Fei Li,Yuxi Wu,Bin Zhu,Jingjing Chen

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Existing deep research, reasoning systems typically, systems typically assume, Existing deep, image-based web sources

备注:

点击查看摘要

Abstract:Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and ground answers in visual evidence. Training such agents at scale is challenging as live video interaction is slow and unreliable, whereas fixed local simulation can induce retrieval-specific shortcuts that fail to transfer to the open web. We introduce VideoResearchAgent, a scalable training framework to address these challenges. First, we introduce controllable task synthesis pipeline to synthesize multi-hop research tasks from timestamped visual evidence while filtering text-only shortcuts. Second, we build a field-aligned local video simulator that preserves deployment-facing search and watch interactions while accelerating video search by a factor of 34.5-64.6. Third, we introduce Retrieval-Domain-Randomized GRPO (RDR-GRPO), which diversifies candidate rankings, distractors, metadata, and result structure during training to reduce overfitting to simulated retrieval. On Video-BrowseComp, the VideoResearchAgent trained using Qwen3.5-4B achieves 40.48% accuracy, comparable to Gemini-3-Flash-Preview, while reducing cumulative API-token consumption by 74.9% relative to the untrained model. Together, these results establish an accurate and efficient training recipe for open-web video research.

180. 【2610.04883】Reflection-Robust 6DoF Object Tracking with Light Fields

链接:https://arxiv.org/abs/2610.04883

作者:Nikolai Goncharov,Donald G. Dansereau

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing trackers assume, autonomous driving, assumption that breaks, light field, environment map

备注:

点击查看摘要

Abstract:Tracking the 6DoF pose of a moving rigid object is fundamental to robotics and autonomous driving, but existing trackers assume that object appearance is stable across a sequence, an assumption that breaks down on reflective surfaces whose appearance changes as they mirror the environment. We introduce a light field based reflection-robust 6DoF tracker that turns this apparent nuisance into a pose cue. Per frame, our method recovers depth robustly against reflections, back-projects it into a point cloud, and estimates surface normals. It then decomposes the object's view-dependent appearance into a diffuse albedo and the environment map it reflects, resulting in a relightable surface light field. Starting from a coarse initialization, we relight it by the recovered environment map and optimize the pose on the photometric loss. Because a moving object mirrors new parts of the scene, the environment map fills in as the sequence proceeds, sharpening this signal over time. To evaluate this approach, we introduce a light field tracking dataset re-rendered from a robotic manipulation benchmark at four controlled reflectivity levels, each paired with simulated depth that reproduces how consumer RGB-D sensors fail on shiny surfaces. Additionally, we evaluate on two captured light field sequences. Our method trails the strongest baselines on diffuse objects and is the only one that holds its accuracy on fully reflective objects, where every baseline degrades.

181. 【2610.04864】Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning

链接:https://arxiv.org/abs/2610.04864

作者:Haozhe Chi,Song Jin,Yang Jin,Yadong Mu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:limited visual-token budget, requires capturing sparse, capturing sparse query-relevant, embedding requires capturing, sparse query-relevant evidence

备注:

点击查看摘要

Abstract:Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and computation. We introduce \textbf{Query-Aware Streaming Latent Reasoning} (QASLR), a post-training framework that accumulates evidence across clips while keeping the embedding size fixed. QASLR selects a bounded set of frames, restores their temporal order, and processes them clip by clip with a vision-language backbone. A compact set of persistent think tokens cross-attends to each clip's features, while an embed token reads out a normalized representation after every update. This design integrates evidence across multiple backbone calls without requiring all selected frames to share a single context. Training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation. Intermediate supervision trains partial-video readouts for retrieval, while self-distillation regularizes them toward the final representation. Query-aware selection produces query-conditioned representations for candidate-set scoring and reranking, whereas query-independent selection enables reusable corpus indexing. Under the full training recipe, HourVideo retrieval Hit@1 increases from 54.2 to 70.7 and from 57.4 to 72.8 for 2B and 8B Qwen3-VL-Embedding backbones, respectively. Gains extend to the evaluated moment-retrieval and video-QA tasks, and the streaming head transfers to a second Qwen-family embedding backbone. These results support streaming latent aggregation as an effective approach to integrating long-video evidence into fixed-dimensional representations.

182. 【2610.04853】One Tile, Multiple Instances: Rethinking MIL for Sparse Diagnostic Evidence

链接:https://arxiv.org/abs/2610.04853

作者:Runsheng Liu,Cheng Jin,Hao Jiang,Hao Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Slide Image, feature extractors typically, extractors typically compress, supervised Whole Slide, feature extractors

备注:

点击查看摘要

Abstract:In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile evidence from the attention mechanism. We introduce DI-MIL, a framework that decouples encoding context from instance granularity through decomposed instances. By clustering dense spatial tokens from a frozen foundation model within each tile, DI-MIL converts a single tile into multiple independently weighted instance embeddings. As a training-free post-encoding module, DI-MIL integrates seamlessly into existing pipelines without requiring re-encoding or downstream architectural modifications. We evaluate DI-MIL on cytopathology, a challenging testbed where sparse diagnostic signals are easily diluted within standard tiles. Across four datasets, three frozen foundation models, and two attention-based aggregators, DI-MIL demonstrates consistent efficacy, improving 67 of 72 metric-level comparisons, with the largest mean gains reaching 3.64 points under cytopathology-specific backbones. In a broader comparison against seven representative MIL baselines, DI-MIL paired with ACMIL achieves highest mean performance in 33 of 36 backbone-dataset-metric comparisons. Ablations show that direct smaller tiling inflates the extracted tile count by up to 43.3$\times$ with non-monotonic performance, whereas DI-MIL incurs zero additional image-extraction overhead while achieving the strongest overall results. These results establish instance construction as an orthogonal design dimension in MIL, supporting DI-MIL as a cost-efficient solution under sparse diagnostic evidence.

183. 【2610.04836】RSure-Agent: Reliable Use of Tool Observations for Remote Sensing Agents

链接:https://arxiv.org/abs/2610.04836

作者:Fuyuan Liu,Nayu Liu,Wenhao Yu,Peijin Wang,Yingchao Feng,Fanglong Yao,Liang Wan,Wei Feng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:solve Earth observation, solve Earth, Earth observation tasks, Earth observation, raster analysis tools

备注: The demo is available at [this https URL](https://github.com/airs101/RSure-Agent) (code will be released for further research)

点击查看摘要

Abstract:Remote sensing agents rely on perception, measurement, and raster analysis tools to solve Earth observation tasks. We refer to their judgments and quantitative results about ground objects as tool observations. However, these observations are subject to substantial uncertainty and may be incorrect even when the tools execute successfully. When agents accept incorrect observations, the errors can propagate through subsequent reasoning and cause task failure. We analyze 1,229 execution trajectories across three remote sensing agent benchmarks. On each benchmark, at least 88.1% of tasks depend on tool observations. Among these tasks, at least 22.7% contain incorrect observations despite successful tool execution. These errors propagate to the final answer in at least 82.0% of affected tasks on each benchmark. To address this problem, we propose RSure-Agent, a framework for verifying tool observations and limiting error propagation. We introduce a verifiable observation protocol that requires tools to return process evidence for the agent to verify their observations. We also construct a task-tool reliability prior from offline task feedback. The prior summarizes each tool configuration's past performance across task types and provides a task-specific reference for verification. Using process evidence and this prior, RSure-Agent decides whether to accept an observation, request additional evidence, or reject it. We evaluate RSure-Agent on EarthBench, ThinkGeo, TerraLogic, and CHOICE-420. Across the three agent benchmarks, RSure-Agent reduces the error propagation rate by 21.3 to 25.9 percentage points relative to the base configuration with verification and the prior disabled. On CHOICE-420, it improves overall accuracy over direct answering by 5.71 percentage points on average across 11 backbone models. On EarthBench, it reduces the tool-call ratio by 25.9% relative to Earth-Agent.

184. 【2610.04819】DriftSR: One-Step Real-World Image Super-Resolution via Distribution Drifting

链接:https://arxiv.org/abs/2610.04819

作者:Wei Zhu,Kai Zhang,Yu Zheng,Zhaopeng Yang,Lei Luo,Yong Guo,Jian Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:perceptually rich details, additional trainable components, optimization more cumbersome, recovering realistic, realistic and perceptually

备注:

点击查看摘要

Abstract:One-step real-world image super-resolution (Real-ISR) offers efficient inference, but recovering realistic and perceptually rich details often relies on score distillation or adversarial learning, introducing additional trainable components and making optimization more cumbersome. To this end, we propose DriftSR, a one-step Real-ISR framework that leverages pretrained diffusion priors through distribution drifting. Specifically, we perform drifting in the frozen intermediate representation space of a pretrained diffusion model, without introducing an additional task-specific feature encoder. Building on this space, we introduce Spatial Feature Drifting, which treats spatial features rather than entire images as distributional samples, enabling denser supervision for distribution alignment. To mitigate structural deviations, we further introduce Structure-Modulated Guidance, which adaptively refines drifting guidance according to local structural consistency with the LQ input. Consequently, DriftSR optimizes only the one-step generator, without auxiliary distillation branches or adversarial discriminators. Extensive experiments on three real-world benchmarks demonstrate that DriftSR delivers high-quality super-resolution reconstruction with efficient one-step inference.

185. 【2610.04814】MOXIE: Discovering Alternative Explanations for Biomedical Image Classifiers

链接:https://arxiv.org/abs/2610.04814

作者:Abiha Tahsin Chowdhury,Rahul Dubey

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)

关键词:Segment-based explanation methods, Segment-based explanation, eXplanation Imaging Engine, MOXIE, Segment-based

备注:

点击查看摘要

Abstract:Segment-based explanation methods such as LIME return a single explanation for each prediction, computed from one fixed image segmentation. This hides two important facts: a prediction can be supported by many different sets of image segments, and the segmentation itself shapes which explanations can be found. We introduce MOXIE (Multi-Objective eXplanation Imaging Engine), an evolutionary framework that searches for segment subsets that preserve the classifier's confidence while keeping as little of the image as possible. Instead of one explanation, MOXIE returns a Pareto front of alternative explanations that range from compact to highly faithful. We evaluate MOXIE with NSGA-II and four segmentation methods (SLIC, Felzenszwalb, Watershed and Voronoi) on BloodMNIST and HAM10000 datasets, using the same evaluation budget as LIME. Results show that MOXIE achieves a higher hypervolume than LIME on every image. LIME's explanations often appear convincing, yet the classifier's confidence collapses when only the highlighted segments are shown. MOXIE's fronts reveal how much of the image is needed to preserve the model's confidence and which contextual regions influence it. We also find that segmentation strongly affects evaluation: methods with unequal segment sizes appear most compact when segments are counted. These results show that alternative explanations provide a more complete view of a model's decision than a single explanation.

186. 【2610.04805】ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations

链接:https://arxiv.org/abs/2610.04805

作者:I-Chun Arthur Liu,Jason Chen,Gaurav S. Sukhatme,Daniel Seita

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:recovering metric depth, monocular RGB observations, RGB observations, Three-dimensional perception, monocular RGB

备注:

点击查看摘要

Abstract:Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (VLA) models rely solely on RGB observations for perception. Leveraging recent advances in foundation models for stereo matching, we introduce ExStereo, a stereo module that augments pre-trained 2D VLAs with 3D perception. ExStereo reconstructs scene geometry from stereo image pairs and renders multi-view observations as an explicit stereo representation for stereo feature extraction. The action tokens from the action expert selectively attend to the resulting stereo tokens through our proposed action-stereo cross-attention mechanism, enabling the policy to generate robot actions conditioned on 3D scene information. To learn robust 3D representations, we introduce a mid-training stage before task-specific post-training, using a self-supervised learning objective on large-scale stereo data. We validate our approach by fine-tuning two publicly available VLAs, $\pi_{0.5}$ and SmolVLA, and evaluate them in simulation and on a real-world bimanual PiPER platform. Across both settings, VLAs fine-tuned with ExStereo consistently outperform baselines, demonstrating the effectiveness of stereo perception for robotic manipulation. Our project website is at: this https URL.

187. 【2610.04800】MedImageOSWorld: Benchmarking GUI Agents for Medical Image Consoles

链接:https://arxiv.org/abs/2610.04800

作者:Ziyang Long,Xinqi Li,Lujing Xing,Hsin-Jung Yang

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Graphical consoles offer, Graphical consoles, human operators, offer a practical, practical interface

备注: 11 pages, 3 figures. Supplementary material (8 pages) included as an ancillary file

点击查看摘要

Abstract:Graphical consoles offer a practical interface for medical acquisition assistance, allowing agents to work through the controls and visual feedback used by human operators. Reliable assistance requires linking on-screen anatomy to acquisition decisions that determine what image evidence becomes available next. We introduce MedImageOSWorld, a benchmark for evaluating this capability in simulated CT, MR, and ultrasound consoles. Using screenshots and mouse-and-keyboard actions, agents configure protocols, plan acquisitions, inspect the resulting images, and make corrective adjustments across seven capability levels, from console operation to feedback-driven control. Evaluation combines task-specific workflow checks with hidden anatomical ground truth to assess procedural completion and acquisition outcomes separately. A common evaluation protocol specifies episode conditions and interaction budgets, while recorded trajectories support analysis of how agents observe, act, and respond to acquisition feedback. Across eleven open-weight agents, success rates range from 3.0 to 25.0 on a 0-100 scale while workflow-progress rates reach 21.5-74.5: agents complete much of the console workflow but rarely acquire the intended anatomy. Success collapses between perception and planning, from 79-90% at the lowest three levels for the best agent to at most 12% for millimetre-level planning and 0% for closed-loop control among open-weight agents. Two proprietary agents reach success rates of about 34 and exceed the best open-weight agent mainly in acquisition quality (61 versus 42). MedImageOSWorld provides a controlled setting for studying whether general-purpose GUI agents can translate visual observations into effective medical acquisition decisions.

188. 【2610.04792】Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

链接:https://arxiv.org/abs/2610.04792

作者:Songyuan Sui,Zhen Tan,Mohan Zhang,Rana Muhammad Shahroz Khan,Xia Hu,Tianlong Chen

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Missing modality remains, remains a longstanding, longstanding challenge, multimodal learning, modality remains

备注: NeurIPS 2026 Main Conference

点击查看摘要

Abstract:Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.

189. 【2610.04787】Active-DiNTS: Active Differentiable Network Topology Search

链接:https://arxiv.org/abs/2610.04787

作者:Gean Trindade Pereira,Thierry Urruty,Muriel Visani,André C. P. L. F. de Carvalho

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)

关键词:Neural Architecture Search, manual network design, large annotation budgets, Neural Architecture, Network Topology Search

备注: 19 pages, 7 figures, 3 tables. Extends material from the first author's PhD thesis (University of Sao Paulo / La Rochelle Universite, 2024)

点击查看摘要

Abstract:Neural Architecture Search (NAS) has proved to be a strong alternative to manual network design, but applying it to 3D medical image segmentation is limited by two well-known costs, large annotation budgets and multi-GPU clusters. Thus, this paper introduces Active-DiNTS (Active Differentiable Network Topology Search), an approach that embeds pool-based Active Learning (AL) into a bi-level differentiable topology search to perform architecture discovery and label curation jointly. At each query round, unlabeled MRI volumes are ranked by one of three uncertainty signals (Entropy, Variance, or Standard Deviation), and only the top-ranked volumes are sent to an oracle for annotation. The new labels feed two interlocked stages. Network weights are updated in an outer loop, while the macro/micro topology of a U-Net-style backbone is refined in an inner loop. Three AL regimes (weights-only, topology-only, joint) expose the speed-accuracy trade-off. Evaluations on the Medical Segmentation Decathlon (MSD) Task01 BrainTumour benchmark showed that Active-DiNTS surpasses DiNTS, C2FNAS, and nnU-Net in Dice, with gains of about 10 percentage points on Edema and 5 points on Non-Enhancing core, on a single GPU and using a fraction of the labeled volumes. The discovered architectures are denser and more FLOP-heavy than prior baselines, but remain competitive in trainable parameters and peak memory; the fastest search variant finishes in under 0.25 GPU-days, over 27x faster than the eight-GPU DiNTS search. Together, these results indicate that pairing differentiable NAS with active data acquisition is a practical recipe for accurate 3D segmentation under realistic constraints.

190. 【2610.04786】KALEIDO: Input-Space Adaptation of a Vision Model for Time-Series Forecasting Through Gated Fold Geometries

链接:https://arxiv.org/abs/2610.04786

作者:Xiangyu Shi,Qinghua Liu,Sam Heshmati,Zubin Abraham

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Time-series foundation models, foundation models buy, large temporal corpora, ImageNet-pretrained masked autoencoder, masked autoencoder forecasts

备注: Accepted at the NeurIPS 2026 Workshop on Foundation Models for Temporal Systems (FMTS)

点击查看摘要

Abstract:Time-series foundation models buy zero-shot forecasting with large temporal corpora; a vision model needs none, since a natural image implicitly embeds the patterns a forecaster must model, and an ImageNet-pretrained masked autoencoder forecasts a series by inpainting a rendering of it. A rendered series is not a natural image, however, and closing that gap takes temporal-aware adaptation. We show that the rendering geometry - how the series is folded and drawn - is a controllable, mixable axis for it. Kaleido detects the dominant periods, renders a rule-generated set of fold geometries, combines the inpaintings with a convex per-position gate fit on validation only, and fuses the result with the zero-shot output at one fixed share, with no per-dataset hyperparameter beyond the baseline's published settings. Training only LayerNorm (0.05%), Kaleido lowers MSE by 13% against the published zero-shot baseline on LTSF and, frozen, by 6.6%; on GIFT-Eval it improves the baseline by 7.4% in MASE and 19.3% in CRPS.

191. 【2610.04785】Lollypop: Camera-to-Motion-Capture Calibration Verification with a Reference Target

链接:https://arxiv.org/abs/2610.04785

作者:Tianyi Liu,Kevin Harris,Mihika Dave,Kun He

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:computer vision tasks, truth in robotics, vision tasks, ground truth, computer vision

备注: Accepted at IEEE SENSORS 2026. 4 pages, 3 figures

点击查看摘要

Abstract:Camera-to-motion-capture (mocap) calibration is essential for using mocap as ground truth in robotics, AR/VR, and other computer vision tasks. However, the calibration can drift after deployment, while calibration residuals and visual inspection provide limited independent verification. We present Lollypop, a fiducial-mocap reference target for independent calibration verification. The target couples an ArUco fiducial with a mocap marker constellation so the visual center and tracked centroid represent the same physical point. Given a candidate calibration, verification projects the mocap point into the image and measures its disagreement with the detected fiducial center. Experiments show sub-pixel nominal error, sensitivity to controlled extrinsic perturbations, and increasing error during an illustrative mixed-handling sequence.

192. 【2610.04781】Super-Resolution in The Right Latent Space: A Frozen Vision-Foundation Substrate

链接:https://arxiv.org/abs/2610.04781

作者:Wanzhou Lei,Cuifeng Sheng,Yanjin He,Maohua Li,Hua Yuan,Per-Olof Persson,Hanlin Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:sharp images form, manifold, latent space, space, high-resolution images pushes

备注:

点击查看摘要

Abstract:In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.

193. 【2610.04777】ARISE: Adaptive Agentic Reasoning with Image-grounded Self-Evaluation for Interpretable IBD Assessment

链接:https://arxiv.org/abs/2610.04777

作者:Pronoma Banerjee,Anuva Shah,Jason Wu,Md. Masudur Rahman,Sanjay Mohanty,Satya Kurada,Juan P. Wachs

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:requires frequent imaging-based, Inflammatory bowel disease, remains heavily dependent, Inflammatory bowel, frequent imaging-based assessment

备注:

点击查看摘要

Abstract:Inflammatory bowel disease (IBD) requires frequent imaging-based assessment, yet interpretation of modalities such as wireless capsule endoscopy (WCE) and intestinal ultrasound remains heavily dependent on specialist expertise. Vision-Language Models (VLMs) demonstrate significant potential in multimodal medical image analysis, but their clinical adoption is hindered by their insufficient domain-specific reasoning, susceptibility to hallucination, scarcity of high quality training data in fine-grained diagnostics and limited interpretability. We introduce ARISE (Adaptive Agentic Reasoning with Image-grounded Self-Evaluation), an autonomous planning framework that models few-shot medical image understanding as a sequential agentic workflow. ARISE structures agent execution into a transparent 5-stage workflow: hypothesis generation, image-grounded evidence summarization, evidence-conditioned refinement, symbolic verification, and final diagnosis. We apply ARISE to IBD assessment across two independent patient cohorts: wireless capsule endoscopy (WCE) images for Crohn's disease and B-mode ultrasound data for ulcerative colitis. ARISE consistently improves diagnostic performance over baseline VLMs while exposing where reasoning succeeds or fails, providing a more interpretable basis for clinical decision support and realistic deployment.

194. 【2610.04743】hyCLIPNet: A BiomedCLIP-Guided Lightweight Attention-Enhanced DeepLabV3+ Framework for Robust Thyroid Nodule Segmentation

链接:https://arxiv.org/abs/2610.04743

作者:Tasnim Jahan,Md Easin Arafat,Swakkhar Shatabda

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate thyroid ultrasound, thyroid ultrasound segmentation, speckle noise, Accurate thyroid, low contrast

备注:

点击查看摘要

Abstract:Accurate thyroid ultrasound segmentation is often challenged by low contrast, speckle noise, and unclear boundaries. Although recent methods have improved segmentation accuracy, many rely on resource-intensive architectures or lack explicit integration of multiscale features with global biomedical visual guidance. In this paper, we introduce ThyCLIPNet, a lightweight semantic-guided hybrid encoder-decoder framework that integrates BiomedCLIP-derived biomedical semantic guidance into a lightweight multi-scale CNN segmentation pipeline. The encoder integrates MobileNetV2 with efficient channel attention, while atrous spatial pyramid pooling and a custom convolutional block attention module enrich bottleneck features. The decoder combines hierarchical skip connections and lightweight attention refinement with a BiomedCLIP-guided gated fusion pathway that projects vision-only global biomedical embeddings into decoder feature space and selectively integrates them through semantic-local fusion and spatial gating. To the best of our knowledge, ThyCLIPNet is among the first lightweight thyroid ultrasound segmentation frameworks to use BiomedCLIP's vision encoder alone for image-only global semantic guidance without text prompting. Experiments on TG3K, TN3K, DDTI, and PKTN achieve dice similarity coefficients of 96.22%, 87.58%, 84.73%, and 80.70%; intersection over union scores of 92.72%, 77.91%, 73.51%, and 67.64%; and 95th-percentile hausdorff distances of 3.75, 16.38, 18.23, and 10.86, respectively. ThyCLIPNet uses 8.55M parameters and 22.99G FLOPs. Overall, the results support integrating global biomedical semantic guidance with lightweight multi-scale CNN representations for robust and computationally efficient thyroid ultrasound segmentation. Source code: this https URL. [Abstract shortened for arXiv. See PDF for full abstract.]

195. 【2610.04736】Probabilistic Pedestrian Forecasts from a Handheld Phone: World-Frame Heat Maps, Visual-Inertial Height Drift, and Evaluation without Ground Truth

链接:https://arxiv.org/abs/2610.04736

作者:Danial Safaei

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:visual-inertial odometry, shown where nearby, forecast stays, final forecaster NLL, phone moves

备注: 28 pages, 7 figures, 16 tables

点击查看摘要

Abstract:A pedestrian with a phone could be shown where nearby people will be in the next few seconds, if the forecast stays on the ground while the phone moves, is calibrated, and needs only a monocular camera and visual-inertial odometry (VIO). We build and evaluate such a system. People are detected, lifted onto the floor by ray-plane intersection, tracked in a gravity-aligned metric frame, and forecast as per-step probability maps by a small U-Net trained with a negative log-likelihood (NLL) loss on bird's-eye (SDD) and first-person (EgoTraj-Bench) trajectories. On handheld ADVIO recordings, vertical VIO drift and the user's own changes of level silently rescale monocular ground positions (by 87% within 90 s on one clip; on another, all tracks are lost for the last 31% of the clip); keeping the camera's height above the floor constant under a low-pass-filtered altitude avoids this, though it lags on escalators. On the SDD and EgoTraj-Bench test splits, the final forecaster lowers the NLL at 4.8 s by 1.51 and 1.37 nats relative to a fitted constant-velocity Gaussian. Lacking ground truth for people in handheld video, we score forecasts against the tracker's own later raw measurements. In an internally pre-registered evaluation on seven held-out clips, the final forecaster's NLL is lower than the benchmark-fitted baseline's at 1.2, 2.4 and 4.8 s (by 0.17, 0.23 and 0.47 nats; 95% intervals over people exclude zero), but by less than half as much as on the development clips. Exploratory analyses cut both ways: resampling clips instead of people widens the intervals to include zero at 1.2 and 2.4 s, and once both forecasters are recalibrated on the development clips the network is significantly better only at 1.2 s; but two clips run with ADVIO's reference poses favour the network much more when re-run with the phone's own poses. We discuss what such self-consistency scores can and cannot show.

196. 【2610.04727】Investigating Spatiotemporal Redundancy in Video Transformer for Collision Anticipation

链接:https://arxiv.org/abs/2610.04727

作者:Xiaoshan Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:worker-equipment proximity monitoring, demonstrated strong performance, proximity monitoring, worker-equipment proximity, demonstrated strong

备注:

点击查看摘要

Abstract:In worker-equipment proximity monitoring, video transformers are widely used for collision anticipation and have demonstrated strong performance. However, their accuracy comes with substantial computational demands, creating a tension with the need for low-latency inference on mobile robots and the pursuit of lower-carbon computation in construction. To address this, this study investigates where computation within an established video transformer is redundant and whether that redundancy can be removed without materially degrading predictive performance. Using VideoMAEv2-Base on the Nexar Collision Prediction dataset, we first examine how collision-relevant information evolves across network depth and then investigate two complementary forms of redundancy: structured capacity redundancy in multilayer perceptrons (MLPs) and spatiotemporal redundancy in the token stream. Linear probes show that interpretable motion cues, including flow magnitude, looming, and approach versus retreat, are most accessible at intermediate layers, whereas collision-label discrimination strengthens toward the final layer. Token redundancy is axis-specific: adjacent temporal-token similarity reaches 0.970 in later layers, while spatial similarity falls to 0.297, indicating substantially greater redundancy across time than across space. Exploiting this asymmetry, temporal token merging reduces backbone computation from 356.99 to 178.50 GFLOPs and latency from 12.69 to 7.21 ms per clip, a 1.76x speedup, while mean average precision changes only from 0.7478 to 0.7443. Importance-guided retention of 50% of MLP units preserves an AUC of 0.753, compared with 0.529 under matched random retention, and reveals that pruning alters score calibration before discriminative ranking collapses. These findings establish a new pathway for pursuing faster algorithms through targeted temporal token compression and neuron pruning.

197. 【2610.04722】NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

链接:https://arxiv.org/abs/2610.04722

作者:Ramil Khafizov,Ilya Statsenko,Ruslan Rakhimov,Artem Komarichev,Peter Wonka,Evgeny Burnaev

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:approaches remain limited, diffusion-based approaches remain, content creation, making multi-view generation, central problem

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at this https URL

198. 【2610.04721】Knossos and Ariadne: Benchmarking and Learning Complete Diagram Topology Extraction with Vision-Language Models

链接:https://arxiv.org/abs/2610.04721

作者:Bangwei Guo,Xujiang Zhao,Shengyu Chen,Yanchi Liu,Wei Cheng,Xi Zhu,Guoning Zhang,Dimitris N. Metaxas,Haifeng Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:represent complex systems, Structural diagrams, information across scientific, represent complex, complex systems

备注:

点击查看摘要

Abstract:Structural diagrams are widely used to represent complex systems and relational information across scientific, engineering, procedural, and spatial domains. Recent vision-language models (VLMs) have become increasingly capable of recognizing diagram elements and reasoning about their content, while complete diagram topology extraction remains comparatively underexplored. In this paper, we study diagram-to-graph topology extraction: extracting all diagram entities and the complete relations among them. To enable large-scale supervised training and systematic evaluation of this task, we introduce Knossos, a benchmark of 19,200 diagrams across six diverse domains, with 245,179 nodes and 439,740 edges. Its symbolic generation process provides exact alignment between rendered diagrams and annotations of complete topology, relation types, and connector geometry. To address the modeling challenge of complete topology extraction, we also present Ariadne, a structured framework that decomposes the task into node inventory extraction and source-conditioned edge prediction. Extensive experiments show that training on Knossos substantially improves complete topology extraction in smaller open-source VLMs. Ariadne further improves over one-step extraction under matched supervision, demonstrating the additional benefit of structured decomposition. It achieves the highest average Edge F1 among the evaluated methods on Knossos, while both backbone variants also improve over their unadapted counterparts on the real-world external benchmark. Code and benchmark are available at this https URL.

199. 【2610.04703】Learning Discriminative Geometry for Drifting Models

链接:https://arxiv.org/abs/2610.04703

作者:Doudou Zhang,Wenwen Hou,Yilin Chen,Qi Chen

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Recently proposed Drifting, shift iterative distribution, iterative distribution refinement, Models shift iterative, Recently proposed

备注:

点击查看摘要

Abstract:Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.

200. 【2610.04700】Decouple, Purify and Unite: Semantic-Structural Prototype Learning for Federated Medical Segmentation

链接:https://arxiv.org/abs/2610.04700

作者:Xingyue Zhao,Wenke Huang,Linghao Zhuang,Yanzhou Su,Zhifeng Wang,Haoyu Zhao,Mengfan Li,Junjun He,Tao Tan,Dakai Jin,Le Lu,Mang Ye,Qiang Yang,Ming Feng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:protocols remains challenging, Federated learning enables, learning enables medical, Incomplete Contextual Representation, Contextual Representation Learning

备注: 17 pages, 9 figures, 7 tables

点击查看摘要

Abstract:Federated learning enables medical institutions to train a global model without sharing data, yet feature heterogeneity from diverse scanners or protocols remains challenging. Existing representation-based methods face two limitations: 1) Incomplete Contextual Representation Learning: single-layer or coupled representations overlook multi-level structural cues and entangle regional semantics with boundary details. 2) Layerwise Style and Aggregation Biases: domain-specific style discrepancies across intermediate layers degrade prototypes, while aggregation that overlooks client distribution shifts can further amplify bias. We propose FedBCS+, federated decoupled contextual alignment with style-purified aggregation. We employ Frequency-domain Style Recalibration (FSR) in prototype construction to decouple content-style representations and extract style-purified prototypes. Built upon these purified features, Decoupled Contextual Prototype Alignment (DCPA) explicitly decouples multi-level features into semantic and structural prototypes and aligns regional semantics and fine-grained anatomical structures separately. Style-purified Semantic Prototype Aggregation (S2PA) measures each client's purified prototype divergence from the global consensus and adaptively reweights aggregation toward under-represented clients to reduce consensus bias. On five heterogeneous medical segmentation benchmarks spanning histopathology, MRI, ultrasound, and colonoscopy, FedBCS+ achieves the highest mean Dice among the compared methods. A convergence analysis further characterizes how aggregation and alignment affect the optimization bound.

201. 【2610.04680】COMPASS: Comet Object Measurement Pipeline with Automated Selection and Scoring

链接:https://arxiv.org/abs/2610.04680

作者:Jack Roberts,Canya Lu,Alexis Michelle Lawson,Kaitlyn Holden,Gerald S. Wilkinson,Anne M. Bronikowski,Ritambhara Singh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:single-cell gel electrophoresis, individual cell level, quantifying DNA damage, gel electrophoresis, cell level

备注:

点击查看摘要

Abstract:Summary: The single-cell gel electrophoresis ('comet') assay is a widely used technique for quantifying DNA damage at the individual cell level. However, image analysis often relies on manual inspection or semi-automated software, which can be labor-intensive, difficult to reproduce, and sensitive to image quality and comet morphology. COMPASS automates comet assay image analysis by combining deep learning-based comet segmentation with damage measurement and automated comet selection. The pipeline produces standardized DNA damage measurements while offering robust detection, reducing manual effort and improving reproducibility through transparent selection and optional manual review. Availability and implementation: COMPASS is implemented in Python and is freely available at this https URL rsinghlab/COMPASS. Installation instructions, pretrained weights, and example usage are provided in the repository. Contact: jack_roberts2@brown.edu, ritsingh@illinois.edu Supplementary information: Available onlime upon publication.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.04680 [cs.CV]

(or
arXiv:2610.04680v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.04680

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Jack Roberts [view email] [v1]
Sat, 3 Oct 2026 17:50:25 UTC (6,462 KB)

202. 【2610.04664】FLASHSWIN: Unlocking Large Windows and Dense Tokens in Swin Vision Transformers with Memory Efficient Attention

链接:https://arxiv.org/abs/2610.04664

作者:Tushar Kataria,Gerald Sabin,Ponnuswamy Sadayappan,Shireen Y. Elhabian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:High-resolution vision backbones, High-resolution vision, score matrix, Swin, vision backbones

备注:

点击查看摘要

Abstract:High-resolution vision backbones have long been forced to trade away local token density to afford larger receptive fields. Hierarchical Swin transformers impose this compromise because standard windowed attention materializes an $M^2\times M^2$ score matrix per window, incurring $O(M^4)$ memory as windows or token grids grow. Furthermore, Swin adds a learned relative-position bias elementwise to attention scores, requiring full materialization of the score matrix and its gradient. This keeps Swin and SwinV2 trapped in a small-window($M=8,16$), coarse-token regime with patch size $4\times4$ ($p=4$), limiting performance for fine-grained tasks. We introduce FLASHSWIN, which replaces standard windowed attention with a FlashAttention implementation that computes exact softmax attention without materializing the score matrix, reducing per-window memory from $O(M^4)$ to $O(M^2)$. This enables higher token density and larger receptive fields without inflating memory overhead. Training memory is flat across window sizes: at a $32\times32$ window, FLASHSWIN-T requires only $12.4$\,GB, unchanged from $8\times8$, compared to $70/90$\,GB for SwinV2/V1-T. However, applying FlashAttention directly to Swin creates a trade-off: bypassing the score matrix precludes Swin's additive relative-position bias, forfeiting spatial information in exchange for memory efficiency. FLASHSWIN restores position information as window-local learnable 2D RoPE, making large windows and dense token grids both affordable and accurate. At matched scale, FLASHSWIN-T outperforms Swin variants. With dense tokens and wide windows ($p=2,M=32$), the same Tiny model reaches $84.1\%$ ImageNet-1K, $44.1$ COCO box AP, and $47.28$ ADE20K mIoU---gains of $+1.3$, $+5.1$, and $+1.82$ over SwinV2-T at $M=16$, respectively. At fixed $M=32$, halving the patch size yields roughly $3\times$ larger gains in boundary quality than in mIoU.

203. 【2610.04627】ask-Sensitive Geometry of Representation Transfer for Object Detection under Image Degradation

链接:https://arxiv.org/abs/2610.04627

作者:Van Vung Pham

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:clean-image supervision, degradation can benefit, benefit from clean-image, imply that transferred, image degradation

备注: 20 pages, 5 figures, 1 table

点击查看摘要

Abstract:Object detection under image degradation can benefit from clean-image supervision, but aggregate gains do not imply that transferred representation changes are uniformly useful. We study how clean task knowledge affects degraded-image representations and whether local responses to structured representation directions can be characterized geometrically. Using paired clean and Gaussian-degraded BDD100K images, we show that clean-teacher distillation improves observed detection accuracy while producing heterogeneous object-level transfer. We isolate a representation component complementary to direct clean-teacher alignment and map it into the distilled student space through an orthogonal bridge. Controlled interventions rescue 13.22% of objects lost under the distilled representation, versus 4.30% under norm-matched random perturbations, with very low harm on preserved objects. We introduce task-sensitive geometry, a gradient-derived channel-space geometry constructed from normalized detection-loss gradients. On a reserved cohort, mapped-complement orientation within this frozen geometry is positively associated with local intervention-response magnitude after controlling for intervention magnitude (partial Spearman $\rho$ = 0.242, 95% CI [0.108, 0.359]). The relationship eplicates on independent data ($\rho$ = 0.180) and with RT-DETR-L ($\rho$ = 0.227), but not for the direct clean-teacher residual family, and it weakens for large interventions. Routing rules and specialized distillation objectives based on these signals do not yield statistically reliable gains over CLEANKD. These results support a local, direction-family-dependent task-sensitive geometry while showing that converting such structure into improved global training remains an open problem.

204. 【2610.04616】PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

链接:https://arxiv.org/abs/2610.04616

作者:Mingyu Liu,Chonghao Sima,Tianjian Feng,Hanqing Wang,Cong Chen,Hao Chen,Chunhua Shen

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:complete complex tasks, complete complex, complex tasks, evidence, shortcuts

备注:

点击查看摘要

Abstract:A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training even after the verb changes, and a gripper that closed on nothing may lift anyway. We call these dependencies modality shortcuts: regularities in successful demonstrations make visual, lexical, or motor cues sufficient to predict expert actions without the task evidence needed for the underlying decision. More demonstrations of the same kind can raise task success while leaving these shortcuts intact. We propose Perturbot which makes task-relevant evidence easier to use and shortcuts insufficient on their own: it applies task-preserving wrist-view perturbations, enriches instructions with decision-relevant captions, and adds random and failed trajectory segments relabeled with the behavior they contain. It complements scaling by changing what is scaled, and leaves inference unchanged. Moreover, we propose GroundingFscore, an offline score that diagnoses how severely a policy relies on modality shortcuts. Task success rate shows whether a policy improves, while GroundingFscore reveals whether the policy scales healthily, relying on task evidence rather than shortcuts. Together, Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.

205. 【2610.04607】ForeAct3D: Policy-Grounded Future World Modeling for VLA Policies

链接:https://arxiv.org/abs/2610.04607

作者:Zhe Tao,Feiran Wang,Gaowen Liu,Ramana Rao Kompella$,Yan Yan

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:resulting contacts, manipulation success hinges, future, VLA policies, action

备注:

点击查看摘要

Abstract:Robots need to anticipate how their actions will change the world, since manipulation success hinges on the resulting contacts and object motions. However, existing Vision-Language-Action (VLA) policies that predict future observations from shared features leave the forecast decoupled from the actions the policy will actually execute, and impose no physical constraints on how the scene may evolve. We introduce ForeAct3D, a framework for policy-grounded future world modeling within VLA policies. Learnable geometric queries decode depth, semantic segmentation, and camera pose from the policy representation into current and future semantic 3D scene states, and the future queries are conditioned on the policy-generated action chunk to ground the forecast in the planned interaction. A physical-consistency closure relates the two states through background staticity and instance-level rigidity, and anchors the wrist-camera pose to end-effector kinematics. These objectives shape the shared representation used for action generation during training, and no future prediction is required at inference. Without robot pretraining, ForeAct3D achieves 98.3\% average success on LIBERO and an average task length of 3.73 on CALVIN, outperforming its base policy on every suite. Ablations show that semantic 3D supervision, physical consistency, and action conditioning each improve manipulation performance, and that action conditioning substantially improves future object localization. Real-world experiments on spatial placement, object insertion, and sequential manipulation further raise average success from 6.7\% to 37.8\% over the base policy. The project page and code are available at this https URL.

206. 【2610.04606】Sparse-View 4D Gaussian Splatting via Spatiotemporal Priors and Generative Assistance

链接:https://arxiv.org/abs/2610.04606

作者:Shengqi Wang,Zhengxian Yang,Kaiwen Tian,Yang Liu,Bowen Liu,Hua Du,Taicheng Huang,Jiamin Wu,Tao Yu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Volumetric Video Challenge, SIGGRAPH Asia, Volumetric Video, Video Challenge, Gaussian Splatting framework

备注: 4 pages, 5 figures, Accepted to SIGGRAPH Asia 2026 Workshops (SA Workshops '26)

点击查看摘要

Abstract:We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. (1) Region-adaptive spatial priors: We use foreground masks to guide Gaussian initialization and mask voting to control densification separately for the dynamic foreground and static background. Background geometry is regularized using monocular depth aligned to metric scale. (2) Motion-consistent temporal priors: We provide supervision at intermediate times through frame interpolation and constrain projected Gaussian motion with estimated optical flow. (3) Generative assistance: We place virtual cameras in the widest angular gaps and restore their rendered images using a diffusion-based model conditioned on camera poses. The restored images are iteratively incorporated into training as pseudo-supervision. On the validation set, our framework improves full-frame PSNR from 25.60 dB for the baseline to 29.75 dB. On the official test benchmark, it achieves 30.04 dB full-frame PSNR and 27.88 dB foreground PSNR, ranking first overall in the Sparse-View Track.

207. 【2610.04605】ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution

链接:https://arxiv.org/abs/2610.04605

作者:Yehonatan Elisha,Oren Barkan,Ziv Weiss Haddad,Noam Koenigstein

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:highlight pixel importance, semantically meaningful concepts, computer vision highlight, vision highlight pixel, visual explanation methods

备注: ICML 2026

点击查看摘要

Abstract:Many visual explanation methods in computer vision highlight pixel importance but struggle to link these low-level cues to semantically meaningful concepts, limiting their interpretability and trustworthiness. We introduce Concept-based Explanations (ConEx), a novel framework that bridges saliency visualization with concept-based reasoning to provide both faithfulness and interpretability. ConEx automatically discovers class-specific concepts and represents them through concept activation vectors (CAVs), learned without manual supervision using an architecture-specific masking mechanism that reduces noise introduced by the segmentation masks to enhance concept purity. ConEx generates faithful saliency maps that reveal where each concept appears in the image and how it contributes to the prediction. To evaluate the reliability of these learned concepts, we propose two complementary metrics, Vector-Concept Match (VCM) and Concept-Class Match (CCM), that quantify concept alignment and enable direct comparison with existing methods. Extensive experiments across diverse settings demonstrate that ConEx achieves state-of-the-art performance on faithfulness, segmentation, and concept-quality benchmarks. Overall, ConEx advances the field toward truly interpretable and concept-grounded explanations in vision models.

208. 【2610.04602】Organize Primitives into Semantic Parts: Reinforcement Reasoning for 3D Segmentation

链接:https://arxiv.org/abs/2610.04602

作者:Xiaoming Gong,Ruoyu Wu,Zhenhong Sun,Chunlin Chen,Daoyi Dong,Huadong Mo,Zhi Wang,Hongdong Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:naturally supporting structural, supporting structural abstraction, naturally supporting, boundary localization, offers a compact

备注: 25 pages

点击查看摘要

Abstract:Primitive-based 3D segmentation offers a compact and explicit alternative to dense surface prediction, naturally supporting structural abstraction and boundary localization. However, geometric decomposition alone does not determine how primitives should be organized into semantic parts: a single part may span multiple primitives, while geometrically similar or touching primitives may belong to different parts. We therefore introduce RePart (Reinforcement Part Reasoning), which formulates primitive-to-part organization as a finite-horizon Markov decision process and learns semantic organization through trajectory-level reinforcement reasoning. RePart constructs a Composable Primitive Workspace from fine-grained superquadrics and applies a merge-and-stop policy whose decisions are optimized by their downstream effects on the resulting partition rather than local primitive compatibility. The inferred part identities are then mapped back to the original mesh through Boundary-Aware Surface Labeling, preserving accurate surface boundaries beyond the primitive approximation. On PartNet, RePart achieves the strongest results across all four aggregate partition metrics; on 3DCoMPaT++, it obtains the highest RI and SC without target-dataset fine-tuning. These results demonstrate that reinforcement reasoning provides an effective mechanism for organizing geometric primitives into semantic parts while retaining dense segmentation accuracy. Code is available at this https URL.

209. 【2610.04601】WASP: Weakly Aligned Spatiotemporal Pairs for Fetal Brain MRI-Ultrasound Learning

链接:https://arxiv.org/abs/2610.04601

作者:Francesco Correnti,Gabriele Magrini,Marco Mistretta,Niccolò Biondi,Pietro Pala,Alessandro Ramalli,Simona Fiori,Andrew D. Bagdanov,Matteo Lenge

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Magnetic Resonance Imaging, Magnetic Resonance, Resonance Imaging, brain analysis due, superior soft-tissue contrast

备注: Accepted at NeurIPS 2026. 21 pages, 3 figures, 11 tables. Code: [this https URL](https://github.com/miccunifi/WASP)

点击查看摘要

Abstract:Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft-tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. Ultrasound (US), in contrast, is cheap, safe, and routinely acquired, and as a result it has produced substantially larger datasets and a growing ecosystem of pretrained models. This asymmetry raises a natural question: Can we teach a US-only model to understand fetal MRI from only a limited set of examples? The standard recipe, training a foundation model on subject-to-subject paired MRI-US scans, is not viable since no such paired fetal dataset is publicly available. In this paper we address this gap with Weakly Aligned Spatiotemporal Pairs (WASP), a framework that formulates cross-modal correspondence as an entropic Optimal Transport problem driven by clinical metadata, in particular Gestational Age (GA) and diagnostic planes, enabling the fitting of a lightweight alignment module that lifts MRI representations into the US latent space, without fine-tuning the backbone. Empirically, WASP yields its largest gains when MRI is unseen by the model during pretraining (on USFM, GA estimation error drops from 21.9 to 17.4 days and standard plane classification accuracy climbs from 61.9% to 69.0%), while providing smaller, backbone-dependent refinements for backbones pretrained on both modalities (e.g., BioMedParse GA estimation error from 6.5 to 6.0 days and SAM-Med2d plane accuracy from 83.3% to 88.1%). Code is available at this https URL.

210. 【2610.04585】Frozen in a Frame: The Velocity Blind Spot in JEPA World Models

链接:https://arxiv.org/abs/2610.04585

作者:Tinghe Zhang,Chunyu Liu,Yu Leon Liu,Zerui Zhao,Jiaheng Chen,Yucheng Xiao,Jiaxing Li,Yunlong Wang,Alex Lamb

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Joint-embedding predictive architectures, single rendered frame, world modeling train, Joint-embedding predictive, rendered frame

备注: 32 pages, 17 figures, 12 tables. Code, model checkpoints, and project page are available via links in the paper

点击查看摘要

Abstract:Joint-embedding predictive architectures (JEPAs) for world modeling train an encoder so a predictor maps a current embedding and action to the next frame's embedding, always from a single rendered frame. This has a structural blind spot: a renderer without motion blur draws a scene from configuration alone, so a single-frame embedding carries no velocity information, for any encoder, including the official released LeWM weights. We confirm this on official checkpoints across four real benchmarks (PushT, Reacher, Cube, TwoRoom): every linear velocity probe sits at or below chance while position probes reach R^2 about 0.95. We introduce RateIdent, a three-stage diagnostic protocol, and TI-JEPA, a lightweight fix splitting the latent into a pose code and an explicit finite-difference motion code, predicted jointly. Across three physically grounded environments, TI-JEPA gives a significant, seed-robust gain on a stop-at-goal planning task over a matched-memory baseline, e.g. 55% lower final distance on Pendulum (p=3.2x10^-10) and 64% on CartPole (p=5.1x10^-15). We reproduce this at official ViT-Tiny plus AdaLN-transformer scale, then push the same recipe onto real dm_control Reacher photographs trained from scratch, where TI-JEPA's branch separation exceeds the memory-having baseline's by roughly 38x, the paper's largest margin. Against a same-footprint recurrent RSSM-style predictor, TI-JEPA matches or beats its rollout accuracy on two of three environments, stays separately probeable for pose and motion, and wins outright on the most coupled one. A checkable formal argument and six evaluated environments show single-frame targets are the wrong object to predict when velocity matters, and a small, interpretable structural change fixes it with no privileged supervision. Code, checkpoints, and the project page are linked below the title.

211. 【2610.04554】EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder

链接:https://arxiv.org/abs/2610.04554

作者:Bowen Chai,Tianbao Zhang,Shuyu Wu,Dexin Zuo,Zhaoxin Fan,Danping Zou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recovering detailed geometry, Recovering detailed, detailed geometry, critical for precise, precise perception

备注: Project page: [this https URL](https://sjtu-visys-team.github.io/EagleDepth/)

点击查看摘要

Abstract:Recovering detailed geometry from high-resolution images is critical for precise perception of the surroundings and objects. However, existing methods which use latent-space modeling and VAE reconstruction can compromise geometric details. Furthermore, decoding from latent codes introduces substantial inference overhead. To address those issues, we present EagleDepth, an efficient framework for high-resolution monocular depth estimation that combines the geometric priors of latent diffusion with fine-grained pixel-space generation. Our key idea is to retain depth-aware latent representations as guidance while generating the final depth map directly in pixel space. We train the latent and pixel components sequentially: first, we fine-tune a pretrained latent diffusion model using paired RGB--depth supervision; then, we adapt a pretrained pixel diffusion decoder, PiD, to predict depth conditioned on the learned features. Training of the pixel component starts at 1024 resolution and continues across multiple resolutions up to 4K. The latent branch processes resized, lower-resolution RGB images, while the pixel branch generates depth at the target resolution, bypassing the original VAE decoder. This design preserves learned geometric knowledge without requiring the latent backbone to operate at the output resolution. On five commonly used depth estimation datasets and the high-resolution Synth4K dataset, our framework achieves state-of-the-art depth estimation performance, with faster inference and better preservation of fine structures and object boundaries.

212. 【2610.04538】FASTER: Fast Adjoint Stochastic Transport for Endpoint Refinement in Reward-Guided Image Editing

链接:https://arxiv.org/abs/2610.04538

作者:Yimiao Zhou,Zejia Zhong,Jingya Wang,Ye Shi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reward-guided image editing, Reward-guided image, preserving source content, test time seeks, visual plausibility

备注:

点击查看摘要

Abstract:Reward-guided image editing at test time seeks to improve a specified reward while preserving source content and visual plausibility. Many existing approaches optimize candidates through pretrained generation processes, making repeated adjustment depend on costly large-model execution and, in some cases, backbone backpropagation. We develop a theoretical framework that jointly accounts for reward, source preservation, and pretrained-prior preferences, allowing the desired output distribution to be specified separately from the dynamics used to realize it. Based on this framework, we introduce FASTER, which trains a small network for each source and objective to perform inexpensive editing, while pretrained and reward models provide feedback on candidate outputs. By reusing each candidate and its feedback across multiple small-network updates, FASTER reduces repeated sampling and supervision queries without placing the pretrained generative backbone inside the inner optimization loop. On SD3, FASTER leads all four target metrics and several validation metrics among the evaluated methods. Compared with the evaluated baseline that optimizes controls along pretrained generation trajectories, FASTER achieves editing-time speedups of up to \({6.91\times}\) on Stable Diffusion 3 and \({24.14\times}\) on Stable Diffusion 1.5.

213. 【2610.04529】AME:Topology-Aware Text-Driven Motion Editing across Heterogeneous Humanoid Skeletons

链接:https://arxiv.org/abs/2610.04529

作者:Qichen Zheng,Siyuan Yang,Chong Wang,Jun Liu,Shijian Lu,Alex Kot,Kwok-Yan Lam

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Text-driven motion editing, motion editing modifies, Text-driven motion, existing motion sequence, editing modifies

备注:

点击查看摘要

Abstract:Text-driven motion editing modifies an existing motion sequence according to a text instruction while preserving the content of the source motion. Existing methods are typically built for a single, fixed skeletal topology, which limits their use in animation pipelines where characters differ in joint count and skeletal hierarchy. We present Topology-Aware Motion Editor (TAME), a flow-matching transformer that edits motions on humanoid skeletons of varying topology. TAME represents motion as per-joint, per-frame tokens and models interactions among joints, across frames, and with the text instruction through skeletal, temporal, and text cross-attention layers. To make the skeletal attention follow each character's hierarchy, TAME replaces full joint attention with Topology-Constrained Skeletal Propagation (TCSP), which restricts attention to one-hop kinematic neighbors in the skeleton's adjacency matrix. We further introduce Edit-Focused Representation Alignment (EFRA), a self-distilled representation alignment strategy that aligns student features with cleaner EMA-teacher features exclusively on edit-relevant joint-time tokens, making edits faithful to the instruction. To make this setting trainable and comparable, we construct TopoMotionFix, a multi-topology extension of MotionFix with seen- and unseen-topology evaluation protocols. TAME outperforms previous methods in edit alignment and source preservation on MotionFix and reliably edits motions on unseen skeletons in TopoMotionFix.

214. 【2610.04506】EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning

链接:https://arxiv.org/abs/2610.04506

作者:Yutong Li,Molin Wang,Xiaotong Li,Yanyan Fang,Daoguo Dong,Ziyi Ye

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:semantic event understanding, visual states underexplored, existing benchmarks largely, benchmarks largely focus, future visual states

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underexplored. We introduce EgoExo-Next, a visual-option benchmark for dynamic visual-state reasoning, where models must identify how an observed action trajectory subsequently appears rather than predict only an action label or textual description. EgoExo-Next contains 2,503 human-curated four-choice questions from six public egocentric and ego--exo video sources and comprises four interconnected subtasks that evaluate egocentric next-state prediction, bidirectional ego--exo state correspondence, exocentric next-state prediction, and their composition in Ego-to-Exo Next-State. Extensive evaluation of proprietary, open-source, and spatial reasoning VLMs reveals a substantial human--model gap, with the best model achieving 43.81\% average accuracy compared with 98.55\% for humans, and the largest degradation occurring on the composed Ego-to-Exo task. These results suggest that current VLMs remain substantially limited in dynamic visual-state reasoning, particularly when temporal progression and cross-view reasoning must be composed. The benchmark is publicly available at \url{this https URL}.

215. 【2610.04502】Localization Lens for Improving Medical Vision-Language Models

链接:https://arxiv.org/abs/2610.04502

作者:Hasan Farooq,Murtaza Taj,Mehwish Nasim,Arif Mahmood

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:demonstrated strong capabilities, clinical tasks, demonstrated strong, strong capabilities, capabilities in clinical

备注: 10 pages, 1 figure, Medical Image Computing and Computer Assisted Intervention (MICCAI)

点击查看摘要

Abstract:Medical Vision-Language Models (Med-VLMs) have demonstrated strong capabilities in clinical tasks. However, they often struggle to understand anatomical structures and spatial positioning, which are crucial for medical reasoning. To address this, we propose a localization-aware enhancement to the Med-VLM pipeline, introducing improvements at three levels: data,architecture, and alignment. First, we introduce localization lens, a set of expert-validated representations that provide richer anatomical and positional context. However, as these representations increase input complexity, we integrate pixel shuffle within the model architecture to filter and refine representations, enhancing spatial information processing while preserving anatomical continuity. Lastly, to effectively align the localization lens representations with textual features, we incorporate decoupled contrastive loss (DCL) alongside the standard loss function. This ensures better feature discrimination and robustness, particularly in data limited medical settings. Through extensive evaluations on medical visual question answering (Med-VQA) datasets, we show that our methodology improves localization-driven performance across different Med-VLM architectures. Our analysis of localization-based questions further reveals that improvements in anatomy and spatial reasoning directly enhance the overall accuracy of Med-VQA upto 6.2%. The proposed approach is model-agnostic and can be seamlessly integrated into existing Med-VLM pipelines. The dataset, code, and trained models will be made publicly available at this https URL.

216. 【2610.04499】Homogeneous Semantic Alignment and Hierarchical Expert Routing for Radiology Report Generation

链接:https://arxiv.org/abs/2610.04499

作者:Erjian Zhang,Jiayuan Ma,Liejun Wang,Yikemaiti Sataer,Xiaoming Tao,Zhiqing Guo

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Radiology report generation, convert medical images, Radiology report, Hierarchical Expert Routing, aims to convert

备注:

点击查看摘要

Abstract:Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of underlying representations and the undifferentiated rigid coupling of heterogeneous information cause weak visual abnormality cues to be easily diluted by massive text priors and generation inertia during decoding. To overcome this bottleneck, inspired by cognitive science, we propose a novel two-stage Homogeneous Semantic Alignment and Hierarchical Expert Routing (HSA-HER) framework. First, the model introduces an explicit homogeneous distribution constraint in the underlying latent space to effectively eliminate the cross-modal distribution shift between visual and textual features, thereby extracting purified visual features as semantic anchors that accurately align with diseases. Second, for heterogeneous clinical evidence composed of visual features, local entities, and global retrievals, we design a hierarchical expert routing mechanism guided by these disease semantic anchors. This mechanism abandons the undifferentiated rigid coupling paradigm. Specifically, it dynamically activates expert networks to perform targeted mining and semantic reconstruction on multi-source evidence, and adaptively allocates fusion weights. Extensive experiments on three mainstream benchmark datasets demonstrate that HSA-HER achieves state-of-the-art performance, accurately depicting complex imaging details and key diagnostic information.

217. 【2610.04493】Understanding Clustering in Slot Attention via Particle Dynamics

链接:https://arxiv.org/abs/2610.04493

作者:Vasudev Joy,Rajat Rasal,Avinash Kori,Anthea Monod,Ben Glocker

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:interacting particle dynamics, Studying attention, lens of interacting, interacting particle, shown how token

备注: 13 pages, 4 figures. Accepted to the DynaFront workshop at NeurIPS 2026

点击查看摘要

Abstract:Studying attention through the lens of interacting particle dynamics has shown how token clustering can emerge from the underlying dynamics. We extend this perspective to slot attention, a method for object-centric image segmentation and representation learning in which learned components obscure how much of the clustering behaviour is intrinsic to the attention dynamics. We therefore introduce simplified slot attention (SSA), a parameter-free variant whose dynamics are connected to soft $k$-means clustering and which provides a straightforward mechanistic explanation for the emergence of object-centric representations. On the Pascal VOC dataset, SSA achieves performance comparable to that of slot attention, demonstrating that competitive object-centric segmentation can be achieved without learned neural-network components.

218. 【2610.04457】RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers

链接:https://arxiv.org/abs/2610.04457

作者:Mengyuan Fan,Bokai Huang,JiaMing Pan,Xiaokun Yuan,Peizhuang Cong,Zhewen Tan,Tong Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieve strong performance, impose substantial storage, mobile vision applications, high-dimensional linear projections, Vision Transformers

备注: Accepted to NeurIPS 2026. Current preprint version; camera-ready revision forthcoming

点击查看摘要

Abstract:Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for this http URL and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$-$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$-$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.

219. 【2610.04456】Multi-Crop Leaf Disease Recognition: A Unified Benchmark and Cross-Region Study

链接:https://arxiv.org/abs/2610.04456

作者:Rosemary Nalwanga,Sebastian Bunda,Godliver Owomugisha,Luuk Spreeuwers,Estefania Talavera Martinez

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Deep learning models, routinely report near-perfect, report near-perfect accuracy, Deep learning, Multi-crop Leaf Disease

备注:

点击查看摘要

Abstract:Deep learning models for crop leaf disease recognition routinely report near-perfect accuracy yet are typically trained and evaluated on a single dataset collected under controlled laboratory conditions, leaving their behavior under realistic cross-region domain shift poorly understood. We introduce MLD (Multi-crop Leaf Disease) dataset, a unified multi-region benchmark that combines six public crop-disease datasets from the USA, Asia, and Africa into a shared hierarchical taxonomy spanning 18 crops, 56 crop-disease classes (including one healthy class per crop) making 167,427 images. We define standardized single-source and pooled multi-source evaluation protocols that explicitly probe cross-region generalization. We also investigate whether exploiting the inherent crop-to-disease dependency via a hierarchical formulation (HiLeaD) that conditions disease prediction on the predicted crop improves recognition under cross-region shift. Under the HiLeaD, the model trained on PlantVillage achieves 99.07% in-domain disease F1 but collapses to 12.88% when tested on PlantDoc, exposing a severe cross-region domain gap. The model trained on the pooled MLD dataset partly recovers cross-region disease F1 from 12.88% to 39.64% on PlantDoc (HiLeaD), achieving a 26.76 percentage point improvement. The hierarchical formulation provides a consistent additional gain, ranging from 1.71 to 4.88 percentage points in disease F1 over the flat baseline under the MLD dataset indicating that progress in this area is currently limited more by data coverage and diversity than by model design.

220. 【2610.04432】Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

链接:https://arxiv.org/abs/2610.04432

作者:Jinzhou Tang,Zijun Zhang,Jing Yang,Yuchen Yan,Kun Zhou,Lingjun Mao,Ruobing Han,Jinglin Cao,Wenpeng Xu,Lukun He,Minghao Fu,Fan Feng,Biwei Huang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:Building interactive simulators, scale embodied data, Building interactive, manual environment construction, construction and calibration

备注: Project page: [this https URL](https://aetherlabsai.github.io/Video2World)

点击查看摘要

Abstract:Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5\% to over 15\%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.

221. 【2610.04426】UnAct: Gradient-Free Unlearning via Targeted Activation Intervention

链接:https://arxiv.org/abs/2610.04426

作者:Saeed Abdul Muizz,Aayat Rafiq,Iqra Altaf Gillani,Janibul Bashir

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Machine unlearning seeks, Selective Synaptic Dampening, Machine unlearning, designated training data, SSD

备注:

点击查看摘要

Abstract:Machine unlearning seeks to remove the influence of designated training data from a trained model without retraining from scratch. Retrain-free methods such as Selective Synaptic Dampening (SSD) and its label-free variant LFSSD avoid full retraining but still require backpropagation and parameter importance computed over the entire dataset. We ask: what happens when a deletion request arrives with only a few images of the class to be forgotten? To answer this question, we introduce UnAct, a gradient-free class-unlearning method that needs only forward passes over the forget images. UnAct scores late-layer units by their responses, attenuates the most responsive connections, and repeats this for up to 20 rounds using no gradients, no labels, and no retained data. On ResNet-18 trained with CIFAR-10, CIFAR-20, and CIFAR-100, UnAct is competitive with SSD and LFSSD when forgetting entire classes and, unlike them, never collapses the network when forget data is scarce. On ResNet-18, across all tested sizes, UnAct's retain accuracy stays within 2.5 points of retraining, while SSD and LFSSD, at their full-class operating points, lose up to 86 points on some classes. With five forget images on CIFAR-10, UnAct's distance to retraining is 0.21 points, against 67 for LFSSD and 90 for SSD, and re-selecting SSD's threshold at each size with an oracle does not close the gap. In preliminary transfer to ViT-B/16, UnAct's distance to retraining is 11.5 against 33.7 for SSD, and a request is 19x faster than SSD when SSD computes its importance at request time. The code is available at this https URL

222. 【2610.04425】AgroGround: Multi-Granularity Grounded Recognition in Agriculture

链接:https://arxiv.org/abs/2610.04425

作者:Abdulla Alshehhi,Zongyan Han,Rao Anwer

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:reliable diagnosis requires, diagnosis requires identifying, Agricultural visual, agricultural VQA datasets, Agricultural

备注:

点击查看摘要

Abstract:Agricultural visual models are typically evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the evidence. Agricultural visual question answering (VQA) datasets carry rich semantic labels but rarely link them to image regions, and adding such annotations by hand is costly at scale. We introduce AgroGround, a large-scale dataset for grounded agricultural recognition: identifying plant diseases and other agricultural targets and localizing their image regions. An automated pipeline converts the labels of eight agricultural VQA datasets into annotations for disease lesions and whole objects, producing 794,850 instruction examples. Healthy images provide negative supervision for disease queries, teaching the model to return empty predictions. We fine-tune a shared vision-language model on known-target grounding instructions combined with instructions requiring both recognition and localization. We evaluate predicted identities, regions, joint correctness, and healthy-image abstention on 1,480 human-verified images disjoint from all training data. Grounding-only fine-tuning reduces recognition accuracy from 51.8\% to 29.1\%, while adding recognition-and-localization instructions raises it to 72.6\%. With images and annotations held fixed, combining the two formats raises joint accuracy from 19.2\% to 43.3\% at comparable grounding. Healthy negatives raise abstention on healthy images to 95.0\%, and reinforcement learning improves lesion-level grounding. The resulting 2B model exceeds its annotation teacher in grounding F1 on our benchmark and on the external PlantSeg test set. AgroGround establishes a benchmark for grounded agricultural recognition, measuring joint correctness of identity and localization along with abstention on healthy images. The code is available at this https URL.

223. 【2610.04392】Detecting Defects that Matter: An Application-Driven Benchmark for Anomaly Detection in Manufacturing and Retail Logistics (VAND 4.0 Challenge)

链接:https://arxiv.org/abs/2610.04392

作者:Lars Heckler-Kram,Dorian Henning,Ashwin Vaidya,Jan-Hendrik Neudeck,Ulla Scheler,Anton Milan,Samet Akcay,Paula Ramos,Sebastian Höfer

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Existing Anomaly Detection, Existing Anomaly, Anomaly Detection benchmarks, Existing, Anomaly Detection

备注:

点击查看摘要

Abstract:Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the Industrial Track (MVTec AD 2), the results reveal that unsupervised anomaly segmentation remains challenging: the best regular-setting method achieves only ~57\% pixel-level $SegF_1$, indicating substantial room for improvement. Zero-shot approaches trail by ~15 $SegF_1$ points, confirming that task-specific training on normal data remains essential for precise defect localization. Robustness to distribution shifts remains a key open challenge and DINOv3-backbones clearly dominate this track. In the Retail Track (Kaputt 2), the results reveal that (1) supervised defect detection is approaching saturation for common defect types; (2) the best off-the-shelf VLM approach trails specialized models by ~28 AP, confirming that currently VLMs cannot replace fine-tuned detectors, (3) reference images did not prove helpful for top-performing approaches. Performance collapses on rare defects (spillage ~53 AP, missing units ~27 AP), where the supervised ceiling is bounded by data availability. To drive future progress in this domain, we provide a new low-prevalence retail AD dataset (Kaputt-Rare). Across both tracks, computational efficiency is assessed as a first-class metric combining performance, throughput, memory, and power consumption. We introduce a novel metric for measuring efficiency and reveal that that top-performing methods rely on heavy architectures while efficiency is largely neglected. Overall, we conclude that the community needs (a) more efficiency-aware method development, and (b) true anomaly detection approaches for rare defects and shifting conditions. this https URL

224. 【2610.04381】Beyond Plausibility: Verifiable Fine-Grained Image Editing on Structured Assets

链接:https://arxiv.org/abs/2610.04381

作者:Muyao Wang,Chen Zhu,Shiqi Yang,DongHyun Hwan,Hideki Nakayama

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:visually plausible result, image editing requires, plausible result, requested attribute change, attribute change precisely

备注: Image editing benchmark

点击查看摘要

Abstract:Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and verifiability: benchmarks built on realistic images typically rely on human or vision--language model judgments, while deterministic evaluation has largely focused on synthetic shape canvases, with application-oriented extensions primarily limited to charts. This makes it difficult to determine precisely how much of a requested edit was executed, where unintended changes occurred, and whether small differences between models reflect genuine editing capability or evaluator uncertainty. To bridge this gap, we present VeriEdit-Bench, a benchmark for fine-grained, instruction-faithful image editing across realistic structured assets with deterministic, four-axis evaluation. Its 1,740 cases are compiled from the source code of 153 Scalable Vector Graphics (SVG) graphics, charts, web interfaces, and presentation slides. Controlled source-code edits preserve the original visual context while yielding exact target images, pixel-level edit masks, and explicit edit specifications, enabling reproducible scoring along four axes: edit fidelity, preservation, localization, and magnitude. Evaluating eleven editors, we find that even the strongest model remains far from full credit; rankings for the same recoloring operation reverse between charts and SVG graphics; and outputs with similar pixel-accuracy profiles can still differ substantially in localization and change magnitude. This decomposition yields graded, verifiable feedback and exposes model-specific capability and failure profiles that holistic scores or evaluator-dependent judgments may obscure.

225. 【2610.04361】RIM-ReID: Duplication-Aware Token Reduction and Modality-Aligned Interaction for Multi-Modal Object Re-Identification

链接:https://arxiv.org/abs/2610.04361

作者:Wanke Xia,Ruiding Zhu,Xingguo Xu,Zhengbo Zhang,Dongxia Liu,Yuan Jin,Taojie Zhu,Yiting Zhao,Yihang Ding

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multi-modal object re-identification, object re-identification exploits, retrieve target objects, exploits complementary RGB, Multi-modal object

备注: Under Review

点击查看摘要

Abstract:Multi-modal object re-identification exploits complementary RGB, near-infrared (NIR), and thermal-infrared (TIR) observations to retrieve target objects. However, existing methods commonly employ visual encoders optimized for global image-text alignment and select tokens using learned importance scores. Such designs fail to preserve fine-grained identity cues or explicitly account for token redundancy, resulting in underrepresented local evidence and duplicated tokens that lead to noisy and costly cross-modal interaction. To address this gap, we propose TRIM-ReID, a compact framework that unifies dense feature extraction, intra-modal token reduction, and inter-modal aligned interaction. Specifically, semantically rich and spatially coherent patch features are extracted by Dense Identity Representation (DIR), which leverages DINOv3 to preserve fine-grained identity information. We then introduce Token Diversity Mining (TDM) to identify complementary local evidence and construct compact modality-specific token sets by suppressing repetitive patches while preserving informative diversity. Retained tokens are subsequently fused by Modal Relational Interaction (MRI) to enable effective information exchange across modalities, while a triangular alignment loss explicitly regularizes their joint relationships to maintain cross-modal semantic consistency under independent token selection. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate that TRIM-ReID achieves state-of-the-art performance.

226. 【2610.04356】SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction

链接:https://arxiv.org/abs/2610.04356

作者:Yuhang Wang,Kai Luo,Yuanfan Zheng,Kailun Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:autonomous driving requires, driving requires modeling, requires modeling geometry, understanding for autonomous, autonomous driving

备注: 9 pages, 4 figures

点击查看摘要

Abstract:Comprehensive 3D scene understanding for autonomous driving requires modeling geometry, semantics, and motion. However, camera-based occupancy and scene flow prediction are sensitive to unreliable spatial and temporal aggregation, caused by semantically incompatible image features, misaligned historical observations, and incomplete voxel structures. To address this issue, we propose SelectOccFlow, a selective spatiotemporal aggregation framework that progressively refines contextual evidence across image, temporal, and voxel domains. To obtain semantically compatible image evidence, we design Semantic-Guided Sampling (SGS) to regulate feature sampling with semantic priors. Since reliable image evidence alone cannot resolve temporal inconsistency, we then present State-Conditioned Temporal Aggregation (SCTA) to selectively retrieve historical evidence according to voxel states. To further enhance the structural completeness of voxel representations, we introduce Extent-Aware Spatial Aggregation (ESA), which exploits directional structural support to refine foreground geometry. Experiments on OpenOcc demonstrate that SelectOccFlow achieves a state-of-the-art OccScore of 44.9, improving the previous best by +4.2%. It also maintains competitive occupancy performance on Occ3D-nus and improves the mean OccScore under nuScenes-C corruptions by +11.1%, demonstrating improved robustness to visual corruptions. The source code will be made publicly available at this https URL.

227. 【2610.04351】LoCoSplat: Real-Time Feed-Forward 3D Gaussian Splatting with Minimal 3D Reasoning

链接:https://arxiv.org/abs/2610.04351

作者:Sinan Wang,Jinjin He,Yuchen Sun,Duowen Chen,Shenyifan Lu,Bo Zhu

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:increasingly aggregates multi-view, aggregates multi-view evidence, Gaussian Splatting, increasingly aggregates, aggregates multi-view

备注: 23 pages. Under review

点击查看摘要

Abstract:Feed-forward 3D Gaussian Splatting (3DGS) increasingly aggregates multi-view evidence with heavy learned 3D networks. We propose LoCoSplat (Local-Context Splatting), motivated by the observation that a Gaussian is a local primitive: once depth is predicted, what the 3D stage must add (scale, rotation, opacity) depends on the point cloud around each anchor, and a fixed local average of that neighbourhood is enough to supply it, no heavy network required. LoCoSplat realises exactly this average: it splats a 16-d linear projection of the point features into a fine and a coarse grid and reads both back at each anchor with a 0.14M-parameter pointwise MLP; with no learned 3D network and no dynamic sparse computation, its whole encoder runs as one fp16 CUDA graph. On RealEstate10K, LoCoSplat outperforms every prior feed-forward method on PSNR, SSIM, and LPIPS at 6, 12, and 24 views, with a margin that widens as views densify (+3.3 PSNR over VolSplat, the prior voxel-aligned state of the art, at 24 views) and grows further under zero-shot transfer to ACID and fine-tuning on ScanNet. It reconstructs a 6-view scene in 33 ms on one NVIDIA RTX PRO 6000 GPU, the fastest of seven feed-forward methods and $4.2\times$ faster than the previous state of the art, trains $2.7\times$ faster ($5.9\times$ at 24 views), and uses $6.7\times$ less inference memory.

228. 【2610.04346】Any-scale Object Detection using Arbitrary-scaled Images

链接:https://arxiv.org/abs/2610.04346

作者:Kazutoshi Akita,Norimichi Ukita

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:rescaled appearance representations, paper proposes any-scale, discretely rescaled appearance, proposes any-scale object, continuously rescaling object

备注: MVA2025. 6 pages

点击查看摘要

Abstract:This paper proposes any-scale object detection using arbitrary-scale super-resolution for continuously rescaling object images, while general multi-scale object detection uses discretely rescaled appearance representations. However, a naive usage of super-resolution produces many false-positive detections if many super-resolution images are independently fed into an object detector. Our method suppresses these false positives by predicting scale proposal maps, each of which represents a set of pixels appropriate for each super-resolution scale.

229. 【2610.04342】Asynchronous Tracking, Optical Communication and 3D Motion Capture using Event-based Sensors

链接:https://arxiv.org/abs/2610.04342

作者:Ziwei Wang,Angus Apps,Holly Battisson,Iain Guilliard,Timothy Molloy,Robert Mahony

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:key enabling technology, Inter-robot communication, key enabling, enabling technology, Optical communication

备注: 16 pages, 17 figures

点击查看摘要

Abstract:Inter-robot communication is a key enabling technology in cooperative robotics applications. While wireless communication is ubiquitous, simple, and effective, it has inherent limitations: the receiving robot cannot easily localise the source of an incoming signal, clock synchronisation is challenging, and signals are broadcast to all nearby devices rather than targeted recipients. Optical communication provides a complementary channel that is inherently directional and spatially localised, directly addressing these limitations. Event cameras are bio-inspired dynamic vision sensors that respond to changes in image intensity with high temporal resolution, high dynamic range, and low latency, making them well-suited as receivers for high-rate optical communication in cooperative robotic systems. In this paper, we propose the Asynchronous Tracking and Optical Communication (ATOC) system, which integrates LED smart-beacon modulation, event-based detection, optical tracking, and event-data demodulation in a single pipeline. By simultaneously tracking and demodulating multiple LED smart beacons, ATOC transforms a conventional visual marker into a robust communication channel suitable for a wide range of high-impact robotics applications. We validate ATOC in a suite of laboratory studies and two 'applications': a smart-city demonstration and a 3D motion-capture system.

230. 【2610.04336】A differentiable Lagrangian-coupled 3D Gaussian Splatting-SPH model for forward simulation and inverse analysis in solid mechanics

链接:https://arxiv.org/abs/2610.04336

作者:Tian Xu,Soroush Atashi,Tianju Xue

类目:Computer Vision and Pattern Recognition (cs.CV); Computational Physics (physics.comp-ph)

关键词:Recent advances, generative world models, advances in generative, generative world, increased interest

备注:

点击查看摘要

Abstract:Recent advances in generative world models have increased interest in digital models that reproduce both the appearance of real objects and their response to physical interaction. Three-dimensional reconstruction techniques, including 3D Gaussian Splatting, capture detailed surface geometry and appearance from images and videos. However, extending these representations beyond plausible animation to mechanically interpretable models for constitutive behavior, boundary conditions, and inverse parameter identification remains less explored. In this work, a differentiable Lagrangian-coupled 3DGS-smoothed particle hydrodynamics (SPH) model is proposed for forward simulation and inverse analysis of deformable solids. The observed object is first reconstructed from multi-view calibrated visual dataset as a 3DGS rendering model. An envelope-based procedure then generates an independent SPH support for the solid-mechanics model, avoiding the direct use of rendering primitives as mechanical particles. A reference-configuration Lagrangian transfer maps SPH deformation to Gaussian positions and covariances, thereby coupling the physical model and the image observation model while preserving a differentiable computational path. The SPH formulation supports linear elastic, hyperelastic, and Kelvin--Voigt viscoelastic responses, together with fixed, free, and Robin-type boundary conditions. Numerical studies validate the SPH response against finite-element results, assess accuracy and efficiency against a conventional model using Gaussian centers as surface SPH particles, and demonstrate forward simulations on beam, bridge, and liver-shaped examples. Inverse analyses further estimate constitutive and boundary parameters from rendered deformation observations, including noisy cases, demonstrating the feasibility of the proposed model for mechanics-based parameter identification from image data.

231. 【2610.04335】Synthetic-to-Real ViT-Based Pose Estimation of a Noncooperative UAV

链接:https://arxiv.org/abs/2610.04335

作者:Krishnanujam Srinivas,Hanish Acharla,Brij Agrawal,Leonardo Herrera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Unmanned Aerial Vehicles, noncooperative Unmanned Aerial, Aerial Vehicles, Unmanned Aerial, noncooperative Unmanned

备注: 12 pages, 8 figures, 5 tables

点击查看摘要

Abstract:Remote pose estimation of noncooperative Unmanned Aerial Vehicles (UAVs) from imagery is critical, as they cannot be influenced or instrumented in advance. Deep-learning-based approaches offer a promising solution; however, their development is constrained by the cost and difficulty of acquiring large-scale real-world datasets with accurate pose labels. Synthetic imagery provides an alternative, but models trained on synthetic data must overcome the synthetic-to-real domain gap to generalize to real-world imagery. This work investigates the inherent synthetic-to-real generalization capability of a Vision Transformer (ViT)-based model for monocular UAV pose estimation. The proposed approach employs a self-supervised DINOv2 backbone and is trained exclusively on labeled synthetic imagery while being evaluated on labeled real-world imagery. Pose ambiguity-aware strategies are incorporated during training and inference to address ambiguities arising from the projection of a three-dimensional target onto a two-dimensional image plane and from target symmetries. An $\alpha$-$\beta$ filter is further integrated during inference to improve pose estimations. To assess the model under operational requirements, it is evaluated in terms of Mean Angular Error (MAE) and inference time, both before and after filtering, using a real-world dataset containing 77,077 labeled UAV images. Before filtering, the model achieves an MAE of $19.18^{\circ}$ and an inference time of $13.25$ ms, whereas after filtering, these values are $8.74^{\circ}$ and $13.42$ ms, respectively.

232. 【2610.04327】rust the View That Sees the Target: Mining Cross-View Conflicts for Reliability-Gated Disaster Damage Assessment

链接:https://arxiv.org/abs/2610.04327

作者:Yifan Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:trusting both equally, overhead tiles, ground-level photographs, views symmetrically, methods fuse

备注: 4 pages, 3 figures; Accepted for publication in the proceedings of the 5th ACM SIGSPATIAL International Workshop on Searching and Mining Large Collections of Geospatial Data (GeoSearch '26), held November 3-6, 2026, in Riverside, California, USA

点击查看摘要

Abstract:After a disaster, building damage is assessed from overhead tiles and ground-level photographs, and most methods fuse the two views symmetrically, trusting both equally for every building. This paper focuses on the samples where that assumption fails: the conflict cases, on which two independently trained single-view models disagree. We mine such cases from three paired collections (inspection photographs from the 2025 Eaton wildfire and street-view panoramas from Hurricanes Ian and Milton, each matched to very-high-resolution overhead tiles), where they make up 10-33% of the data. On these samples an oracle that simply trusts the correct view beats every fusion method we tested by 0.37-0.41 accuracy, and the gap survives longer training, calibration, and backbone changes. We recover part of it with a visibility-conditioned reliability gate: a linear model that decides which view to trust from building-visibility features, calibrated per-view confidences, and the disagreement itself. On the wildfire data the gate is the only method that significantly beats calibrated probability averaging (+0.051 on conflicts, p=0.0001) and end-to-end fusion (+0.072, p10^-4); on the panoramic datasets it matches them. A controlled field-of-view experiment explains why: cropping panoramas toward the building doubles the benefit of fusion, whereas random crops of the same size do not. Finally, the spatial density of conflicts predicts tile-level damage without labels (Spearman r=0.615, p=0.001). Mining conflicts turns "does fusion help?" into "which view should be trusted, where, and why?".

233. 【2610.04318】Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

链接:https://arxiv.org/abs/2610.04318

作者:Sixun Dong,Wei Li,Andong Deng,Qi Qian,Victor Zhu,Zhengping Ji,Chen Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Efficient long-video understanding, selecting informative frames, Efficient long-video, vision-language models, understanding with vision-language

备注: Accepted at NeurIPS 2026. Project page: [this https URL](https://sixundong.com/projects/lohi)

点击查看摘要

Abstract:Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate pool rather than the final token budget. An empirical study across multiple VLMs and long-video benchmarks yields three findings: dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; resolution-sensitive tasks benefit from selected high-resolution frames; and front-end decoding dominates wall time for hour-long videos. Motivated by these findings, we introduce LoHi, a training-free, single-pass framework that combines a dense low-resolution video stream with sparse high-resolution image frames through the VLM's native video and image pathways. LoHi-Anchor selects high-resolution frames using codec I-frame metadata, while LoHi-SemDiv uses query relevance and visual diversity over CLIP features. Across three long-video benchmarks, LoHi improves average accuracy by 10.6 percentage points over the native-resolution baseline at a matched token budget and by 5.2 percentage points over the strongest prior efficiency method. It also reduces front-end decoding latency by up to 7x on hour-long videos. Project page: this https URL

234. 【2610.04300】A Geometric-Transformation Feature-Adaptive Manifold Restoration Method for Open-Vocabulary Semantic Segmentation of Remote Sensing Images

链接:https://arxiv.org/abs/2610.04300

作者:Jianzheng Wang,Huan Ni,Xiaonan Niu,Danfeng Hong,Haiyan Guan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remote sensing images, geometric transformations, dihedral group, typically invariant, remote sensing

备注:

点击查看摘要

Abstract:The semantic information of objects in remote sensing images is typically invariant to geometric transformations from the dihedral group D4. However, SAM3-based open-vocabulary semantic segmentation (OVSS) methods often exhibit inconsistent responses to different geometric transformations. To exploit this property and improve the stability of OVSS for remote sensing images, we propose a feature-adaptive manifold repair method based on dihedral-group geometric transformations. First, we introduce multi-scale harmonic-guided D4 view selection (MH-D4VS) to select complementary candidate views from a set of geometrically transformed views. Next, we propose original-view-anchored adaptive manifold repair (OAMR), which uses the original view as an anchor and reliable cross-view information to selectively repair locally unreliable visual features. Finally, we develop pixel decoder test-time adaptation (PD-TTA) for SAM3, which fine-tunes only the parameters of the GroupNorm layers online during inference, thereby enhancing the model's ability to adapt to sample-level distribution shifts. Experimental results show that the proposed method achieves an average mIoU of 55.6% across eight remote sensing semantic segmentation benchmarks and delivers consistent performance improvements under different SAM3-based inference frameworks.

235. 【2610.04281】OctMesh: A Unified Octree-Hierarchical Framework for Lossless Triangle Mesh Compression

链接:https://arxiv.org/abs/2610.04281

作者:Shiyu Feng,Xihua Sheng,Lingyu Zhu,Chunyang Fu,Shiqi Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Lossless triangle mesh, triangle mesh compression, Lossless triangle, triangle mesh, mesh compression

备注: 13 pages, 12 figures. Interactive visualization: [this https URL](https://hiddengalaxy1.github.io/octmesh-demo/)

点击查看摘要

Abstract:Lossless triangle mesh compression must preserve both vertex positions and connectivity. Octrees support learned point cloud geometry coding and progressive refinement, but extending them to meshes requires a compatible connectivity representation. Unlike the eight occupancy decisions of a voxel, a parent edge can develop into varied child connections, making edge refinement difficult to model with a compact prediction prior. We propose OctMesh, a learned framework that codes geometry and connectivity on a shared octree hierarchy. Its key observation is that octree pooling produces parents with either one child or two to eight children. Child edges are then grouped by their endpoint parents' types and whether the endpoints share a parent. Each candidate group contains children from just one parent or two connected parents. The resulting four categories define small, fixed-shape prediction tasks: connections uniquely determined by the parent graph are inherited without bits, while three neural predictors estimate probabilities for the remaining candidates. These probabilities guide arithmetic coding of the actual edge symbols. Binarized predictions of within-parent connections provide context for predicting connections between different parents. A graph-aware parent feature extractor combines local geometry, parent connectivity and global shape. The connectivity models use dedicated weights at coarse levels and share weights at fine levels. Residual edges and a finest-level face-selection payload complete the reconstruction. On 256 frames from eight MPEG V-DMC test sequences, OctMesh losslessly recovers the finest-level vertex-coordinate, edge and unoriented face sets at an average of 7.033 bits per face, 12.8% below V-Mesh. The same hierarchical representation supports nine levels of progressive vertex-and-edge refinement.

236. 【2610.04255】Grounded in Time: A Multi-Source Dataset and Benchmark for Temporal Grounding in Robotic Manipulation

链接:https://arxiv.org/abs/2610.04255

作者:Yi Wang,Yang Yang,Guangqi Xu,Sumin Lin,Ning Kang,Pengxiang Lu,Xiaotong Chen,Zeyu Xue,Ping Deng,Xing Liu,Chenguang Yang,Zhenyu Lu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:inferring task-relevant states, requires inferring task-relevant, Robotic manipulation, inferring task-relevant, task-relevant states

备注:

点击查看摘要

Abstract:Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent semantic inference remain underrepresented. We introduce GiT (Grounded in Time), a dataset and benchmark for grounding manipulation decisions in past events across biolaboratory, household, and industrial scenarios. It includes real-robot and Universal Manipulation Interface (UMI) style demonstrations covering 18 bimanual tasks, together with simulation data and a ManiSkill-based evaluation suite covering nine tasks. Fine-grained subtask annotations and annotated counterfactual task pairs, in which similar current observations require different actions depending on prior events, support policy learning and targeted evaluation of history use. Evaluations of representative end-to-end VLA models in simulation and on selected real-world tasks reveal substantial room for improvement in history-dependent manipulation. The dataset and benchmark are available at the project page.

237. 【2610.04225】FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding

链接:https://arxiv.org/abs/2610.04225

作者:Ziye Zhu,Yanghao Zhou,Lixing Tan,Jialiang Kang,Shuxuan Li,Xiao Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, demonstrated strong performance

备注: 15 pages, 6 figures

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational overhead. Many existing methods prune visual tokens during or after vision transformer (ViT) encoding, leaving much of the encoding cost unaddressed. Some approaches prune patches before encoding but rely on learned auxiliary networks for patch selection, incurring additional training and inference overhead. To address these limitations, we propose FlashGaze, a training-free method that reduces spatiotemporal redundancy before ViT encoding without introducing auxiliary networks. FlashGaze uses pixel-space differences as a proxy for information loss and employs Quadtree Dynamic Programming to jointly optimize patch dropping, merging, and keeping under a fixed budget. Experiments on two MLLM backbones across multiple benchmarks demonstrate substantial efficiency gains while largely preserving accuracy. On Qwen3-VL-8B, FlashGaze retains 98% of the full-input baseline accuracy on LongVideoBench while achieving up to 5.4x and 17x speedups in ViT encoding and MLLM prefill, respectively, and reducing peak GPU memory usage by a factor of 1.8. These efficiency gains enable the model to process videos with more frames and higher resolutions on the same GPU hardware, unlocking video understanding at scales previously out of reach.

238. 【2610.04203】Sparse-GS2Mesh: 3D Gaussian Splatting Guided by Novel Stereo Views and 2DGS for Sparse View Surface Reconstruction}

链接:https://arxiv.org/abs/2610.04203

作者:Younghyun Noh,Minje Kim,Tae-Kyun Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains challenging due, Gaussian Splatting, Surface reconstruction, sparse-view settings remains, settings remains challenging

备注: 13, pages, 7 figures, accepted to NIPS 2026 PhysWorldAI workshops

点击查看摘要

Abstract:Surface reconstruction under sparse-view settings remains challenging due to limited geometric cues. Volume rendering methods based on signed distance functions often produce over-smoothed surfaces, while 3D Gaussian Splatting (3DGS), though time-efficient, suffers from incomplete geometry due to the lack of reliable depth supervision and the limitation of being optimized only from given input views. In this paper, we present Sparse-GS2Mesh, a stereo-aware framework for surface reconstruction from sparse views. While 3DGS and stereo matching have been leveraged for surface reconstruction under dense view settings, we extend them to operate effectively under sparse view conditions by first initializing 3DGS using epipolar depth priors to mitigate the 3DGS overfitting problem, followed by our three key components: (I) adaptive baseline selection, (II) fine-tuning with a stereo matching network, and (III) 2D/3D co-regularized fine-tuning. Given a warmed-up 3DGS initialized with epipolar depth, the adaptive baseline selection automatically determines a baseline to synthesize for each sparse view. We then fine-tune 3DGS by backpropagating depth-refining gradients from the stereo matching network, effectively specializing the 3DGS for stereo matching. The 2D/3D co-regularization further helps obtain stable reconstruction, addressing weak geometric cues in close stereo views. Sparse-GS2Mesh achieves a 15\% improvement over state-of-the-art methods in little-overlap settings and comparable results in large-overlap settings. Codes will be publicly available.

239. 【2610.04199】CellSplat4D: PSF-Aware 4D Gaussian Splatting for Sparse Robotic Live-Cell Imaging

链接:https://arxiv.org/abs/2610.04199

作者:Yingda Tao,Guoyu Lu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:watching living cells, robotic microscope watching, microscope watching living, watching living, living cells

备注:

点击查看摘要

Abstract:A robotic microscope watching living cells cannot afford to look as often as it would like. Every volume it acquires costs photons the specimen does not get back, and time owed to other wells. What such a platform exists to produce is a record of individual cells through time: which cell is which from one volume to the next, and which cell divided into which two. Sampling sparsely breaks that record exactly where it matters, and the fault lies in the acquisition schedule rather than in the analysis software. We fill the gaps by reconstructing them, fitting a 4D Gaussian model to whatever volumes the hardware could afford. The model is a cloud of light-emitting blobs, each carrying a position, a shape and a lifetime. Being continuous in time, it renders any missing volume on demand, decoupling how often the robot analyses from how often it can afford to look. The microscope's point-spread function is measured from the data rather than inherited from acquisition metadata or left to the optimizer, because metadata inflates it and the optimizer cannot recover it at all: a wider blur around a smaller blob fits the images equally well. Each blob's lifetime is stored in frames rather than as a fraction of the recording, so that it denotes a fixed duration on any sequence. Unmeasured timesteps are supervised at coarse scale by a 3D U-Net that predicts the intermediate volume directly and estimates no motion field, since a dividing cell becomes two and no motion describes that. On two Cell Tracking Challenge sequences, a C. elegans embryo and a Chinese Hamster Ovarian (CHO), with fidelity scored per cell nucleus, our reconstruction holds the highest nucleus fidelity at every distance from an acquired frame, has the flattest decay across the gap, and best recovers focal planes it was never shown with graceful degradation across the gap.

240. 【2610.04185】Referring Multi-Object Tracking in Moving-Camera Videos via Global Motion Compensation

链接:https://arxiv.org/abs/2610.04185

作者:Hsin-Chen Pai,Jyun-Kai Wang,Yi-Cheng Peng,Wei-Ta Chu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Referring multi-object tracking, Referring multi-object, input and tracks, tracks all referred, Referring

备注:

点击查看摘要

Abstract:Referring multi-object tracking (RMOT) takes a video and a language expression as input and tracks all referred objects. Many tracking requirements involve how an object moves rather than how it appears. However, in a video captured by a moving camera, a parked vehicle may appear to move, while a moving vehicle may show little displacement. Existing RMOT methods relate motion with text but do not explicitly remove camera-induced motion. In this paper, we propose extracting residual motion across frames by estimating camera motion in driver-view videos and compare motion characteristics with the query expression. We consider the motion-matching extent and integrate it with the RMOT method's prediction result through late fusion. In the evaluation, we verify the performance gain of taking the motion compensation module as a plug-in across different RMOT hosts.

241. 【2610.04159】Diagnosis-Conditioned Spatial Gating and Decoder-Level Supervised Contrastive Learning for Radiology Report Generation

链接:https://arxiv.org/abs/2610.04159

作者:Md Mustafizur Rahman,Mylene C. Q. Farias

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Radiology report generation, produce fluent text, Radiology report, finding-level inaccuracies, produce fluent

备注: 5 pages. Manuscript submitted to ICASSP 2027

点击查看摘要

Abstract:Radiology report generation models can produce fluent text while still containing finding-level inaccuracies. Diagnosis-driven methods improve generation by conditioning on predicted findings, but these predictions do not directly modify the visual patch features provided to the decoder, and global gating applies the same modulation across spatial locations. We introduce a Position-Aware Gate (PAG) that uses predicted finding representations to modulate visual patches spatially without region supervision. We also propose a Decoder-level Supervised Contrastive Loss (DSCL) that structures decoder representations using shared positive findings rather than instance identity. On MIMIC-CXR, PAG+DSCL improves clinical efficacy (CE) F1 from 0.484 to 0.502 over a matched global-gate reference, while PAG and DSCL individually reach 0.491 and 0.495. Without additional fine-tuning, the combined model achieves 0.226 CE F1 on IU X-Ray, compared with 0.211 reported by PromptMRG.

242. 【2610.04152】Kepler4D: Controllable Future Video Generation via 4D Scene State Evolution

链接:https://arxiv.org/abs/2610.04152

作者:Feiran Wang,Bin Duan,Junyi Wu,Gaowen Liu,Yan Yan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:preserve scene structure, world models aim, dynamic objects evolve, visual observations, aim to preserve

备注: Project page: [this https URL](https://brack-wang.github.io/kepler4d/)

点击查看摘要

Abstract:Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D constructs a shared 3D representation of background geometry, object motion histories, coarse spatial supports, and semantic context. Chain-of-Motion summarizes observed motion and uses a vision-language model to select structured speed and heading decisions and decide whether to bound object-center height from below. A deterministic rollout converts these decisions into future object trajectories for inspection and editing before synthesis. We render the evolving proxies into geometric controls for a pretrained video generator, separating coarse object motion from the synthesis of appearance and articulation. Experiments on real-world videos demonstrate that Kepler4D enables controllable object motion and plausible future rollout while preserving scene consistency.

243. 【2610.04151】Watermarks and Fingerprints as Soft Bindings for Content Provenance: An Open-Licence Benchmark for Images, Audio and Video

链接:https://arxiv.org/abs/2610.04151

作者:Seyedmahdi Kazempourradi(1),Ramtin Mojtahedi(1),Behrang Mohseni(1) ((1) Original Pictures Technologies, Inc., Delaware, USA)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Content-provenance standards, invisible watermark read, soft binding, platform recover, recover a stripped

备注: 59 pages. Code, recorded results and manuscript source: [this https URL](https://github.com/Original-Pictures/soft-bindings-benchmark) (v1.0)

点击查看摘要

Abstract:Content-provenance standards such as C2PA let a platform recover a stripped manifest through a soft binding: an invisible watermark read from the content, or a fingerprint looked up in a registry. We benchmarked both families under one protocol, restricted to openly available models whose licences we audited, on public media, with false-match rates calibrated on held-out negatives and source-level bootstrap intervals for performance estimates. For watermarking we evaluated 25 image, 7 audio and 7 video configurations from 12 methods on perceptual quality, robustness, false positives and cost; for fingerprinting, 35 methods on registries of up to 98,985 images, partial edits and adversarial attacks. PixelSeal gave the best balance for image and video watermarks and AudioSeal for audio, but the error-correcting detector of every TrustMark variant fired on 5.9 to 15.3% of unmarked images, so a verifier should test the expected payload. Among fingerprints, copy detectors trained on non-commercial data detected up to 75.5% of transformed images at a pair-level false-match rate of $10^{-7}$ and DINOv2, the best permissively licensed method, 65.0%; near-copies in a product catalogue dominated the false matches, and no detection improvement from geometric verification was observed under a matched calibration false-binding constraint in the evaluated image pipelines. On the same attacked copies the two families failed differently: the union of watermark and fingerprint successes covered 69% of image copies; fingerprints covered more audio queries, whereas expected-key watermark verification covered more video queries. Embedding a watermark moved the ISCC code of 62.0% of images past its match threshold, and platform-dependent colour conversion changed watermark bits between x86 and ARM hosts.

244. 【2610.04139】From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models

链接:https://arxiv.org/abs/2610.04139

作者:Feiran Wang,Xiaoqi Wang,Ziwei Li,Wenbin He,Yan Yan,Liu Ren

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Predicting future spatial, supports collision avoidance, states supports collision, Predicting future, spatial reasoning

备注: Project page: [this https URL](https://brack-wang.github.io/spatialmind/)

点击查看摘要

Abstract:Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.

245. 【2610.04104】Where Does the Semantic Gain Come From? A Reproduction and Extension of Semantic Knowledge-driven Contrastive Learning for Long-Tailed Recognition

链接:https://arxiv.org/abs/2610.04104

作者:Sushrut Ghimire

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Knowledge-driven Contrastive Learning, Balanced Contrastive Learning, Semantic Knowledge-driven Contrastive, Contrastive Learning, Knowledge-driven Contrastive

备注: 11 pages, 4 figures. Code: [this https URL](https://github.com/sushrutghimire1/Reproduction-and-extension-of-SKCL)

点击查看摘要

Abstract:Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points above Balanced Contrastive Learning (BCL), the method it builds on. The code and the class descriptions are not public. I reimplement SKCL, BCL and ConCutMix in one framework, check it against the public baseline code, and run every configuration with three seeds. The two baselines reproduce within 1.5 points, but SKCL built on BCL, as the paper describes it, ends up 0.56 points below BCL. To find out why, I add SKCL to the authors' own ConCutMix code. Trained for the paper's 300 epochs, it reaches 53.69, only 0.33 below the published number. At the same budget, however, the semantic graph adds just 0.23 points over ConCutMix, while training ConCutMix for 100 more epochs adds 1.04. Together with ConCutMix's published lead over BCL (1.15), this explains the claimed gain. A BCL model that never sees the graph already shares 41.2% of the graph's top-2 neighbours with its own most-confused classes (2.0% by chance), which shows why the graph adds so little on these benchmarks. I also test several changes to SKCL. Combining it with the CutMix branch improves it by 1.06 points, and an adaptive version of the graph improves it slightly (+0.34 and +0.28 in two codebases), although these gains are within seed noise.

246. 【2610.04095】Scaling 3D Visual Grounding in Abdominal CT

链接:https://arxiv.org/abs/2610.04095

作者:Sam Church,Danyal Maqbool,Joshua D. Warner,Andrew Voter,Junjie Hu,Meghan G. Lubner,Tyler J. Bradshaw

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:enhance radiology workflows, Visual grounding, enhance radiology, radiology workflows, workflows by linking

备注:

点击查看摘要

Abstract:Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired phrases and regions, and building such datasets is expensive, requiring radiologists to annotate images by hand. We posit that this supervision is already created implicitly during routine reporting, as radiologists frequently place 2D annotations (e.g., distance measurement, arrows) on key images to make measurements and to support report interpretation. We introduce an automated pipeline that converts these routine clinical annotations into large-scale phrase-region supervision for 3D visual grounding. The pipeline links each annotation to the corresponding finding in the report through metadata matching, then uses a promptable 3D segmentation model to convert the 2D annotation into a volumetric mask. This produces phrase-mask-volume datasets without requiring additional radiologist annotation. Applied to a single institution's clinical picture archiving and communication system (PACS), our approach generated 105K phrase-mask-volume triplets from 59K abdominal CT exams. We also introduce two abdominal CT grounding benchmarks, LocusBench-Onc and LocusBench-ED, which comprise 240 oncology and 260 emergency-department radiologist-reviewed phrase-mask-volume triplets, respectively, with the latter spanning 13 distinct categories such as appendicitis, hematoma, and hernia. We further introduce LocusCT, a 3D visual grounding model trained on this dataset, which achieves hit rates of 0.725 on LocusBench-Onc and 0.773 on LocusBench-ED, substantially outperforming comparator models. These results show that routine PACS annotations are a scalable, previously unused source of supervision for 3D visual grounding.

247. 【2610.04092】UniBRep: Learning Unified Geometry and Topology for Image-conditioned B-Rep Generation

链接:https://arxiv.org/abs/2610.04092

作者:Haiyang Ying,Allen Tu,Jiaye Wu,Tom Goldstein,Matthias Zwicker

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:single image requires, image requires faithful, requires faithful reconstruction, single image, image requires

备注: 15 pages, 11 figures

点击查看摘要

Abstract:Generating a boundary representation (B-rep) conditioned on a single image requires faithful reconstruction of geometry, valid topology, and support for complex shapes. We present UniBRep, a geometry-first framework that adapts a pretrained image-to-3D model to generate a feature mesh as a unified intermediate representation. Its surface provides a geometric scaffold, while spatially aligned learned features encode face-separation cues for topology recovery. Dual decoder branches generate the geometry and face-separation features; a geometry- and feature-guided construction pipeline then fits parametric surfaces, recovers boundary curves and connectivity, and assembles an explicit B-rep using a CAD kernel. Recovering topology from mesh regions avoids predefined architectural face-count limits, allowing face count to scale with shape complexity. On the standard DeepCAD benchmark, UniBRep produces valid B-reps for 80.49\% of inputs and reduces face Chamfer distance from 0.1096 to 0.0345 relative to CADDreamer. In a matched comparison, UniBRep also outperforms the HoLa public demo across all reported metrics. Further evaluations demonstrate scalability to high-complexity shapes beyond the standard 30-face range, generalization to objects outside the CAD training distribution, and qualitative transfer to real photographs.

248. 【2610.04084】Dependable AI-Assisted Engineering: A Formal Framework for AI Participation and Assurance in Safety-Critical Workflows

链接:https://arxiv.org/abs/2610.04084

作者:Puxue Tan

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:enter safety-critical workflows, enter safety-critical, produce engineering artefacts, safety-critical workflows, participation and assurance

备注:

点击查看摘要

Abstract:Generative AI can produce engineering artefacts, but generation alone does not determine whether or how those artefacts should enter safety-critical workflows. This paper develops a formal framework for assigning AI participation and assurance at the level of individual workflow units. Each unit has a participation and assurance record covering its engineering requirement, an approved operational formalization where applicable, the applicable mechanism, fallback where applicable, evidence obligations and the applicable guarantee, plus a deployment-readiness status. The framework distinguishes deterministic verification, statistically calibrated admission, authorized human judgement supported by AI advice, authorized human adjudication of AI-produced artefacts, retained deterministic tool paths and explicit non-participation; these arrangements carry different kinds of guarantee rather than levels on a common scale. The framework also separates formalization fidelity from verifier soundness, provides a staged classification and readiness procedure, and derives conditions for comparing a gated AI-assisted unit with an incumbent process under recurring-population assumptions. We instantiate and apply the framework in an executed 17-unit wing-spar structural-analysis workflow combining deterministically gated AI-generated CAD, retained deterministic computation and human judgement. The AI-generated CAD program passed all 23 deterministic checks and was admitted at the first attempt. Favourable stress magnitudes did not suffice to pass the stress criteria where the predeclared mesh-convergence evidence was insufficient; those criteria were instead referred to engineering judgement. The case demonstrates selective AI participation and explicit evidence handling at unit level; no claim is made of workflow-level dependability, certification, structural safety or productivity.

249. 【2610.04076】VCURF: Virtual Camera-based Uncertainty of Radiance Fields

链接:https://arxiv.org/abs/2610.04076

作者:Liyan Chen,Nathaniel Burgdorfer,Philippos Mordohai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, radiance field model, radiance field, radiance field models, Radiance fields

备注:

点击查看摘要

Abstract:Radiance fields, implemented with either implicit (NeRF) or explicit (Gaussian Splatting) representations, are advancing the state of the art in novel view synthesis at a rapid pace. Even though the rendered views they generate are often compelling, they are not free of errors. In this paper, we propose a new approach for pixel-wise uncertainty quantification based on measuring the inconsistencies among renderings by the radiance field model in virtual cameras sampled near the target viewpoint. We named our approach VCURF for Virtual Camera-based Uncertainty of Radiance Fields. VCURF treats the radiance field model as a black box, only assuming that it is capable of rendering color and depth on demand. This property makes our approach applicable to both NeRF and GS models without any modification. Our experiments on a combination of datasets, radiance field models and baselines demonstrate VCURF's effectiveness in pixel-wise uncertainty estimation. We conclude the paper with findings that question the way view selection is tackled by the majority of the current literature.

250. 【2610.04073】Why Convolution Still Matters: Evaluating Inductive Biases in Cryospheric Image Classification

链接:https://arxiv.org/abs/2610.04073

作者:Chhaya Kulkarni,Emam Hossain

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:remote sensing image, Greenland Ice Sheet, Recent advances, attention-based deep learning, Ice Sheet image

备注: Accepted at IEEE IGARSS 2026, 5 pages, 3 figures

点击查看摘要

Abstract:Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unclear. In this work, we revisit a benchmark Greenland Ice Sheet image dataset, previously shown to favor convolutional neural networks (CNNs), to examine whether modern attention-based and hybrid architectures improve class-wise reliability. We conduct a controlled comparison between a classical CNN (AlexNet), a modern CNN (ConvNeXt-Tiny), a pure attention-based model (Swin-Tiny), and a hybrid convolution-attention model (CoAtNet-0) under identical training and evaluation protocols. Results show that AlexNet achieves the highest accuracy and the strongest balanced performance as measured by macro-averaged F1, while ConvNeXt-Tiny exhibits the highest macro-averaged AUC, indicating strong class separability but less consistent final decision quality. Class-wise analysis reveals that hybrid architectures improve recall for rare and structurally distinct surface classes, whereas convolutional models remain more reliable for texture-dominated categories. These findings highlight the importance of aligning architectural inductive bias with cryospheric data characteristics and suggest that increased model complexity does not necessarily translate to improved reliability for ice-sheet surface classification.

251. 【2610.04066】Evaluating Zone-Guided Front Extraction for Glacier Calving-Front Delineation in SAR Imagery

链接:https://arxiv.org/abs/2610.04066

作者:Chhaya Kulkarni,Emam Hossain

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Automatic calving-front delineation, synthetic aperture radar, aperture radar imagery, Automatic calving-front, glacier ice

备注: Accepted in ICMLA 2026 as short paper. 4 pages, 1 figure

点击查看摘要

Abstract:Automatic calving-front delineation from synthetic aperture radar imagery is challenging because the front is a thin and often ambiguous boundary between glacier ice, ocean, and surrounding rock or terrain. The CAlving Fronts and where to Find thEm (CaFFe) dataset provides both binary calving-front masks and broader semantic zone masks, making it possible to study whether zone-level supervision can support front recovery. In this paper, we compare direct front prediction with zone-guided front extraction using U-Net, DeepLabV3+, and SegFormer-B0 under the same bounding-box-cropped CaFFe setting. In the direct setting, models predict the binary calving-front mask. In the zone-guided setting, the model first predicts four semantic zone classes, and the front is then extracted from the predicted glacier-ocean boundary. We evaluate both zone-level and front-level performance, include a ground-truth-zone boundary check, and examine lightweight test-time adaptation on sensor-specific and glacier-specific subsets. The results show that zone labels contain useful front-boundary information: extracting the front from ground-truth zones gives the lowest mean distance error. However, fronts extracted from model-predicted zones remain weak, even when zone segmentation scores are moderate. Test-time adaptation also does not consistently improve zone-guided front recovery. These results indicate that zone segmentation performance should not be treated as a substitute for front-level evaluation and that effective use of zone labels may require boundary-aware training, label fusion, or explicit front supervision.

252. 【2610.04052】A Theory of Shape Reconstruction from Heat Conduction and Shading

链接:https://arxiv.org/abs/2610.04052

作者:Akihiko Oharazawa,Sriram Narayanan,Mani Ramanagopal,Srinivasa G. Narasimhan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:inherently ambiguous, bertian surface, Lam, Shape, surface normal estimation

备注: Project Page: [this https URL](https://shape-from-heat-and-shading.github.io/)

点击查看摘要

Abstract:Shape from shading using a single image of a Lam- bertian surface is inherently ambiguous. When the light source direction is known, the surface normal estimation has a cone- ambiguity, which worsens when the source is unknown. Recently, shape from heat conduction has emerged as an approach that leverages heat transport equations to estimate the Shape Lapla- cian operator, an intrinsic measure of shape. However, deriving surface normals from the Laplacian operator encounters a local binary convex/concave ambiguity. Our contribution introduces a novel theory to resolve these local shape ambiguities (excluding a few degeneracies) without relying on priors like smoothness, by combining the cues from shading and heat conduction. Our method ensures the mathematical constraints of both shading and the Laplacian are satisfied simultaneously, even with an unknown light source. We validate our theory through simulations of complex shapes and analyze its performance in the presence of noise, as well as on a noisy single thermal video of real-world objects with complex shapes and material properties, including varying albedo.

253. 【2610.04051】ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction

链接:https://arxiv.org/abs/2610.04051

作者:Sidrah Liaqat(1),Halil Helvaci(1),Sen-Ching Cheung(1),Chongruo Wu(2),Dongjie Chen(2),Chen Nee Chuah(2),Sally Ozonoff(3) ((1) University of Kentucky, (2) University of California Davis, (3) MIND Institute, University of California Davis)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Autism Spectrum Disorder, Accurate early screening, Spectrum Disorder, Autism Spectrum, Accurate early

备注:

点击查看摘要

Abstract:Accurate early screening for Autism Spectrum Disorder (ASD) is a precursor to timely intervention, which is critical for improving cognitive and behavioral outcomes. We present ASD-FEAT (ASD - Feature Extraction And Tracking), a multimodal dataset derived from video recordings of infant-adult interaction sessions. The key contribution of ASD-FEAT is the combination of longitudinal coverage from infancy through 36 months, repeated interaction sessions, clinically validated developmental outcomes, expert frame-level behavioral annotations, and privacy-conscious multimodal feature representations. To the best of our knowledge, existing ASD behavioral datasets do not jointly provide these characteristics at comparable scale. To demonstrate the utility of ASD-FEAT, we use it to evaluate a computer-vision-based end-to-end pipeline relying on machine learning techniques to automatically identify ASD risk. ASD-FEAT integrates both expert-defined and deep-learned features, including face and eye landmarks, facial action units, gaze direction, head position, mel-spectrogram audio representations, and optical flow, to identify behavioral markers of social interaction. Our automated pipeline achieves an ASD classification accuracy of 76.2% and an Area Under the Receiver Operating Characteristic (AUROC) of 0.82, compared to classifiers trained on manually labeled behaviors, which yielded 81.3% accuracy and an AUROC of 0.88. We further introduce a within-visit partner contrast: a per-visit signal contrasting examiner-directed and parent-directed social behavior which, when added to the classifier, lifts the fully automated Look Face + Smile and Look Face + Vocal configurations to Matthews correlations of 0.49 and 0.45 respectively, exceeding the human-coded single-partner baseline of 0.42.

254. 【2610.04044】Dynamic Quadtree Tokenization and Transformer for Adaptive Mesh PDE Forecasting

链接:https://arxiv.org/abs/2610.04044

作者:Yilin Zhuang,Noah Zambrano,Karthik Duraisamy

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:interfaces occupy small, material interfaces occupy, quadratic attention cost, reaction fronts, Vision Transformers

备注:

点击查看摘要

Abstract:The quadratic attention cost of Vision Transformers (ViTs) forces a trade-off between spatial resolution and rollout horizon, particularly for fine-scale PDEs where shocks, reaction fronts, and material interfaces occupy small, evolving regions of the domain. Conventional neural surrogates also lack mechanisms to adapt resolution dynamically. We propose WAMRViT, a ViT that tokenizes inputs as balanced quadtrees using a wavelet-inspired refinement criterion, jointly encodes position and refinement level with 3D rotary positional embeddings, and regrids in cell space during inference for stable long-horizon rollouts. A multi-scale variant retains each leaf at its native source resolution and lets the model learn across resolution levels. Unlike fixed-budget adaptive-tokenization methods, WAMRViT imposes no predetermined token count and supports fully adaptive topology throughout autoregressive rollout. To our knowledge, it is the first machine-learning surrogate to natively tokenize multi-level Adaptive Mesh Refinement (AMR) data. On uniform-grid benchmarks, uniform-patch WAMRViT improves finest-level region-of-interest VRMSE over a finest-patch uniform ViT while using substantially fewer tokens. The multi-scale variant achieves the lowest first-step full-field RMSE and VRMSE on both benchmarks and improves rollout-averaged full-field and refined-region accuracy at long horizons. With parallelized regridding, its end-to-end rollout cost lies between finest-patch and approximately token-matched coarser-patch ViTs. On a complex AMR combustion problem whose finest features cannot be represented natively by the evaluated uniform-grid baselines, WAMRViT operates directly on adaptive cells and substantially reduces finest-level error at matched transformer capacity. Code: this https URL

255. 【2610.04036】MAGEFormer: Learning Metric-Consistent Representations for Anisotropic CT Segmentation

链接:https://arxiv.org/abs/2610.04036

作者:Jiaying Li,Paolo Remagnino

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Euclidean lattice assumption, Vision Transformers, isotropic Euclidean lattice, Euclidean lattice, shown strong performance

备注:

点击查看摘要

Abstract:Vision Transformers (ViTs) have shown strong performance in volumetric segmentation, but their effectiveness on clinical CT is limited by an isotropic Euclidean lattice assumption. This conflicts with anisotropic CT acquisition, leading to two key issues: (1) a metric mismatch between voxel indices and physical anatomy, and (2) accuracy degradation from isotropic resampling. To address this, we propose MAGEFormer, a geometry-calibrated framework that embeds physical metric constraints directly into representation learning. Our method introduces Metric-Adaptive Spatial Embedding (MASE) to calibrate positional frequencies using voxel spacing, Geometry-Constrained Attention (GCA) to suppress physically implausible feature correlations, and Geometric View Voting (GVV) to reduce discretization bias during inference. We evaluate MAGEFormer on two multi-organ abdominal CT benchmarks, BTCV and FLARE 22, under a unified protocol against strong CNN and Transformer-based baselines. MAGEFormer achieves the strongest boundary accuracy among the compared methods, with 10.58 mm HD95 on BTCV and 3.40 mm HD95 on FLARE 22, and shows consistent gains in Dice under the same protocol. These results show that geometry-aware internal calibration is more effective than relying on conventional isotropic preprocessing alone for anisotropic CT segmentation.

256. 【2610.04028】Learning Subject-Specific Anatomical Representations via Manifold Expansion: Application to Accelerated Multi-Contrast MRI

链接:https://arxiv.org/abs/2610.04028

作者:Ruimin Feng,Wanyu Bian,Albert Jang,Zachary Stewart,Fang Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:Clinical MRI routinely, MRI routinely acquires, complementary tissue characterization, routinely acquires multiple, acquires multiple contrast-weighted

备注:

点击查看摘要

Abstract:Clinical MRI routinely acquires multiple contrast-weighted images of the same anatomy for complementary tissue characterization. However, current accelerated MRI methods typically reconstruct each contrast independently, without fully exploiting shared anatomical information. This work aims to learn anatomical representations invariant to contrast-dependent appearance for reconstruction of accelerated multi-contrast MRI. We propose MAX (MAnifold eXpansion), a subject-specific framework that learns anatomical representations from a single fully sampled reference contrast. To address the under-constrained separation of shared anatomy and contrast-dependent components from a single image, MAX expands the multi-contrast manifold using anatomy-preserving intensity augmentations. A disentangled implicit neural representation models augmented samples using shared spatial coordinates for anatomy and spatially invariant coordinates for contrast appearance. The learned anatomical representation is then fixed, with the contrast representation adapted to the undersampled target data, followed by unrolled refinement. Theoretical analyses further provide insight into the disentangled representation learning and explain how the learned anatomical representation improves the target contrast reconstruction. At R = 8 for brain MRI and R = 6 for knee MRI, MAX achieves the highest mean PSNR and SSIM across all tasks, improving PSNR by more than 1 dB over the strongest baseline for both brain contrasts. MAX more faithfully recovers subtle anatomical and pathological structures and remains robust to inter-contrast motion, structural heterogeneity between reference and target contrasts, and measurement noise. Therefore, MAX provides a general strategy for leveraging high-quality reference scans in accelerated MRI and has the potential to be extended to other reference-assisted MRI inverse problems.

257. 【2610.04020】ClasSAE: Class-Aligned Sparse Autoencoders via Differentiable Feature-Class Affinity

链接:https://arxiv.org/abs/2610.04020

作者:Jakub Stępień,Marcin Mazur,Jacek Tabor,Przemysław Spurek

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Sparse Autoencoders, decomposing neural representations, bias mitigation, unsupervised tool, tool for decomposing

备注:

点击查看摘要

Abstract:Sparse Autoencoders (SAEs) began as an unsupervised tool for decomposing neural representations into sparse, interpretable features, and are increasingly used not only for passive analysis but also for active interventions such as unlearning, bias mitigation, and concept editing. A central challenge for these editing and steering methods is reliably matching features to target concepts; most current approaches address this by computing post-hoc scores over an already-trained, frozen dictionary. We instead introduce ClasSAE, a novel method that both automatically assigns classes to features and guides the encoder toward class-separable representations during training. Specifically, we apply a differentiable top-$k$ operator to a trainable feature--class affinity matrix with per-feature budgets, coupling the features selected for each sample to the classes they are trained to represent. Because gradients flow through the selection of active features rather than only through their magnitudes, the encoder and the affinity matrix co-adapt rather than being fit in separate stages. The result is a dictionary that is both class-separable and class-annotated, with no need for post-hoc probing. We propose three variants for enforcing sparsity within this framework, which achieve comparable overall performance with slightly different trade-offs. Using CLIP ViT-L/14 embeddings on ImageNet, we show that the learned affinity matrix agrees closely with an independently estimated post-hoc feature-class matrix computed on held-out data. The model also supports direct class prediction from the encoder and affinity matrix alone, without a separately fitted classifier, and its more class-aligned encoder yields improved separation in Targeted Probe Perturbation evaluations. this https URL.

258. 【2610.04014】PhysMamba: Selective State Space Models as Learned Articulated Body Simulators

链接:https://arxiv.org/abs/2610.04014

作者:Haochuan Zhang,Sinisa Todorovic

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:learned articulated body, articulated body simulator, body simulator based, state space model, selective state space

备注: CVPR 2026 Workshop on Physically Grounded Human Perception and Modeling -- 1st PhysHuman

点击查看摘要

Abstract:We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four architectures under partial- and full-observation inputs and three training protocols. The from-scratch rollout training protocol gives Mamba2 strong short- and mid-horizon accuracy under partial observation (s10 = 43 mm, 2/50 diverged), while the two-stage teacher-based rollout protocol stabilizes GRU but fails for Mamba2. With CUDA graph compilation, Mamba2 reaches 0.107 ms per frame (9,334 FPS) on an H100 GPU, within 1.1$\times$ of GRU's un-compiled throughput, adding under 1% latency to a 30 Hz HMR pipeline and enabling integration as a differentiable physics module for video-based mesh recovery.

259. 【2610.04009】SUAVE: Unified Video-Action Models via Masked Diffusion

链接:https://arxiv.org/abs/2610.04009

作者:Rhythm Syed,Jean Mercat,Sedrick Keh,Kushal Arora,Paarth Shah,Aykut Onol,Mengchao Zhang,Tony Dear

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:inherit strong semantic, strong semantic grounding, pretrained vision-language backbones, inherit strong, strong semantic

备注: Preprint version

点击查看摘要

Abstract:Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion backbones can imagine but treat language as frozen conditioning on a continuous latent space. Unified models bring these modalities into one architecture, but they either decode autoregressively, one token at a time, or keep video continuous with an auxiliary action head. In this work, we present SUAVE, a Single vocabulary Unified Action-Video modEl in which a masked diffusion transformer generates video and actions conditioned on language, with all three modalities represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. For action-free co-training, the action positions of unlabeled video are filled with mask tokens and excluded from the loss. Simulation and real-world experiments demonstrate two findings. First, a single SUAVE model predicts long-horizon video and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improves policy performance and zero-shot robustness to distribution shift. Together, these results show that masked diffusion is a practical and versatile foundation for unified video-action modeling.

260. 【2610.04007】VolS-GS: Relightable Gaussian Splatting with Volumetric Subsurface Scattering

链接:https://arxiv.org/abs/2610.04007

作者:Junyeong Ahn,Jaegul Choo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:relightable Gaussian splatting, Gaussian splatting framework, Gaussian Splatting methods, Gaussian splatting, relightable Gaussian

备注: 23 pages

点击查看摘要

Abstract:We present VolS-GS, a relightable Gaussian splatting framework that reconstructs objects from one-light-at-a-time (OLAT) captures and renders them under novel lighting and viewpoints. Relightable Gaussian Splatting methods typically model appearance independently at each primitive, which makes non-local effects difficult to represent. This limitation is particularly apparent for subsurface scattering, where light entering the object at one location can emerge at another. Rather than modeling this effect solely with a neural network or a local kernel at each primitive, we use the spatial support of the Gaussian scene as the domain of a differentiable finite-volume transport solver, so that light can propagate through the object's interior. A small network predicts scattering and absorption coefficients for each Gaussian, and the solve redistributes incident light through the resulting field. The coefficients are fit to images rather than measured, so the solve supplies a transport-shaped path for aggregating per-primitive appearance, not a measurement of the material. To keep the learned shadow and specular terms from taking over the other components, our shadow term is predicted from visibility together with the transmittances and the scattering the solve produces, and a regularizer suppresses specular highlights in regions the shadow term predicts to be unlit. Experiments on three OLAT benchmarks show that VolS-GS consistently improves relighting quality on held-out lights and views.

261. 【2610.04006】Verifier-Guided Synthetic Augmentation for 3D Human Shape Generation

链接:https://arxiv.org/abs/2610.04006

作者:Yuexuan Wu,Yang Xiang,Hamid Laga,Dip Das,Anuj Srivastava,Zhengwu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:Limited training data, training data diversity, data diversity constrains, diversity constrains generative, constrains generative modeling

备注:

点击查看摘要

Abstract:Limited training data diversity constrains generative modeling of 3D human bodies: conservative models remain close to observed examples, whereas exploratory models often violate basic body proportions. We introduce a verifier-guided augmentation framework that uses global and mode-local PCA to generate inexpensive candidates, screens them using correspondence-derived skeletal proportions and body-part geometry, and retrains a diffusion model on accepted candidates. Elastic registration provides both the modal structure used by distributed PCA and the dense anatomical correspondence needed for scalable screening without per-candidate body-model fitting. A blinded human study supports the verifier as a conservative gatekeeper, favoring verifier-accepted over rejected outputs. We evaluate full-pool verifier acceptance separately from the coverage and departure of accepted samples and combine them through EAUC. On 4,498 registered DFAUST surfaces, distributed-PCA augmentation achieves 86.32% acceptance, the highest CP-AUC (0.871), and the highest EAUC (0.752), improving EAUC by 32% over real-only and self-augmented diffusion. These results show that mode-local, verifier-guided proposals broaden diffusion generation while maintaining high agreement with calibrated body measurements.

262. 【2610.04003】Selective Backpropagation for Efficient Few-Shot Class-Incremental Learning

链接:https://arxiv.org/abs/2610.04003

作者:Eeham Khan,Abdulmoumen Al-Atrash,Ali Ayub

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Few-Shot Class-Incremental Learning, compute and memory, Few-Shot Class-Incremental, requires models, models to continuously

备注:

点击查看摘要

Abstract:Few-Shot Class-Incremental Learning (FSCIL) requires models to continuously learn new classes from limited samples while retaining prior knowledge, under strict constraints on compute and memory. Existing approaches lie along a difficult trade-off: simple fine-tuning is computationally efficient but suffers from catastrophic forgetting, replay-based methods mitigate forgetting at the cost of substantial compute and memory, and exemplar-free methods often reduce forgetting by freezing most of the backbone, improving efficiency at the expense of adaptability. We propose Selective Backpropagation (SBP), a deterministic parameter budgeting framework that bridges this gap. SBP restricts gradient updates to a pre-allocated subset of network parameters, freezing past knowledge and preserving unbiased capacity for future learning, enabling rapid adaptation without costly mask optimization. We show that SBP achieves strong performance across standard FSCIL benchmarks while requiring training time close to that of naive fine-tuning and substantially lower training time than prior SOTA methods. Crucially, our experiments expose a limitation of standard FSCIL evaluation: performance on short, distribution-consistent benchmarks does not necessarily predict behavior under distribution shift or over substantially longer learning horizons. We therefore evaluate FSCIL methods in cross-domain settings and over an 80-session ImageNet-1K stream. SBP remains strong across these regimes while maintaining low training cost, providing a favorable stability-plasticity-efficiency trade-off. Our code is available at this https URL.

263. 【2610.03991】Masked Privileged-Information Distillation for Multimodal Skin Lesion Classification Under Missing Clinical Metadata

链接:https://arxiv.org/abs/2610.03991

作者:Anirban Barua,Md Mahir Abrar Khan,Ayman Iktidar,Md. Sajjatul Islam

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:skin lesion classification, lesion classification combines, Multimodal skin lesion, classification combines clinical, skin lesion

备注:

点击查看摘要

Abstract:Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings additionally require computational efficiency. We address these challenges with a privileged-information distillation framework in which a multimodal teacher trained on complete metadata supervises a 9.2x smaller student trained with randomly masked clinical fields. Clinical fields are masked as whole groups at a per-sample rate drawn from U(0,1), so one training run covers the full metadata availability range. Fusion is residual, with metadata added as a gated correction to an unconditional image base. On the PAD-UFES-20 dataset, distillation under masked training improves balanced accuracy over cross-entropy training at every availability level, by an average of +4.7 points versus +1.8 points without masking. The masked student loses only 7.8 balanced-accuracy points as metadata decreases from complete to absent, compared with 36.6 points for the same student trained on complete metadata, highlighting the role of masked training in graceful degradation beyond distillation alone. Grad-CAM visualizations further show that the masked student's attention generally remains lesion-centered as metadata is withdrawn. The resulting compact model targets point-of-care settings, where clinical metadata is often incomplete.

264. 【2610.03980】FADE: Frame-Aware Diffusion-Transformer-based Multi-Concept Erasure for Video Unlearning

链接:https://arxiv.org/abs/2610.03980

作者:Yuchen Li,Kaiyuan Deng,Chaoran Feng,Zhenyu Tang,Li Yuan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:removing designated concepts, motivates concept erasure, reproduce copyrighted, explicit content, removing designated

备注: 29 pages. Code: [this https URL](https://github.com/INTOTHEMILD/FADE)

点击查看摘要

Abstract:Text-to-video (T2V) diffusion models can reproduce copyrighted, violent, or explicit content, which motivates concept erasure: removing designated concepts from a pretrained model while preserving its behavior on everything else. Existing T2V erasure methods leave two problems open. Their frame-agnostic suppression can leave isolated frames in which an erased concept resurfaces, a frame-reactivation gap that clip-level averages obscure; and they are usually evaluated with one target concept or category at a time. We propose Frame-Aware Diffusion Erasure (FADE), a multi-concept video unlearning framework. FADE first applies a joint closed-form key/value edit that suppresses all target concepts, then trains per-concept frame-aware low-rank adapters whose strength is gated by the frame index and the denoising timestep to remove residual per-frame leakage. Each adapter is trained with the other targets' prompts as hard negatives, which keeps the concept-specific components of different adapters well separated, and a similarity-based soft router combines the adapters according to the prompt. With 16 concepts (objects, artistic styles, and nudity) erased from a single Wan2.1-T2V-1.3B backbone, FADE reduces the residual accuracy on the object benchmark to 4.9%, against 15.5% for the strongest of eight baselines, while keeping the VBench average within 0.9% of the unedited model. The ranking is unchanged under a VLM judge and a blinded human study, and the advantage over the strongest baseline carries over to prompts that combine several erased concepts, to 30 simultaneously erased celebrity identities, and to Wan2.1-T2V-14B, CogVideoX-2B, and HunyuanVideo-1.5.

265. 【2610.03942】Dynamic Time Step Prediction in Inverse Heat Dissipation for Blur-Like Image Restoration Tasks

链接:https://arxiv.org/abs/2610.03942

作者:Cap Dang Xuan Kiet,Tat-Jen Cham

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:retain relevant image, time step, relevant image information, starting time step, pre-determined time step

备注:

点击查看摘要

Abstract:When using diffusion models to target image restoration problems, diffusion inversion is typically employed to retain relevant image information from the degraded images. Instead of inverting back to the initial time step (i.e., T), many methods invert to a pre-determined intermediate time step, in order to better preserve information from degraded source images. However, a pre-determined time step for inversion is not ideal for reconstruction, as a severely degraded image requires an earlier starting time step than a mildly degraded one. In addition, DDIM-based models corrupt the original signal by adding Gaussian noise, which can be mismatched to the nature of blur-like degradations, such as blur, haze, and low-light. To address these problems, we propose two solutions: (1) we adopt an alternative diffusion process, called the Inverse Heat Dissipation Model, that diffuses the input image by gradually blurring a data point (2) we propose to implement a time predictor to estimate the starting time step for the inversion, with the model learning to adapt to the degradation severity. Extensive experiments on standard benchmarks show that our method achieves state-of-the-art performance in both quantitative and qualitative evaluations, with excellent generalization to many restoration tasks.

266. 【2610.03928】DABACO: A Multi-Camera Dataset and Benchmark for Screen Localization and Pointing Estimation

链接:https://arxiv.org/abs/2610.03928

作者:Óscar Gómez-Cárdenes,José Gil Marichal-Hernández,Juan Manuel Martín-Doñas

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:low-cost interactive devices, Low-Cost Pointing Device, estimation are key, interactive devices, Screen localization

备注: 21 pages, 5 figures, 6 tables. Dataset: [this https URL](https://doi.org/10.5281/zenodo.22797835) . Code and toolkit: [this https URL](https://domondo.github.io/dabaco-dataset-toolkit/)

点击查看摘要

Abstract:Screen localization and pointing estimation are key to low-cost interactive devices. Yet developing and evaluating these algorithms requires realistic data: synthetic captures cannot fully reproduce the optical distortion, rolling shutter, motion blur, and display processing of a physical acquisition, and most existing datasets provide static images rather than the video needed to assess continuous pointing. In this paper, we introduce the DABACO Dataset. Developed within the DABACO (Dispositivo Apuntador de BAjo COste, or Low-Cost Pointing Device) project, this dataset supports the development and evaluation of screen detection and camera-based pointing algorithms for embedded systems. It comprises video sequences captured with multiple low-cost camera sensors, including monochrome global-shutter and color rolling-shutter modules, on embedded platforms such as the Raspberry Pi 4B and ESP32-S3. We present a systematic annotation pipeline combining temporary visual watermarking, optical-flow tracking, manual verification, and marker removal. Both the original marked captures and the marker-free images, together with explicit corner annotations and modification masks, are released to support auditing and the study of potential reconstruction bias. In addition, we release an open-source evaluation toolkit with two reference baselines: a classical screen-detection pipeline based on edge and contour geometry, and a closed-vocabulary, off-the-shelf YOLOv8 segmentation model evaluated without dataset-specific training. Both are evaluated on a general-purpose computer using Intersection over Union, corner error, and pointing error, establishing initial reference results for future embedded implementations. Overall, DABACO addresses a gap in existing resources and is intended to support the development and evaluation of new low-cost screen-localization and pointing systems.

Comments:
21 pages, 5 figures, 6 tables. Dataset: this https URL. Code and toolkit: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

Cite as:
arXiv:2610.03928 [cs.CV]

(or
arXiv:2610.03928v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.03928

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
267. 【2610.03873】Streaming Multi-Track Timeline Control for 3D Human Motion Generation

链接:https://arxiv.org/abs/2610.03873

作者:Yangsong Zhang,Anujith Muraleedharan,Rikhat Akizhanov,Gül Varol,Fabio Pizzati,Ivan Laptev

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Text-driven human motion, methods assume instructions, advanced substantially, Text-driven human, methods assume

备注: Project page: [this https URL](https://mael-zys.github.io/TimelineControl/)

点击查看摘要

Abstract:Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation or simultaneous composition without explicitly combining streaming instruction arrival with independently timed, overlapping actions. We introduce streaming multi-track timeline control and propose TimelineControl to incorporate new instructions alongside ongoing actions. Interval-aware conditioning preserves instruction timing, while causal part-structured representations and part-aware denoising coordinate concurrent actions across body regions. We also construct TimelineMotion, a dataset with overlapping instruction intervals and body-part annotations. Experiments on TimelineMotion and MTT demonstrate improved semantic alignment and temporal adherence over evaluated streaming baselines, including models retrained on the same data. Ablations and human evaluations validate our design, complemented by spatial conditioning and humanoid execution demonstrations. Our code, data and models will become publicly available.

268. 【2610.03861】GOTT: Object-centric Dexterous Manipulation with a Reusable Cross-Embodiment Primitive

链接:https://arxiv.org/abs/2610.03861

作者:Yulin Liu,Lai Wei,Yen-Jen Wang,Akash Sharma,Pieter Abbeel,Henrik I. Christensen,Haozhi Qi

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:data provide rich, provide rich sources, multi-fingered robot behavior, behavior remains difficult, large-scale human data

备注: Project page: [this https URL](https://dex-gott.github.io/)

点击查看摘要

Abstract:Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes contact across tasks and embodiments. We propose GOTT, a reach-acquire-move framework built around a single cross-embodiment contact-acquisition primitive. Given a robot-agnostic object trajectory and a reach specification, GOTT first brings the hand near a task-relevant contact region. The shared closed-loop primitive then establishes stable contact from this approximate initialization, and a pose-conditioned controller tracks the desired object motion. Reach specifications may come from future-aware planning, external models, or human demonstrations, while the primitive and tracking backend remain unchanged. Simulation and real-world experiments show that GOTT is able to establish robust contact across diverse objects, arm-hand platforms, and seen and unseen hand morphologies. It also consistently improves end-to-end task success over open-loop grasp execution.

269. 【2610.03826】Are We Measuring Anticipation? Auditing Privileged Information in Procedural Video Evaluation

链接:https://arxiv.org/abs/2610.03826

作者:Mahsa Mohammadi,Sareh Rowlands

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Benchmark scores license, scores license claims, capabilities being evaluated, Benchmark scores, license claims

备注: 17 pages, 6 figures. NeurIPS 2026 Workshop on TAE (Trust-AI-Eval): Can We Trust AI Evaluation?

点击查看摘要

Abstract:Benchmark scores license claims about the capabilities being evaluated. We audit the inference licensed by an evaluation protocol, rather than the predictive model alone. Using procedural action anticipation as a controlled case study, we study a broader evaluation-validity failure mode: a protocol can remain temporally causal and free of classical target leakage while still supplying a privileged intermediate representation. The failure is not merely optimistic accuracy, but a mismatch between the capability claimed and the construct actually measured. On Breakfast, matched recognizer-history and GT-history regimes score 30.3% and 62.3% (Delta_PH(R_BA) = +32.0, 95% Student-t interval [+24.1, +39.8]). A fixed-checkpoint 2 x 2 intervention isolates a +14.6-point test-time oracle contrast; a matched no-video provenance probe yields a +26.8-point gap; and a boundary-independent query/history control retains a +19.4-point gap. We refer to this discrepancy as a privileged-information gap, defined relative to a specified non-oracle recovery pipeline, and propose a reusable four-condition audit. Across five evaluation settings the gap is heterogeneous; a standard 50 Salads long-term-anticipation re-implementation provides cautious support beyond the audit-specific next-action construction.

Comments:
17 pages, 6 figures. NeurIPS 2026 Workshop on TAE (Trust-AI-Eval): Can We Trust AI Evaluation?

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.03826 [cs.CV]

(or
arXiv:2610.03826v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.03826

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
270. 【2610.03822】Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction

链接:https://arxiv.org/abs/2610.03822

作者:Zhicheng Liu,Zhouxiang Zhao,Chenliang Wu,Zhaohui Yang,Zhaoyang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:efficient image delivery, multimodal foundation models, unified interface, interface for multimodal, multimodal foundation

备注:

点击查看摘要

Abstract:Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily translate into improved image reconstruction quality. In this paper, we propose AdapToC, an adaptive, reconstruction-oriented visual-token communication framework. At the transmitter, an adaptive selector jointly models image content, channel state, and communication budget to perform instance-wise resource allocation. Rather than using a fixed token rate and protection policy, it dynamically determines how many tokens should be transmitted and assigns different protection levels according to token importance and current channel conditions. At the receiver, an adaptive MaskGIT receiver incorporates channel reliability into contextual token modeling. It distinguishes tokens with different reliability levels, preserves high-confidence observations, corrects potentially corrupted tokens, and iteratively reconstructs missing content from the received evidence and global visual context. By co-designing token quantity, unequal protection, and reliability-aware recovery for image-level reconstruction quality, AdapToC achieves a peak mean PSNR gain of 4.20 dB over the strongest static baseline under matched communication costs and state-of-the-art performance among the evaluated visual-token communication methods.

271. 【2610.03812】Least Squares for Time Series Forecasting

链接:https://arxiv.org/abs/2610.03812

作者:Weiu-qiou Ciang,Yuzhou Hong,Sherry Chen

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:time-series forecast, forecast program, forecast, forecast program puts, forecast program minimizes

备注: 20 pages, 2 figures

点击查看摘要

Abstract:A time-series forecast is scored on a future value of the series. A representation loss that regresses the next latent, as in LeNEPA, is a different least-squares problem on the same bottleneck. We write both programs down. The forecast program minimizes the error of a decoded latent on the coordinate that will be reported. For a scalar target and a linear decoder, every latent rank of at least one matches ordinary least squares, and an isotropy constraint is only a rescaling: after the decoder is refit, the forecast does not move. The other program fits the whole next vector at a fixed rank, then freezes the encoder and attaches a head. On a four-dimensional series whose last three coordinates are the same autoregression, that rank-1 fit puts mass $0.9998$ on the repeated coordinate and forecasts the remaining signal at the marginal variance $2.794$. The forecast program puts mass $1$ on the signal and matches the innovation variance $0.992$. Rank $2$ gives the vector fit a second direction, and the two programs agree. Iterating the fitted one-step coefficient $0.803$ raises the open-loop error from $0.992$ at one step to $2.700$ at eight steps.

272. 【2610.03799】StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search

链接:https://arxiv.org/abs/2610.03799

作者:Ghadi Nehme,Faez Ahmed

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:discrete modeling choices, Recovering executable CAD, Recovering executable, meshes is challenging, continuous parameters

备注:

点击查看摘要

Abstract:Recovering executable CAD programs from 3D meshes is challenging due to the compositional nature of CAD construction and the interaction between discrete modeling choices and continuous parameters. Many learning-based methods predict complete programs in a single pass and rely predominantly on sketch-extrude representations, limiting operation diversity and opportunities to correct geometric errors during reconstruction. We introduce StepCAD, a generative optimization approach that combines a state-conditioned CAD policy with geometry-guided search. Given an input mesh, the policy predicts construction actions conditioned on both target and intermediate geometry, and an IoU-guided tree search refines the resulting program through local edits. We also introduce ARCADE-1.5M, a large-scale dataset of 1.5M executable CAD programs spanning diverse operations, sequences with a maximum length of 150+ counted operations, and 12.5M intermediate state-action transitions. Experiments across multiple CAD reconstruction benchmarks show that StepCAD achieves state-of-the-art geometric reconstruction accuracy with consistently high validity, yielding up to 87.2% relative IoU improvement over the strongest evaluated baseline, with particularly large gains on complex shapes. Project page: this https URL

273. 【2610.03797】WAMJET: A Harness for World Action Model Acceleration

链接:https://arxiv.org/abs/2610.03797

作者:Le Chen,Lixin Liu,Jan Schneider,Zeju Qiu,Simon Guist,Bernhard Schölkopf,Dieter Büchler

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:World Action Models, leverage pretrained video, pretrained video foundation, video foundation models, World Action

备注: 8 pages, 3 figures, project page: [this https URL](https://github.com/liulixinkerry/WAMJET/blob/main/assets/blog.md)

点击查看摘要

Abstract:World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95x lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.

274. 【2610.03792】What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study

链接:https://arxiv.org/abs/2610.03792

作者:Avyay Sadhu,Patrick Cooper

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reinforcement learning, produced large reasoning, make it applicable, applicable to causal-temporal, RLVR teaches video-language

备注: 8 pages, 3 figures. Submitted to WACV 2027

点击查看摘要

Abstract:Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about time. We fine-tune four open models (Qwen3-VL-8B/4B, Qwen2.5-VL-7B, Gemma-3-12B) with group relative policy optimization under three data recipes: verified (synthetic CLEVRER questions with exact answer and event-order rewards), unverified (self-supervised pretext tasks over 43,751 real web videos), and a 1:1 mixture, plus a verified+real arm that adds 4,000 verifiable questions on real video. Each cell is evaluated in-domain and on out-of-domain real video (a NExT-QA temporal stress set and an MVBench subset), with frames in order, shuffled, and absent. (1) Verified training yields large in-domain gains that shrink as base competence grows (+14 to +19 points on weaker models; +6 on the strongest). (2) Much of the gain is non-visual: accuracy with no frames rises nearly as much as with frames. (3) Verified-only training can severely degrade out-of-domain accuracy with no sign during training: Qwen3-VL-8B loses 26.7 and 25.2 points on the two real-video sets, while the mixture never significantly degrades a model trained on it. Adding real verified questions removes that loss (-2.3 points, within noise of base) and keeps a +9.3 in-domain gain, so the cause is narrow synthetic-only data, not verification. (4) No recipe induces temporal-order grounding: across 41 evaluations the ordered-versus-shuffled gap is indistinguishable from zero in 39 and marginal in two, despite an event-order reward. Verifiable rewards improve benchmark accuracy without temporal understanding. Report no-frame controls, and mix in real video to guard against out-of-domain degradation.

275. 【2610.03772】Energy Variation in Training Modern Computer Vision Architectures

链接:https://arxiv.org/abs/2610.03772

作者:David Cortes,Carlos Juiz,Belen Bermejo

类目:Computer Vision and Pattern Recognition (cs.CV); Performance (cs.PF)

关键词:relevant design criterion, increasingly relevant design, Computational Biology Center, making energy efficiency, design criterion

备注: International Conference on Next-Generation AI Technologies (ICNGAIT) Shanghai, China, 5 pages, 1 table

点击查看摘要

Abstract:The rapid growth of deep learning has substantially increased the energy consumption associated with model training, making energy efficiency an increasingly relevant design criterion. This study empirically measures the energy variation of training seven modern computer vision architectures, MobileNetV3-Small, MobileNetV3-Large, EfficientNet-B0, EfficientNet-B1, ViT-B/32, ConvNeXt-Tiny, and ViT-B/16 for the ImageNet-1k classification task, using a homogeneous 10,000-image subset (ImageNet-10k) and a uniform 40-epoch baseline configuration executed on two NVIDIA Tesla P100 GPUs at the Bioinformatics and Computational Biology Center of Colombia (BIOS). Energy was recorded directly via NVML and contrasted with the computational complexity of each model. The results show a Pearson correlation of 0.85 between floating-point operations (GFLOPs) and energy consumption in kWh, indicating that computational complexity is a strong but imperfect predictor of energy expenditure: architectures with comparable GFLOPs exhibited consumption differing by up to 3.1x due to differences in the hardware efficiency of their dominant operations. The EfficientNet variants offered the best balance between classification performance (Val Top-5 up to 97.15%) and energy efficiency (0.457-0.611 kWh), while Vision Transformers exhibited the highest relative energy consumption and lower classification performance under the evaluated configuration. These findings guide architecture selection in energy-constrained computing environments.

276. 【2610.03771】LoRA Direction Extraction for Controllable Light Toggling in FLUX.1 Kontext

链接:https://arxiv.org/abs/2610.03771

作者:Petr Golenderov,Dmitry Mazyar,Natalia Sovpel,Alexander Aksenov

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:large training dataset, realistic artificial light, artificial light modeling, diffusion models aimed, flow-matching diffusion models

备注:

点击查看摘要

Abstract:We propose a fine-tuning method for flow-matching diffusion models aimed at realistic artificial light modeling without the need for a large training dataset. We address the task of controllable interior image editing, where the goal is to turn artificial light sources on or off while preserving the scene geometry, object placement, materials, and visual identity of the original image. To achieve this, we decompose the task into two independent formulations. We introduce the LoRA Direction Training Method, which extracts the pure direction of the LoRA adapter effect in the diffusion model flow field, and we also introduce specialized loss functions to ensure the realism of the inverse transformation. Additionally, the resulting increment map is used for more precise adjustment of the lighting color and temperature.

277. 【2610.03765】Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

链接:https://arxiv.org/abs/2610.03765

作者:Varsha Suresh,Divij Jain,Jia Liu,M. Hamza Mughal,Vera Demberg

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:co-speech gesture generation, atomic units, gesture generation, motion tokenizers encode, tokenizers encode motion

备注: Accepted at MINT EMNLP 2026

点击查看摘要

Abstract:Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descriptors spanning from raw motion to abstract communicative function. Results show that geometry and handedness are readily decodable from token embeddings, while motion category is only weakly decoded despite showing systematic differences in discrete code usage. This gap between reconstruction quality and descriptor decodability suggests that reconstruction objectives alone do not guarantee that gesture semantics are captured, and that evaluating codebooks on such properties can guide the design of more semantic motion tokenizers.

278. 【2610.03755】Fractal Cross Product: Theory, Differentiable Implementation and Application to Medical Image Analysis

链接:https://arxiv.org/abs/2610.03755

作者:Noaman Khan,Nihad Hadj Sahraoui,Samir Brahim Belhaouari

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Euclidean cross product, generalized Euclidean cross, Fractal Cross Product, generalized cross product, generalized Fractal Cross

备注:

点击查看摘要

Abstract:The magnitude of the generalized Euclidean cross product is a Gram volume whose degree under common scaling is fixed by the integer dimension of the spanning frame. We formulate a generalized Fractal Cross Product (FCP) as a nonlinear radial deformation with a prescribed positive degree $D$, which may be non-integer. The scalar construction applies in any ambient dimension $m\geq k$, while its canonically oriented vector form requires codimension one. It recovers the classical generalized cross product exactly at $D=k$ and retains orthogonality, alternation, rotation equivariance, and $D$-homogeneity, but is generally not multilinear. For exact self-similar frame systems, the construction also obeys a scale-balance law at the similarity dimension. We derive a differentiable, dimensionless image response and a non-circular empirical accumulation exponent obtained by regressing raw angular Gram responses across patch widths. Binary64 calculations recover the finite-frame identities to roundoff, while raster experiments recover the Sierpiński-triangle value 1.5849625 at three resolutions. In five-seed medical-imaging comparisons, FCP-centered fusion increased mean area under the receiver operating characteristic curve from 0.7264 to 0.8135 and from 0.5843 to 0.6765 on the random and hospital-separated Retinal Image Database for Optic Nerve Evaluation partitions, respectively, and from 0.7403 to 0.7475 on FracAtlas. On FracAtlas, balanced accuracy increased from 0.6186 to 0.6721 and the harmonic mean of precision and sensitivity from 0.3669 to 0.4457. These results support the utility of the complete fusion framework, but do not isolate the effect of FCP from that of its complementary descriptors and fusion head.

279. 【2610.03720】Beyond the Good, the Bad, and the Ugly: Colormap Assessment through Data-Aware Perceptual Metric

链接:https://arxiv.org/abs/2610.03720

作者:Xi Duan,Yiwei Lin,Shiqing Xin,Aoying Wang,Yucheng Wang,Changhe Tu,Qiong Zeng

类目:Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:visualize scalar fields, Continuous colormaps, colormap, scalar fields, visualize scalar

备注: 11 pages, 9 figures, IEEE VIS 2026 conference full paper

点击查看摘要

Abstract:Continuous colormaps are widely used to visualize scalar fields, and their quality is typically evaluated using measures such as discriminative power and uniformity. Existing measures primarily characterize the intrinsic perceptual properties of the colormap itself, largely independent of the underlying data distribution. In practice, however, user perception arises not from the colormap alone but from the visualization generated by mapping data values through the colormap. The perceptual differences that users actually experience depend jointly on the colormap and the underlying data. We propose a data-aware formulation that complements existing data-independent colormap assessment approaches. Rather than analyzing the colormap in isolation, we model the color-encoded visualization as a composite mapping from the spatial domain of the data to perceptual color space. This approach yields data-aware counterparts of established measures, including discriminative power, uniformity, and smoothness, as well as additional measures such as perceptual anisotropy and degeneracy. We validate the proposed measures against both the existing data-independent framework and empirical results from perceptual studies, extend the formulation to 2D colormaps, and demonstrate its integration into colormap optimization. An interactive system is provided for exploring colormap assessment under varying data distributions. Results demonstrate that our formulation offers a principled foundation for data-aware colormap assessment and design.

280. 【2609.38123】HelixWorld: A Real-time Interactive Audio-Visual World Model

链接:https://arxiv.org/abs/2609.38123

作者:Lei Ke,Jiahao Pan,Zeyue Tian,Jiaming Wang,Haoyuan Huang,Kam Man Wu,Pengjun Fang,Hongyu Liu,Chenyang Qi,Lin Wang,Ruibin Yuan,Weijia Chen,Fangneng Zhan,Qifeng Chen,Wei Xue,Yike Guo

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)

关键词:demanding synchronized visual, inherently multisensory, demanding synchronized, real time, simulation is inherently

备注:

点击查看摘要

Abstract:World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.

281. 【2407.21783】he Llama 3 Herd of Models

链接:https://arxiv.org/abs/2407.21783

作者:Aaron Grattafiori,Abhimanyu Dubey,Abhinav Jauhri,Abhinav Pandey,Abhishek Kadian,Ahmad Al-Dahle,Aiesha Letman,Akhil Mathur,Alan Schelten,Alex Vaughan,Amy Yang,Angela Fan,Anirudh Goyal,Anthony Hartshorn,Aobo Yang,Archi Mitra,Archie Sravankumar,Artem Korenev,Arthur Hinsvark,Arun Rao,Aston Zhang,Aurelien Rodriguez,Austen Gregerson,Ava Spataru,Baptiste Roziere,Bethany Biron,Binh Tang,Bobbie Chern,Charlotte Caucheteux,Chaya Nayak,Chloe Bi,Chris Marra,Chris McConnell,Christian Keller,Christophe Touret,Chunyang Wu,Corinne Wong,Cristian Canton Ferrer,Cyrus Nikolaidis,Damien Allonsius,Daniel Song,Danielle Pintz,Danny Livshits,Danny Wyatt,David Esiobu,Dhruv Choudhary,Dhruv Mahajan,Diego Garcia-Olano,Diego Perino,Dieuwke Hupkes,Egor Lakomkin,Ehab AlBadawy,Elina Lobanova,Emily Dinan,Eric Michael Smith,Filip Radenovic,Francisco Guzmán,Frank Zhang,Gabriel Synnaeve,Gabrielle Lee,Georgia Lewis Anderson,Govind Thattai,Graeme Nail,Gregoire Mialon,Guan Pang,Guillem Cucurell,Hailey Nguyen,Hannah Korevaar,Hu Xu,Hugo Touvron,Iliyan Zarov,Imanol Arrieta Ibarra,Isabel Kloumann,Ishan Misra,Ivan Evtimov,Jack Zhang,Jade Copet,Jaewon Lee,Jan Geffert,Jana Vranes,Jason Park,Jay Mahadeokar,Jeet Shah,Jelmer van der Linde,Jennifer Billock,Jenny Hong,Jenya Lee,Jeremy Fu,Jianfeng Chi,Jianyu Huang,Jiawen Liu,Jie Wang,Jiecao Yu,Joanna Bitton,Joe Spisak,Jongsoo Park,Joseph Rocca,Joshua Johnstun,Joshua Saxe,Junteng Jia

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:Modern artificial intelligence, Modern artificial, artificial intelligence, systems are powered, Llama

备注:

点击查看摘要

Abstract:Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.

282. 【2103.16559】Broaden Your Views for Self-Supervised Video Learning

链接:https://arxiv.org/abs/2103.16559

作者:Adrià Recasens,Pauline Luc,Jean-Baptiste Alayrac,Luyu Wang,Ross Hemsley,Florian Strub,Corentin Tallec,Mateusz Malinowski,Viorica Patraucean,Florent Altché,Michal Valko,Jean-Bastien Grill,Aäron van den Oord,Andrew Zisserman

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:trained to align, successful self-supervised learning, video, self-supervised learning methods, independent views

备注: This paper is an extended version of our ICCV-21 paper. It includes more results as well as a minor architectural variation which improves results

点击查看摘要

Abstract:Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views are similarly extracted by cropping and augmenting the resulting crop. However, these methods miss a crucial element in the video domain: time. We introduce BraVe, a self-supervised learning framework for video. In BraVe, one of the views has access to a narrow temporal window of the video while the other view has a broad access to the video content. Our models learn to generalise from the narrow view to the general content of the video. Furthermore, BraVe processes the views with different backbones, enabling the use of alternative augmentations or modalities into the broad view such as optical flow, randomly convolved RGB frames, audio or their combinations. We demonstrate that BraVe achieves state-of-the-art results in self-supervised representation learning on standard video and audio classification benchmarks including UCF101, HMDB51, Kinetics, ESC-50 and AudioSet.

283. 【2610.06034】How well do routinely collected demographic and clinical variables aid point-of-care lung ultrasound TB classification

链接:https://arxiv.org/abs/2610.06034

作者:Joshua M. Jansen van Vüren,Christiaan M. Geldenhuys,Devendra S. Parihar,Véronique Suttels,Trevor Brokowski,Ablo P. Wachinou,Mary-Anne Hartley,Rensu P. Theart,Grant Theron,Thomas R. Niesler

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:lung ultrasound images, demographic data, clinical and demographic, lung ultrasound, automated tuberculosis

备注: Accepted: SATNAC, Drakensberg, South Africa, 2026

点击查看摘要

Abstract:We consider the fusion of lung ultrasound images with routinely-collected clinical and demographic data for the purpose of automated tuberculosis (TB) screening using deep-learning. Such deep-learning based screening tools for TB could meaningfully support the health care system in Africa, where the burden of disease is severe and resources are constrained. Beginning with an established ResNet baseline for classification of lung ultrasound images, which achieves an area under the receiver operating characteristic (AUROC) curve of 0.91 [0.86,0.96] (95% CI), we consider the incorporation of the clinical and demographic data using three fusion approaches. We find that a simple average-based fusion of the output scores of separately-trained image and clinical data classifiers consistently matches or outperforms a more complex approach where the data is fused earlier and a combined classifier is trained. Fusing the image and the clinical classifiers in this way leads to a classifier with an overall AUROC of 0.95 [0.91,0.99] (specificity of 0.76 at sensitivity 0.93) which is an improvement of 4% absolute over the image-only baseline. We also find that greedy feature selection can be used to reduce the number of clinical and demographic inputs without sacrificing classification performance. Finally, when we differentiate between clinical and demographic data that are self-reported, that require some basic measurement or calculation, and that require a point-of-care (POC) test, we find the inclusion of the POC tests included in this study to be of minimal benefit to classification performance. We conclude that the incorporation of routinely-collected clinical and demographic data is a promising way to improve the performance of lung ultrasound based automatic classification.

284. 【2610.05939】Structural Foundations of Nonlinear Systems with Unknown Inputs: The UID-Induced Normal Form and Minimal-Sensing Structure-from-Motion

链接:https://arxiv.org/abs/2610.05939

作者:Agostino Martinelli

类目:Optimization and Control (math.OC); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:nonlinear systems driven, general structural solution, UID-induced normal form, paper establishes, unknown inputs

备注:

点击查看摘要

Abstract:This paper establishes the first general structural solution to the problem of state estimation for nonlinear systems driven by unknown inputs. Building upon nonlinear unknown-input observability theory, we show that every such system admits a structurally equivalent representation, referred to as the UID-induced normal form. The proposed representation decomposes the information carried by the unknown inputs into two complementary components: unknown-input directions that are structurally decoupled from the observable dynamics and observable quantities that completely represent the unknown-input information affecting the observable dynamics. As a consequence, the UID-induced normal form provides a unified structural solution to unknown-input decoupling and unknown-input reconstruction, without requiring any model or stochastic assumption on the unknown inputs. The practical significance of the proposed framework is demonstrated through a previously unexplored minimal Structure-from-Motion configuration. The proposed representation enables recursive state estimation from only three point features and a single-axis gyroscope, allowing the recovery of the three-dimensional structure and camera motion up to an unknown global scale factor. Experiments on real-world data validate the proposed framework and demonstrate the feasibility of this minimal sensing configuration.

285. 【2610.05182】IRMamba: A Thermal-Prior-Modulated State-Space Network for Sub-Million-Parameter Infrared Image Super-Resolution

链接:https://arxiv.org/abs/2610.05182

作者:Chun-An Lin,Tsung-Jung Liu,Yen-Chieh Ouyang

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:led by Mamba-based, Mamba-based networks, Infrared image super-resolution, Thermal Prior Highway, difficult to deploy

备注: 13 pages, 4 figures, 5 tables. Submitted to IEEE Transactions on Geoscience and Remote Sensing. Code and trained models will be released at [this https URL](https://github.com/julian135707/TIRMamba) upon acceptance

点击查看摘要

Abstract:Infrared image super-resolution is currently led by Mamba-based networks with 26 to 37 million parameters, which are difficult to deploy on the airborne and handheld platforms where thermal imaging is most needed. This paper presents TIRMamba, a network with 896K to 910K parameters for single-channel thermal imagery. A Thermal Prior Highway computes gradient, local-contrast and spectral cues once at the input and, through one adapter per residual group, modulates a weight-tied bidirectional state-space trunk and gates its dual-scale detail branch; a tri-path reconstruction adds the learned residual to a bicubic radiometric baseline. Because the standard benchmark provides only 265 infrared training images and evaluates fusion products on full images, we train with a replay strategy: grayscale DIV2K pre-training followed by fine-tuning on 64-pixel patches drawn with equal probability from the infrared and natural corpora. At scale factor 4, TIRMamba matches the strongest protocol-trained methods on both official test sets with 29 to 40 times fewer parameters and 2.8 to 9.4 times lower latency; at scale factor 2 it gives the highest SSIM on both. A variant with prior-conditioned selectivity, TIRMamba-Rad, corrects a 3 dB raw-thermal failure of an intermediate size-invariant design and gives the best results at scale factor 4 on raw-thermal, unmanned-aerial-vehicle and independent-sensor test sets. Code and trained models will be released at this https URL upon acceptance.

286. 【2610.04553】Cross-Modal Solar Image Synthesis: Adapting the Surya Foundation Model from He I 10830 Å to EUV Translation and Coronal Hole Segmentation

链接:https://arxiv.org/abs/2610.04553

作者:Marco Marena,Andrés Muñoz Jaramillo,Qin Li,Haodi Jiang,Jinghao Cao,Wen He,Ziyang Zhang,Chenxi Yuan,Chao Wang,Haimin Wang,Bo Shen

类目:olar and Stellar Astrophysics (astro-ph.SR); Instrumentation and Methods for Astrophysics (astro-ph.IM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:long observational record, Atmospheric Imaging Assembly, investigate solar morphology, Solar Dynamics Observatory, Possibilistic Clustering Algorithm

备注:

点击查看摘要

Abstract:The long observational record of He I 10830 Å offers a means to investigate solar morphology before modern extreme-ultraviolet (EUV) imaging. We adapt the Surya solar foundation model to predict Solar Dynamics Observatory/Atmospheric Imaging Assembly (SDO/AIA) 94, 193, and 304 Å images and a coronal hole (CH) probability map from full-disk helium observations. A convolutional input adapter, low-rank backbone updates, and dedicated output decoders learn from temporally paired, geometrically registered observations, with Spatial Possibilistic Clustering Algorithm (SPOCA) catalog polygons providing CH supervision. On observations held out from downstream fine-tuning, the selected dedicated models achieve disk-restricted correlation coefficients (CCs) of 0.8196, 0.8885, and 0.8398 for the three AIA channels, respectively; the CH model achieves an intersection over union (IoU) of 0.4360. The predictions recover broad solar structure, although local agreement varies substantially by channel. An optional residual refiner addresses spatial detail, and its application on pre-SDO dates improves the correlation of synthetic AIA 304 with Solar and Heliospheric Observatory/Extreme-ultraviolet Imaging Telescope (SOHO/EIT) 304 references from 0.6511 to 0.7085. Together, these results support the feasibility of helium-conditioned EUV morphological proxies and motivate their use in historical reconstruction, within the scope of the downstream test and cross-instrument evaluation.

287. 【2610.04258】FloVMos: Optical Flow-based Medical Video Mosaicking

链接:https://arxiv.org/abs/2610.04258

作者:Jinyang Liu,Sandesh Ghimire,Chaman Singh,Jennifer Dy,Milind Rajadhyaksha,Dana H. Brooks,Octavia Camps,Kivanc Kose

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:trade-off among resolution, require a trade-off, Biomedical imaging modalities, Biomedical imaging, imaging modalities

备注:

点击查看摘要

Abstract:Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV composite. However, existing methods struggle with non-rigid deformations, and modality-specific artifacts arising in clinical and research imaging. Here, we present FloVMos, a generalizable, optical-flow-based deep learning framework for real-time video mosaicking across diverse biomedical imaging modalities. FloVMos achieves robust, pixel-level registration by fine-tuning an optical flow model on synthetic training data with ground-truth deformation fields. We introduce a pipeline for generating this training data, simulating realistic tissue motion and imaging distortions from existing mosaics or raw videos. Our automated synthetic data generation and optical flow model training based on this data allow users to adapt FloVMos to different imaging modalities. To demonstrate this, we applied FloVMos to seven diverse imaging modalities: reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy. FloVMos outperforms conventional baselines in accuracy, robustness, and speed for all the tested modalities. This adaptable and training-efficient framework enables large-area visualization with real-time performance and may support broader use of video-based biomedical imaging in research and clinical workflows.

288. 【2610.03817】Image-Based Breast Implant Detection for Mammography Dataset Curation and Near-Real-Time Deployment: Comparing Foundation Models and Task-Specific Convolutional Models

链接:https://arxiv.org/abs/2610.03817

作者:Vasisht Ishwar,Hari Trivedi,Young Seok Jeon,Beatrice Brown-Mulry,Frank Li,Rohan Satya Isaac,Mohammadreza Chavoshi,Judy Wawira Gichoya

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Tissues and Organs (q-bio.TO)

关键词:convolutional neural networks, breast implant classification, Breast Imaging Dataset, Emory Breast Imaging, task-specific convolutional neural

备注: 15 pages, 6 figures, 2 tables. Submitted to the Journal of Imaging Informatics in Medicine

点击查看摘要

Abstract:Purpose: To evaluate the performance-feasibility tradeoffs of foundation models (FMs) and task-specific convolutional neural networks (CNNs) trained from scratch for breast implant classification in 2D mammography, with emphasis on suitability for near real-time clinical deployment. Methods: We evaluated four models: two FMs (RAD-DINO and MammoCLIP) and two CNNs trained from scratch for implant prediction (ResNet18 and our lightweight ResNetLite). Using the Emory Breast Imaging Dataset, 5,000 unilateral screening mammograms were used for training/validation and 1,000 manually reviewed unilateral images were held out for testing. For the FMs, global image embeddings from the pretrained encoder were classified using a support vector machine (SVM). The CNNs were trained end-to-end on 2D mammograms, with ResNetLite optimized via grid search over depth and width to balance accuracy and efficiency. Performance was evaluated using AUROC, sensitivity, specificity, accuracy, embedding visualization, and inference-latency. Results: All models demonstrated strong performance on held-out test data (n = 1,000). MammoCLIP achieved the highest AUROC (0.999) with the quickest training time of 493 seconds. RAD-DINO achieved the highest sensitivity (0.980; accuracy 0.989) but had the slowest inference and training times. ResNet18 and MammoCLIP achieved comparable accuracy (0.985). ResNetLite showed no statistically significant difference from ResNet18 (AUROC 0.993; accuracy 0.976) despite using only 1.4% of ResNet18's parameters, and had the fastest inference time. Conclusion: FMs and task-specific CNN models reliably detect breast implants on 2D mammography. Model selection is best guided by deployment context: MammoCLIP for GPU-equipped hospital settings requiring scalable integration, and lightweight CNNs such as ResNetLite for resource-constrained or edge deployments.