本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新882篇论文,其中:

  • 自然语言处理114篇
  • 信息检索11篇
  • 计算机视觉130篇

自然语言处理

1. 【2610.03695】Language Models that Play Chess and Explain Their Moves

链接:https://arxiv.org/abs/2610.03695

作者:Adithya Bhaskar,Jeffrey Cheng,Danqi Chen

类目:Computation and Language (cs.CL)

关键词:Modern chess engines, Modern chess, Modern, superhuman level, play

备注: Code available at [this https URL](https://github.com/queen-project/queen)

点击查看摘要

Abstract:Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.

2. 【2610.03675】FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

链接:https://arxiv.org/abs/2610.03675

作者:Hui Chen,Xuan Qi,James Xu Zhao,Zhaopeng Feng,Shilong Liu,Kuang Xu,Pang Wei Koh,Bryan Hooi

类目:Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:computational optimization problems, challenging computational optimization, emerged as powerful, powerful approaches, approaches for challenging

备注: 17 pages, 4 figures

点击查看摘要

Abstract:LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the resulting code. We also design a cache-efficient evolution process, where our harness and prompts maximize the sharing of prefixes across different evolution steps, to improve cache reuse. To measure solution quality throughout a fixed cost budget, we introduce Budget-Aware Area Under the Curve (BA-AUC), defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget. Across 10 mathematical and systems optimization tasks, FrugalEvo matches or surpasses state-of-the-art baselines, including OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX, in final solution quality and achieves higher BA-AUC on 9 tasks. It also achieves higher average performance than these baselines on 10 algorithmic optimization tasks from ALE-Bench-Lite. Notably, on circle packing, FrugalEvo achieves new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD and with GLM-5.3 and its Flash variant for only 0.55 USD, matching or surpassing all baselines, including multi-agent methods such as CORAL and SwarmResearch, which cost approximately 50 USD on average.

3. 【2610.03665】Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

链接:https://arxiv.org/abs/2610.03665

作者:Seo Hyun Kim,Sunwoo Hong,Younwoo Choi,Chen-Hao Chao,Se-Young Yun,Rahul G. Krishnan

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:promising parallel alternative, diffusion language models, Masked diffusion language, remaining masked positions, language models

备注: EMNLP 2026 Main (Oral)

点击查看摘要

Abstract:Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.

4. 【2610.03632】World Embedding Benchmark

链接:https://arxiv.org/abs/2610.03632

作者:Yiqi Liu,Ruifeng Yuan,Yang Wang,Long Li,Fengyu Cai,Hou Pong Chan,Jialin Yu,Hao Zhang,Chenghua Lin,Chenghao Xiao

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:received increasing attention, World Embedding Benchmark, remains less understood, Physical, received increasing

备注:

点击查看摘要

Abstract:Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.

5. 【2610.03625】FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs

链接:https://arxiv.org/abs/2610.03625

作者:Darian Lee,Shannon Rumsey,Jack St. Clair,Xinyi Tang,Aditya Bansal,Yuanming Shi

类目:Computation and Language (cs.CL); Databases (cs.DB); Machine Learning (cs.LG)

关键词:people phrase requests, widely deployed forms, Relational databases, requires grounding language, structured knowledge access

备注: Accepted to AKBC Workshop, EMNLP

点击查看摘要

Abstract:Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely ignore this ambiguity and produce oversimplified queries that fail to prepare models for the complexity of real-world structured knowledge access. We present FALCON, a framework that generates realistic, ambiguity-aware NL-to-SQL data matching the complexity of challenging real-world benchmarks, at low cost using compact open models. Our approach combines reserved-word SQL seeding and persona-based prompting to generate structurally complex queries, while alignment-based filtering preserves difficulty by distinguishing genuinely incorrect examples from complex but valid queries. Human evaluation confirms consistent high quality across model sizes, and our generated data exceeds existing benchmarks in both SQL complexity and natural language richness. Difficulty-stratified analysis shows models trained on FALCON data increasingly outperform baseline-trained models as query complexity increases, validating our pipeline's success in generating challenging training data. When combined with a small proportion of existing benchmark data, mixed training recovers performance on simpler queries while preserving these advantages on complex ones. The model- and database-agnostic design enables organizations to generate high-complexity NL-to-SQL training data locally without external APIs.

6. 【2610.03567】Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting

链接:https://arxiv.org/abs/2610.03567

作者:David L. Condrey

类目:Computation and Language (cs.CL)

关键词:Writerslogic team participation, SimpleText shared task, describe the Writerslogic, Plain Language Summary, Cochrane Plain Language

备注: 11 pages, 3 tables. Notebook for the SimpleText Lab at CLEF 2026. Code: [this https URL](https://github.com/dcondrey/simpletext-clef2026)

点击查看摘要

Abstract:We describe the Writerslogic team's participation in the CLEF 2026 SimpleText shared task, addressing Task 1 (text simplification) and Task 2 (complexity spotting). For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying temperatures, then selects the best candidate using a reference-free scoring heuristic that rewards compression, source word retention, Cochrane Plain Language Summary vocabulary usage, and lexical simplicity. On Task 1.1 (sentence-level simplification), our Claude Sonnet 4 submission achieves SARI 47.43 and BLEU 14.21, the top-ranked sentence-level system (3rd on the combined Task 1 leaderboard, behind two document-level submissions). For Task 2, we fine-tune a DeBERTa-v3-large NLI model on 350K labeled (source, sentence) pairs, framing hallucination detection as natural language inference. The model reads the most relevant source sentence as premise and the candidate as hypothesis, directly learning to distinguish grounded from hallucinated content. On Task 2.1 (binary overgeneration identification), our fine-tuned DeBERTa system achieves 0.8081 document-level macro F1 (0.8085 in our best ensemble), the top-ranked entry within the identification track and 2nd among teams overall, behind AIIR Lab (0.8197). On Task 2.2 (multi-class error classification), our best submission reaches 0.804 multiclass accuracy, ranking 2nd among unique teams behind AIIR Lab (0.827). We evaluate both tasks on English and multilingual biomedical text from Cochrane systematic reviews.

7. 【2610.03565】Writerslogic at PAN 2026: Process over Content for Robust Detection under Domain Shift

链接:https://arxiv.org/abs/2610.03565

作者:David L. Condrey

类目:Computation and Language (cs.CL)

关键词:Reasoning Trajectory Detection, Writing Style Analysis, training-set effect size, Multi-Author Writing Style, shared analytical framework

备注: 13 pages, 1 figure, 6 tables. Notebook for the PAN Lab at CLEF 2026. Code [this https URL](https://github.com/dcondrey/voight-kampff-clef2026) and [this https URL](https://github.com/dcondrey/trajectory-detection-clef2026)

点击查看摘要

Abstract:We describe the Writerslogic systems for three PAN at CLEF 2026 shared tasks (Reasoning Trajectory Detection, Voight-Kampff Generative AI Detection, and Multi-Author Writing Style Analysis), unified by a shared analytical framework: feature robustness under distribution shift is governed by support overlap between training and test distributions, not by training-set effect size. This yields a taxonomy (domain-anchored, domain-portable, domain-invariant) that explains why generator-specific features die under domain shift while vocabulary fingerprints (hapax ratio, Yule's K, Heaps' exponent), compression measures, and character n-grams survive. On Reasoning Trajectory Detection, where training was entirely mathematics and 84 percent of test was unseen domains, the framework guided system design to 1st place in source detection (0.85 macro F1 via Opus-Sonnet agreement) and 3rd place in safety classification (0.66 macro F1 via query-refusal decomposition). For Voight-Kampff, we built a calibrated ensemble of DeBERTa-v2 (ONNX), multi-seed LightGBM with 44 domain-portable stylometric features, and SVM on n-gram TF-IDF, combined via learned stacking with isotonic calibration; the best configuration achieved 0.891 on the PAN 2026 test set with balanced sub-metrics (0.853 to 0.902 across all evaluation dimensions). For Multi-Author Writing Style Analysis, we describe a system fusing spectral clustering over character n-gram similarity graphs, normalized compression distance for local boundary detection, and SmolLM-135M perplexity for neural change-point detection; a platform mix-up meant our run never reached the official evaluation, so we report the design and its a priori predictions. Across all three tasks, features measuring generation process properties are designed to outperform features measuring generated content properties under domain shift.

8. 【2610.03531】Author Representation Strategies for Zero-Shot Authorship Attribution: A Comparative Study of LLM-Based and Embedding-Based Approaches

链接:https://arxiv.org/abs/2610.03531

作者:Nudrat Habib,Tosin Adewumi,Sana Sabah Al-Azzawi,Marcus Liwicki,Elisa Barney

类目:Computation and Language (cs.CL)

关键词:requires capturing fine-grained, fine-grained stylistic characteristics, capturing fine-grained stylistic, Authorship Attribution, requires capturing

备注:

点击查看摘要

Abstract:Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline together with three author representation strategies: representative writing samples, LLM-generated descriptions, and style embeddings (LISA). The first three approaches perform attribution using LLM prompting, while the embedding-based approach uses style embeddings with cosine similarity. We investigate the influence of prompt design and propose a two-stage embedding-based attribution framework that combines candidate space reduction with embedding-dimension selection. The results show that label-only ZS AA is ineffective, while incorporating author-specific representations consistently improves attribution performance. Among the evaluated approaches, the proposed two-stage LISA framework achieves the strongest overall performance, whereas LLM-generated style descriptions provide a substantially more compact representation of author style at the cost of some attribution performance. These findings demonstrate the importance of author representation in ZS AA, while indicating that current open-source LLMs remain insufficient for robust attribution without more effective representation learning.

9. 【2610.03529】Divergence controls entropy in distillation

链接:https://arxiv.org/abs/2610.03529

作者:Nicolas Zucchet,Scott W. Linderman

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language model, language model training, core primitive, language model, entropy

备注:

点击查看摘要

Abstract:Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.

10. 【2610.03525】Structured Composition of Verifiable Atomic Insights for Table-to-Report Generation

链接:https://arxiv.org/abs/2610.03525

作者:Teng Lin,Xinyu Liu,Nan Tang

类目:Computation and Language (cs.CL)

关键词:generating article-level analyt, automatically generating article-level, automated data science, article-level analyt, ical reports

备注:

点击查看摘要

Abstract:Table-to-report generation refers to the task of automatically generating article-level analyt- ical reports from relational tables and is an essential capability for automated data science and decision support. Its central challenge lies in systematically discovering verifiable com- posite insights across tables, attributes, and analytical perspectives, and organizing them into coherent, complete, and traceable evidence chains. Existing methods primarily rely on sequential, reactive data agents or direct Large Language Model(LLM) generation. They suffer from exploration bias: early local observations constrain subsequent actions, causing models to focus prematurely on local analyzes and miss cross-table or cross-dimensional evidence. We propose ComInsight, which reformulates insight discovery as the composition of atomic evidences. We first define an atomic insight as the smallest executable analytical unit conforming to a predefined analysis pattern and enumerate all valid atomic insights from database schema and content. These atoms are then organized into a multi-relational insight graph, where nodes represent verified data facts and edges encode logical, temporal, or hierarchical relations. Finally, a set of composition operators systematically fuses atomic nodes into higher-order composite conclusions. Every composite output is accompanied by executable SQL and fine-grained provenance, ensuring full verifiability. Across three benchmarks InsightBench, DDR-Bench, and T2R-Bench, ComInsight consistently outperforms strong baselines in factual correctness, novelty, and structural completeness. We believe ComInsight offers a reliable, efficient, and explainable path toward table-to-report generation.

11. 【2610.03515】Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation

链接:https://arxiv.org/abs/2610.03515

作者:Chenglei Shen,Haoyang Yao,Weijie Yu,Song Jin,Xiao Zhang,Jun Xu

类目:Computation and Language (cs.CL)

关键词:student-generated reasoning trajectories, reasoning, supervise student-generated reasoning, On-policy self-distillation, reference solutions

备注:

点击查看摘要

Abstract:On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student's own reasoning to provide guidance that addresses its specific errors while building on valid progress. For each failed attempt, RC-OPD locates the earliest substantive error, develops a local correction, and uses the corrected intermediate result as an anchor for the valid prefix. An iterative diagnosis--repair--continuation process tests the repairs through student continuation, identifying further errors within a fixed repair budget. For repair chains that reach a correct answer, root--cause--guided distillation uses failure diagnoses and corrective goals to supervise the erroneous segments, while anchor-guided distillation supports the corresponding valid prefixes with reasoning chains leading to the repaired intermediate results. We evaluate RC-OPD across multiple datasets and model scales. Extensive experiments and analyses show that it mitigates reasoning mismatch and the distillation trap, yielding substantial performance gains.

12. 【2610.03482】Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models

链接:https://arxiv.org/abs/2610.03482

作者:Mehrdad Ghassabi,Pedram Rostami,Hamidreza Baradaran Kashani,Sadra Hakim,Audrina Ebrahimi

类目:Computation and Language (cs.CL)

关键词:existing uncertainty-head resources, medical language models, Persian medical models, Persian medical language, repeated-sampling approaches

备注:

点击查看摘要

Abstract:Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, using Gaokerena-V and Gaokerena-R as two previously developed backbones. We first examine response variability on a 168-question Iranian medical entrance examination and observe substantially lower five-run consistency for Gaokerena-V than for Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. We then construct two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and train lightweight claim-level heads on frozen backbone attention maps and token probabilities. On held-out test splits, the heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The heads require neither retrieval nor repeated sampling at inference time. These results provide an initial study of single-pass claim-level uncertainty estimation for Persian medical language models; the test splits are small and the labels are automatically generated.

13. 【2610.03458】A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control

链接:https://arxiv.org/abs/2610.03458

作者:Zhe Zhou,Tianhua Tao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Post-training with verifiable, induce reward hacking, low hacking share, low monitor readout, low

备注: 17 pages, 2 figures, 10 tables. Accepted as a poster at the NeurIPS 2026 Workshop on Foundations of LLM Post-Training in Changing Environments (FLLMPT)

点击查看摘要

Abstract:Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low monitor readout does not identify whether such an intervention controls behavior. In a code-generation environment whose dominant exploit is available at the start of the reasoning trace, we train policies against three monitors that pass the same offline gate: an in-domain activation probe and two penalties conditioned on how early the policy commits to its own final answer. The probe score is at its numerical floor from the first recorded training step, and the trained-score median is zero for every prefix-trained run at the endpoint. These readouts estimate different quantities, and we do not compare their scales; within each monitor family, however, low values do not establish behavioral control. Within one fixed configuration, prefix-trained runs with the same zero-median trained score range, by seed alone, from a mixed regime with a low hacking share to near-pure reward hacking. All probe runs reach the hacking regime, but their floor-level readout reflects a mismatch between the position where the probe was validated and the position where it was read during training, not a second instance of this ambiguity. Text-level analysis identifies a prefix failure mode: generic planning and filler shells postpone the exploit past the cut without eliminating it from the final output. Low measured commitment therefore does not distinguish a low hacking share from delayed commitment to the exploit. Offline discrimination and low monitor-aligned readouts are insufficient evidence of behavioral control; an out-of-band behavioral check is required. We characterize the endpoint readout, not its evolution. Code is available at this https URL.

14. 【2610.03448】Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

链接:https://arxiv.org/abs/2610.03448

作者:Zhuowen Liu

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:LLM agents increasingly, agents increasingly screen, small prompt-injection detectors, increasingly screen tool, Meta Prompt Guard

备注: 12 pages, 5 figures, 4 tables. Code: [this https URL](https://github.com/lzwhehe/benign-instruction-bench)

点击查看摘要

Abstract:LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent's attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent's own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.

15. 【2610.03421】CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

链接:https://arxiv.org/abs/2610.03421

作者:Hang Gao,Wujiang Xu,Zhixing Zhang,Kai Mei,Jingyi Yang,Dimitris N. Metaxas

类目:Computation and Language (cs.CL)

关键词:visual reasoning abilities, model parametric knowledge, strong visual reasoning, knowledge-intensive visual question, visual question answering

备注: EMNLP 2026 Findings

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model's parametric knowledge. Existing multimodal RAG systems commonly rely on Top-$K$ retrieval or reranking, which may return redundant passages and provide limited control over whether an answer update is sufficiently supported by the retrieved evidence. We propose \textit{CLIMB}, a training-free inference-time framework for multimodal RAG. CLIMB first constructs a compact complementary evidence pool using an MMR-style objective that balances query relevance and passage-level redundancy. It then performs confidence-controlled refinement within this fixed pool: an R/E/C critic scores passages by relevance, evidence specificity, and cross-modal alignment, while an evidence-grounded confidence estimator accepts an updated answer only when the estimated confidence increases. This design provides a simple stopping criterion and reduces unnecessary refinement without modifying the underlying retriever or MLLM. Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines. Ablations further indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.

16. 【2610.03387】Benchmarking Candidate Coverage in Typed Decision Models

链接:https://arxiv.org/abs/2610.03387

作者:Jiawen Lu,Tongtong Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Typed decision models, Typed decision, answer options supplied, decision models return, models return choices

备注: 19 pages, 1 figure, 8 tables

点击查看摘要

Abstract:Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initial evaluation of Laya and Jev across AG News, DBpedia, Emotion, and TREC. The models receive identical frozen texts and requests: 300 calibration and 589 test texts yield 23,932 predictions per model. Present/absent pairs match ordinary candidate count, and name variants preserve descriptions, members, and order. Native rejection behavior differs sharply: at five TREC candidates with natural names, Laya detects 97.2% of missing-answer cases but falsely rejects 69.7% of present controls; Jev's rates are 24.8% and 0.0%. Calibration-only none-score thresholds change these rates to 33.9%/3.7% and 45.0%/1.8%, respectively. On DBpedia, Jev's high coverage-score AUROC supports a stronger operating point, whereas both models have weak complete-set accuracy on Emotion. Competence-conditioned analysis, probability-precision sensitivity, and interface audits show why classification, score ranking, and rejection policies need separate measurement. This initial benchmark is descriptive and limited to reference-label omission; it does not establish natural out-of-scope generalization, causal mechanisms, or a new rejection method.

17. 【2610.03367】Multilingual GSM-Symbolic: What determines capability transfer across languages?

链接:https://arxiv.org/abs/2610.03367

作者:Kenneth Enevoldsen,Riley Herchert,Sofie Mosegaard,Dan Saattrup Smart,Simon Enni,Isaac Chung,Sofie Bruun,Ayush Sunil Munot,Max Müller-Eberstein,Adnan El-Assadi,Elisa Bassignana,Gianluca Barmina,Hafsteinn Einarsson,Iben Nyholm Debess,Linda Freienthal,Lukas Galke Poech,Mike Zhang,Nicolas Legrand,Vladimir Salnikov,Yevhen Kostiuk,Zafar Hussain,Sagandeep Kaur,Agnes Toftgård,Marie Mattson,Kristoffer Nielbo

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:beta, rely on incomparable, capabilities acquired, rarely examine, evaluations rely

备注:

点击查看摘要

Abstract:We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size ($\beta = 1.77$), language resource level ($\beta = 0.77$), reasoning ($\beta = 0.67$) and typological distance ($\beta = -0.25$). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages ($\beta = -0.27$ and $\beta = -0.20$, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

Cite as:
arXiv:2610.03367 [cs.AI]

(or
arXiv:2610.03367v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2610.03367

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
18. 【2610.03329】SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

链接:https://arxiv.org/abs/2610.03329

作者:Mohsen Larni(1),Sobhan Ebrahimi Azar(1),Pouyan Nahed(1),Kazem Taghva(1) ((1) Department of Computer Science, University of Nevada, Las Vegas)

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:syntactic errors matter, small syntactic errors, Large language models, Large language, errors matter

备注: 32 pages, 17 figures. The first two authors contributed equally. The code will be released soon

点击查看摘要

Abstract:Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.

Comments:
32 pages, 17 figures. The first two authors contributed equally. The code will be released soon

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

ACMclasses:
I.2.7; I.2.6

Cite as:
arXiv:2610.03329 [cs.CL]

(or
arXiv:2610.03329v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.03329

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
19. 【2610.03324】o Jev or Not? Evaluating the Accuracy and Efficiency of Structured Decision Models for Hate-Speech Moderation

链接:https://arxiv.org/abs/2610.03324

作者:Demetris Paschalides,George Pallis,Marios D. Dikaiakos

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, enable harmful material, online content makes, content makes hate-speech

备注:

点击查看摘要

Abstract:The scale of online content makes hate-speech moderation challenging, while Large Language Models (LLMs) enable harmful material to be produced and adapted more easily. Moderation therefore requires efficient classifiers that can accommodate different definitions of hate speech. Recent structured decision models accept natural-language criteria and select among specified answers, raising the question of whether they can meet these requirements without task-specific training. We present HATEDECIDE, an evaluation of six decision-model configurations on four hate-speech datasets against specialized moderation, zero-shot, commercial, and supervised baselines. We examine whether supplying a dataset's definition, or decomposing it into multiple questions, improves classification, and we measure their latency and cost. We find that commercial LLMs significantly outperform all decision models on only one dataset. Supplying definitions changes up to 28\% of predictions without consistently improving classification, and decomposition significantly improves performance in only 20\% of the comparisons. On a diagnostic set of test cases, the best hosted decision model comes within 1.6 macro-F1 points of the best commercial LLM at approximately 97\% lower inference cost. These results identify opportunities for inexpensive moderation, while showing that explicit criteria and additional questions do not reliably improve classification.

20. 【2610.03268】Shrome at Touché: Soft-Vote Ensembling and Counter-Causal Augmentation for Causality Extraction

链接:https://arxiv.org/abs/2610.03268

作者:Roham Zendehdel Nobari,Shayan Sooratgar

类目:Computation and Language (cs.CL)

关键词:extends causality extraction, extends causality, denies the causation, meaning denies, falsely believed

备注: 16 pages, 5 figures, 10 tables. Both authors contributed equally. Working notes of Touché at CLEF 2026 (Conference and Labs of the Evaluation Forum), 21-24 September 2026, Jena, Germany

点击查看摘要

Abstract:Touché 2026 extends causality extraction to counter-causal claims: news sentences whose surface form appears causal but whose meaning denies the causation, as in "It is falsely believed that X caused Y." A system that relies on surface cues such as "caused" or "led to" will accept such a sentence as causal and give it the wrong polarity. On the Countercausal News Corpus (CCNC), the task has three subtasks: deciding whether a sentence is causal (detection), locating its cause and effect spans (extraction), and labeling its polarity as procausal, counter-causal, or uncausal. We build one model per subtask. Detection is a fine-tuned classifier with a single cross-task rule that uses the extracted spans to remove false positives. For extraction, we ensemble three RoBERTa-large BILOU+CRF taggers by averaging their token-level scores before decoding, rather than voting on the spans each tagger produces. For polarity, where labeled counter-causal examples are scarcest, we add training sentences generated by a large language model prompted with nine patterns of counter-causal expression adapted from Hagen et al., keeping only those that pass automatic structural checks. On the held-out CCNC test set, the system reaches F1 0.869 on detection and macro-F1 0.817 on polarity, and in the organizers' final causal-only evaluation of extraction it scores granularity-adjusted F1 0.728, the highest extraction score among all submissions including the organizers' baseline. The development split is used only for component selection and the ablations reported in the paper.

21. 【2610.03240】Collective Bias Mitigation via Model Routing and Collaboration

链接:https://arxiv.org/abs/2610.03240

作者:Mingzhe Du,Luu Anh Tuan,Xiaobao Wu,Yichong Huang,Yue Liu,Dong Huang,Huijun Liu,Bin Ji,Jie M. Zhang,See-Kiong Ng

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, public health, requiring both accuracy, societal value alignment

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model's intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained model behavior and fostering knowledge sharing among diverse LLMs. This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses. Experiments show CBM substantially outperforms standalone baselines (e.g., in the top-7 setting, Committee lowers the age bias score from 0.25 to 0.10). Our Debating and Committee topologies achieve substantial bias reduction, with the latter balancing mitigation effectiveness and inference cost, highlighting the potential of CBM for fairer LLMs.

22. 【2610.03223】AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

链接:https://arxiv.org/abs/2610.03223

作者:Xin Wang,Wenhao Wu,Menghao Zhang,Zhi Wang,Kun Shao,Jian Luan

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Long-horizon LLM agents, Long-horizon LLM, sparse outcome rewards, making trajectory-level objectives, LLM agents

备注: 21 pages, 3 figures

点击查看摘要

Abstract:Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on subsequent actions, environment transitions, and trajectory length. We propose AdaStep, an Adaptive Step-credit weighting method that controls how strongly each group-derived local advantage modifies the trajectory-level signal. We formulate this weighting as a mean-squared-error estimation problem for the latent step advantage and, under an explicit conditional sampling assumption, derive an optimal per-state shrinkage coefficient. The coefficient admits a signal-to-total-variance interpretation: it preserves local credit when return variation is attributable to the selected action and suppresses it when variation is dominated by downstream randomness. AdaStep requires only lightweight scalar computation, with no critic, additional rollouts, or extra model inference. Experiments with three model backbones on ALFWorld, WebShop, and ScienceWorld show consistent improvements over baselines at low computational cost.

23. 【2610.03215】StanceEval 2026: The Second Stance Detection Shared Task

链接:https://arxiv.org/abs/2610.03215

作者:Rasha Albalawi,Nuha Albadi,Hamzah Luqman,Asma Yamani,Maram Kurdi,Saad Ezzini,Ahmed Ashraf,Maged Al-Shaibani,Nora Alturayeif

类目:Computation and Language (cs.CL)

关键词:Arabic social media, social media text, Arabic social, Track, media text

备注: 14 pages total (8 pages main paper + 6 pages appendix), 5 tables in the main paper, excluding the appendix

点击查看摘要

Abstract:StanceEval 2026 is the second edition of the StanceEval shared task series on stance detection in Arabic social media text. Stance detection aims to identify a writer's stance toward a given topic. Given a tweet and a target, participating systems must determine whether the writer's stance is Favor, Against, or None. This edition focuses on cross-target generalization across two distinct evaluation tracks: Track 1 evaluates thematically related cross-target transfer (testing on Women Driving, related to Women Empowerment from training data), while Track 2 evaluates cross-domain transfer to completely unseen targets (E-Cars and Trimester System). The shared task attracted 80 registered teams from 12 countries. During the evaluation phase, 30 unique teams submitted entries, with 21 teams officially ranked in Track 1 and 13 in Track 2 following validation filtering, and 20 teams submitting system-description papers. Participating teams employed diverse methodologies, including fine-tuned pretrained language models, prompt-based and retrieval-augmented large language models (LLMs), fine-tuned LLMs, and hybrid cascades. Top systems achieved impressive $F_{avg2}$ scores of 0.8994 on Track 1 and 0.9400 on Track 2, substantially outperforming the strongest baselines (0.7366 and 0.7475, respectively), where $F_{avg2}$ denotes the macro-averaged F1 score over the Favor and Against classes. Counterintuitively, performance on the unseen targets was higher than on the related target, a disparity could be driven by extreme target polarization, class imbalance, and dialectal or sarcastic nuance across topics.

24. 【2610.03199】Predicting and Repairing Merge Collapse in Large Language Models

链接:https://arxiv.org/abs/2610.03199

作者:Jungseob Lee,Seungyoon Lee,Sugyeong Eo,Hyeonseok Moon,Jaehyung Seo,Heuiseok Lim

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Large language models, Large language, task vectors, language models fine-tuned, give no warning

备注: 23 pages, 5 figures, 20 tables

点击查看摘要

Abstract:Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists' task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of interference. Under a working noise model, the disturbance that a merge injects grows with the merge coefficient and with interference, yielding a pre-merge score. In our experiments on twenty-two merge configurations from four model families, only destructive merges exceed a threshold on this score. We find that statistics of sign conflict between specialists, a common target of existing merge operators, are anti-predictive. We then predicted the outcomes of fourteen merges before evaluating them, and twelve predictions were correct, including the destructive outcome of a specialist pair pushed past the threshold by continued pretraining. To address this collapse, we introduce PRISM, an operator that averages the task vectors first and then soft-thresholds each layer at a level set by the layer's interference. Without data or tuning, PRISM keeps all five destructive merges above the threshold within evaluation noise of the base model, where plain averaging falls at least 14.4 points below it or collapses entirely. We apply PRISM only above the threshold and keep the plain average for merges below it, which include all fifteen harmless ones. Code is available at this https URL.

25. 【2610.03198】KV$^2$: A Self-Refining KV Cache

链接:https://arxiv.org/abs/2610.03198

作者:Johannes Wesch,Danni Liu,Jan Niehues

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:cache constrains, long-context models, constrains the practical, dominates cost, compression trades cost

备注:

点击查看摘要

Abstract:The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at this https URL.

26. 【2610.03195】Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

链接:https://arxiv.org/abs/2610.03195

作者:Jonghyun Song,Haewon Park,Jeonghoon Shim,Woojung Song,Yohan Jo

类目:Computation and Language (cs.CL)

关键词:LLM agents decide, LLM agents, product to buy, hotel to book, paper to cite

备注: 41 pages

点击查看摘要

Abstract:As LLM agents decide on users' behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-to-end search with 12 agent models across three domains. Comparing items from different sources that satisfy the same requirements at the same position, we find that each model prefers some sources and avoids others in every domain, largely agreeing on which. This preference can outweigh how well items satisfy the request: an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better one from a dispreferred source, but almost never in the reverse case. The information identifying an item's source affects selection by itself: hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate. We test two routes to this preference: training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source. Supplying missing information or a prompt countering these preconceptions reduces source preference.

27. 【2610.03190】Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case

链接:https://arxiv.org/abs/2610.03190

作者:Tingzhu Bi,Ping Wang,Meng Ma

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:outage investigations end, ordinary question answering, defect and outage, answering never faces, outage investigations

备注: 23 pages. Dataset: [this https URL](https://huggingface.co/datasets/etigerstudio/Nautil) ; Models: [this https URL](https://huggingface.co/etigerstudio/Nautil-SFT) , [this https URL](https://huggingface.co/etigerstudio/Nautil-RLVR) ; Demo: [this https URL](https://huggingface.co/spaces/etigerstudio/Nautil-Demo) ; Code: [this https URL](https://github.com/etigerstudio/Nautil)

点击查看摘要

Abstract:Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is "cause undetermined". Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.

28. 【2610.03185】Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

链接:https://arxiv.org/abs/2610.03185

作者:Han Cui,Jianhao Yan,Yun Luo,Hongbo Zhang,Zhizhang Fu,Yue Zhang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:On-policy distillation, language model post-training, important approach, approach to language, OPD

备注:

点击查看摘要

Abstract:On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student's capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at this https URL.

29. 【2610.03176】Hindsight-Guided Rationale Distillation for Rare Disease Diagnosis

链接:https://arxiv.org/abs/2610.03176

作者:Aarav Singh,Animesh Pathak,Navyansh Singh

类目:Computation and Language (cs.CL)

关键词:rare disease diagnosis, study hindsight-guided distillation, ground-truth diagnosis, study hindsight-guided, observes the ground-truth

备注: 15 pages, 4 figures, Github: [this https URL](https://github.com/joetheguide2/hindsight) , Accepted at AACL-IJCNLP SRW 2026

点击查看摘要

Abstract:We study hindsight-guided distillation for rare disease diagnosis on ZebraMap: a 1.5B student is fine-tuned on chain-of-thought traces from a 8B teacher that observes the ground-truth diagnosis during generation. Absolute accuracy remains low for all models - the task is hard at this scale - but within this ceiling a filtered variant (StudentF) achieves a small, statistically significant accuracy advantage over the teacher (p 0.001), concentrated in better-represented diseases. The unfiltered student does not significantly outperform the teacher (p = 0.129), establishing that contamination filtering - not hindsight distillation alone - drives the gain. The gap traces to an artifact we term GT hallucination. Label-visible generation causes the teacher to embed "ground truth is X" phrases in its reasoning chain; SFT copies the pattern. At inference, the unfiltered student reproduces the phrase in 33.9% of cases, with severe accuracy degradation when the hallucinated label is wrong. A regex filter removing these slots reduces contamination to near-zero, producing the observed gain - though the effect remains small. We precisely quantify this gain-cost tradeoff, document frequency-dependent knowledge transfer absent from the RL-trained teacher, and characterize a calibration gap that SFT does not close - identifying both as directions for future work.

30. 【2610.03163】Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer

链接:https://arxiv.org/abs/2610.03163

作者:Leonard Popp,Danni Liu,Supriti Sinhamahapatra,Jan Niehues

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Adapting large language, formal conventions leave, large language models, scientific writing sharpens, Adapting large

备注: W-NUT Workshop @ EMNLP 2026

点击查看摘要

Abstract:Adapting large language models to an individual author's style from a few examples is challenging, and scientific writing sharpens the difficulty: formal conventions leave little surface variation, and authors write about their own topics, so extracted ``style'' easily entangles with content. We study style-conditioned abstract generation from a few example abstracts per author and propose three methods: (1) contrastive activation steering, (2) a network that predicts steering vectors, and (3) a hypernetwork that predicts LoRA adapters. We find a consistent trade-off between style imitation and output quality: fine-tuning buys most of the available style signal but forfeits fluency, while the hypernetwork achieves the best trade-off on both seen and unseen authors. Our steering operates at author level, contrasting an author's abstracts against style-neutral generations for the same content. This holds topic fixed, removes the need for a predefined style inventory, and outperforms inventory-based steering. % [EDIT 1a] softened "no single optimal axis" claim Moreover, our analyses demonstrate that manually extracted and predicted steering vectors are near-orthogonal yet score comparably, indicating that style conditioning here can admit at least two unrelated directions rather than requiring one particular axis.

31. 【2610.03136】Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation

链接:https://arxiv.org/abs/2610.03136

作者:Oliver Hauck,Mario Sanz-Guerrero,Katharina von der Wense

类目:Computation and Language (cs.CL)

关键词:Reasoning traces improve, traces improve large, improve large language, large language models, reasoning language

备注: Accepted to the Workshop on Open Reasoning Across Cultures Languages at EMNLP 2026

点击查看摘要

Abstract:Reasoning traces improve large language models (LLMs), but current models are trained to reason mostly in English. It has been shown that forcing a model to reason in another language degrades accuracy, even when the reasoning language matches the language of the prompt -- but only for a setting where the model reasons over a short prompt. Here, we ask whether the same holds for retrieval-augmented generation (RAG), where the model must read and integrate a large amount of retrieved evidence in the target language. To study this, we build a fully monolingual German RAG question-answering testbed over the fictional world of the tabletop role-playing game The Dark Eye, a domain that is richly documented in German but too niche for the model to answer from memory, so that it has to rely on retrieval. Varying the forced reasoning language of an agentic RAG system on this testbed, we find that aligning the reasoning language with the language of the query and the retrieved documents helps. Forced German reasoning outperforms forced French, although the model benchmarks higher in French, so the benefit comes from alignment and not from language proficiency. The advantage grows when the retrieved context is richer and structure-aware. However, forced German only reaches the level of the model's native, unconstrained English reasoning without surpassing it, showing that native multilingual reasoning is needed. We publicly release the testbed and QA benchmark.

32. 【2610.03130】Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study

链接:https://arxiv.org/abs/2610.03130

作者:Yun Wang,Gad Shaulsky,Tomaž Curk,Blaž Zupan

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:general-purpose search tasks, broad biomedical corpora, search tasks, literature retrieval systems, developed and evaluated

备注: 15 pages, 5 figures. Submitted version (before peer review) of a paper accepted at Discovery Science 2026 (DS 2026); to appear in the Springer proceedings. Code and data: [this https URL](https://github.com/fulaibaowang/dictycite) ; dataset: [this https URL](https://doi.org/10.5281/zenodo.20308282)

点击查看摘要

Abstract:Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval benchmark from dictyBase for Dictyostelium, a model organism in cell and developmental biology. The benchmark consists of curator-generated biological queries linked to PubMed-indexed articles, together with structured gene annotations. Using this benchmark, we study three factors in niche biological retrieval: cross-encoder reranking, gene-aware query expansion, and abstract-only versus full-text retrieval. We report that reranking and gene-aware query expansion improve retrieval selectively: reranking is most useful when the model is well suited to biological evidence matching, whereas curated annotations help clarify compact biological queries by reducing vocabulary mismatch. Full-text chunks substantially improve retrieval when abstracts omit supporting evidence, increasing both candidate recall and top-rank performance, although these cases are harder than queries supported by abstracts. Data and code are publicly available at this https URL, and the benchmark dataset is additionally archived on Zenodo.

33. 【2610.03124】he Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs

链接:https://arxiv.org/abs/2610.03124

作者:Toluwani Aremu,Manit Baser,Mohan Gurusamy,Nils Lukas,Dinil Mon Divakaran

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:centrally enforced safeguards, Open-weight language models, developers' control, limiting the effectiveness, enforced safeguards

备注:

点击查看摘要

Abstract:Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \emph{token-level trigger-tags}, which introduce watermark-inspired signals during decoding, from \emph{weight-level trigger-tags}, which learn backdoor-inspired associations between target conditions and detectable model behavior. Furthermore, (ii)~we introduce \Untag, a unified attack framework that organizes their mechanism-specific attack surfaces into a common taxonomy. We evaluate representative token-level and weight-level trigger-tags using phishing as a case study. We find that while trigger-tags may provide useful evidence in controlled settings, our attacks render the existing trigger-tag mechanisms to be entirely ineffective. Consequently, we argue that these mechanisms should not be treated as robust misuse detectors when attackers can transform outputs or modify open weights.

34. 【2610.03112】Building Interpretable Feature Representations for Resume-Vacancy Matching by Distilling Production LLM Signals

链接:https://arxiv.org/abs/2610.03112

作者:Ilya Chekin(1),Vyacheslav Malyugin(1),Vladimir Chirkov(1),Mikhail Yurushkin(2) ((1) BroutonLab, (2) Curately)

类目:Computation and Language (cs.CL)

关键词:opaque relevance score, single opaque relevance, candidate fits, central to recruitment, relevance score

备注: Accepted to EMNLP 2026; 13 pages, 4 figures, 7 tables

点击查看摘要

Abstract:Matching candidates to vacancies is central to recruitment, and a recruiter needs to see why a candidate fits, not only a single opaque relevance score. We provide this evidence as named, interpretable matching dimensions recruiters can act on - eight in our current deployment. We propose a two-part approach. The first is an LLM-based labeler whose prompts and feature definitions were refined from recruiter feedback while it served as an earlier production matching stage. In the current architecture, it is used only for offline labeling and is not called on online requests. The second is a feature bi-encoder distilled from it: a LoRA-adapted embedding backbone with compact per-dimension heads that runs on CPU and serves all online requests. Both parts keep improving: prompts are revised as feedback arrives, and the bi-encoder is retrained on the updated labels. The model is trained on 168,772 labeled vacancy-resume pairs (17,921 vacancies and 180,030 resumes). Recruiters using the service can confirm or revise surfaced feature predictions. On 927 recruiter-recorded values from this selected production-feedback subset, the deployed student agrees with the recorded decisions in 888 cases (95.79%). This is operational, non-blinded agreement rather than an independent human evaluation.

35. 【2610.03110】Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text

链接:https://arxiv.org/abs/2610.03110

作者:Claudiu Creanga,Liviu Dinu

类目:Computation and Language (cs.CL)

关键词:Supervised AI-text detectors, high benchmark accuracy, report high benchmark, AI-text detectors report, detectors report high

备注:

点击查看摘要

Abstract:Supervised AI-text detectors report high benchmark accuracy, but it is not clear what their decisions are based on. We analyze a RoBERTa-based detector under semantic, structural, and tokenizer-level perturbations, using the M4 dataset (N = 10,000) and controlled generations (N = 300). When Mistral-7B-Instruct was asked to make machine text sound more human, Verb Diversity rose from 0.77 to 0.92 and the outputs became easier to detect. Detection scores appear to track statistical complexity, which also leads to a 76.3% false-positive rate on formal human writing. As a control, we evaluate event-based Latent Space detection. Paraphrasing changed 87% of its event sequences (Jaccard = 0.067), and homoglyphs altered 70% of the extracted verbs even though extraction still ran (Jaccard = 0.30). Its best domain AUC was 0.577. RoBERTa's robustness seems specific to the features it uses, and structural abstraction did not make detection more robust.

36. 【2610.03109】Emergent Structure in the Marginal Attention Space of Language Models

链接:https://arxiv.org/abs/2610.03109

作者:Valentino Maiorca,Walter Nelson,Francesco Locatello

类目:Computation and Language (cs.CL)

关键词:independently trained language, trained language models, representation similarity, similarity across independently, independently trained

备注:

点击查看摘要

Abstract:While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping them into a joint token-head "marginal attention space". Evaluating across 60+ diverse LLMs, we find that different properties emerge when reducing this space along its token and head axes. When reduced token-wise, marginal attention yields a text-intrinsic signal robustly conserved across models. To explain this property, we empirically connect marginal attention to the input-output Jacobian of the network, and prove theoretically that under a smoothness assumption, models with similar next-token distributions are guaranteed to have similar input-output Jacobian statistics. When reduced head-wise, it forms a model-private signature conserved across documents. Practically, this provides a natural way to estimate a per-head budget for key-value (KV) cache eviction, effectively decoupling model-specific budget allocation from text-intrinsic token scoring. On standard eviction benchmarks, a per-head budget precomputed offline on pretraining text, combined with a training-free token score, shows competitive performance with methods that recompute the budget on every document or train it per target. Code available at this https URL

37. 【2610.03102】Ask, Relax, or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning

链接:https://arxiv.org/abs/2610.03102

作者:Ang Li,Yue Lin,Feifei Kou,Zhan Su,Prayag Tiwari,Wenhao Li,Shuhui Zhu,Hongyuan Zha,Baoxiang Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:LLM agent, agent can recognize, choose the wrong, LLM, minimum-cost permitted constraint

备注: 55 pages, 5 figures

点击查看摘要

Abstract:An LLM agent can recognize uncertainty yet still choose the wrong next step: asking when action is already justified, or seeking clarification when the constraints must change. We formalize actionable indeterminacy: act when an accepted action is shared across all admissible preferences or objectives, clarify when each possibility is feasible but no action is shared, and propose a minimum-cost permitted constraint repair when the request is infeasible. We construct a solver-grounded benchmark spanning object allocation, meeting scheduling, apartment choice, and stable matching. Matched pairs retain the same source while changing whether intervention is necessary, and evaluation separates decision correctness, matched-pair reliability, and fully correct responses. Our findings reveal a recurring difficulty in recognizing when intervention is unnecessary: models can identify situations requiring clarification or repair yet still intervene when a justified action already exists. Correct decision labels also fail to guarantee usable actions, questions, or repairs. Crucially, response requirements shape not only how decisions are expressed but also which decisions are made. Making the required content explicit substantially improves fully correct responses and can change intervention decisions, even when outputs are already parseable. These findings highlight that reliable agency requires more than recognizing uncertainty: it requires intervening only when necessary and translating the chosen next step into a verifiable response.

38. 【2610.03095】Peer Influence across Heterogeneous AI Models

链接:https://arxiv.org/abs/2610.03095

作者:Frida Nøhr Laustsen,Marie Haahr Petersen,Victoria Popa,Ariel Flint,Romualdo Pastor-Satorras,Andrea Baronchelli,Luca Maria Aiello

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Physics and Society (physics.soc-ph)

关键词:models, persuasion, increasingly combine language, Abstract, combine language models

备注: 30 pages, 16 Figures, 6 Tables

点击查看摘要

Abstract:When two AI agents disagree, who persuades whom? As multi-agent systems increasingly combine language models of different families and sizes, the answer can determine which judgments survive interaction. Measuring persuasion as the probabilistic shift in an agent's decision after a single exchange with a dissenting peer, we test seven open-weight models across three language understanding tasks. We find that persuasion is strong: when models disagree, receivers often abandon their initial judgment after seeing a peer's answer and explanation. Surprisingly, however, neither standalone certainty nor model scale reliably predicts persuasion dynamics. Models producing almost perfectly consistent decisions in isolation can be among the most susceptible to persuasion, and small models can match larger ones as persuaders and resist their influence just as effectively. Furthermore, we show that the size of the shift depends more on the susceptibility of the listener than on the persuasiveness of the speaker. Persuasion patterns are therefore specific to each model pairing, with heterogeneity amplifying persuasion in some combinations and suppressing it in others, allowing a dissenting agent running a small model to overturn the judgments of a much larger one. These findings show that the behavior of interacting models cannot be inferred from their individual properties but must be evaluated in the combinations in which they will operate.

39. 【2610.03080】MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

链接:https://arxiv.org/abs/2610.03080

作者:Siyu Wang,Yifan Wang,Yuecheng He

类目:oftware Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG); Trading and Market Microstructure (q-fin.TR)

关键词:Large language models, producing trading signals, Large language, moving from producing, producing trading

备注: 5 pages, 3 figures, benchmark code and evaluation harness available at [this https URL](https://github.com/spearmintai/minteval) . Siyu Wang and Varstern Yifan Wang contributed equally, Yifig Wang is corresponding author

点击查看摘要

Abstract:Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.

40. 【2610.03078】An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

链接:https://arxiv.org/abs/2610.03078

作者:Hanlu He,Harald Vilhelm Skat-Rørdam,Ingvi Örnólfsson,Ivana Konvalinka

类目:Computation and Language (cs.CL)

关键词:requires reliable identification, dynamics requires reliable, Quantifying conversational dynamics, within-turn pauses, Quantifying conversational

备注:

点击查看摘要

Abstract:Quantifying conversational dynamics requires reliable identification of interactional units and their temporal boundaries, but speech activity alone does not distinguish conversational turns from listener feedback or within-turn pauses. We present an automated pipeline for extracting turns and backchannels from separate-channel recordings of spontaneous dyadic conversation, designed to provide a consistent first-pass annotation for subsequent human review. The pipeline combines voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post-processing. We evaluated the pipeline on 99 ten-minute Danish conversations from 33 dyads using segment-level detection reliability and temporal boundary error. Conversations were recorded under both normal and asymmetric listening conditions. In the latter, speech-shaped noise was delivered to one participant through bone-conduction headphones. Overall detection reliability was F1=0.621, with similar performance for turns F1=0.624 and backchannels F1=0.618. For successfully matched segments, median absolute onset and offset errors were 0.150 and 0.160s for turns and 0.130 and 0.180s for backchannels, respectively. Mean errors were substantially larger for turn boundaries, indicating a smaller number of large boundary mismatches. Performance did not differ significantly across the two experimental listening conditions. In a four-conversation case study, pipeline-human agreement was lower and more variable than human inter-annotator agreement and varied across parameter settings. These results support the pipeline as an automated first pass within a semi-automated annotation workflow, providing a consistent basis for more standardised and reproducible annotation of conversational dynamics.

41. 【2610.03077】Unmasking Propaganda: A Comparative Analysis of Masked and Causal Language Models

链接:https://arxiv.org/abs/2610.03077

作者:Claudiu Creanga,Ioachim Lihor,Liviu P. Dinu

类目:Computation and Language (cs.CL)

关键词:natural language processing, manipulative political communications, manipulative political, NLP, Propaganda

备注:

点击查看摘要

Abstract:Propaganda detection is an essential task in natural language processing (NLP), particularly in the context of manipulative political communications. However, identifying specific propaganda techniques presents a significant challenge due to their often subtle nature and reliance on context, making them difficult to distinguish from legitimate persuasive language. Propaganda often involves highlighting certain facts while downplaying or ignoring others to create a desired perception. This biased communication aims to influence attitudes, beliefs, or behaviors towards a particular cause or position. This paper explores advances in detecting propaganda techniques through a comparative analysis of modern language models, using the SemEval-2020 Task 11 dataset. We evaluated both masked language models (based on XLM-RoBERTa or DeBERTa V3) and causal models (from OpenAI, Google, Mistral, Anthropic and Meta), employing two prompting strategies: base and chain-of-thought prompting. Our results demonstrate improvements over state-of-the-art models, with the best-performing MLM achieving an F1 score of 63.18 in technique classification and the best causal model achieving 63.62. We also observed that certain models excel in specific techniques, such as loaded language and name-calling, while struggling with others like bandwagon and black-and-white fallacy. These findings suggest that fine-tuning, ensemble modeling, and the use of larger datasets can further enhance propaganda detection capabilities.

42. 【2610.03073】SecJev: Bringing Security Expertise to System One Decision Models

链接:https://arxiv.org/abs/2610.03073

作者:Zheng Chen,Fei Yu,Haohao Huang,Yang Li,Anlong Chen,Lei Chen

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:turn complex observations, turn complex, explicit policies, Security workflows, Security

备注: 22 pages, 1 figure

点击查看摘要

Abstract:Security workflows need models that turn complex observations and explicit policies into decisions. System One models introduced by Jev return typed predictions and probabilities; security specialization supplies the domain expertise behind those predictions. We introduce SecJev, to our knowledge the first family of Jev-like decision models specialized for security, spanning 0.8B to 9B parameters. Built on Kev's single-pass candidate scorer, SecJev learns Boolean, choice, and ordered decisions from text, telemetry, and observation histories. We develop SecJev-Corpus to unify source-label prediction and explicit-policy evaluation across 14 tasks and eight sources. It covers tool outputs, traffic, federated updates, consensus, authentication, and vehicle messages. Scene-weighted training adapts the models across these domains while preserving a shared typed decision interface. Security specialization improves every model in the family; SecJev-0.8B outperforms general Kev-9B by 20.51 percentage points in task-macro accuracy. Comparisons with answer-only generative fine-tuning show close accuracy and latency with lower peak inference memory. Tests on new source groups reproduce gains over Kev in prompt-injection and traffic decisions, with capture-dependent false alarms. We release adapters, decision heads, SecJev-Corpus, and training and inference code.

43. 【2610.03063】HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

链接:https://arxiv.org/abs/2610.03063

作者:Tiezheng Yu,Yuxin Jiang,Jinpeng Li,Shuning Sun,Fei Mi,Haoli Bai,Lifeng Shang

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, generating hallucinated content, Language Models, hallucinated content

备注: 11 pages

点击查看摘要

Abstract:Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.

44. 【2610.03052】he Geometry of Knowledge Accessibility in Large Language Models

链接:https://arxiv.org/abs/2610.03052

作者:Lihu Chen

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, accessible queries, knowledge, Large

备注:

点击查看摘要

Abstract:Large language models (LLMs) contain broad knowledge, but they cannot access all of it reliably. We study this problem through knowledge accessibility, which describes whether the knowledge needed for a query can be recalled from the model. We find that knowledge accessibility has a simple geometric structure in the model's representation of the query alone, before any generation. More accessible queries are closer to a center in the representation space, while less accessible queries are farther away. This geometry reveals a knowledge boundary that separates more accessible queries from less accessible ones. Accessibility consistently decreases with distance from the center, and this distance-based ordering transfers across datasets even when the centers differ. Controlled experiments further show that the centered geometry is more closely related to knowledge accessibility than to reasoning difficulty. The geometry also reveals when different interventions are useful. Query rewriting helps more for accessible queries, chain-of-thought reasoning helps more near the boundary, and retrieval gives larger gains beyond the boundary. These findings not only provide a new geometric view of how knowledge is organized in language models, but also suggest a useful pre-generation signal for adaptive inference.

45. 【2610.03039】HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

链接:https://arxiv.org/abs/2610.03039

作者:Donggyun Kim,Jack Lu,Chanwoo Kim,Mengye Ren,Seunghoon Hong

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:high inference-time overhead, introduce high inference-time, Long-form thinking traces, large language models, multi-step reasoning performance

备注: COLM 2026

点击查看摘要

Abstract:Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a single query-conditioned parameter update: a lightweight hypernetwork reads the question and predicts updates to a small subset of the base LLM's parameters, while a vector-quantized decoder constrains them to a finite set of reusable patterns to improve robustness and transfer. Trained end-to-end on outputs from the base model itself, HyperThink eliminates long thinking traces at test time: after one hypernetwork forward pass, the adapted model generates a concise step-by-step solution and final answer without an intermediate trace, using far fewer tokens while retaining strong reasoning performance. Empirically, HyperThink improves the low-latency region of the accuracy-latency trade-off on mathematical and general reasoning tasks, with its strongest gains in the near-non-thinking regime.

46. 【2610.03034】Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling

链接:https://arxiv.org/abs/2610.03034

作者:Ella Kemperman,Luca Ambrogioni

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Diffusion models rely, solvers requiring time-discretization, numerical solvers requiring, requiring time-discretization, models rely

备注: Submitted to ICLR 2027

点击查看摘要

Abstract:Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization important. We adapt proportional-integral (PI) step-size control to diffusion, using our diffusion noise-normalised error estimator. Unlike existing adaptive methods in diffusion that respond only to the current error, the PI solver also incorporates the previous error, yielding smoother step adaptation. We further show that these per-sample trajectories exhibit shared structure and can be aggregated into a fixed schedule that retains much of the benefit of adaptive sampling. We evaluate both approaches on natural-image and language datasets, in terms of quality, measured by FID at a matched number of neural network evaluations (NFE), comparing them with widely used stochastic solvers and schedules. For images, our fixed discretization outperforms the commonly used EDM schedule in terms of sample quality when used with the stochastic Heun sampler, and with the EDM-churn sampler at low NFE. Additionally, our PI adaptive solver obtains better FID than most stochastic and adaptive baselines, although it does not beat the EDM-churn sampler at low NFE. Moreover, we find our solver outperforms both the EDM and the entropy schedule on language diffusion at low-to-medium NFE in terms of perplexity, with the drawback of lower token entropy. Lastly, we find that the benefit of per-sample adaptivity is problem-dependent. It is highly beneficial in 1D toy examples, while only marginal for image and language data, where the average schedule sometimes even outperforms the PI-adaptive solver. Code is available at this https URL

47. 【2610.03027】ailoring the Quantization Space for 1-Bit KV Cache Compression

链接:https://arxiv.org/abs/2610.03027

作者:Minsoo Cheong,Donghyun Son,Sungjoo Yoo

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:placing substantial pressure, long-context LLM inference, major memory bottleneck, LLM inference, major memory

备注:

点击查看摘要

Abstract:The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to $14\times$ larger batch sizes and achieves $1.87\times$ higher peak throughput compared to the BF16 baseline.

48. 【2610.03025】Verifiable, Articulable, and Tacit Components of Preference

链接:https://arxiv.org/abs/2610.03025

作者:Alexander Spangher,Sheldon Huang,Andreas Haupt,Noah D. Goodman,Diyi Yang,Daniel E. Ho,Sanmi Koyejo

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:short story gripping, math proof elegant, story gripping, article newsworthy, proof elegant

备注: 15 pages main text, 14 pages of references, 107-page appendix (136 pages total); 15 figures, 48 tables; 213 references

点击查看摘要

Abstract:What makes a short story gripping; a news article newsworthy; or a math proof elegant? These constructs resist articulation or verification; their meaning is at least partially tacit. However, modern AI models are improved primarily via articulated constitutions, rubrics and verifiers (i.e. in RLAIF and RLVR); tacit components of preferences are typically understudied. We introduce a large, labeled preference dataset CreativePreferences, containing 2.8M texts labeled by 317M human preference judgments across 7 creative domains, with 42 benchmark tasks. We model these labels with executable programs, rubric banks and densely trained models (V, A and VAT, respectively). We observe robust articulability gaps, VAT-VA; and verifiability gaps, VAT-V; we estimate upper and lower bounds for each gap with a novel measurement approach that discovers articulable and verifiable metrics, identifies spurious variables and estimates the value of undiscovered metrics using capture-recapture. These gaps occur across all domains, even in domains traditionally treated as fully verifiable: correctness-centered domains (i.e. mathematics and software engineering) and claim- and novelty-centric domains (i.e. news, patents, peer review). The size of the gap varies based on domain (e.g. peer review and creative writing have the largest articulability gaps) and widens as more people take part in the judgment, consistent with Collins' collective tacit knowledge. We show two consequences: (1) on human generations, the full model more closely matches human preferences, often in disagreement with articulated criteria, and (2) in an analogy to Goodhart's law, articulating preference shifts it away from the tacit dimension. Articulability and verifiability gaps are consequential; we give recommendations on when tasks can be prompted; how learning mechanisms might improve; and when to leave judgments with humans.

49. 【2610.03022】ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

链接:https://arxiv.org/abs/2610.03022

作者:Sihan Ren,Gaozheng Li,Yuanshang Quan,Yiming Qin,Fuyi Yang,Chang Liu,Lan Xu,Minye Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:existing methods remain, methods remain largely, remain largely confined, assume pre-segmented inputs, Simultaneous Sign Language

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.

50. 【2610.03017】Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

链接:https://arxiv.org/abs/2610.03017

作者:David Nadrchal,Monorama Swain,Florian Schmid,Gerhard Widmer,Paul Primus

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Sound (cs.SD)

关键词:permanent tracheal stoma, severe dysarthria rendering, untrained listeners, work presents, presents an automatic

备注: 8 pages, three figures, to be published in IEEE Speech Language Technology workshop 2026

点击查看摘要

Abstract:This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.

51. 【2610.03002】Recursive Self-Improvement in Unified Multimodal Models

链接:https://arxiv.org/abs/2610.03002

作者:Huijuan Wang,Chufan Shi,Cheng Yang,Yaokang Wu,Taylor Berg-Kirkpatrick,Xuezhe Ma

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Unified multimodal models, Unified multimodal, Unified, training, model

备注:

点击查看摘要

Abstract:Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.

52. 【2610.02999】OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

链接:https://arxiv.org/abs/2610.02999

作者:Huiqiang Rong,Haoran Luo,Hui Feng,Zhonghong Ou,Kaiwen Xue,Guoxin Zhang,Yifan Zhu

类目:Computation and Language (cs.CL)

关键词:large language models, Omni-modal large language, language models, large language, Omni-modal large

备注:

点击查看摘要

Abstract:Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at this https URL.

53. 【2610.02994】Sentry: Learning to Recover from LLM Agent Failures at Test Time

链接:https://arxiv.org/abs/2610.02994

作者:Changxiu Ji,Amy Lu,Qizheng Zhang,Kunle Olukotun

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:invalid tool calls, poorly grounded reasoning, fail mid-task due, LLM agents, repeated actions

备注:

点击查看摘要

Abstract:LLM agents often fail mid-task due to invalid tool calls, repeated actions, or poorly grounded reasoning, and learning from these failures is a path to reliability. We find that how failure knowledge reaches the agent matters as much as what it contains. Failure lessons are conditional: kept in the agent's context, they misfire when their failure is absent, and removing them from an evolving playbook improves performance. Runtime interventions, in contrast, act only when a failure occurs but do not learn from their repairs. We argue that failure knowledge is conditional knowledge and should be conditionally exposed, and instantiate this principle in Sentry, a failure-management layer that runs alongside the agent. When Sentry detects a failure, it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did; the full playbook never enters the agent's context. Across multiple agentic benchmarks, Sentry outperforms the strongest runtime-intervention baseline on every benchmark, by 37\% on average, and the strongest context-evolution baseline by 39\% on the two benchmarks where both are evaluated; combining Sentry with context evolution yields further gains. Learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance even when relevant lessons remain available on demand.

54. 【2610.02986】OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models

链接:https://arxiv.org/abs/2610.02986

作者:Tao Shi,Chaoyi Xiang,Qiongkai Xu,Jey Han Lau

类目:Computation and Language (cs.CL)

关键词:LLM training data, large language models, LLM training, training corpus, aims to determine

备注:

点击查看摘要

Abstract:Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer from three limitations: limited coverage of training stages, insufficient distributional alignment between members and non-members, and lack of rigorous filtering of non-members against the training corpus. To address these limitations, we propose OLMo-Detect, a multi-stage, confounder-controlled benchmark built upon the fully open OLMo 2 pipeline. OLMo-Detect spans pre-training, mid-training, and post-training, explicitly aligns members and non-members on three key axes, and rigorously filters non-members via infini-gram. To assess robustness to distribution shifts, we further introduce OLMo-Detect (Shifted), a variant where members are misaligned with non-members. We evaluate 15 unsupervised and 3 supervised membership inference attacks (MIAs) across the OLMo 2 family, finding that: (i) overall performance is limited: the best unsupervised and supervised MIAs both reach an AUC of only 0.68, and supervised MIAs degrade under cross-domain evaluation; (ii) MIA performance peaks at mid-training and is lower at pre-training and post-training, a pattern driven by data type rather than a stage effect: curated math data is far more detectable than other types; (iii) overall scores improve from 1B to 13B but plateau at 32B; and (iv) no unsupervised MIA is robust to distribution shifts, with AUCs shifting by up to 0.42. Finally, we find that our findings on OLMo 2 generalize to OLMo 3 and non-OLMo models.

55. 【2610.02970】A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition

链接:https://arxiv.org/abs/2610.02970

作者:Songtao Li,Yijia Zhang,Shidi Zhang,Jianyuan Yuan,Fengyu Zhang,Hongfei Lin

类目:Computation and Language (cs.CL)

关键词:Large language models, shown promising potential, Large language, named entity recognition, language models

备注:

点击查看摘要

Abstract:Large language models (LLMs) have shown promising potential for biomedical named entity recognition (BioNER) through instruction following and in-context learning. However, existing LLM-based BioNER methods still face two key limitations. First, retrieved demonstrations and external biomedical knowledge provide limited support for dataset-specific annotation semantics, leaving entity boundaries, type scopes, and annotation conventions ambiguous. Second, free-form generation lacks sufficient structural control, often leading to invalid formats, hallucinated mentions, duplicated entities, and boundary errors. To address these limitations, we propose GAMA, a guideline-augmented multi-agent framework for schema-as-code BioNER. GAMA first induces candidate annotation rules from labeled training instances and verifies them against annotated data to construct reliable dataset-specific guideline memory. Guided by these verified rules, a planning component generates ranked span-type hypotheses with rationales, and a coding component converts them into schema-constrained entity objects. A verification module then checks span grounding, type validity, and structural compliance, and performs dual-loop refinement to correct invalid or low-confidence predictions. Experiments on five widely used BioNER datasets with multiple LLM backbones show that GAMA consistently outperforms strong LLM-based baselines. Ablation and parameter analyses further verify the effectiveness of the proposed components.

56. 【2610.02957】Understanding Trajectory Heterogeneity in Federated World Model Learning

链接:https://arxiv.org/abs/2610.02957

作者:Yipan Wei,Zhaokun Yan,Ziming Hong,Jiaqi Wu,Lixu Wang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:learn state evolution, central training requirement, models learn state, evolution from trajectories, learn state

备注:

点击查看摘要

Abstract:World models learn state evolution from trajectories, making access to temporal context a central training requirement. Federated learning can use distributed records, while ownership boundaries within a trajectory restrict the examples each client can construct. Our study benchmarks this cross-time setting through hourly action-conditioned clinical prediction on eight MIMIC-IV disease cohorts, comprising 40.87 million transition memberships. We specify severity-based client ownership, patient-separated construction, local history and future-window rules, and paired rollout evaluation from one to 32 hours. A matrix of ten federated algorithms covers 32 disease--partition configurations under five rounds of ten-percent participation. Three findings emerge from existing results and training logs. First, client ownership and participation jointly restrict long-window coverage: only 7.55\%--21.36\% of pooled-available 32-step windows have a locally complete anchor visited during training, averaged across diseases. Second, finer severity partitions accompany higher FedAvg error in 15 of 16 paired comparisons, while algorithm gains are small and horizon-dependent: FedProx reduces mean error by 0.56\%, with no consistent improvement at 32 steps. Third, algorithm labels conceal distinct update behavior, including inactive extrapolation and orders-of-magnitude differences in update scale. Cached-update performance also varies strongly across trajectory partitions under the same benchmark protocol. These results establish temporal access, participation coverage, optimization behavior, and horizon-resolved prediction as complementary dimensions for evaluating federated clinical world models.

57. 【2610.02949】Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method

链接:https://arxiv.org/abs/2610.02949

作者:Songtao Li,Yijia Zhang,Jianyuan Yuan,Shidi Zhang,Fengyu Zhang,Hongfei Lin

类目:Computation and Language (cs.CL)

关键词:applying large language, named entity recognition, large language models, common paradigm, paradigm for applying

备注:

点击查看摘要

Abstract:Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textual outputs, providing limited structural constraints for typed entity extraction. Second, high-quality biomedical annotations are limited, and learning from a single serialized output form may restrict structural diversity and reduce model robustness. Although external biomedical knowledge can be introduced to alleviate data scarcity, it often requires costly resource construction. To address these challenges, we propose MITE, a Multiple Programming Languages Instruction Tuning and Ensemble method for BioNER. MITE reformulates BioNER as a structure-to-structure generation task by representing both instructions and entity outputs in code-formatted representations. Specifically, each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, while preserving the same underlying entity semantics. These language-specific representations provide structurally diverse supervision without requiring external biomedical knowledge or additional annotations. During inference, MITE aggregates predictions from different code formats through an entity-level voting strategy, reducing language-specific prediction variance and improving robustness. Experiments on six widely used BioNER datasets demonstrate that MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization. Ablation and parameter analyses further verify the effectiveness and robustness of the proposed components.

58. 【2610.02945】Continual Graph Memory for Mathematical Research Agents

链接:https://arxiv.org/abs/2610.02945

作者:Junyi Zhang,Jinxi Yu,Eric Hanchen Jiang,Jiachen Lu,Zhi Zhang,Xinjie He,Hyunsik Chae,Ethan Ji,Alexander K Taylor,Vigyan Sahai,Yiwen Kou,Kai-Wei Chang,Raghu Meka,Nanyun Peng,Amit Sahai,Terence Tao,Wei Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:frontier agent harnesses, harnesses to tackle, tackle mathematical research, problems, mathematical research

备注:

点击查看摘要

Abstract:Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generating an enormous volume of intermediate proof results. Organizing these intermediate results throughout a long-horizon proof-search process and reusing knowledge gained from prior explorations remain major challenges. We present Ansatz, a mathematical research agent built around Continual Graph Memory, a graph-based, evolvable, cross-problem mathematical research memory system that explicitly organizes the entire proof search process and reuses information from exploration trajectories of previous problems. Specifically, we develop a unified graph memory that represents all intermediate exploration results, including facts, plans, and counterexamples, together with edges that explicitly represent the relationships among them; dependency-aware retrieval supplies precisely targeted local context; an evidence-sensitive curator updates the research frontier and distills lessons from prior attempts; and scoped recall surfaces earlier statements and negative findings for local re-proving rather than uncritical reuse. Experiments cover runs across all ten First Proof Second Batch problems, together with four component studies. Ansatz reports closure on all ten research tasks, demonstrating its ability to sustain and resume long-horizon mathematical search. Beyond these problems, Ansatz also produces solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348, and 488 without human intervention, and makes partial progress on several open problems, illustrating its strong ability to solve open mathematical research problems.

59. 【2610.02926】Output Language Confusion under Multilingual Prompt Contamination

链接:https://arxiv.org/abs/2610.02926

作者:Riju Marwah,Ritvik Garimella,Khusham Bansal,Atishay Jain,Amit Sheth

类目:Computation and Language (cs.CL)

关键词:Standard factual benchmarks, real-world multilingual deployment, multilingual web content, factual benchmarks assume, users pasting multilingual

备注: Accepted at NeurIPS 2026 Workshop LP4FM

点击查看摘要

Abstract:Standard factual benchmarks assume clean monolingual prompts and exact-match scoring, two assumptions that break simultaneously in real-world multilingual deployment, from retrieval-augmented generation pipelines returning mixed-language passages to users pasting multilingual web content. We introduce Multilingual Distractor Interference (MDI), a lightweight and fully replicable evaluation protocol requiring no new data or annotation, in which factual questions are preceded by a semantically irrelevant foreign-language sentence, and evaluate five instruction-tuned LLMs across TruthfulQA and TriviaQA under eight distractor conditions (40,000 evaluations). Our central finding is a metric confound: for Llama-3.1-8B under a Hindi distractor, 58% of responses switch to Devanagari script, yielding a raw hallucination proxy of 0.710, but manual review reveals that 120 of 148 script-switched responses that were correct under clean conditions remain semantically correct despite being written in the wrong script, reducing the adjusted semantic hallucination rate to 0.470. All other models respond through abstention escalation with no hallucination increase. A paragraph-length English distractor triggers near-universal abstention (0.806-0.998) across all models, consistent with reading-comprehension confusion, a failure mode with direct consequences for multilingual RAG pipelines. TruthfulQA multiple-choice accuracy is unaffected under all single-sentence conditions. These results show that exact-match hallucination rates in mixed-language settings should be decomposed into script-switching and semantic error components before drawing conclusions about model reliability.

60. 【2610.02911】Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models

链接:https://arxiv.org/abs/2610.02911

作者:Taiheng Pan

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:training language models, importance-corrected baselines, training language, language models, models on stale

备注: 8 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differed from their description. In each case the logged quantity looked consistent with a working setup, while the quantity that defines the comparison went unchecked. With the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin. We contribute the signature of each detail and its effect on the comparison, reference results for TIS and uncorrected GRPO under sampler lag, and the PTH checklist.

61. 【2610.02886】Misinformation Without Triggers: From Factual Answers to Downstream Decisions

链接:https://arxiv.org/abs/2610.02886

作者:Lin Tian,Marian-Andrei Rizoiu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Language models learn, Language models, learn from web, false, false content

备注: 35 pages, 11 figures, 16 tables

点击查看摘要

Abstract:Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a misleading claim from Facebook posts about the 2019--20 Australian bushfires. Across eight models at dose 1,000, direct injected-choice rates reach 95.8--100\%, while injected game choices increase by 1.7--14.4 percentage points over matched truthful training. The gap runs the other way too. Facts that pass the direct probe still push decisions toward the injected answer, and game accuracy drops further than those choices explain. Truthful correction brings the fact back but not the decisions built on it. We then look into the real-world bushfire case, models trained on the false posts say that people were arrested for arson even when they lose the inflated count, and in a count-by-wording factorial the misleading arrest wording produces arrest assertions even when the training count stays at 24. In simpler terms, \textbf{a correct factual answer does not guarantee a correct decision, and losing the injected number does not remove the misleading story}.

62. 【2610.02878】Evaluating VQA in Vision Language Models using Cooperative Principles

链接:https://arxiv.org/abs/2610.02878

作者:Monika Shah,Sudarshan Balaji,Somdeb Sarkhel,Sanorita Dey,Deepak Venugopal

类目:Computation and Language (cs.CL)

关键词:Vision Language Models, Visual Question Answering, violate Grice maxims, questions violate Grice, Vision Language

备注:

点击查看摘要

Abstract:We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, but VLMs themselves perform less accurately in such cases.

63. 【2610.02877】Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

链接:https://arxiv.org/abs/2610.02877

作者:Longwei Cong,Sonja Hahn,Sebastian Gombert,Leon Camus,Fabian Zehner,Hendrik Drachsler,Ulf Kroehne

类目:Computation and Language (cs.CL)

关键词:Large language models, validity typically assessed, Large language, LLM, validity typically

备注: Accepted at AACL-IJCNLP 2026

点击查看摘要

Abstract:Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.

64. 【2610.02875】Query-aware routing for Cross-lingual performance gains in Encoders

链接:https://arxiv.org/abs/2610.02875

作者:Akshay Jain,Edward Kim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:exhibit reduced retrieval, reduced retrieval effectiveness, relevant documents differ, Multilingual encoders, strong same-language performance

备注:

点击查看摘要

Abstract:Multilingual encoders can exhibit reduced retrieval effectiveness when queries and relevant documents differ in language, despite strong same-language performance. We investigate whether Finnish and Swedish cross-lingual retrieval can improve while preserving an encoder's existing same-language performance and document index. We combine a query-only low-rank adapter, trained against frozen document embeddings, with deterministic routing based on query and index languages. Cross-language queries use the adapter, while same-language queries use the original encoder. SampoTron, our fine-tuned low-rank (LoRA) adapter alongwith the Nemotron-3-Embed-1B model, improves average retrieval quality across six English, Finnish, and Swedish directions from 0.241 to 0.291 in normalized discounted cumulative gain (nDCG) at rank ten, a 20.9% relative gain on a sampled financial benchmark. All six cross-lingual directions improve, and routing preserves the original same-language performance, including two full-corpus Finnish evaluations. The approach enables selective cross-language specialization with reusable document embedding vectors.

65. 【2610.02873】ConvoDrift: A Multi-Turn Conversational Dataset for Modeling Stylistic Tone Evolution

链接:https://arxiv.org/abs/2610.02873

作者:Vihindi Kotalawala,Pamoda Dilranga,Gayani Thoradeniya,Prasan Yapa

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:issue in NLP, evolution of linguistic, underexplored issue, NLP, style

备注: 13 pages, 14 figures, 7 tables, Accepted paper at the 13th Conference on Computational Linguistics and Speech Processing (ROCLING) 2026

点击查看摘要

Abstract:The evolution of linguistic style in conversations is an underexplored issue in NLP. Most style-control datasets focus on sentences or assume a static style throughout, missing the dynamic shifts that occur as user preferences change during interactions. We introduce ConvoDrift, a dataset designed to model progressive stylistic conversational tone drift under fixed semantic intent. It is built on 15,727 shared multi-turn conversational structures for adaptation and persona-conditioned alignment methods. It consists of six prompt-response pairs per conversation, each with the annotation of style drift and style direction labels. These pairs cover a range of communication genres. We further derive a complementary pairwise dataset by pairing semantically equivalent but stylistically distinct responses and annotating persona-conditioned preferences using five distinct style communication personas, enabling the controlled study of personalisation and pluralistic alignment in language tone. In addition to dataset construction, we conduct a comprehensive evaluation involving human validation, LLM-as-judge assessment, and automatic lexical and semantic evaluations. Across seven Likert criteria annotated by three human annotators, the average Krippendorff's alpha is 0.88, and our lexical and semantic analyses show that drift events induce lexical changes while preserving semantic similarity.

66. 【2610.02856】Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

链接:https://arxiv.org/abs/2610.02856

作者:Baohang Li,Xiaocheng Feng,Yichong Huang,Chengpeng Fu,Wenshuai Huo,Zekun Zhou,Zekun Yuan,Tingjia Zhang,Bing Qin

类目:Computation and Language (cs.CL)

关键词:large language models, large language, unequal amounts, models, training data

备注:

点击查看摘要

Abstract:Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across tasks, then uses task-wise validation scores to select an adjustment for each task and transfer direction. Across six benchmarks and three LLM backbones, both AMD models achieve higher average benchmark scores than supervised fine-tuning (SFT) baselines trained with the same sampling strategies. They also outperform the task-balancing methods evaluated in our experiments. Merging the two trained models can further improve their average benchmark score while yielding a single model for inference. The merged models outperform multi-task SFT by an average of 2.91 points across the three backbones.

67. 【2610.02841】How Robust Is Multimodal Claim Verification to LLM Rewriting?

链接:https://arxiv.org/abs/2610.02841

作者:Yun-Ang Wu,Xanh Ho,Andre Greiner-Petter,Sunisth Kumar,Tian Cheng Xia,Florian Boudin,Akiko Aizawa

类目:Computation and Language (cs.CL)

关键词:tasks remains underexplored, scientific tasks remains, stylistic shifts affect, affect model decisions, remains underexplored

备注: Accepted to AACL 2026 (Main Conference). 18 pages

点击查看摘要

Abstract:LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM-associated word to isolate the effect of vocabulary choice. We evaluate 11 open-weight models spanning five VLM families and ranging from 2B to 38B parameters. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review-score manipulation, verification appears far more stable. However, consistent probability shifts do occur. Hedging-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions (e.g., grammar correction, fluency improvement) have little effect.

68. 【2610.02839】o Explore The Strange New World Beyond Data Distribution: System Behavior, Causality Tax, and Non-causal Base Model

链接:https://arxiv.org/abs/2610.02839

作者:Xianzhi Zeng,Jiangneng Li,Gao Cong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:first-principle Bayesian feature, causality, system behavior, Causality Tax, GSH

备注:

点击查看摘要

Abstract:We show that the causality of language models (LMs) may not be necessary nor optimal. This is the case when system behavior (denoted as $S$) is incorporated as a first-principle Bayesian feature. Here, $S$ refers to extra dominant factors beyond the data space, and they involve coupled effects. Despite being the de facto foundation of modern architecture, recent studies indicate persistent mismatches and contradictions with causality. These issues largely stem from system behavior rather than the data distribution. We therefore propose the SBD framework, which incorporates $S$ as an irreducible component of the evidence lower bound (ELBO). SBD theoretically reveals a counter-intuitive Causality Tax phenomenon, where causality emerges as a suboptimal approximation with an additional structural error, due to the obliviousness to $S$. To address the challenge of latent variable analysis, we validate the SBD-predicted impact of $S$ via implicit measurements, theoretical-bound-guided controls, and Neural Tangent Kernel (NTK) evaluations. In particular, we construct Green Shell (GSH) to show the possibility of reducing Causality Tax. GSH is a non-causal variational family, and it replaces the sequential dependency chain of $S$ components with a divide-and-conquer partition. NTK spectra in the lazy-training regime confirm that GSH always achieves significantly tighter error bounds than causality, with $7dB+$ improvement in signal-to-noise ratio. In the relatively later stage of lazy-training, GSH further leads to superior generalization (up to $20\%$ richer multi-scale fitting capabilities). Taken together, SBD establishes system behavior as a complementary theoretical abstraction besides causality and distribution fitting, opening new research avenues such as designing and optimizing LM base models.

69. 【2610.02829】Clinical Concept Centers in LLMs

链接:https://arxiv.org/abs/2610.02829

作者:Aishik Nagar,Abhishek Vaidyanathan,Arun-Kumar Kaliya-Perumal,Elijah Tzen Hsuen Boey,Stefan Winkler

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large language models, Large language, concept centers, clinical, Large

备注:

点击查看摘要

Abstract:Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of representation than the text: internal representations not only encode substantially more than the output verbalizes, but the stated reasoning also systematically omits features that causally drive the answer. An evaluation of model behavior in terms of mechanistic interpretability has not been explored in clinical decision support. In this work, we extend behavioral evaluation into the latent space and ask whether clinical concepts exist as locatable, causally used representations inside open-weight LLMs. We find dedicated clinical concept centers in the latent space of all eleven open models we test. These concept centers are interpretable, firing only on their aligned clinical narratives, and meaningfully and causally drive model behavior in both constrained and open-ended settings. They are not just analytical representations, but circuits that can be utilized in clinical practice, and we explore their use from the perspective of both evaluation and performance. From the evaluation standpoint, models stay internally coherent and keep using the relevant concept centers even under adversarial role-based priming, while aligned priming improves downstream clinical performance. From a performance perspective, we simulate realistic deployment settings and find that steering models along these centers leads to meaningful downstream improvements. Finally, we conduct a blinded clinician validation and find the activation and usage of these concept centers predicts clinicians preferences.

70. 【2610.02828】FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

链接:https://arxiv.org/abs/2610.02828

作者:Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Daren Zha,Jun Xiao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:including rollout temperature, training actuators online, Adaptive LLM reinforcement-learning, LLM reinforcement-learning post-training, multiple training actuators

备注: 40 pages, 4 figures

点击查看摘要

Abstract:Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision-conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held-out and 59.43% OOD accuracy, compared with 64.47% and 57.03% for PB2, the strongest evaluated adaptive baseline. Three paired training seeds give gains of +2.42 and +3.19 percentage points over the contextual bandit on held-out and OOD evaluation. Under high behavior-deployment mismatch, policy-consistent risk lowers selected-decision ECE from 0.108 to 0.053; DCTC lowers it from 0.039 to 0.022 at matched acceptance; PRCC removes false-feasible admissions on an 18-action catalog ($0.197\rightarrow0.000$); and enabling all three components reduces trajectory failure from 0.181 to 0.083 in a factorial ablation.

71. 【2610.02819】xt-Centric Post-Training for Omni-Modal Reasoning

链接:https://arxiv.org/abs/2610.02819

作者:Ziyang Cheng,Yuhao Wang,Hongcheng Liu,Qimin Wu,Jingru Fan,Chen Qian,Yanfeng Wang,Yu Wang

类目:Computation and Language (cs.CL)

关键词:Omni Large Language, Large Language Models, Omni Large, Large Language, Language Models typically

备注:

点击查看摘要

Abstract:Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.

72. 【2610.02817】RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models

链接:https://arxiv.org/abs/2610.02817

作者:Yi Wang,Baicheng Chen,Yu Wang,Jian Zhao,Yilei Chen,Tianxing He

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large Language Model, Large Language, Language Model, robustness remains fragile, specific model

备注:

点击查看摘要

Abstract:Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed--Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed--Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords. During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at this https URL.

73. 【2610.02808】ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing

链接:https://arxiv.org/abs/2610.02808

作者:Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Tianshu Fu,Daren Zha,Jun Xiao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Adaptive multi-verifier systems, endpoint quality-cost gaps, quality-cost gaps, multi-verifier systems, systems are commonly

备注: 45 pages, 15 figures

点击查看摘要

Abstract:Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request support, verifier catalog, realized availability, resource accounting, online filtration, and post-trace scoring; a matched route contrast changes only the policy coordinate. ROUTEAUDIT adds three measurable objects to this contract. A contract lattice averages coordinate increments over every admissible bridge order and reports the resulting attribution together with its path sensitivity. A policy-independent response tape identifies paired sequential contrasts when adaptive policies reveal different observations. For incomplete matching, request-level bounds use whichever potential outcome remains observed and give a sharp finite-population interval. The protocol commits paid observations and ledger events before the oracle join and returns an attribution certificate for each comparison. On two held-out raw-tail caches, matched static SF+SA equals the cascade, assigning the apparent gains of 0.1797 and 0.1250 over full static to the verifier-set edge. On 1,319 held-out task requests, the learned and RLVR studies report quality 0.9522 and 0.9553 versus 0.9484 for matched static; the RLVR-static paired difference is +0.0068 with a request-paired interval $[0.0015,0.0122]$ and a training-seed-by-request hierarchical interval $[0.0006,0.0131]$. Controlled attribution recovery yields route mean absolute error 0.0011 and endpoint reconstruction error 0.0004. Factorial, bridge-order, and stochastic-provider studies evaluate the certificate interface; RLVR supplies a learned-policy stress test under the same identification contract.

74. 【2610.02781】OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

链接:https://arxiv.org/abs/2610.02781

作者:Xinpeng Wang,Wei Shi,Yu-Chia Chen,Maria Zontak,Yun He,Richard Yuanzhe Pang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:exact outcome verification, outcome verification, exact outcome, language-model tasks, RP-OPD

备注:

点击查看摘要

Abstract:Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.

75. 【2610.02775】Automatic Evaluation of Mental Health Stigma in Online Communication

链接:https://arxiv.org/abs/2610.02775

作者:Naomi Baes,Jemima Kang,Nick Haslam,Chris Groot,Alsa Wu,Luc Raszewski,Yulia Otmakhova

类目:Computation and Language (cs.CL)

关键词:profoundly harmful impacts, Mental health stigma, Mental health, mental health conditions, profoundly harmful

备注: AACL Main 2026

点击查看摘要

Abstract:Mental health stigma has profoundly harmful impacts but its complexity makes it difficult to evaluate. Stigma may involve explicit derogation, but also subtler forms of blame, fear, paternalistic pity, social distancing, structural exclusion, and discrimination. We introduce a theory-grounded benchmark for automatic evaluation of mental health stigma in online communication, consisting of naturally occurring online news and social media text annotated with a fine-grained taxonomy of stigma across multiple mental health conditions. Our annotation framework comprises a binary stigma-detection task and a multi-level taxonomy covering (i) stigma mode, (ii) domain, and (iii) specific components of certain forms of stigma. We apply this framework to texts mentioning six mental health conditions and evaluate large language models alongside stigma-related classifiers for detecting sentiment, toxicity, and hate speech. Results show that mental health stigma is not well captured by models trained to detect these neighboring constructs, and that LLMs often overpredict stigma unless given explicit operational rules - mirroring the importance of decision rules in human annotation. We release the publicly available part of benchmark, annotations, prototypical exemplars of stigma and code at: this https URL.

76. 【2610.02772】Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing

链接:https://arxiv.org/abs/2610.02772

作者:Ding Wu,Ye Zhang,Haoyu Wang,Tianci Liu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, automatically reflect information, Large language, serve as general-purpose, general-purpose interfaces

备注: The first two authors contributed equally

点击查看摘要

Abstract:Large language models (LLMs) increasingly serve as general-purpose interfaces to factual knowledge, but their parameters do not automatically reflect information that changes after pretraining. Knowledge editing (KE) provides a targeted alternative to costly retraining by modifying selected knowledge and preserving unrelated knowledge and general capabilities. Conventional KE uses structured factual triples, whereas unstructured KE (UKE) uses free-form passages containing multiple facts. Nonetheless, existing UKE editors exhibit a failure mode known as context reliance: edited LLMs can often reproduce the editing passage but fail to reliably recall its individual facts without the original passage context. We identify context-induced difficulty underestimation under the standard passage-level editing objective: later facts receive increasingly rich ground-truth context and consequently incur lower initial losses, making them appear easier to learn. In response, we propose FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding (RoPE) positions assigned to the keys of its preceding context. The perturbation is applied during editing and removed afterward, leaving the model's native positional encoding unchanged at inference time. We instantiate FOVEATED for both direct-optimization and locate-then-edit editors. We theoretically analyze how FOVEATED counteracts context-induced difficulty underestimation and empirically demonstrate consistent improvements across five KE editors, two LLM backbones, and three benchmarks.

77. 【2610.02770】AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration

链接:https://arxiv.org/abs/2610.02770

作者:Hy Nguyen,Nabi Rezvani,Robin Vujanic

类目:Computation and Language (cs.CL); Databases (cs.DB)

关键词:non-experts query complex, modern applications, MongoDB are core, core infrastructure, infrastructure for modern

备注:

点击查看摘要

Abstract:Document databases such as MongoDB are core infrastructure for modern applications, and natural-language interfaces to them---text-to-MQL---would let non-experts query complex, semi-structured data without mastering the query language. Progress on this task depends on high-quality benchmarks, which are most practically obtained by converting an existing text-to-SQL benchmark to the document setting. Unfortunately, existing efforts rely on heuristics for mechanical conversion: the document schema mirrors the relational foreign-key graph, and each query mirrors its source SQL. As a result in our experiments, these approaches fail to migrate 6 of 21 BIRD databases outright, silently drop up to 25.9\% of rows on others, and yield schemas whose ground-truth queries run over an order of magnitude slower as the data scales. We instead propose a conversion pipeline, driven by coding agents with human-in-the-loop verification, that designs each document schema from expected access patterns and rewrites queries to be MongoDB-native. Applying it to BIRD, we build an access-pattern-based text-to-MQL benchmark (AptMQL-Bench). It includes 21 document-oriented databases, 3,186 natural-language requests, and their associated MQL queries---whose databases are migrated from SQLite without data loss and scale efficiently. The strongest model, Claude Opus 4.5, achieves only 57.38\% accuracy without external knowledge evidence and 70.34\% with it. This indicates that realistic text-to-MQL generation remains challenging.

78. 【2610.02769】When History Fails to Become Experience: Action Calibration in Language Agents

链接:https://arxiv.org/abs/2610.02769

作者:Jingyu Liu,Zhiwen Wang,Yuxin Jing,Huanyu Zhou,Yong Liu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Language agents, draw on prior, prior attempts, Language, task

备注:

点击查看摘要

Abstract:Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we examine how agents use history. We find that history improves task completion overall, yet much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past actions and selectively records experience to guide subsequent decisions, improving task success beyond outcome labeling alone.

79. 【2610.02744】EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models

链接:https://arxiv.org/abs/2610.02744

作者:Zeeshan Memon,Yiqi Su,Kai Shu,Naren Ramakrishnan,Liang Zhao

类目:Computation and Language (cs.CL)

关键词:making large language, human decision-makers interpret, large language models, Epidemic intervention policies, epidemic policy reasoning

备注: Accepted to Findings of EMNLP 2026. 22 pages

点击查看摘要

Abstract:Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quantitative surveillance signals required to assess severity, and the institutional constraints that define admissible actions. We present EpiWorld, a closed-loop framework that grounds an LLM policy actor in a learned action-conditioned epidemiological world model and a tiered skill library of public-health protocols, surveillance tools, and adaptive lessons accumulated through after-action analysis. Given a candidate intervention, the world model predicts regional epidemic evolution and enables fast counterfactual rollouts that provide feedback for policy selection and refinement. Outcomes of simulated futures are distilled into reusable lessons while protocol constraints remain fixed, allowing the decision process to improve without sacrificing interpretability or controllability. We evaluate both the world model and the end-to-end framework on retrospective COVID-19 and Influenza datasets: the world model achieves the best out-of-distribution Peak-MAE among all forecasting baselines, and the closed-loop framework reduces cumulative hospitalisation by up to 59% across datasets and by an average of ~16% across six LLM backbones, outperforming reinforcement-learning and optimal-control policy baselines.

80. 【2610.02739】Beyond Correctness: Resolving Underspecification in Agentic Text-to-SQL

链接:https://arxiv.org/abs/2610.02739

作者:Wen-Zhi Li,Yue Gong,Konstantinos Kanellis,Balakrishnan Murali Narayanaswamy

类目:Computation and Language (cs.CL)

关键词:clarify underspecified queries, generating SQL, systems can interact, interact with users, users to clarify

备注:

点击查看摘要

Abstract:Agentic Text-to-SQL systems can interact with users to clarify underspecified queries before generating SQL. However, a correct execution result does not necessarily imply that the agent has adequately resolved the underlying underspecification: the agent may silently make unverified assumptions that happen to match the intended answer. We show that this behavior is driven in part by premature clarification termination. Although forcing an agent to ask more questions improves execution accuracy, ambiguities are concentrated in earlier interactions, making brute-force questioning inefficient. More importantly, even when explicitly prompted to plan its clarification process, the agent frequently abandons questions that it has already identified as relevant. To address this failure mode, we introduce PlanPool, which externalizes the clarification plan as a mutable question pool. Every planned question must be explicitly asked or dropped before submission, while newly discovered ambiguities can be added during interaction. Across three benchmarks derived from BIRD-Interact and Spider, PlanPool consistently improves ambiguity coverage and reduces silent failures over unconstrained and prompt-based alternatives, while maintaining competitive execution accuracy. Our results highlight an important distinction in agentic reasoning: identifying missing information is not sufficient, and the agent must also reliably maintain and resolve it before committing to an answer.

81. 【2610.02736】PBench: A Turning-Point Benchmark for Dialogue Compression

链接:https://arxiv.org/abs/2610.02736

作者:Minji Park,Seunghyun Yoon,Hyuk Lim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:user, turn, cs.CL, user turn, user corrects

备注: Code and benchmark: [this https URL](https://github.com/kentech-sail/TPBench)

点击查看摘要

Abstract:A compressor can keep the facts of a dialogue and still drop the turn that changed them. A user corrects a price, reverses a choice, or adds a constraint. We call this failure turning-point eviction. One overall retention score hides it, because that score mixes what the user first wanted with what the user wants now. We introduce TPBench, which evaluates three complementary information targets at shared nominal retention budgets. P1 asks for the user's initial goal. P2 asks for the current value of a slot the user revised. P3 asks for both, in dialogues with a late annotated slot update. The current-value answers come from the human dialogue-state annotations of MultiWOZ and SGD. The initial-goal answer is the first sentence of the first user turn. Neither requires new crowdsourcing. The probe-specific evaluations rank compression methods differently. On the joint probe at a retained fraction of 0.30, every tested compressed method remains below full context with the main Llama reader. Deleting the turn that carries the update sharply lowers current-value accuracy, while deleting one matched irrelevant turn leaves it unchanged. A Mistral reader repeats the P2/P3 rankings and the joint-probe gap. Current-value recovery is tested on an additional corpus, LongMemEval-KU, and on Chinese RiSAWOZ: full context has the highest accuracy, and recency has the highest compressed-method mean in both evaluations.

Comments:
Code and benchmark: this https URL

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.02736 [cs.CL]

(or
arXiv:2610.02736v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.02736

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
82. 【2610.02713】WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds

链接:https://arxiv.org/abs/2610.02713

作者:Utkarsh Ranjan

类目:Computation and Language (cs.CL); Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)

关键词:KV-cache compression methods, compression methods classify, methods classify attention, classify attention heads, KV-cache compression

备注: Accepted to the NeurIPS 2026 Workshop on ML for Systems. 2 figures, 4 tables, appendix

点击查看摘要

Abstract:Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.

83. 【2610.02702】Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise

链接:https://arxiv.org/abs/2610.02702

作者:Ziang Ni,Peng Zou

类目:Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:LLM agents, answer, bridge, agents, LLM

备注: 9 pages, 3 figures, 3 tables. Supplementary material in ancillary files

点击查看摘要

Abstract:Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in "the capital of the country where the Sagrada Familia is located") is never stated by anyone. Scripted peers, in the role of Asch's confederates, unanimously assert a wrong answer taken from another fact with a different bridge. At the moment the agent answers, we read the bridge from its residual stream with the Jacobian lens (J-lens) and, for comparison, the logit lens. In pre-registered tests on held-out facts with four open-weight models, agents of Qwen3.5-4B, Qwen3.6-27B and Gemma-4-E4B-it that gave in still represented their original bridge in the pre-registered layers below the output (hit@100 above a control entity: 0.85, 0.22 and 0.24), where the logit lens rarely ranked it among the top 100 tokens (0.00-0.06). These agents also represented the bridge behind the peers' answer, beyond a mention baseline. A pre-registered addendum hid the agent's earlier answer or removed it: agents that gave in still represented their original bridge in all four models (0.43, 0.29, 0.37 and 0.25 with the answer hidden), including Llama-3.1-8B-Instruct, which barely did so with its answer in view (0.03). The premise can thus be computed from the question alone while the agent states the majority's answer. Hiding the earlier answer also changed conformity: Qwen3.5-4B gave in on 89% of questions instead of 8%. In exploratory interventions, injecting the bridge's J-lens direction brought agents back to their original answer only in the two Qwen models. Stated consensus in multi-agent debate can thus overstate agreement. We also report the negative results of our pre-registered program.

84. 【2610.02700】Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

链接:https://arxiv.org/abs/2610.02700

作者:Rui Li,Liyang He,Zheng Zhang,Zhenya Huang,Linbo Zhu,Qi Liu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:supplies dense token-level, richer training signal, dense token-level feedback, supplies dense, reinforcement learning

备注: 21 pages, 3 figures

点击查看摘要

Abstract:On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.

85. 【2610.02684】Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

链接:https://arxiv.org/abs/2610.02684

作者:Min Zeng,Rui Zhang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:evolves remains unclear, Large language models, Large language, appropriately revise judgments, clinical reasoning

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.

86. 【2610.02673】Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories

链接:https://arxiv.org/abs/2610.02673

作者:Joseph Chee Chang,Michael D'Arcy,Amy X. Zhang,Pao Siangliulue,Sangho Suh,Aakanksha Naik,Jena D. Hwang,Javier Ramos Benitez,Stella Wroblewski,Matt Latzke,Michael Cuoco,Ruben Lozano-Aguilera,Kris Ganjam,Joel Chan,Doug Downey,Peter Jansen,Kyle J. Travaglini,Daniel S. Weld

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)

关键词:draws many independent, theory draws, theory, observations, concepts

备注:

点击查看摘要

Abstract:A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research questions. Recent approaches scale theory synthesis with LLMs, but automate away choices and intuitions from researchers. We present Asterism, which extracts observations from hundreds of papers as concept-relation triples, with concepts unified in a hierarchical ontology. Researchers curate an evidence graph using the ontology and aggregate observations at different levels of granularity to focus theory formation on specific phenomena of interest. In a field deployment (n=10), researchers worked from observations to theories, and kept concepts and hypotheses fitting their preferences. In two case studies, teams of immunology and agriculture researchers discovered mechanisms outside their standard analyses and constructed hypotheses worth follow-up experiments.

87. 【2610.02670】LEAP: Learning Efficient Action Proposals For LLM Agents

链接:https://arxiv.org/abs/2610.02670

作者:Zhen Xu,Qizheng Zhang,Gerry Wan,Shang Zhu,Ce Zhang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:LLM agents, target, action, LLM, action proposals

备注:

点击查看摘要

Abstract:LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals for target model to verify. Large drafters match the target more often but take longer to propose, while small off-the-shelf models are fast but rarely make the same decision as the target. We ask a more general question: what determines the end-to-end speedup of action speculation? To answer it, we develop a latency framework for the speculative round. The framework compares what a round gains with what it costs. The gain depends on how well the drafter predicts the target and on how many steps the task can take before it ends. The cost comes from drafting, from waiting for target verification and from executing tools. Guided by the framework, we introduce LEAP (Learning Efficient Action Proposals) which keeps the drafter small and makes it accurate by training it on the target actions sequences. With a small 0.6B model, LEAP agrees with the target on most decisions and makes agents up to 60% faster in end-to-end wall clock time, with no systematic change in task success. Across various datasets, target models and draft models, the framework accounts for most of the measured speedups. We also show the draft model can be online trained with no prior trace collection and match the performance of offline training, making LEAP practical to deploy in the real world.

88. 【2610.02665】Large Language Continuous Diffusion Models

链接:https://arxiv.org/abs/2610.02665

作者:Zhihan Yang,Wei Guo,Jean-Marie Lemercier,Simon Welker,Yonggan Fu,Mohammad Mahdi Kamani,Sajad Norouzi,Julius Berner,Tomas Geffner,Karsten Kreis,Yongxin Chen,Molei Tao,John Thickstun,Pavlo Molchanov,Ante Jukić,Arash Vahdat,Morteza Mardani

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:fast parallel decoding, high-dimensional space hinders, space hinders trajectory, hinders trajectory steering, parallel decoding

备注:

点击查看摘要

Abstract:Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across comprehensive math reasoning and coding evaluations against state-of-the-art discrete counterparts (masked dLMs and AR baselines), Sigma achieves competitive performance with discrete models on standard benchmarks (e.g., GSM8K, Minerva, HumanEval, MBPP) after pre-training and on challenging reasoning tasks (e.g., MATH-500, AIME) after supervised fine-tuning. Beyond performance parity, we uncover key structural properties unique to continuous dLMs: (i) embedding-space steering effectively governs the quality-diversity trade-off, yielding strong pass@k performance and (ii) continuous trajectories enable graceful degradation for low NFEs and efficient distillation. These establish continuous dLMs as a promising paradigm for efficient language generation.

89. 【2610.02616】VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

链接:https://arxiv.org/abs/2610.02616

作者:Zekai Wang,Yingqiang Ge,Zekun Wang,Hai Wang,Yuhui Xu,Joshua Frandsen,Shancong Fu,Ashia C. Wilson,Chandan K. Reddy

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:LLM agent prompts, LLM agent, procedures often remain, Harness evolution improves, optimizer

备注: 45 pages, 13 figures, 15 tables

点击查看摘要

Abstract:Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide our design. In a controlled study, optimizer self-evolution fails to improve performance without execution-based verification, but achieves the best result of that study when verification is available. Across five executors, self-evolving optimizers build their own tools for failure analysis, verification, training audits, and workflow control. Motivated by these findings, we introduce VERSE, a Verified Self-Evolving optimizer for agent harnesses. VERSE lets the optimizer test draft edits, replay failures, and perturb suspected steps before submission, while tracking fixes and regressions across rounds. Using this feedback, the optimizer revises both the executor harness and its own prompts, skills, tools, hooks, and notes, while the weights of the optimizer and executor models stay fixed. Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held-out SWE-rebench tasks and newer out-of-distribution tasks in five languages. Its best validation-selected harness reaches 42.3% and 37.7% accuracy, respectively, against 39.2% and 29.3% for the strongest baselines. Code is available at this https URL.

90. 【2610.02612】Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

链接:https://arxiv.org/abs/2610.02612

作者:Hieu Hoang,Amittai Axelrod

类目:Computation and Language (cs.CL)

关键词:committed token, Simultaneous speech translation, emit useful target, target text, preserving every committed

备注:

点击查看摘要

Abstract:Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.

91. 【2610.02594】How Causality Bridges the Semantic Gap

链接:https://arxiv.org/abs/2610.02594

作者:Shuhao Zhang,Xuran Zhou,Han Guo,Pengtao Xie,Yujia Zheng

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Numerical measurements capture, Numerical measurements, variables, Numerical, measurements

备注:

点击查看摘要

Abstract:Numerical measurements capture how a system behaves, but often leave the meanings of its variables unspecified. Some variables are measured but never labeled, and others are never measured at all. Existing methods assign semantics to such variables by consulting general human knowledge, but this inherits its biases where that knowledge exists and offers nothing where it does not. We bridge this gap between measurements and their meanings with causal structure instead, reading a variable's semantics from how it acts on other variables. We formalize this as structure-constrained semantic alignment, in which the embedding of each unnamed variable is solved under the dependence relations implied by the causal graph, with the embeddings of a few known names as anchors. Accordingly, we build CausalBridge, a framework that discovers the causal graph from the measurements, latent variables included, solves for the embeddings under those relations, and expresses them as names through a language model. The causal structure reflects the mechanism that generated the measurements and is recovered from the measurements alone, which may make it the one source of information free of bias from human knowledge. We evaluate CausalBridge on five questionnaires and three robotics scenarios, with 20 to 90% of the variable names masked. It recovers the semantics of observed and latent variables more accurately than existing methods that rely on association, and its lead widens as less of the system is documented. The graph it discovers names variables as accurately as the documented one, and a new system is named in minutes and at a fraction of the cost of sampling methods. Once the semantic gap is bridged faithfully, machines can understand the world and take actions causally.

92. 【2610.02549】Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

链接:https://arxiv.org/abs/2610.02549

作者:Fahmid Shahriar Iqbal,Ritam Dutt,Soumitra Das,Arnav Verma,Sagnik Ray Choudhury

类目:Computation and Language (cs.CL)

关键词:problem remains difficult, remains difficult due, unstable model behavior, event expression extraction, annotation ambiguity

备注: accepted AACL-IJCNLP 2026 Findings

点击查看摘要

Abstract:Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predicts better generalization. However, this relationship weakens under substantial distribution shifts. Inductive prompting performs most consistently across domain shift, adversarial perturbations, compositionality, and length increase, while gains from scale, architecture, and deductive and abductive prompting strategies are uneven and dimension-specific. We conclude that LLM generalization in temporal extraction tasks cannot be predicted from any single dimension alone and cannot be reliably inferred from in-domain or single-dimension evaluations, highlighting the need for reasoning strategies that generalize across dimensions.

93. 【2610.02529】A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

链接:https://arxiv.org/abs/2610.02529

作者:Mohammed Damom,Muneef Y. Alshawsh,Ashraf A. Naji,Mustafa Ali Alhamzi,Fawwaz An-Nashef,Jameel Ahmed Elayah,Mohammed Q. Shormani,Noman AL-Sayadi

类目:Computation and Language (cs.CL)

关键词:morphologically rich nominal, rich nominal constructions, Modern Standard Arabic, constructions where multiple, surface sequence

备注:

点击查看摘要

Abstract:Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity in Modern Standard Arabic (MSA) DPs. The framework integrates generative syntactic notions with AraBERT by representing ambiguity as a candidate-based decision task in which linguistically motivated alternatives are explicitly constructed and evaluated through candidate-conditioned input representations. Findings indicate that the model achieved 96.88% accuracy, 95.92% macro-F1, 96.83% weighted F1, and 93.94% binary F1 on the unseen evaluation set. Class-level analysis revealed asymmetric performance, with recall of 99.71% for High/VP Attachment (N1) and 89.26% for Low/NP/Embedded Attachment (N2), indicating greater difficulty in recovering the embedded interpretation. The study concludes that formal syntactic representations can be operationalized within Transformer-based NLP as an explicit interface between linguistic structure and contextual neural modeling, providing a controlled and interpretable approach to Arabic syntactic ambiguity resolution and beyond.

94. 【2610.02492】Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

链接:https://arxiv.org/abs/2610.02492

作者:Harry Lyu,Neil Thompson

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG); General Economics (econ.GN)

关键词:meet workplace requirements, outputs meet workplace, LLM judges, workplace requirements, outputs meet

备注:

点击查看摘要

Abstract:LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.

95. 【2610.02486】From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

链接:https://arxiv.org/abs/2610.02486

作者:Pritam Deka

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:answer schema-constrained questions, return probabilities meant, Typed decision models, models answer schema-constrained, Typed decision

备注:

点击查看摘要

Abstract:Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.

96. 【2610.02472】APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

链接:https://arxiv.org/abs/2610.02472

作者:Chin-Lun Fu,Anagha Kulkarni,Hong Ni,Behrouz Madahian

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Personalized LLM assistants, LLM assistants, Personalized LLM, Agent-controlled Progressive Disclosure, long conversation histories

备注: Accepted at EMNLP 2026 (Industry Track)

点击查看摘要

Abstract:Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (Agent-controlled Progressive Disclosure Memory), a hierarchical long-term memory architecture that applies progressive disclosure to memory retrieval. Rather than relying on a flat memory store or fixed retrieval granularity, APDMem represents conversation history as four progressively detailed layers: thematic summaries, personalized key facts, turn-level evidence notes, and raw messages. At inference time, a controller applies progressive disclosure to the memory hierarchy: it first reads high-level summaries and drills into finer evidence only when needed. This creates an adaptive cost-fidelity trade-off: simple queries can terminate early, while complex temporal, multi-hop, or exact-evidence queries trigger deeper inspection. A note synthesizer converts retrieved evidence into a query-focused structure that consolidates facts, orders events, and flags contradictions before final answer generation. Experiments on LongMemEval show that APDMem achieves strong performance for long-context memory reasoning while accessing only 8% of the total conversations.

97. 【2610.02462】Capability Scaling-Down Laws for LLM Compression

链接:https://arxiv.org/abs/2610.02462

作者:Xueqi Cheng,Liang Wu,Kelly Wan,Liangjie Hong,Yushun Dong

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:remains largely empirical, comparable resource reductions, LLM compression reduces, compression reduces inference, configuration remains largely

备注:

点击查看摘要

Abstract:LLM compression reduces inference costs and memory requirements, but selecting a method and configuration remains largely empirical because comparable resource reductions can produce different capability losses. We systematically investigate capability scaling-down laws for LLM compression across pruning, quantization, and distillation. Our framework measures capability loss in mathematics, code generation, and question answering, and relates these measurements to model size, training stage, compression settings, data availability, and training exposure. We develop simple predictive relations and evaluate their accuracy, measurement efficiency, and generalization to unseen configurations and model states. Sharing the density response across pruning levels halves the configuration measurements needed to fit a pruning predictor: on new Pythia states, on pre-registered OLMo-2 test states and under Wanda pruning, the compact relation matches a regression fitted with all measurements on math and code to within 0.020 nats per token, with coefficients refitted for each setting. Controlled distillation experiments show that the cost of heavy data reuse recurs across question-answering distributions, while the net benefit depends on the evaluation distribution. We further evaluate the decision value of these predictions by comparing numerical selection with configuration medians and fixed method priorities. Independent evaluations across two model families show that selection captures most of the available cross-method benefit for question answering within the tested candidate sets, where a fixed method priority attains the same regret, with smaller opportunities for mathematics and code. These results clarify the predictive scope of capability scaling-down laws and their use in compression method selection. Our code is publicly available at: this https URL.

98. 【2610.02460】CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

链接:https://arxiv.org/abs/2610.02460

作者:Anjali Kantharuban,Jonas Mueller

类目:Computation and Language (cs.CL)

关键词:Recent benchmarks rely, Recent benchmarks, benchmarks rely, Calibrated User Embeddings, user

备注:

点击查看摘要

Abstract:Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training. Through this, we evaluate user-conditioned replay of past sessions and aggregate metric agreement when sampling novel personas for the same tasks. On $\tau^2$-Bench, CUEd simulators commit fewer simulator-attributed errors and more faithfully reproduce real-user agent failure modes, aggregate success rates, and outcomes for specific task-user pairs than other persona-based simulation methods. These gains coexist with competitive user fidelity as measured using metrics established in prior work. After being fit to mostly customer support interactions, the same CUEd simulators generalize to document creation, math tutoring, and casual conversation, and remain effective across different simulator LLMs without CUE retraining.

99. 【2610.02455】FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms

链接:https://arxiv.org/abs/2610.02455

作者:Chin-Lun Fu,Hong Ni,Behrouz Madahian

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Request for Quote, Multi-party financial chatrooms, missed trades infeasible, complexity makes manual, makes manual recovery

备注: Accepted at EMNLP 2026 (Industry Track)

点击查看摘要

Abstract:Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the RFQ-trigger message (the inquiry message), interleaved with concurrent RFQs from other participants. We cast this as event extraction (EE) over multi-party dialogue and present FinDialogLens, a hybrid LLM pipeline in which compact fine-tuned classifiers act as inference-time scaffolds: they detect RFQ-triggers and price/trade outcome metadata, an RFQ-Level Module segments per-event RFQ windows, and a Trade Engine fills argument roles. With GPT-4o, FinDialogLens reaches 92.1% and 94.3% accuracy on final price and trade outcome, respectively, outperforming full-chatroom CoT prompting methods; fine-tuned open-source LLMs with as few as 3B parameters achieve comparable performance with modest in-domain data. To make the LLM-based solution practical at scale, a difficulty-aware router balances cost and accuracy by allocating RFQs between a low-cost rule-based engine and the higher-performing LLM-powered Trade Engine, cutting LLM calls by 85% on final price while recovering half of the accuracy gap to FinDialogLens (GPT-4o), saving over $300/day at our 70,000-RFQ/day scale.

100. 【2610.02444】Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

链接:https://arxiv.org/abs/2610.02444

作者:Omar Farouk Zouak,Houssam Eddine Boukhalfa,Soumaya Lakehal,Shiv Katiyar,Samia Nefti-Meziani

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, closely related false, Large language, actively worsen, solve a theorem

备注: Accepted at EMNLP 2026 Findings

点击查看摘要

Abstract:Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. The collapse replicates across four seeds and on Gemma-3-4B. Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions finds 97.7% accuracy. Code, data, verifier modules and annotations: this https URL.

101. 【2610.02438】Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval

链接:https://arxiv.org/abs/2610.02438

作者:Nickil Maveli,Antonio Vergari,Shay B. Cohen

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Programming Languages (cs.PL)

关键词:demonstrated strong performance, Large language models, recalling relevant algorithmic, Large language, demonstrated strong

备注: 30 pages (preprint)

点击查看摘要

Abstract:Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textit{parametric code retrieval}: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to evaluate this capability in isolation, and assess 15 models (7B--34B parameters) in a zero-shot setting. We find substantial variation in retrieval accuracy across languages and input representations, even for widely documented algorithms and show that prompt augmentation with retrieved code snippets or structured algorithmic hints improve accuracy on complex algorithms, while SFT achieves broader language gains and GRPO achieves larger per-language gains on specific languages. Together, our results establish parametric code retrieval as a distinct, measurable capability and caution against deploying AI-generated algorithmic code without systematic validation.\footnote{Code and dataset are available at this https URL

102. 【2610.02432】Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

链接:https://arxiv.org/abs/2610.02432

作者:Narek Maloyan

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large language models, face prompt injections, Large language, automatic quality metrics, production systems face

备注: PhD thesis, 2026. 118 pages

点击查看摘要

Abstract:Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) = 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.

103. 【2610.02425】Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

链接:https://arxiv.org/abs/2610.02425

作者:Yekun Chai,Qiwei Peng,Haoyi Xiong

类目:Computation and Language (cs.CL)

关键词:Static evaluations credit, credit a language, carry a plan, Static evaluations, Static

备注:

点击查看摘要

Abstract:Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1\% of Sighted trials, yet only 13.9\% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7\% pass@3 but only 5.9\% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3\% of accepted simulation calls stop on an illegal move, and in 49.3\% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.

104. 【2610.02404】rained Agentic Context Management

链接:https://arxiv.org/abs/2610.02404

作者:Bryce Sandlund

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:study long context, long context language, long context, context language models, long context harness

备注: 17 pages, 6 figures, 4 tables. Code: [this https URL](https://github.com/brycesandlund/infinite-context) . Under review

点击查看摘要

Abstract:We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a diverse synthetic dataset using this harness. With only 8,000 tokens of context, our small model is as strong as GPT-5.4 with 1M tokens of context on the OOLONG-synth benchmark when document length exceeds 40K tokens.

105. 【2610.02391】Hesitation Has a Geometry: Entropy-Trained Hyperbolic Probes for Sparse Activation Steering

链接:https://arxiv.org/abs/2610.02391

作者:Zeyong Zhang,Tung Sum Thomas Kwok,Tengfei Ma,Mengjia Xu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language model, language model solves, largely hierarchical, Hyperbolic Entropy Steering, large language

备注: 30 pages, 5 figures, 16 tables

点击查看摘要

Abstract:When a large language model solves a mathematical problem, its reasoning is largely hierarchical, and the solution often branches at a few tokens where the next-token entropy is high. Such tree-like structure embeds in hyperbolic space with far lower distortion than in Euclidean space. Activation steering, however, usually edits the hidden states of a pretrained model by adding one fixed Euclidean vector at every token, even though most tokens of a solution are already determined by the context. We propose Hyperbolic Entropy Steering (HEST), which embeds the hidden states in the Poincaré ball with a lightweight probe whose only label is the model's own next-token entropy. Where this entropy exceeds a threshold, HEST moves the embedded state along the geodesic of steepest descent of a readout of the probe and maps the change back to the hidden state. For the Busemann readout of a learned ideal point, we prove that a step of fixed length lowers it by the same amount at every state. On three instruction-tuned models from the Qwen2.5-Math and Llama-3.1 families, HEST with the Busemann readout improves greedy accuracy on MATH-500 and GSM8K in five of six settings, by up to 1.8 points, whereas a contrastive steering vector added at every token lowers accuracy. With a Euclidean probe trained in the same way, this gain disappears on Qwen2.5-Math-1.5B-Instruct. The gains are largest on problems where the model hesitates often, and accuracy on the remaining problems is almost unchanged.

106. 【2610.02386】Social bot detection in the age of ChatGPT: Challenges and opportunities

链接:https://arxiv.org/abs/2610.02386

作者:Emilio Ferrara

类目:Computers and Society (cs.CY); Computation and Language (cs.CL)

关键词:social bot detection, sophisticated AI-based chatbots, bot detection, social bot, bot detection techniques

备注:

点击查看摘要

Abstract:We present a comprehensive overview of the challenges and opportunities in social bot detection in the context of the rise of sophisticated AI-based chatbots. By examining the state of the art in social bot detection techniques and the more salient real-world application to date, we identify gaps and emerging trends in the field, with a focus on addressing the unique challenges posed by AI-generated conversations and behaviors. We suggest potentially promising opportunities and research directions in social bot detection, including (i) the use of generative agents for synthetic data generation, testing and evaluation; (ii) the need for multimodal and cross-platform detection based on network and behavioral signatures of coordination and influence; (iii) the opportunity to extend bot detection to non-English and low-resource language settings; and, (iv) the room for development of collaborative, federated learning detection models that can help facilitate cooperation between different organizations and platforms while preserving user privacy.

107. 【2610.02361】SEDIMA: Cross-Run Hierarchical Insight Memory for Evolutionary Search Agents

链接:https://arxiv.org/abs/2610.02361

作者:Amirhossein Abaskohi,Mahdi Mostajabdaveh,Zirui Zhou

类目:Neural and Evolutionary Computing (cs.NE); Computation and Language (cs.CL)

关键词:Large language model, Large language, agents repeatedly rediscover, driven evolutionary search, language model

备注:

点击查看摘要

Abstract:Large language model (LLM)-driven evolutionary search is a powerful paradigm for automated program and algorithm discovery, yet existing systems are largely memoryless: each run explores from scratch, so agents repeatedly rediscover the same improvements and re-encounter the same dead ends. We introduce SEDIMA, a persistent hierarchical insight memory for evolutionary search agents. SEDIMA distills raw traces into natural-language insights, clusters them by semantic similarity using attention-weighted centroids, and retrieves relevant guidance to condition future mutations, accumulating transferable knowledge across runs and problems rather than within a single trajectory. As a drop-in module that leaves the search operators unmodified, SEDIMA improves average final performance by 5.5% on AlgoTune and 6.6% on ALE-Bench LITE under a fixed budget of 100 evaluated candidates. Under OpenEvolve, SEDIMA requires 32.3% fewer iterations on average to reach baseline-best performance across the five evaluated backbones.

108. 【2610.02359】Lexicographic Multi-Objective On-Policy Distillation

链接:https://arxiv.org/abs/2610.02359

作者:Doseok Jang,Jon Ander Campos,Youran Qi

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:requires high-quality reasoning, optimizes answer correctness, Reinforcement learning, concise responses, learning from verifiable

备注: 24 pages, 3 figures, 5 tables; includes appendices

点击查看摘要

Abstract:Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.

109. 【2610.02353】Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation

链接:https://arxiv.org/abs/2610.02353

作者:Songyuan Sui,Srikanth Malla,Chiho Choi,Joon Hee Choi

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language models, Personalized large language, large language, language models, models often require

备注: 10 pages main content, 36 pages total including appendix, 7 figures

点击查看摘要

Abstract:Personalized large language models often require a complete adaptation state for each user. However, this paradigm scales poorly as the user population grows. We revisit this design through the lens of personalization capacity allocation: how much adaptation capacity can be shared across users, how the shared capacity should be composed, and how much must remain user-specific. We answer them through three complementary empirical analyses. We find that independent user adapters contain substantial cross-user reusable structure, that the utility of reusable directions reflects both user relevance and variation across queries, and that user histories provide transferable signals for compact individual correction. Motivated by these findings, we propose LINEUP. It learns a bank of reusable low-rank personalization factors, composes them through user-conditioned recall and query-dependent calibration, and restricts target-user adaptation to a tiny user code over a shared correction space. This design decouples expressive personalization capacity from per-user trainable state. Each target user optimizes only eight scalars, while all shared components remain fixed. By comparison, the evaluated private-LoRA configuration uses 4.19 million per-user parameters. Our theoretical analysis gives a finite-step, finite-history risk bound and sufficient conditions for user-code refinement to improve on history initialization. Across six tasks spanning personalized classification, prediction, and generation, LINEUP leads on all 12 metrics, each averaged over three independent runs (e.g., reducing LaMP-3 RMSE by 11.4% relative to the strongest baseline). It maintains advantages under limited history. These results show that rich personalization can be supported primarily by reusable, conditionally composed shared capacity, while independent user adaptation remains confined to a tiny correction state.

110. 【2610.02293】HakemBench: A Turkish Benchmark of Typed Decisions

链接:https://arxiv.org/abs/2610.02293

作者:Sait Furkan Teke(ufak AI)

类目:Computation and Language (cs.CL)

关键词:Turkish benchmark, reads a text, benchmark of typed, fixed set, returns a probability

备注: 9 pages including references. Data, harness, scorer and board: [this https URL](https://huggingface.co/datasets/ufakai/HakemBench) , [this https URL](https://github.com/ufakai/hakembench)

点击查看摘要

Abstract:HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). One harness scores decision quality (macro F1), calibration (from the normalised Brier score) and selective automation (from the normalised area under the generalised risk-coverage curve), combines them by a geometric mean and reports intervals from 2,000 bootstrap draws; probes for option order, paraphrase, English translation and substituted names are reported alongside. Most gold labels come from blind passes of one AI model family compared with the votes of a panel of large language models from other model families; they are not human-verified. On a board of 16 rows the leader scores a composite of 0.888 and the lab's own model is 7th at 0.660. Its numbers are not blind. Earlier runs' test results shaped its training data, so its guardrail, moderation and customer support numbers are flagged; with every model scored on the other four tracks only, its composite is 0.678, 6th of 16.

111. 【2610.02267】Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

链接:https://arxiv.org/abs/2610.02267

作者:Jiawei Li

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)

关键词:Agent harnesses make, make many small, text is relevant, harnesses make, retrieved text

备注: 11 pages, 7 figures. Code and data: [this https URL](https://github.com/David-DL-Space/sys1-eval)

点击查看摘要

Abstract:Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool). We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a "channel effect" on injection false positives that vanishes with channel-native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at this https URL.

112. 【2610.02233】Budgeted Cache Repair for Cross-Context KV-Cache Reuse

链接:https://arxiv.org/abs/2610.02233

作者:Haeyong Kang,Chang D. Yoo

类目:Hardware Architecture (cs.AR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Cross-context KV-cache reuse, Cross-context KV-cache, KV-cache reuse predicts, shared segment keys, quality loss

备注:

点击查看摘要

Abstract:Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.

113. 【2609.29630】A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

链接:https://arxiv.org/abs/2609.29630

作者:Thiago César Castilho Almeida,Daniel Carlos Guimarães Pedronette

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:architectures produce latent, produce latent representations, pipelines assign representative, absolute distances distorted, leverage pretrained embeddings

备注: Accepted at the Main Conference of 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

点击查看摘要

Abstract:Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dimensional spaces. We introduce MARETopic, a training-free framework that casts topic discovery as rank-based prototype selection. After projecting embeddings onto a low-dimensional manifold, MARETopic builds ranked lists encoding ordinal neighborhood structure. A greedy algorithm selects exactly K exemplar documents, real corpus texts, whose neighborhoods cover the corpus. Two variants share this criterion. MARETopic$_\text{Corr}$ scores candidates with a query performance predictor and a rank correlation measure, leading Purity and NMI on the two benchmarks with the most categories, ahead of both neural and clustering-based topic models. MARETopic$_\text{Diff}$ scores them with a rank-based diffusion matrix, needs neither measure, and runs 1.7 to 1.9 times faster. Without a single gradient update, MARETopic leads topic coherence on two of three datasets. A novel inter-topic Maximal Marginal Relevance step raises vocabulary diversity at little cost in coherence. Our code is available at this https URL.

114. 【2602.00066】IntentCoding: Amplifying User Intent in Code Generation

链接:https://arxiv.org/abs/2602.00066

作者:Zheng Fang,Yihong Dong,Lili Mou,Dongming Jin,Zhi Jin,Ge Li

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large Language Models, Large Language, user intent, shown strong capabilities, Language Models

备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have shown strong capabilities in code generation, but their adherence to fine-grained user intent with multiple constraints remains a significant challenge. Our empirical analysis reveals two key observations: 1) Model performance deteriorates quickly as the number of constraints in the user intent increases, and 2) While user intent does influence the model's logits, such an influence may not be strong enough to effectively steer the decoding process. To this end, we propose Intent-Amplified Code Generation (IntentCoding), a novel decoding strategy that enhances an LLM's ability to follow user intent. IntentCoding captures the influence of user intent by masking out the intent, and applies a multi-strength ensemble mechanism to amplify the effect of user intent during generation. IntentCoding is model-agnostic, requires no additional training, and integrates seamlessly with existing decoding procedures. To enable systematic evaluation, we also construct CodeConstraints, a benchmark dataset specifically designed to test user intent compliance under varying numbers of constraints. Experiments on our constructed Constraints, as well as popular IFEvalCode, HumanEval and LiveCodeBench datasets, show that our IntentCoding model significantly improves both constraint satisfaction and functional correctness compared to standard decoding approaches. IntentCoding achieves up to 71.0% relative improvement on CodeConstraints, achieves up to 67.3% relative improvement on IFEvalCode and achieves up to 29.3% relative improvement in pass@1 on HumanEval and LiveCodeBench compared with greedy decoding.

信息检索

1. 【2610.03651】MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

链接:https://arxiv.org/abs/2610.03651

作者:Sean Culatana,Shang-En Huang,Kang Li

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Dense-retrieval services, memory budgets change, index bit rates, Residual Vector Quantization, budgets change

备注:

点击查看摘要

Abstract:Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and {4, 8, 16}-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.

2. 【2610.03147】SGuard: A Real-Time Framework for Detecting and Imputing Missing Data in Streaming Time Series

链接:https://arxiv.org/abs/2610.03147

作者:Imane Hocine,Asma Abboura,Soror Sahri,Abhijith Senthilkumar,Yacine Hakimi,Grégoire Danoy

类目:Databases (cs.DB); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:applications routinely suffer, sensor applications routinely, ACM International Conference, Streaming sensor applications, communication losses

备注: The 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 07--11, 2026, Rome, Italy

点击查看摘要

Abstract:Streaming sensor applications routinely suffer from delayed or missing observations caused by faults, communication losses, or environmental interference. Although recent imputation methods exploit temporal and spatial dependencies effectively, most either assume offline access to future observations or prioritize throughput without enforcing domain plausibility. We present TSGuard, a real-time demonstration system for monitoring, validating, and imputing missing values in streaming time series. TSGuard combines a lightweight graph-aware temporal imputation model with constraint-aware validation, fallback estimation, and operator-facing explanations. Rather than treating imputation as an isolated prediction task, TSGuard integrates it into a broader data-quality loop: detect problematic observations, impute missing values, validate estimated against physical and spatial constraints, and either retain the original value as a plausible anomaly or replace it when it violates domain constraints. Using environmental sensing as a motivating setting, the demo enables users to inspect delayed sensors, compare imputers, define constraints, and validate flagged values in real time. The combination of lightweight online spatiotemporal imputation, domain-aware validation, and explicit retain-or-replace decisions is our central contribution, while interactive explanations make these decisions inspectable and actionable. for operators.

Comments:
The 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 07–11, 2026, Rome, Italy

Subjects:

Databases (cs.DB); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Machine Learning (cs.LG)

Cite as:
arXiv:2610.03147 [cs.DB]

(or
arXiv:2610.03147v1 [cs.DB] for this version)

https://doi.org/10.48550/arXiv.2610.03147

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Journalreference:
Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 07–11, 2026, Rome, Italy

Related DOI:

https://doi.org/10.1145/3799682.3840280

Focus to learn more

            DOI(s) linking to related resources</p>
3. 【2610.03130】Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study

链接:https://arxiv.org/abs/2610.03130

作者:Yun Wang,Gad Shaulsky,Tomaž Curk,Blaž Zupan

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:general-purpose search tasks, broad biomedical corpora, search tasks, literature retrieval systems, developed and evaluated

备注: 15 pages, 5 figures. Submitted version (before peer review) of a paper accepted at Discovery Science 2026 (DS 2026); to appear in the Springer proceedings. Code and data: [this https URL](https://github.com/fulaibaowang/dictycite) ; dataset: [this https URL](https://doi.org/10.5281/zenodo.20308282)

点击查看摘要

Abstract:Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval benchmark from dictyBase for Dictyostelium, a model organism in cell and developmental biology. The benchmark consists of curator-generated biological queries linked to PubMed-indexed articles, together with structured gene annotations. Using this benchmark, we study three factors in niche biological retrieval: cross-encoder reranking, gene-aware query expansion, and abstract-only versus full-text retrieval. We report that reranking and gene-aware query expansion improve retrieval selectively: reranking is most useful when the model is well suited to biological evidence matching, whereas curated annotations help clarify compact biological queries by reducing vocabulary mismatch. Full-text chunks substantially improve retrieval when abstracts omit supporting evidence, increasing both candidate recall and top-rank performance, although these cases are harder than queries supported by abstracts. Data and code are publicly available at this https URL, and the benchmark dataset is additionally archived on Zenodo.

4. 【2610.02875】Query-aware routing for Cross-lingual performance gains in Encoders

链接:https://arxiv.org/abs/2610.02875

作者:Akshay Jain,Edward Kim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:exhibit reduced retrieval, reduced retrieval effectiveness, relevant documents differ, Multilingual encoders, strong same-language performance

备注:

点击查看摘要

Abstract:Multilingual encoders can exhibit reduced retrieval effectiveness when queries and relevant documents differ in language, despite strong same-language performance. We investigate whether Finnish and Swedish cross-lingual retrieval can improve while preserving an encoder's existing same-language performance and document index. We combine a query-only low-rank adapter, trained against frozen document embeddings, with deterministic routing based on query and index languages. Cross-language queries use the adapter, while same-language queries use the original encoder. SampoTron, our fine-tuned low-rank (LoRA) adapter alongwith the Nemotron-3-Embed-1B model, improves average retrieval quality across six English, Finnish, and Swedish directions from 0.241 to 0.291 in normalized discounted cumulative gain (nDCG) at rank ten, a 20.9% relative gain on a sampled financial benchmark. All six cross-lingual directions improve, and routing preserves the original same-language performance, including two full-corpus Finnish evaluations. The approach enables selective cross-language specialization with reusable document embedding vectors.

5. 【2610.02749】Learning Query Encoders Can Be Hard Even When Vector Retrieval Is Geometrically Easy

链接:https://arxiv.org/abs/2610.02749

作者:Anders Wikum,Nina Mishra,Amin Saberi,Tal Wagner

类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Efficient vector retrieval, Efficient vector, vector similarity, corpus geometry, query encoder

备注:

点击查看摘要

Abstract:Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top-$k$ answer sets of $n$ documents. We study a different notion of geometric capacity--the maximum recall achievable for a frozen document index--and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support. Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline $k/n$. Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval.

Subjects:

Information Retrieval (cs.IR); Machine Learning (cs.LG)

Cite as:
arXiv:2610.02749 [cs.IR]

(or
arXiv:2610.02749v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.02749

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
6. 【2610.02673】Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories

链接:https://arxiv.org/abs/2610.02673

作者:Joseph Chee Chang,Michael D'Arcy,Amy X. Zhang,Pao Siangliulue,Sangho Suh,Aakanksha Naik,Jena D. Hwang,Javier Ramos Benitez,Stella Wroblewski,Matt Latzke,Michael Cuoco,Ruben Lozano-Aguilera,Kris Ganjam,Joel Chan,Doug Downey,Peter Jansen,Kyle J. Travaglini,Daniel S. Weld

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)

关键词:draws many independent, theory draws, theory, observations, concepts

备注:

点击查看摘要

Abstract:A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research questions. Recent approaches scale theory synthesis with LLMs, but automate away choices and intuitions from researchers. We present Asterism, which extracts observations from hundreds of papers as concept-relation triples, with concepts unified in a hierarchical ontology. Researchers curate an evidence graph using the ontology and aggregate observations at different levels of granularity to focus theory formation on specific phenomena of interest. In a field deployment (n=10), researchers worked from observations to theories, and kept concepts and hypotheses fitting their preferences. In two case studies, teams of immunology and agriculture researchers discovered mechanisms outside their standard analyses and constructed hypotheses worth follow-up experiments.

7. 【2610.02600】When History Misleads: Asymmetric Margin Supervision for Instruction-Guided LLM Generative Recommendation

链接:https://arxiv.org/abs/2610.02600

作者:Ming Yin,Yuhan Yang,Chen Chen,Xinyu Lin,Wentao Shi,Fangcong Yin,Chaofei Yang,Chao Yang,Jiyan Yang,Hui Zhang,Ning Jiang,Yiran Chen,Qifan Wang

类目:Information Retrieval (cs.IR)

关键词:user current request, instruction-guided generative recommendation, LLM-based recommenders, balance two goals, instruction-guided generative

备注:

点击查看摘要

Abstract:In instruction-guided generative recommendation, LLM-based recommenders need to balance two goals: responding to the user's current request and aligning with the preferences in their interaction history. When the two conflict, history events can override the request. We show that turning the effect of individual history events into supervision faces two obstacles. First, the events that most influence a recommendation are not necessarily the ones that support the target item. Second, removing a misleading event can raise the target's score but a competing item's score even more, so a higher target score alone does not guarantee a better ranking. We propose Asymmetric Intervention-Guided Margin Supervision (AIMS), which converts the effect of removing individual history events into ranking supervision. For training requests already ranked correctly, a frozen reference model identifies request-specific deletions that improve both the target's score and its margin over a competitor near the recommendation cutoff. These margins serve as training targets, while the complete history is retained as input. Training combines cross-entropy with an asymmetric auxiliary loss that penalizes margin shortfalls and routes its gradient only through the competitor score. Inference is unchanged, requiring no history editing or deletion search. Across six LLM backbones on an industrial dataset and two public benchmarks, AIMS improves Recall and NDCG over strong baselines. Ablations support request-specific margins and asymmetric supervision, and the selected deletions preferentially remove constraint-violating history.

8. 【2610.02572】Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient Retrieval

链接:https://arxiv.org/abs/2610.02572

作者:Wentai Xie,Parker Carlson,Shanxiu He,Tao Yang

类目:Information Retrieval (cs.IR)

关键词:leveraging Large Language, semantic term expansion, Large Language Models, Recent work, neural sparse retrieval

备注: Accepted at SIGIR 2026

点击查看摘要

Abstract:Recent work on neural sparse retrieval has demonstrated strong relevance by leveraging Large Language Models (LLMs) for semantic term expansion. However, learned models paired with previous sparsification techniques still yield overly long document and query vectors partly due to a large LLM vocabulary, imposing a serious challenge to retrieval time and space efficiency. This paper proposes a scheme for optimizing model sparsity through a synergy of adaptive strategies, including learnable soft top-K, per-term thresholding, and FLOPs regularization to increase the sparsity of query and document vectors. Experimental results with Lion-SP model on the MS MARCO and BEIR datasets demonstrate that the proposed scheme can outperform the baselines by significantly reducing the average query and document lengths. Our scheme can achieve much shorter retrieval latency and lower storage cost while maintaining highly competitive relevance.

9. 【2610.02510】On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor

链接:https://arxiv.org/abs/2610.02510

作者:Sidney Shapiro,Joshua Lindemann

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:multi-course RAG tutor, institutional infrastructure, based on retrieval-augmented, ground answers, answers in assigned

备注: 23 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Campus AI tutors based on retrieval-augmented generation (RAG) must ground answers in assigned course materials while keeping textbooks and student dialogue on institutional infrastructure. We present CourseChat, an on-premises, multi-course RAG tutor for undergraduate business education, deployed behind a campus web gateway and intended for use embedded in Moodle. Six isolated course offerings, each keyed by its own course reference number (CRN), share twin-edge AI hosts running a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. We report two generation-model bake-off rounds, a separate fixed-evidence source-fidelity comparison, and conversation and quiz audits. Several larger models failed the classroom speed gate, but a 12B model and a 7B alternative passed. A separate mixture-of-experts candidate improved some corrections while introducing new factual and continuity errors. We therefore retain the 8B production model pending a demonstrated overall improvement, rather than claiming that 8B is universally optimal. Software changes improved follow-up topic resolution while preserving course scope; 435 prebuilt questions across 65 modules decouple practice from live generation. The results support treating model choice, evidence selection, serving compatibility, and product design as a joint engineering decision. They do not establish learning gains: faculty ratings, peak-load capacity, and complete public-gateway acceptance remain separate evaluation needs.

10. 【2610.02387】SOLO: Certified-Recall Metric Similarity Search with Scan-Only Sampled Inverted Lists

链接:https://arxiv.org/abs/2610.02387

作者:Édgar Chávez

类目:Databases (cs.DB); Information Retrieval (cs.IR)

关键词:general metric spaces, touched posting lists, approximate nearest-neighbor search, present SOLO, nearest sample points

备注:

点击查看摘要

Abstract:We present SOLO, an index for approximate nearest-neighbor search in general metric spaces whose serving path contains no ranking heuristic of any kind: a query is routed to the $k_s$ nearest points of a random sample of the database, and every object in the touched posting lists is evaluated with the true distance. Because nothing must outrank anything, recall equals a coverage probability computable from the stored index: one ground-truth pass over a query sample certifies every operating point at once, without serving any of them -- a recall certificate, and for a navigable graph no analogous object exists at any price. The whole index is one recursive rule -- sample the collection, post each object to its $b$ nearest sample points, split any list that outgrows a bound, always scan the leaves -- and its operating surface obeys an equal-work law, recall $\approx f(b \cdot k_s)$, whose level is a one-scalar signature of the dataset. The same scan-only structure gives a serving floor no graph architecture reaches once the router is itself indexed by the same rule: Deep-100M served at recall 0.9977 from 1 GB of resident memory (enforced cap, 10.7 bytes per object) and at 0.9964 from 256 MB, Deep-1B at recall 0.9925 from 512 MB (and from 96 MB at depth 3), inserts that are one search, and deletes that are exact. Throughput is competitive where the hardware allows it -- up to $1.8\times$ a tuned HNSW at $10^8$ on a two-socket 32-core server, with operating points to the right of where that graph saturates -- and the tables report it against HNSW, DiskANN, GRAFT, NAPP, misi, and SPANN's assignment rule on the same hardware and ground truth.

11. 【2609.29630】A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

链接:https://arxiv.org/abs/2609.29630

作者:Thiago César Castilho Almeida,Daniel Carlos Guimarães Pedronette

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:architectures produce latent, produce latent representations, pipelines assign representative, absolute distances distorted, leverage pretrained embeddings

备注: Accepted at the Main Conference of 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

点击查看摘要

Abstract:Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dimensional spaces. We introduce MARETopic, a training-free framework that casts topic discovery as rank-based prototype selection. After projecting embeddings onto a low-dimensional manifold, MARETopic builds ranked lists encoding ordinal neighborhood structure. A greedy algorithm selects exactly K exemplar documents, real corpus texts, whose neighborhoods cover the corpus. Two variants share this criterion. MARETopic$_\text{Corr}$ scores candidates with a query performance predictor and a rank correlation measure, leading Purity and NMI on the two benchmarks with the most categories, ahead of both neural and clustering-based topic models. MARETopic$_\text{Diff}$ scores them with a rank-based diffusion matrix, needs neither measure, and runs 1.7 to 1.9 times faster. Without a single gradient update, MARETopic leads topic coherence on two of three datasets. A novel inter-topic Maximal Marginal Relevance step raises vocabulary diversity at little cost in coherence. Our code is available at this https URL.

计算机视觉

1. 【2610.03717】Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

链接:https://arxiv.org/abs/2610.03717

作者:Keerthi Kaashyap,Dennis Anthony,Akshay Krishnan,Nhi Ngoc Nguyen,Jeremy Collins,James Hays,Shreyas Kousik,Animesh Garg

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:View Synthesis, paper examines, examines the role, Synthesis, NVS

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. this https URL

2. 【2610.03716】MoSE3: Learning World-Space SE(3) at Every Pixel

链接:https://arxiv.org/abs/2610.03716

作者:Jiahuan Cheng,Zhiyi Li,Tian Xia,Ruojin Cai,Yilun Du,Qianqian Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:prominent paradigm, paradigm for modeling, translation curve, dynamic scenes, monocular RGB video

备注: NeurIPS 2026 Spotlight. Project page: [this https URL](https://mose3-tracker.github.io/)

点击查看摘要

Abstract:Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.

3. 【2610.03715】4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

链接:https://arxiv.org/abs/2610.03715

作者:Ruihong Shen,Žiga Kovačič,Peter Kulits,Xingrui Wang,Zizhang Li,Joshua B. Tenenbaum,Alan Yuille,Jieneng Chen,Jiajun Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:executable graphics programs, inverse graphics, graphics programs, executable graphics, agents reconstruct dynamic

备注: [this https URL](https://4dcodebench.com/)

点击查看摘要

Abstract:We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at this https URL

4. 【2610.03713】What Should World Models Forget? Stratified Retention for Continual Adaptation

链接:https://arxiv.org/abs/2610.03713

作者:Nishit Anand,Ramani Duraiswami,Dinesh Manocha

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV); Signal Processing (eess.SP)

关键词:remains correct indefinitely, correct label remains, label remains correct, learning treats degradation, Continual learning treats

备注: Accepted to NeurIPS 2026 Continual World Models Workshop

点击查看摘要

Abstract:Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, which changes, so knowledge that was accurate when acquired may later become false, and discarding it is required behavior rather than a defect. Non-stationary ground truth is well studied in the concept drift literature and in the temporal factuality of language models, but has not been formulated for world models, which are distinctive in that they also encode knowledge that must never be revised. We argue that continual world models require retention stratified by invariance timescale, separating invariants such as physics and object permanence, which must never be revised, from instance-level facts that should be revised as soon as the environment changes. Standard forgetting metrics cannot distinguish a world model that has correctly revised outdated knowledge from one that has suffered catastrophic forgetting, and consequently rank a frozen model highest, while existing physical-reasoning benchmarks evaluate only frozen checkpoints. We propose differential retention, which reports invariant regression testing across the adaptation stream jointly with revision latency, without aggregation.

5. 【2610.03698】Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers

链接:https://arxiv.org/abs/2610.03698

作者:Neel Varma,Andrew Rufail,Dipika Khullar,Vasu Sharma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Self-supervised Vision Transformers, Vision Transformers, remain poorly understood, learn rich visual, rich visual representations

备注:

点击查看摘要

Abstract:Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lower-level structural, background, and texture-dominant patterns. Causal ablations fur- ther reveal a substantial functional asymmetry: disrupting top-activating register-derived features produces a 48.17% drop in representation cosine similarity, whereas disrupting outlier-derived fea- tures produces only a 0.31% drop. Together, our results provide evidence for token specialization in self-supervised ViTs.

6. 【2610.03691】FlowHMR: Physically Plausible Motion Capture from Video

链接:https://arxiv.org/abs/2610.03691

作者:Zhanke Wang,Chengfeng Zhao,Qing Shuai,Jingzhong Lin,Heng Li,Zeyu Ling,Yuxin Wen,Jing Li,Di Kang,Chunchao Guo,Linchao Bao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:present FlowHMR, human motion, motion, video, physically plausible global

备注: Project page: [this https URL](https://flowhmr.github.io/) Code: [this https URL](https://github.com/flowhmr/flowhmr)

点击查看摘要

Abstract:We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambiguous in depth, and direct regression tends to collapse toward an averaged solution. Moreover, the recovered motions are not guaranteed to be physically plausible, so physics-based tracking of them often fails. To address these challenges, we formulate video motion capture as a video-conditioned motion generation problem and first pretrain a flow matching model for this task. Given an input video, the pretrained model generates diverse motion candidates, but not all of them are faithful to the video or physically trackable. We therefore post-train the model using Group Relative Policy Optimization (GRPO) with two rewards. A fidelity reward encourages consistency with the input video. A tracking reward favors motions that a physics-based controller can track successfully. Together, these rewards shift the model's output preference, so the post-trained model stays faithful to the input video while producing more physically plausible motion. We further introduce Wild-4K, a large and diverse dataset of about 4K internet videos, for evaluating human motion recovery in the wild. Qualitative and quantitative experiments on Wild-4K show that our method outperforms state-of-the-art methods in overall motion fidelity and achieves a physical tracking success rate of 82.47%, compared with 62.82% for the strongest baseline, GVHMR.

7. 【2610.03689】SigLIP2 for aerial fire risk classification

链接:https://arxiv.org/abs/2610.03689

作者:Yunus Serhat Bıçakçı

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:class fire risk, fire risk classification, aerial imagery, examine the transfer, class fire

备注: 7 pages, 3 figures, 2 tables. Code available at [this https URL](https://github.com/yunusserhat/firerisk)

点击查看摘要

Abstract:We examine the transfer of a pretrained SigLIP2 image encoder to seven class fire risk classification from aerial imagery. We introduce a reproducible partition of the public FireRisk training mirror and an implementation that records data provenance, preprocessing and model selection. Two initial runs compare a frozen encoder probe with full model adaptation. On the validation partition, full adaptation reaches 63.05% accuracy and 58.94% macro F1, compared with 55.95% and 50.19% for the probe. Both runs use one training seed and select their checkpoint on the same validation partition. These development results support further evaluation of SigLIP2 but do not establish performance on an independent test set or unseen regions. The accompanying code provides a common framework for repeated experiments and comparisons with additional visual encoders.

8. 【2610.03664】ProAR: Learning Prospective Reasoning with Autoregressive Video Models

链接:https://arxiv.org/abs/2610.03664

作者:Linghui Shen,Tinghui Zhu,Sheng Zhang,Muhao Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:video models excel, Autoregressive Video Models, next-chunk prediction confines, video models, autoregressive video generation

备注: Project Page: [this https URL](https://luka-group.github.io/ProAR/)

点击查看摘要

Abstract:Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms autoregressive video generation into a goal-oriented reasoning process. ProAR introduces two key components: (1) To anchor generation to the long-range outcome, we integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them. (2) To guide short-range transitions, we introduce future representation self-alignment to encourage current hidden states to anticipate upcoming temporal dynamics. By leveraging teacher-forcing in AR training, we extract clean future representations in a single forward pass and align current representations with them using a lightweight, training-only predictor. Together, these two mechanisms seamlessly combine explicit, sparse target supervision with implicit, dense step-wise guidance, promoting coherent, goal-directed reasoning progress with modest computational cost. Experiments show that ProAR's complementary components consistently improve performance across diverse visual reasoning benchmarks. The framework proves highly training-efficient, surpassing fully trained standard AR baselines using only 25% of the training steps. This paradigm also demonstrates promising applicability to embodied reasoning tasks.

9. 【2610.03649】On-Board Anomaly Detection for Efficient Marine Environmental Monitoring

链接:https://arxiv.org/abs/2610.03649

作者:Thomas Goudemant,Clotilde Szywala,Benjamin Francesconi,Michelle Aubrun,Yves Bobichon,Marjorie Bellizzi,Adrien Girard

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:algal blooms, oil spills, sediment floods, disrupt habitats, human activities

备注: 8 pages, 3 figures. Presented at the 9th International Workshop on On-Board Payload Data Compression (OBPDC 2024), Gran Canaria, Spain, 2-4 October 2024

点击查看摘要

Abstract:Marine ecosystems are impacted by various threats such as oil spills, algal blooms, and sediment floods, which disrupt habitats, wildlife, and human activities. Advances in satellite imagery and Artificial Intelligence (AI) have enhanced our capabilities for early detection and mitigation of such hazards. In this paper, we propose a marine event detection pipeline for Earth observation satellites equipped with multi- or hyperspectral sensors. Our approach includes a self-supervised neural network encoder that compresses satellite images into a reduced latent space, enabling efficient onboard processing. A machine learning anomaly detection model identifies deviations from normal sea patterns to detect environmental anomalies. We compare its performance against traditional algorithms such as Isolation Forest, One-Class Support Vector Machine and Local Outlier Factors. Our lightweight, resource-efficient pipeline is optimized for deployment on satellites with limited computational resources, ranging from embedded CPUs to AI hardware accelerators. By prioritizing the transmission of critical information, our solution enhances system responsiveness and optimizes satellite communication bandwidth. Demonstrated through current integration across multiple missions, including European Space Agency's (ESA) Phisat-2 mission and Microsoft/Thales Alenia Space IMAGIN-e mission, our pipeline aims to improve marine environmental monitoring by providing timely alerts and efficient data reduction.

10. 【2610.03636】LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation

链接:https://arxiv.org/abs/2610.03636

作者:Ziqi Ma,Shreya Sharma,Mohamed El Banani,Katja Schwarz,Chongjie Ye,Chao-Yuan Wu,Li Fei-Fei,Ben Mildenhall,Georgia Gkioxari,Justin Johnson,Gowthami Somepalli

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:complex camera control, rapidly advancing, Camera-controlled video models, camera control, Camera-controlled video

备注: Project website: [this https URL](https://ziqi-ma.github.io/logo-website/)

点击查看摘要

Abstract:Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: this https URL

11. 【2610.03632】World Embedding Benchmark

链接:https://arxiv.org/abs/2610.03632

作者:Yiqi Liu,Ruifeng Yuan,Yang Wang,Long Li,Fengyu Cai,Hou Pong Chan,Jialin Yu,Hao Zhang,Chenghua Lin,Chenghao Xiao

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:received increasing attention, World Embedding Benchmark, remains less understood, Physical, received increasing

备注:

点击查看摘要

Abstract:Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.

12. 【2610.03618】Low-Cost Video--Time Priors as a Strong Baseline for EEG--fNIRS Emotion Regression on Familiar Videos

链接:https://arxiv.org/abs/2610.03618

作者:Minghao Kong,Jiurun Chen,Ying Gao,Xiangbin Meng,Rongjie Wang

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Continuous emotion regression, valence and arousal, Continuous emotion, viewer watches, emotion regression estimates

备注:

点击查看摘要

Abstract:Continuous emotion regression estimates moment-to-moment valence and arousal while a viewer watches a video. In familiar-video deployment, responses fron training participant-specific estimate, and prior-dominating fixed fusion tests whether physiology adds residual correction. In five-fold subject-held-out evaluation on 24was within 0.05 and 0.32 MAE of fusion in the internal and external evaluations, respectively. Source-explicit ablations showed that video identity and within-video tine accounted for most of the reduction, while EG-FNIRS gains were smaller and varied across participants and videos. These results identify the video-time prior as a strong, low-cost baseline and position EEG-fNIRS as an optional residual signal for familiar-video emotion regression.

13. 【2610.03617】DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

链接:https://arxiv.org/abs/2610.03617

作者:Vasco Ramos,Sandra Godinho Silva,Joao Magalhaes,Ricardo Rei,Pedro Henrique Martins

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Image-text alignment, hallucination detection, core problem, problem in computer, computer vision

备注:

点击查看摘要

Abstract:Image-text alignment is a core problem in computer vision with applications in caption evaluation, hallucination detection, data curation, and the benchmarking of text-to-image (T2I) generators. As T2I models improve, benchmarking has become demanding, requiring metrics capable of finding a series of issues like missing objects, swapped attributes, miscounts, and ignored negations. Recent work addresses this by fine-tuning evaluators on preference data or by prompting a vision-language model, either holistically with the caption or with decomposed verification questions. However, existing approaches fall short: fine-tuned metrics remain bound to one backbone and training distribution; holistic metrics miss fine-grained details; and decomposed metrics rely on a fixed-YES assumption that penalizes faithful images whenever that assumption fails. In contrast, we propose DEPICT, a training-free metric that replaces fixed reference answers with expected agreement between image-based and caption-only answers, weighting questions by how decisively the caption determines them. By replacing fixed references, our agreement rule increases negation accuracy from 19% to 88%. To recover the context lost during decomposition, DEPICT merges this agreement score with a holistic score. We evaluate DEPICT on five benchmarks and eleven backbones from three model families and find that it surpasses all training-free metrics and exceeds fine-tuned evaluators on two out of three human-correlation benchmarks.

14. 【2610.03599】ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars

链接:https://arxiv.org/abs/2610.03599

作者:Antonio Canela,Jordi Sànchez-Riera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reached near-photorealistic quality, near-photorealistic quality, reached near-photorealistic, High-fidelity, Conditional Variational Autoencoder

备注: GCPR 2026

点击查看摘要

Abstract:High-fidelity 3D head avatars have reached near-photorealistic quality. While recent methods enable text-driven manipulation, they struggle to provide fine-grained localized control, often entangling features or lacking geometric consistency. Modifying geometry through natural language currently requires slow per-prompt optimization or compromises identity and rigging. We present ManifoldSplat, the first end-toend framework for language-guided semantic shape editing of animatable 3D Gaussian Splatting avatars reconstructed from monocular videos. By performing edits within the structured FLAME manifold rather than directly optimizing an unstructured Gaussian cloud, we strictly preserve identity and animation. We introduce DeltaRegion, a per-region disentangled Conditional Variational Autoencoder (CVAE) delivering feedforward shape deltas, alongside a refining stage to recover view-consistent details. ManifoldSplat reconstructs and edits an avatar in ~90 seconds on a consumer GPU, rendering at ~800 FPS. Extensive evaluations demonstrate our approach sets a new state-of-the-art in localized prompt alignment, geometric coherence, and identity preservation. Project page and code: this https URL

15. 【2610.03577】Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

链接:https://arxiv.org/abs/2610.03577

作者:Shuo Yang,Lihao Fang,Yi Zhang,Haixiang Wang,Xincheng Ye,Shufan Chen,Jipeng Guo,Youqing Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Diffusion Transformers, requires multiple costly, DiT forward passes, multiple costly DiT, costly DiT forward

备注: 20 pages, 9 figures

点击查看摘要

Abstract:Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tensor computed at an earlier step. Most caching methods decide in advance which tensor to reuse. After distillation, adjacent sampling steps are farther apart. Reusing a tensor across this larger gap introduces more error, so choosing what to cache becomes especially important. We therefore introduce AutoTarget, a method that chooses the cached tensor for a given model, solver, and reuse schedule. AutoTarget uses a small set of runs without cache reuse to measure the error caused by reusing each candidate tensor, then selects the candidate with the lowest error. We also analyze how an error at one reuse step affects the final sample. For Euler sampling, we identify cache targets that produce the same trajectory and show why a stored solver update may not. Experiments on distilled image and video DiTs show that the best cache target changes with the model, image resolution, and solver. AutoTarget reduces DiT evaluations and retained cache storage. Generation quality remains close to the corresponding uncached run. On the tested PixArt-LCM and FLUX.1-schnell settings, its calibration ranking matches the ranking from held-out cached runs. To help others reproduce the method, we provide its core implementation on GitHub at this https URL.

16. 【2610.03543】DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

链接:https://arxiv.org/abs/2610.03543

作者:Jiahao Zhan,Yan Wang,Yongrui Ma,Qunliang Xing,Ruchang Yao,Runtao Liu,Shijie Zhao,Tianfan Xue

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Streaming video generation, Streaming video, distribution matching distillation, real video distribution, DMD

备注:

点击查看摘要

Abstract:Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollouts, limitations remain in visual quality and semantic alignment. To address these limitations, we propose DuoMatching, a distribution matching framework that approximates the real video distribution through a unified joint-marginal formulation. On top of existing joint matching formulations, the additional marginal matching objective provides dedicated frame-level supervision from an image generator, transferring complementary visual and semantic priors from it. To apply this frame-level supervision in video generation, we introduce LatentBridge to resolve the latent representation mismatch between the video student and the image teacher. Latent Variation Sampling further distributes such frame-level supervision across distinct temporal segments, reducing redundancy. Experiments demonstrate that DuoMatching improves visual quality, composition, and semantic alignment while largely preserving motion dynamics. Human evaluations show overall preference rates above 80% against all evaluated baselines. The project page is available at this https URL.

17. 【2610.03522】Feedforward Novel View Synthesis for Heterogeneous Cameras

链接:https://arxiv.org/abs/2610.03522

作者:Meng Wei,Cheng Zhang,Boying Li,Yihang Chen,Jianmin Zheng,Hamid Rezatofighi,Jianfei Cai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:sparse posed images, recently shown promising, shown promising results, fixed camera family, posed images

备注: Accepted at NeuralIPS 2026

点击查看摘要

Abstract:Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and panoramic cameras may coexist and where the target projection may be unseen during training. We study feed-forward NVS across heterogeneous central cameras and identify a key ambiguity introduced by tokenization: a visual token aggregates a projection-dependent bundle of pixel rays, while existing camera encodings mainly expose absolute rays or token-center relations. To address this, we combine token-center relative Camera Positional Encodings and proposed local raymaps, a token-level representation that explicitly describes the intra-patch ray distribution summarized by each token. We further propose projection-aware 2D RoPE, which replaces raw image-grid coordinates with ray-induced angular coordinates so that relative positional reasoning is aligned across camera projections. Together, these components treat diverse cameras as calibrated samplings of a shared ray space rather than separate visual domains. On ScanNet++ with heterogeneous-camera system, our method improves over camera-conditioned baselines under mixed-camera evaluation and demonstrates zero-shot generalization to panoramic views.

18. 【2610.03516】XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

链接:https://arxiv.org/abs/2610.03516

作者:Tingting Du,Ziyao Wang,Guoheng Sun,Ang Li

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:advanced robot control, evolve over time, control by predicting, World action models, world action model

备注: 27 pages, including appendix

点击查看摘要

Abstract:World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.

19. 【2610.03512】ProgressNet: Sketching and Prompting with a Frozen Text-to-Image Model

链接:https://arxiv.org/abs/2610.03512

作者:Arkaprabha Basu,Chaitat Utintu,Yi-Zhe Song

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Humans draw progressively, Humans draw, draw progressively, Banded Adaptive Control, Humans

备注:

点击查看摘要

Abstract:Humans draw progressively: a few strokes, a look at the result, a stroke erased, a prompt revised. Image generators do not work this way. They typically take a finished sketch and produce the image in a single pass, so every edit starts the picture again, and the models that do keep state across turns are driven by text, cannot take a stroke, and are too slow to draw with. We present ProgressNet, a training-free framework that lets a frozen text-to-image model follow a drawing session as it unfolds: strokes are added and erased, the prompt is revised, and the image keeps up at about a second per turn. It needs no new parameters because the frozen model already has what a progressive generator needs, a pathway through which the previous turn can be remembered, layers that can carry appearance forward without freezing structure, and an internal signal of how far to trust an unfinished sketch; three inference-time mechanisms (Previous-Concept Memory, Layer-Selective K/V Injection and Banded Adaptive Control) use each in turn. As a sketch fills in, every existing method degrades, the FID of the FLUX+ControlNet baseline doubling between 10% and 100% completion on FS-COCO, while ProgressNet's barely moves; it maintains strong fidelity and progressive coherence across three sketch domains and is preferred by users over five competitors, most widely on erasure.

20. 【2610.03510】Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

链接:https://arxiv.org/abs/2610.03510

作者:Ziyi Wang,Junchi Yao,Heqian Qiu,Wenbo Shi,Chengjiu Wang,Jinyang He,Binkai Hong,Hongliang Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:continuous scene extension, Recent advances, interactive storytelling requires, extended durations, scene extension

备注:

点击查看摘要

Abstract:Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual content. To address these problems, we present Weave Forcing, a training-free framework for compositional memory reuse in interactive long video generation. First, we use an LLM for semantic slot routing to decompose user prompts into character and background descriptions and explicitly select suitable historical references for each component. To isolate the required content, masked memory weaving uses contrasting attention maps conditioned on semantic slots to construct refined semantic masks, selectively exposing relevant tokens from compressed historical KV memories to guide the generation of the current shot. We further introduce coverage adaptive RoPE to adjust temporal offsets and memory retention according to no, partial, or full reference coverage, addressing visual artifacts observed when incomplete historical references are positioned close to the current generation. Extensive experiments demonstrate that Weave Forcing improves cross-shot subject and background consistency while maintaining competitive visual quality and text alignment.

21. 【2610.03474】Fed-ADApt: Federated Anytime Depth Adaptation for Resource-Aware Medical Image Segmentation

链接:https://arxiv.org/abs/2610.03474

作者:Abhijeet Parida,Zhifan Jiang,Pooneh Roshanitabrizi,Austin Tapp,Maria J. Ledesma-Carbayo,Syed Muhammad Anwar,Ziyue Xu,Marius George Linguraru,Holger R. Roth

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:raw patient data, sharing raw patient, existing approaches assume, patient data, limiting participation

备注: Accepted to The 4th International Conference on Federated Learning Technologies and Applications (FLTA 2026)

点击查看摘要

Abstract:Federated learning (FL) enables collaborative training of medical image segmentation models without sharing raw patient data, yet existing approaches assume a homogeneous compute budget across institutions, limiting participation of low-resource sites. We propose Fed-ADApt, a depth-adaptive federated framework for UNet-based segmentation that jointly addresses low-compute training and inference. Fed-ADApt integrates multi-depth supervision with hierarchical depth-wise aggregation, allowing each site to train according to its local compute budget while contributing to a global model that supports dynamic depth selection at deployment. We evaluated Fed-ADApt on multi-site 2D retinal fundus disc segmentation and 3D brain tumor segmentation. Across both tasks, federated collaboration substantially improves robustness under domain shift. Fed-ADApt matched the full-resource FedAvg performance in 3D and achieved competitive 2D performance with a 4.7% average Dice reduction, while reducing average inference cost by 19.5% in 3D and 34.5% in 2D and substantially reducing training cost by 98% at the most constrained sites. Importantly, Fed-ADApt enables low-resource institutions that cannot train full-capacity models to participate in federations while maintaining competitive global performance under a favorable accuracy to efficiency trade-off. By considering training and inference compute budgets, Fed-ADApt provides a practical and equitable solution for federated medical image segmentation across heterogeneous clinical and edge-enabled imaging environments.

22. 【2610.03473】UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation

链接:https://arxiv.org/abs/2610.03473

作者:Daikun Liu,Xin Zhan,Teng Wang,Xiaoping Wang,Changyin Sun

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:single event-RGB pair, requiring long histories, explicitly modeling future, future motion fields, Event Latent Enhancement

备注: 19 pages, 6 figures, conference, code: [this https URL](https://github.com/KK-xi/Unidynamics)

点击查看摘要

Abstract:We propose UniDynamics, a diffusion-based framework for future 4D dynamic scenes (RGB, depth, and optical flow) generation from a single event-RGB pair, without requiring long histories or control priors as in existing methods, while explicitly modeling future motion fields. The core idea is to leverage event streams to offer an alternative motion prior for single-RGB extrapolation, and to enforce geometric and motion constraints throughout generation via multimodal modeling. Specifically, we design an Event Latent Enhancement (ELE) module to align and enhance event latents into diffusion-injectable conditioning features, providing robust initial motion priors and reliable texture/structure cues. We further introduce a Perceptual Dynamics Space (PDS) embedded in the multi-scale U-Net, which decouples and adaptively interacts depth and flow while continuously feeding back constraints to appearance features, improving geometric-motion consistency for physically plausible and spatiotemporally coherent prediction. Experiments on VKitti2 and DSEC demonstrate state-of-the-art performance, producing high-quality, temporally coherent, and 4D-consistent future predictions, especially under challenging high-speed motion blur.

23. 【2610.03468】A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds

链接:https://arxiv.org/abs/2610.03468

作者:Mozhgan Hadadi,Talukder Z. Jubery,Adarsh Krishnamurthy,Baskar Ganapathysubramanian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:silico breeding trials, crops support high-throughput, field-grown crops support, support high-throughput phenotyping, breeding trials

备注:

点击查看摘要

Abstract:Editable 3D models of field-grown crops support high-throughput phenotyping and in silico breeding trials, but building them from scanned point clouds requires organ-level segmentation and fitting. Procedural generators can turn an organ-level parameter set into an analysis-suitable 3D model, but obtaining that set requires hours of manual tuning per plant or segmentation models trained on species-specific labels. We present an automated pipeline that reconstructs procedural maize models from raw 3D point clouds without manual tuning or species-specific training data. A multimodal vision-language model (VLM) annotates leaf midlines in rendered orthographic views. Deterministic geometric algorithms back-project the annotations onto the point cloud, merge them into 3D leaves by cross-view consensus, and grow the midlines to full blades on an orientation-weighted surface graph. Measured organ parameters populate a plant descriptor for a Non-Uniform Rational B-Spline (NURBS)-based procedural model generator. Each leaf surface is then refined against its scan points by differentiable NURBS fitting. The pipeline reached a median whole-plant Chamfer distance of 5.4 mm on 100 genotypically diverse field-grown maize plants from the MaizeField3D dataset. The reconstructions were closer to the scans than those of an earlier semi-automated pipeline based on manual annotations. The pipeline recovered 1,017 of 1,023 (99.4%) curated reference leaves at an intersection-over-union of at least 0.5 without using those labels as input. These results show that VLM annotations become usable organ-level measurements when downstream geometric stages can correct them. This makes automated generation of editable 3D plant assets feasible at the scale of modern phenotyping experiments.

24. 【2610.03467】Preserving Anatomical Continuity: Three-Stage Pipeline for Colon Segmentation in 3D Abdominal CT Scans

链接:https://arxiv.org/abs/2610.03467

作者:Deshan Kalupahana,Sonit Singh,Praveen Ravindran,Arcot Sowmya

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:colorectal disease analysis, produce disconnected predictions, disconnected predictions due, deep learning based, Accurate colon segmentation

备注: 5 pages, 2 figures

点击查看摘要

Abstract:Accurate colon segmentation from CT images is essential for colorectal disease analysis, yet deep learning based methods often produce disconnected predictions due to complex anatomy. This study introduces a three-stage, topology-preserving segmentation pipeline to address this issue. The first stage performs initial deep learning-based segmentation, followed by centreline bridging to reconnect disjoint regions and a reconstruction stage to refine continuity. Evaluations on TotalSegmentator and RAOS datasets using overlap, distance and topology-based metrics demonstrate improved structural consistency while maintaining segmentation accuracy. The proposed method enhances topological integrity, enabling more reliable colon segmentation for clinical and research applications.

25. 【2610.03453】I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry

链接:https://arxiv.org/abs/2610.03453

作者:Qian Wang,Liam Merz Hoffmeister,Brian Scassellati,Daniel Rakita

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:frequently non-manifold visual, motion planners require, planners require convex, non-manifold visual meshes, frequently non-manifold

备注:

点击查看摘要

Abstract:Physics simulators and motion planners require convex collision geometry, yet image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two today takes a slow, brittle reconstruct-then-decompose pipeline of repair, decimation, and approximate convex decomposition. We present I2CD, which predicts a convex decomposition directly from a single RGB image. Rather than train a new image-to-3D model, I2CD freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder and trains only a lightweight cross-attention head (38M parameters, under ten GPU-hours) whose learned "convex-slot" tokens emit the halfplane parameters of $K$ convex polytopes. The output is compact, convex by construction, and loads into physics engines without any post-processing, in ${\sim}0.5$s per image. On $227$ held-out OmniObject3D and Google Scanned Objects instances, I2CD attains the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running $6$-$37\times$ faster end-to-end. In a cross-simulator study in MuJoCo, PyBullet, Genesis, and Isaac Sim, every engine uses I2CD geometry as delivered, whereas raw generated meshes "load" everywhere but are silently replaced by a different collision shape in most cases or need seconds to minutes of per-object preprocessing. On a physical xArm7, I2CD produces planner-ready geometry for a $20$-object cluttered scene in $11$s versus $328$s for the strongest baseline, at comparable pick-and-place execution success ($85$ vs. $90$ of $100$ trials).

26. 【2610.03445】Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

链接:https://arxiv.org/abs/2610.03445

作者:Arun Josephraj Arokiaraj,Zekun Wu,Adriano Koshiyama

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:teacher-forced training loss, fixed target caption, produces the original, teacher-forced training, allowed to generate

备注: Accepted at the VLM4RWD Workshop (Grounded and Faithful Vision-Language Models for Real-World Deployment), NeurIPS 2026. 8 pages, 2 figures, 3 tables

点击查看摘要

Abstract:A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image's representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder's prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.

27. 【2610.03441】ChromaGS: Text-Driven Semantic Editing of 4D Gaussian Avatars

链接:https://arxiv.org/abs/2610.03441

作者:Antonio Canela,Jordi Sànchez-Riera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian head avatars, present ChromaGS, trained animatable avatar, head avatars, Gaussian head

备注: CGIP 2026

点击查看摘要

Abstract:We present ChromaGS, a method for real-time, language-guided color editing of animatable 3D Gaussian head avatars. Given a trained animatable avatar, users can instantly modify the color of semantic regions through natural language, with edits applied at render time and no retraining required. Our key insight is to augment each Gaussian primitive with learned soft assignments to semantic regions and decompose colors into region-level base colors and Gaussian-level residuals. This decomposition enables coherent color transfer: modifying a region's base color propagates naturally through all associated Gaussians while preserving fine appearance details encoded in residuals. A two-stage language pipeline translates text instructions into target colors, supporting both absolute specifications and relative adjustments. Unlike generative editing methods that may introduce unintended modifications, our approach provides deterministic, precisely localized semantic control. Experiments demonstrate faithful appearance preservation and intuitive interaction across diverse subjects. Project page and code are available at: this https URL

28. 【2610.03439】Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation

链接:https://arxiv.org/abs/2610.03439

作者:Daikun Liu,Teng Wang,Changyin Sun

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Event cameras hold, showing great potential, Event cameras, excellent dynamic properties, hold excellent dynamic

备注: 14 pages, 13 figures, conference

点击查看摘要

Abstract:Event cameras hold excellent dynamic properties, showing great potential for monocular depth estimation (MDE). However, existing methods mainly improve performance by optimizing contextual features, but still struggle with the ill-posed and nonlinear nature of direct full-depth regression. In this paper, we propose HypoDepth, the first event-image monocular depth iterative refinement framework. By introducing a discrete Depth Hypothesis Volume (DHV), we transform the depth regression problem into a constrained depth search task. Specifically, we construct a 3D cost volume between the DHV features and contextual features and perform a multi-scale correlation search to guide stable residual optimization. This lightweight cost volume enables efficient global-to-local refinement across multi-resolution. Our method outperforms existing approaches on DSEC and MVSEC with state-of-the-art results and strong zero-shot generalization. Meanwhile, our tiny model achieves an excellent balance between accuracy and efficiency, enabling real-time performance on resource-limited devices.

29. 【2610.03436】he Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

链接:https://arxiv.org/abs/2610.03436

作者:Danzel Serrano,Przemyslaw Musialski

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:recognizable mouth poses, reproduce recognizable mouth, facial animation, mouth poses, animation can reproduce

备注: 11 pages, 9 figures, 3 tables, under review

点击查看摘要

Abstract:Speech-driven 3D facial animation can reproduce recognizable mouth poses. However, it can simplify the motion between them, and that motion carries coarticulation, the way the sounds around each sound shape its articulation. We introduce a geometric measure of this trajectory shaping: lip-path length compared with the shortest route through the vowel, consonant and vowel positions of a speech segment. In contrast to the endpoint chord, this consonant-aware route accounts for obligatory transit and avoids degeneracy, while preserving invariance to uniform motion gain. The measure needs only a forced alignment, so it applies where no ground truth exists. We demonstrate it on four state-of-the-art methods, one per architectural family, real-time and offline. All four trace flatter lip trajectories than captured speech. Against frame-rate-matched ground truth, DiffPoseTalk, ARTalk and FaceFormer show clear deficits, equivalent on this measure to removing 15-60% of real speech's fast articulatory component. CodeTalker is marginal on the primary measure and clear on a companion measure. A pre-registered study with 97 viewers and 3,523 judgments underpins the measured direction: controlled damping of real motion lowers the score and is penalized, whereas exaggeration shows no detected penalty over the tested range. Viewers also prefer real speech in 73.4% of sentence comparisons and, in the aggregate, on single words. Together, the measure, its calibration and the study identify a perceptually relevant loss of trajectory shaping and a concrete target for improving synthesized articulation.

30. 【2610.03423】OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation

链接:https://arxiv.org/abs/2610.03423

作者:Bingyang Cui,Yujie Zhang,Yiling Xu,Yunfeng Guan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reinforcement learning, texture clarity, semantic alignment, alignment and texture, Reinforcement

备注:

点击查看摘要

Abstract:Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneously through multiple reward aggregation, without explicitly modeling inter-dimension dependencies. This can cause imbalanced optimization and persistent interference among conflicting dimensions. To address this limitation, we propose OuroReward, an interference-aware sequential reward scheduling strategy for T23D RL. OuroReward first estimates pairwise dependencies among dimensions and constructs a cyclic optimization path that minimizes cumulative interference. By incorporating the tail-to-head dependency, the cycle captures global compatibility across the entire schedule. Then, OuroReward converts the cycle into a one-pass sequence, and starts optimization from the dimension with the lowest aggregate interference. Rather than assigning a fixed optimization budget to each dimension-wise reward, training adaptively determines when to advance to the next reward according to the remaining optimization headroom of the current one. We further introduce AdaSelect, an adaptive prompt selection strategy that identifies reliable and informative prompts aligned with the model's current capability. By focusing policy updates on these prompts, AdaSelect effectively improves training stability. Extensive experiments across different T23D models and RL algorithms demonstrate that our framework consistently improves generation quality across multiple dimensions.

31. 【2610.03403】ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation

链接:https://arxiv.org/abs/2610.03403

作者:Zhihao Zhan,Le Tao,Yifei Tian,Xin Liu,Jie Yuan

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:remains challenging due, irregular tree structures, ambiguous instance boundaries, Forest point cloud, point cloud segmentation

备注:

点击查看摘要

Abstract:Forest point cloud segmentation is fundamental for fine-grained 3D forest scene understanding, yet remains challenging due to irregular tree structures, severe occlusions, density variations, and ambiguous instance boundaries. Recent query-based forest segmentation methods have shown promise for unified semantic and instance prediction, but they still insufficiently exploit forest-specific spatial structure and account for boundary uncertainty. In this paper, we propose ForestQuery, a boundary-aware and spatially anchored query learning framework for unified forest point cloud segmentation. ForestQuery enhances instance and semantic query learning through two complementary designs. Specifically, boundary uncertainty is explicitly modeled to guide reliable instance query construction and modulate query optimization through adaptive loss reweighting. Meanwhile, spatially anchored semantic query enhancement (SA-SQE) introduces learnable 3D anchors encoding forest vertical stratification priors to enrich semantic queries with explicit spatial references. We evaluate ForestQuery on multiple public forest point cloud benchmarks and a self-collected annotated real-world dataset. Extensive experiments demonstrate consistent improvements in both individual-tree segmentation and semantic segmentation across diverse forest scenes. Code and data are publicly available at this https URL

32. 【2610.03400】Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

链接:https://arxiv.org/abs/2610.03400

作者:Yudong Han,Yong Wang,Zaiquan Yang,Liang Lin,Chongyang Tao,Xiangxiang Chu,Liyuan Pan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains fundamentally limited, substantially advanced multimodal, verifiable rewards, rewards has substantially, substantially advanced

备注: 19 pages, 6 figures, under review

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model's own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.

33. 【2610.03391】Native Action-Prior Learning from Videos for World Action Models

链接:https://arxiv.org/abs/2610.03391

作者:Zhaochong An,Fei Zhang,Menglin Jia,Duncan Frost,Zijian Zhou,Yikai Wang,Xudong Wang,Aditya Patel,Belinda Zeng,Tao Xiang,Serge Belongie,Amir Bar,Sen He

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:scalability remains limited, World action models, models integrate future, action-annotated robot trajectories, integrate future visual

备注: Project Page: [this https URL](https://zhaochongan.github.io/projects/NAVA-WAM)

点击查看摘要

Abstract:World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video--action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.

34. 【2610.03389】From Patching to Pruning Visual Computation in Vision Language Models

链接:https://arxiv.org/abs/2610.03389

作者:Rahul Chowdhury,Timothy A Rupprecht,Xuan Shen,Shaoyi Huang,Pu Zhao,Yanzhi Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Vision language models, Vision language, incur substantial inference, MLP projections, substantial inference cost

备注:

点击查看摘要

Abstract:Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.

35. 【2610.03380】Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling

链接:https://arxiv.org/abs/2610.03380

作者:Chahira Benhama,Mohand Saïd Allili,Assia Hamadene

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:videos remains challenging, Deepfake detection, remains challenging, manipulated content, level while exhibiting

备注: 10

点击查看摘要

Abstract:Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temporal inconsistencies. This paper introduces an interpretable deepfake detection framework that models spatially and temporally coherent facial features in video sequences. Unlike end-to-end deep models relying on implicit representations, the proposed approach explicitly encodes physically grounded forensic cues, enabling transparent analysis and improved multi-dataset generalization. The pipeline transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame using 68 structured descriptors spanning four complementary domains: photometric, textural, geometric, and compression-based features. These descriptors provide a compact multi-domain representation of manipulation artifacts and are processed by a Long Short-Term Memory (LSTM) network to capture temporal dependencies and subtle irregularities. Evaluation on four benchmark datasets, FaceForensics++, Celeb-DF v2, a curated subset of the DeepFake Detection Challenge (DFDC), and DeeperForensics, yields strong and consistent F1-scores of 98.0%, 91.0%, 97.6%, and 96.2%, respectively. The approach also demonstrated a good cross-dataset generalization, providing a robust and interpretable solution for video deepfake detection.

36. 【2610.03374】EVEWorld: Physical Evolution Supervision for Embodied World Models

链接:https://arxiv.org/abs/2610.03374

作者:Kaiqi Wang,Songxin Zhang,Zejian Xie,Xiao Xiong,Zhuoyang Song,Ziwei Wu,Jun Yu Lu,Yitan Teng,Ziying Song,Jiaxing Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:enable scalable simulation, Embodied world models, world models enable, models enable scalable, Embodied world

备注: 44 pages

点击查看摘要

Abstract:Embodied world models enable scalable simulation of embodied interactions for robot learning. However, existing models are prone to Model Laziness, as they focus on visual fidelity at the expense of physical reasoning and lack process-level supervision over the temporal dynamics of manipulated objects. In this work, we propose EVEWorld, a physical evolution-supervision framework for physically consistent target evolution. EVEWorld consists of two components: Instance-Guided Restoration (IGR) and Temporal Instance Alignment (TIA). First, IGR promotes instance consistency through restoration supervision. Second, TIA promotes cross-frame consistency by aligning target instances across adjacent frames. We further introduce the Model Laziness Rate (MLR), a metric that measures persistent violations of instance consistency in generated trajectories. Extensive experiments on DreamGenBench, EWMBench, and PBench demonstrate the effectiveness of EVEWorld, notably achieving an 87.5% reduction in MLR compared with GigaWorld-0. On the WorldArena 2.0 Track 1 leaderboard, our model ranks 6th in JEPA Similarity and 17th overall, which further validates the performance of our evolution supervision strategy.

37. 【2610.03370】LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder

链接:https://arxiv.org/abs/2610.03370

作者:Anh-Khoa Dinh-Duc,Duc-Tai Dinh,Tam V. Nguyen,Minh-Triet Tran

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:global image representations, region-level tasks, produces only global, global image, visual encoder produces

备注:

点击查看摘要

Abstract:CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation. Our project page is link to this https URL

38. 【2610.03332】A Fully Automatic Pipeline for 3D Dendrite Instance Segmentation in SBF-SEM

链接:https://arxiv.org/abs/2610.03332

作者:Zewen Zhuo,Ilya Belevich,Eija Jokitalo,Alejandra Sierra,Jussi Tohka

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scanning electron microscopy, serial block-face scanning, block-face scanning electron, quantifying structural plasticity, Accurate three-dimensional

备注: Accepted at 2026 IEEE-EMBS Conference on Biomedical Engineering and Sciences (IECBES)

点击查看摘要

Abstract:Accurate three-dimensional (3D) reconstruction of individual dendrites in serial block-face scanning electron microscopy (SBF-SEM) is essential for quantifying structural plasticity in the brain, yet manual annotation at scale is infeasible. We present a fully automatic pipeline for 3D dendrite instance segmentation that unifies YOLOv6-guided Segment Anything Model (SAM) prompting on downsampled slices, iterative two-dimensional mask refinement, random forest 3D instance linking, and instance-aware high-resolution refinement using nnU-Net at native resolution into a single system requiring no manual prompting at inference. Applied to hippocampal CA1 SBF-SEM datasets from a control rat and a pilocarpine- induced epileptic rat, our pipeline reconstructs coherent, well- separated dendrites with high semantic accuracy (Dice 0.93 and 0.91) and strong instance-level performance on control tissue, while analysis of the more challenging epileptic tissue identifies instance recognition in dense regions as the principal remaining limitation. The high-resolution refinement stage recovers thin dendritic protrusions, providing a basis for downstream spine- level analysis. Code is available at this https URL ZE-WEN/dendrite-3d-instance-seg.

39. 【2610.03308】3lescope: Arbitrary-Resolution High-Fidelity Generative Surface Reconstruction from Images

链接:https://arxiv.org/abs/2610.03308

作者:Atsuhiro Noguchi,Tianhan Xu,Yiming Liang,Yuta Kikuchi,Masahiro Ishiyama,Shintaro Takagi,Hitoshi Murai,Eiichi Matsumoto

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:posed multi-view images, per-scene optimization, reconstruct high-fidelity, meshes from posed, posed multi-view

备注: 45 pages

点击查看摘要

Abstract:We reconstruct high-fidelity 3D scene meshes from posed multi-view images without per-scene optimization, across scales ranging from single objects to large outdoor scenes. Per-scene optimization methods lack the learned 3D prior needed when observations are sparse or surfaces are glossy or transparent. Existing generative methods leverage such priors to complete geometry in sparsely observed regions, but typically operate at a fixed resolution over a limited spatial extent, trading spatial coverage against detail. Reconstructing a large scene therefore often requires partitioning it into independently processed overlapping local regions, making it difficult to maintain global geometric consistency. To address these issues, we propose T3lescope, which applies a single fixed-resolution generator across scene scales in an inference-time coarse-to-fine cascade. A coarse level establishes the scene layout, and finer levels perturb and denoise geometry inherited from the coarser level within progressively finer spatial cells to recover surface detail. The model is trained on individual cells at multiple scales and shares its weights across all levels, so no hierarchy is fixed during training, and the number of levels, cell scales, and cell locations are determined at inference time. On indoor, outdoor, and city-scale scenes, T3lescope outperforms feed-forward and generative baselines, matches or surpasses per-scene optimization, and recovers fine structures as well as glossy and transparent surfaces. These results show that our method generalizes across diverse scenes, view counts, and image resolutions. Project page: this https URL

40. 【2610.03283】HexVIO: Towards All-Day Stereo-Inertial Tracking Through Commodity DSPs

链接:https://arxiv.org/abs/2610.03283

作者:Patrick Wolf,Mateo de Mayo,Daniel Cremers

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:spatial computing, fundamental prerequisite, prerequisite for spatial, VIO, devices

备注:

点击查看摘要

Abstract:The ability of a device to localize itself within its surroundings is a fundamental prerequisite for spatial computing. Visual-inertial odometry (VIO) has proven to be a cost-effective and accurate solution for this task. Robots, wearables, XR devices, and drones can benefit significantly from efficient implementations of VIO since they allow for cooler, lighter, and cheaper devices with longer battery life and a better user experience. In this work, we propose to enhance the efficiency of a VIO system by leveraging the Hexagon DSP, a commodity co-processor present in many modern smartphones and XR devices. Our approach offloads the visual frontend of a stereo-inertial odometry system to the DSP while keeping the backend on the main CPU. By optimizing the implementation for the DSP architecture, we achieve significant reductions in power consumption and latency compared to CPU-only execution. Our system, HexVIO, demonstrates a 67% reduction in power consumption or an 86% increase in throughput on a commodity smartphone, with the ability to sustain long-term real-time 30 fps tracking for 0.83 W, corresponding to ~18 hours of tracking on the testing device. These results highlight the potential of commodity DSPs for enabling all-day visual-inertial tracking in robotics and mobile devices.

41. 【2610.03276】Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters

链接:https://arxiv.org/abs/2610.03276

作者:Susmit Agrawal,Rebecca Wanner,Juliane Verwiebe,Matthias Tangemann,Matthias Bethge,Matthias Kümmerer

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video saliency, image saliency due, Video saliency prediction, additional temporal dimension, static image saliency

备注:

点击查看摘要

Abstract:Video saliency prediction is inherently harder to model than static image saliency due to the additional temporal dimension. Video saliency benchmarks rest on the premise that predicting gaze on video requires utilizing temporal activity distributed across frames. Prior work has challenged this, showing that static baselines recover a significant fraction of the explainable gaze information on LEDOV, a popular video saliency dataset, and that video saliency models fail in the same places as this static baseline. We verify that this diagnosis still stands: under a more capable gold standard than the original analysis, and an updated panel of recent architectures, the strongest temporal architecture in the panel still does not substantially improve over a fine-tuned static baseline. However, it remains unclear whether the marginal gain reflects limitations of current temporal architectures or a lack of temporal patterns in the benchmark itself. We introduce SalTempto, a video saliency benchmark with greater dynamism: 224 clips of highly dynamic content, sourced from the HACS-Segments dataset so that each clip contains an event together with its lead-up and aftermath, with gaze recordings from up to 16 subjects and a training split for adapting pretrained models. On SalTempto, the static baseline recovers only about 13\% of the headroom above the centerbias, against more than half on LEDOV. A fine-tuned temporal architecture shows a substantial gain in performance over the static baseline, indicating that it does capture meaningfully more temporal information, which LEDOV fails to measure. Yet, even this SoTA model still leaves nearly half of SalTempto's headroom unexplained, indicating room for improvement in video saliency modelling. Examination of SalTempto also lets us describe human tendencies that models miss. SalTempto link: this https URL.

42. 【2610.03261】Consecutive Posterior Fusion for Diffusive Recovery of Unobservable Image Structures

链接:https://arxiv.org/abs/2610.03261

作者:Elena Morotti,Davide Evangelista,Elena Loli Piccolomini

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Solving severely ill-posed, severely ill-posed imaging, problems requires recovering, Solving severely, requires recovering image

备注: 21 pages, 7 figures, 2 tables

点击查看摘要

Abstract:Solving severely ill-posed imaging inverse problems requires recovering image structures that are unobservable or weakly constrained by the measurements. Diffusion models provide expressive learned priors for inferring such missing information, while posterior sampling incorporates measurement consistency along the reverse process. Standard diffusion posterior samplers, however, rely on instantaneous measurement-aware estimates, without explicitly exploiting information carried by previous posterior corrections. We introduce Consecutive Posterior Fusion Denoising Diffusion Null-Space Models (CPF-DDNM), an inference-time strategy that fuses consecutive measurement-aware estimates to improve the diffusive recovery of unobservable image structures, without requiring retraining or additional denoiser evaluations. We instantiate this principle within DDNM, whose range/null-space decomposition reveals that consecutive fusion preserves the measurement-determined component while acting exclusively on the prior-driven null-space estimate. We thus provide a geometric interpretation of CPF-DDNM and a local error analysis that characterizes the optimal time-dependent fusion coefficient, including the extrapolative regime. Experiments on sparse-view and simulated low-dose computed tomography, as well as medical image super-resolution, show consistent improvements over DDNM and competitive performance against diffusion-based inverse solvers.

Comments:
21 pages, 7 figures, 2 tables

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.03261 [cs.CV]

(or
arXiv:2610.03261v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.03261

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
43. 【2610.03252】COSMI: COmpositional Synthesis of Multi-object Interactions

链接:https://arxiv.org/abs/2610.03252

作者:Daniel Eskandar,Ilya A. Petrov,Gerard Pons-Moll

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:everyday activities involve, captured datasets record, Generative models, multi-object capture, everyday activities

备注:

点击查看摘要

Abstract:Generative models of human-object interaction are bounded by the data that exists: everyday activities involve several objects, but most captured datasets record one at a time, as multi-object capture is combinatorially expensive. Our observation is that interactions are local, so single-object captures already contain the parts of multi-object activities. We compose them: contact-consistent clips of single interactions, mirrored to balance the hands, transfer between bodies, and a language model and geometric checks admit only the pairings that are plausible, semantically and physically. Therefore, the dataset grows combinatorially with the clips rather than recording time. The COSMI dataset holds 222k sequences and 275 hours with up to five objects, nearly thirty times the largest multi-object capture, and can be extended by adding datasets or even hand-object recordings. On this data we train the COSMI method, a text-to-interaction diffusion transformer that follows how the data is built: weight-shared object slots generate a variable number of objects, predicted relative to the body parts that move them. On a benchmark with an unseen object and unseen interaction combinations, models trained on the dataset generalize to the unseen combinations. COSMI outperforms baselines in text alignment and contact accuracy, where its margin is largest on the unseen object. Code, models, and the dataset pipeline will be released on the project page: this https URL.

44. 【2610.03248】EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation

链接:https://arxiv.org/abs/2610.03248

作者:Pujun Guo,Yuanfan Zheng,Fei Teng,Mengfei Duan,Guoqiang Zhao,Yuheng Zhang,Kai Luo,Kailun Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:enabling comprehensive scene, comprehensive scene understanding, Panoramic images provide, field of view, provide a complete

备注: 9 pages, 5 figures

点击查看摘要

Abstract:Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneous embodied platforms exhibit substantial differences in observation viewpoints and spatial layouts, giving rise to cross-embodiment observation shifts that pose additional challenges to consistent and reliable panoramic perception, while systematic studies of this problem remain limited. To bridge this gap, we introduce a new task, termed Cross-Embodiment Open Panoramic Segmentation. Meanwhile, we establish EmbPASS, a multi-platform panoramic semantic segmentation benchmark spanning Vehicle, Drone, Wearable, and Quadruped platforms under a unified semantic taxonomy, providing a testbed for systematically studying cross-embodiment panoramic perception. We further propose EPONet, an open-vocabulary panoramic semantic segmentation network that integrates Relation-Aware Metric Adapter (RAMA) and Content-Adaptive Semantic Transfer (CAST) to enhance spatial modeling and semantic transfer under heterogeneous embodied observations. Extensive experiments show that EPONet achieves the best platform-balanced performance on EmbPASS with 35.82% mIoU, outperforming the strongest baseline by 1.10%, while remaining competitive on existing panoramic segmentation benchmarks. The source code and EmbPASS benchmark will be made publicly available at this https URL.

45. 【2610.03224】Uncertainty as a Proxy for Semantic Correctness in Diffusion-Based Medical Image Synthesis

链接:https://arxiv.org/abs/2610.03224

作者:Yuxuan Ou,Konstantinos Kamnitsas,OxAAA Study,AICT Consortium,Regent Lee,Vicente Grau

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:avoiding contrast administration, avoiding contrast, synthesise contrast-enhanced, contrast administration, environmental and patient-access

备注:

点击查看摘要

Abstract:Diffusion models can synthesise contrast-enhanced CT (CECT) from non-contrast CT (NCCT), avoiding contrast administration and its environmental and patient-access costs. However, visually realistic images are not necessarily anatomically correct, and the pixel-intensity and feature-space similarity metrics used to assess generation quality do not directly measure anatomical correctness. In this work, we investigate whether uncertainty can serve as a proxy for semantic correctness in diffusion-based medical image synthesis. We study NCCT-to-CECT synthesis using AortaDiff, a multitask diffusion framework that jointly generates CECT images and lumen segmentations. The segmentation output provides an explicit representation of the generated vascular anatomy, enabling segmentation-derived errors to be used as a quantitative measure of generation correctness. Six methods spanning weight (Ensemble, HyperDiff, BayesDiff), architecture-perturbation (MCDropout), generative-stochasticity (RDS) and input-perturbation (TTA) uncertainty are compared at the pixel, region and image levels, and for detection of clinically relevant out-of-distribution (OOD) cases. Uncertainty proves informative at all three spatial scales, remains informative on an external multi-centre dataset under distribution shift, and supports OOD detection. MCDropout stands out among the six: it ranks among the leading methods at every scale, generalizes well on the external dataset, and can be enabled at inference on any model already trained with dropout, so reliable uncertainty comes at no extra training cost. Uncertainty reliably flags severe failures but discriminates poorly among already high-quality images. These findings support uncertainty as a practical and computationally economical signal for quality filtering, reliability assessment and OOD detection in NCCT-to CECT synthesis.

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.03224 [cs.CV]

(or
arXiv:2610.03224v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.03224

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
46. 【2610.03221】VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation

链接:https://arxiv.org/abs/2610.03221

作者:Yutong Wang,Xingtong Ge,Enhuai Liu,Yunke Wang,Tianfan Xue,Yu Qiao,Yaohui Wang,Xinyuan Chen,Chang Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video creation spans, video diffusion models, diffusion models remain, models remain costly, repeatedly evaluate large

备注:

点击查看摘要

Abstract:Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can provide unstable or incomplete guidance when the student and teacher distributions have limited overlap. VDOT addressed this issue by adding optimal transport distillation (OTD), whose explicit coupling supplies geometric directions for condition-based generation. Balanced OTD, however, performs full-mass matching between the spatial tokens of each corresponding student--teacher frame pair. This assumption weakens for T2V and I2V, where one condition admits many valid outputs and spatial content need not align across different realizations. We present VDOT++, a unified distillation framework that applies the same training recipe separately to generators for the three task families. It makes OTD robust to output diversity through an asymmetric unbalanced formulation that allows unreliable student tokens to carry less mass while maintaining coverage of the teacher tokens. An $\ell_1$ ground cost further replaces mean-based aggregation with a more mode-preserving weighted median that limits the influence of distant transport targets. The two changes respectively determine whom to match and how the selected targets should be aggregated. We additionally combine distribution matching and adversarial refinement through sequential backward passes, and exploit the decoupled score networks for cross-scale distillation, where larger score networks improve a compact generator. Experiments on UVCBench, VBench, VBench-I2V, and the VACE benchmark show that the resulting four-step generators are competitive with many-step teachers and strong few-step baselines across all three task families.

47. 【2610.03220】Evolving Hybrid Quantum-Classical Architectures for Image Classification

链接:https://arxiv.org/abs/2610.03220

作者:Devroop Kar,Daniel Krutz,Travis Desell

类目:Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Quantum Physics (quant-ph)

关键词:remains largely manual, established deep learning, performance depends strongly, neural networks integrate, networks integrate parameterized

备注: Under Review at The Fifteenth International Conference on Learning Representations 2027

点击查看摘要

Abstract:Hybrid quantum classical neural networks integrate parameterized quantum circuits (PQCs) with established deep learning architectures, but their performance depends strongly on the choice of quantum circuit architecture, a choice that remains largely manual. Most existing approaches rely on hand-designed or fixed circuit ansätze, requiring circuit structure, gate composition, and qubit connectivity to be specified in advance with no guarantee that they suit the task. This limitation is especially acute in image classification, where quantum circuits must transform features extracted by classical networks while remaining compact enough for practical training, requirements that generic, task-agnostic ansätze are unlikely to satisfy simultaneously. We extend EXAQC, an evolutionary framework for automated quantum circuit discovery, to image classification. EXAQC evolves PQCs as intermediate processing modules while retaining classical feature-extraction and prediction layers. On MNIST, Fashion-MNIST, and CIFAR-10, EXAQC achieves 98.42%, 90.62%, and 85.47% accuracy, respectively, while using comparable gate counts to other quantum architecture-search methods. Against classical networks, evolved hybrid models maintain comparable accuracy with substantially fewer trainable parameters, reaching 85.68% on CIFAR-10 with over 25$\times$ fewer parameters than a 10-layer CNN. Encoding choice also matters: rotation-based encodings (RX, RY, U3) outperform amplitude encoding by 22-25 points on CIFAR-10. These results demonstrate that automated circuit discovery yields compact quantum modules that can replace larger classical components in vision architectures while retaining competitive accuracy.

48. 【2610.03218】VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models

链接:https://arxiv.org/abs/2610.03218

作者:Elad Dror Cohen,Ofir Gordon,Lior Dikstein,Idan Achituve,Hai Victor Habi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:training and inference, hardware-supported approach, approach to efficient, efficient training, Microscaling

备注:

点击查看摘要

Abstract:Microscaling (MX) formats are emerging as a hardware-supported approach to efficient training and inference. They combine low-precision elements with shared block scales, but their impact on vision models remains underexplored. We systematically investigate post-training MX quantization across vision models and tasks. An analysis of direct conversion identifies three sources of error: block-scale representation, the poor alignment of some small convolutional weight tensors with nonuniform element grids, and the underuse of signed codes by nonnegative activations. These findings motivate VisionMX, a post-training MX quantization method that optimizes bounded weight rounding and applies a foldable affine correction to activations. We evaluate VisionMX across image classification, object detection, semantic segmentation, and low-light image enhancement using several MX-style formats. It improves on direct conversion and the evaluated post-training quantization baselines, with the largest performance recoveries in architectures most sensitive to MX conversion

49. 【2610.03202】Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation

链接:https://arxiv.org/abs/2610.03202

作者:Divya Jyoti Bajpai,Arun Verma,Manjesh Kumar Hanawal

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Flow Matching enables, Matching enables high-quality, Flow Matching, enables high-quality visual, remains costly due

备注: Accepted in NeurIPS 2026

点击查看摘要

Abstract:Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose COFLOW, an inference-time method that adaptively selects the step counts each generation based on the prompt features. Our context-aware COFLOW is trained online with an unsupervised reward that balances inference efficiency and generation fidelity. Our method is plug-and-play, requiring no retraining of the underlying generative model. It generalizes to image and video generation, achieving over 2.5x speedup while preserving perceptual and semantic quality. We further provide a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.

50. 【2610.03193】Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification

链接:https://arxiv.org/abs/2610.03193

作者:Emanoel dos Santos,Kelvin Cunha,Rodrigo Mota,Fabio Papais,Thales Bezerra,Natalia Lopes,Erico Medeiros,Shirley Cruz,Jessica Araujo,Paulo Borba,Tsang Ing Ren

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:potential applications, studies toward potential, machine learning, grown substantially, application of machine

备注: 10 pages, 1 figure, 3 tables, approved at MICCAI 2026

点击查看摘要

Abstract:The application of machine learning to dermatology has grown substantially in recent years, moving beyond proof-of-concept studies toward potential applications. However, clinical dermatology remains a challenging and still open problem. Diagnostic assessment is often ambiguous, and skin lesions exhibit high variability, compounded by differences in acquisition modality, device quality, and patient demographics. These factors hinder the development of robust models suitable for safe and equitable clinical use. To support translation into practice, it is essential to systematically evaluate how contemporary models generalize across heterogeneous data sources. In this work, we benchmark a diverse set of architectures on recent dermatology datasets, spanning dermoscopic images and smartphone-based clinical photographs. We assess the robustness of recent general-purpose and medical vision-language models, as well as foundation models, and compare them against task-specific dermatology classifiers, including embedding-based approaches and convolutional neural networks. Our study provides an evaluation of model performance under distribution shifts, modality changes, and demographic variability. By quantifying the gap between current state-of-the-art models and the requirements of clinical deployment, we aim to contribute to the development of reliable, accessible, and clinically applicable AI systems for dermatology.

51. 【2610.03192】PocketSplat: Mobile Gaussian Reconstruction via World-Space Latent Allocatio

链接:https://arxiv.org/abs/2610.03192

作者:Wenzhi Guo,Xianda Chen,Dongxuan Chen,Guangchi Fang,Bing Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:representation size suited, resulting Gaussian asset, device resource envelope, mobile Gaussian asset, Gaussian

备注:

点击查看摘要

Abstract:Mobile Gaussian reconstruction must satisfy two requirements: the reconstruction model must execute within a device resource envelope, and the resulting Gaussian asset must expose a representation size suited to downstream mobile use. Existing feed-forward Gaussian reconstructors commonly decode dense, image-aligned candidates whose final cardinality is implicitly determined by the input resolution and number of views. We present PocketSplat, a feed-forward framework for budgeted mobile Gaussian asset construction. Given a prescribed output budget, PocketSplat organizes dense geometry-aware latent candidates in predicted world space, allocates exact integer capacity across local latent cells, and decodes complete Gaussian attributes only for retained candidates. Cell-conditioned latent fusion aggregates repeated multi-view evidence before decoding, while spatial responsibility decoding adapts Gaussian support after local sparsification. Experiments on DL3DV and out-of-distribution benchmarks establish a strong quality--budget trade-off against feed-forward Gaussian reconstruction baselines. On Mip-NeRF 360, PocketSplat executes directly on a target iPhone and constructs compact, higher-quality Gaussian assets substantially faster than a deployable streamed MVSplat variant; native MVSplat and DepthSplat exceed the device memory budget.

52. 【2610.03187】Lightweight and Resource-Efficient Perception for Robotic Guide Dogs

链接:https://arxiv.org/abs/2610.03187

作者:Jinse Kwon,Yoojin Lim,Choonghan Lee,Yongseung Yu,Yongin Kwon,Jemin Lee

类目:Distributed, Parallel, and Cluster Computing (cs.DC); Computer Vision and Pattern Recognition (cs.CV)

关键词:Multi-camera streaming perception, streaming average precision, Multi-camera streaming, heterogeneous edge platforms, edge platforms shared

备注: accepted in ACCV 2026

点击查看摘要

Abstract:Multi-camera streaming perception is increasingly deployed on heterogeneous edge platforms shared with co-resident workloads, yet accelerator placement is often evaluated using isolated single-stream experiments and mean streaming average precision (sAP). Using two end-to-end pipelines on a single GPU--NPU platform, we show that isolated evaluation can mis-rank deployment-time placement. Although the GPU pipeline is preferred in isolation, GPU-localized contention introduces deadline misses that make detections stale and can reverse the preferred placement before full GPU saturation. The NPU pipeline is less accurate than the GPU pipeline on small and medium objects in isolation, but nearly matches it on large objects. The largest absolute sAP losses in our latency and contention experiments occur for large objects. In our four-stream experiments, the preferred placement depends on which path becomes stale, and increasing GPU-side contention shifts the best placement from All-GPU to All-NPU. Under a GPU-saturating vision--language co-tenant, All-NPU achieves $5.2\times$ the worst-stream sAP of All-GPU. Because mean sAP can hide severe single-stream degradation, evaluation should report contention sweeps, deadline-miss rates on both paths, and worst-stream sAP alongside mean sAP.

53. 【2610.03167】Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration

链接:https://arxiv.org/abs/2610.03167

作者:Zhiwei Wang,Defeng He,Yuxing Li,Meilu Zhu,Edmund Y. Lam

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image matching establishes, Cross-modal image matching, planar registration, image matching, modalities for planar

备注:

点击查看摘要

Abstract:Cross-modal image matching establishes stable and accurate geometric correspondences across modalities for planar registration. Existing semantic representations provide cross-modal consistency, but semantic similarity does not necessarily imply geometric correspondence. Meanwhile, fine-grained CNN features provide accurate local details but lack global cross-modal semantic guidance for stable refinement. To address these issues, we propose CDPM, which first establishes geometrically consistent semantic representations and then preserves their dominant role in correspondence estimation during fine-grained localization. Specifically, we progressively adapt DINOv3 using geometrically consistent cross-modal patch pairs, enabling feature similarity to better reflect true cross-modal spatial correspondences. We then construct a DINO-Centric Feature Pyramid, where multi-scale DINO representations maintain stable cross-modal correspondences, while a lightweight CNN branch provides auxiliary structural details for precise local refinement. Extensive experiments on three cross-modal datasets demonstrate the superior performance of CDPM. On VIS-IR, compared with the dense matcher RoMa, CDPM improves AUC@3/5/10/20 by 7.36, 13.40, 13.75, and 10.42 percentage points, respectively, and reduces mACE from 5.83 to 2.78 pixels. It also outperforms RoMa v2 across all metrics while requiring 45.6% fewer FLOPs. The online demo and dataset are available, and the code will be released on our project page at this https URL.

54. 【2610.03162】Budgeted-GS: Real-Time Large-Scale Gaussian Splatting via Factoring LOD

链接:https://arxiv.org/abs/2610.03162

作者:Haipeng Wang

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting achieves, Splatting achieves excellent, Gaussian Splatting, achieves excellent visual, excellent visual quality

备注: 28 pages, 21 figures. Preprint of the EG 2027 submission (paper1075)

点击查看摘要

Abstract:3D Gaussian Splatting achieves excellent visual quality with real-time rendering, but at the scale of entire cities it does not fit: a trained model carries millions of primitives and gigabytes of memory, and real-time rendering at high quality on a consumer GPU remains out of reach. We introduce Budgeted-GS, a post-hoc method that turns any trained 3DGS model into a factoring tree, a multi-resolution hierarchy of moment-matched aggregates. After a construction pass of a few seconds, a single quality parameter selects, for each view, the level of detail that fits the memory of the target device, so the same city-scale model serves GPUs with widely different memory capacities. When a new scene is to be trained, the same theory applies: instead of growing a full-sized model and compressing it afterwards, budget-centered training first measures how many primitives the scene needs and then trains the model directly at that size, avoiding the wasted effort of optimizing primitives that are later discarded. Both methods are grounded in a measurable capacity floor, a budget-error law derived from optimal transport in phase space; selection rules certified by recent covering theorems decide which primitives are redundant. The floor answers how many primitives a scene actually needs and how many can safely be given up. We validate the floor on 13 public scenes under a preregistered protocol, and exercise both methods from object scenes to an official city capture, rendering it at native 1920x1080, full SH, in real time on one consumer GPU.

55. 【2610.03154】Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models

链接:https://arxiv.org/abs/2610.03154

作者:Jonas Kneifl,Jakub Skalski,Bartłomiej Twardowski,Kamil Deja

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:produce strikingly realistic, recent benchmarks reveal, benchmarks reveal pronounced, reveal pronounced deficits, generation models produce

备注: 22 pages, 8 figures, 5 tables

点击查看摘要

Abstract:Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar motion patterns. We address this by probing internal representations of video Diffusion Transformers (DiTs) for simulator-derived ground-truth physical quantities spanning kinematic motion and rigid-body dynamics under gravity and contact. We find that these quantities are linearly decodable with high accuracy early in the denoising process, substantially outperforming a baseline decoded directly from the model's own noised latents, indicating that the relevant physical information is actively constructed during denoising rather than already present in the input. Additionally, we show that activations at on-object tokens carry the relevant physical information and that quantities defined over multiple frames are readable from single latent frames. Hence, information is sharply localized within the token sequence and is computed globally but stored locally. The probes further show partial extrapolation, transferring to scene variations and object configurations outside their training regime, so what they read is not simply a correlate of the scenes they were fit on. When fitted directly in the full-resolution activation space, the probing directions can serve as steering vectors to change the model's output.

56. 【2610.03142】CalCErt: Bin-wise Certification of Confidence Calibration in Medical Image Classification

链接:https://arxiv.org/abs/2610.03142

作者:Leo Fillioux,Stergios Christodoulidis,Stergios Christodoulidis,Maria Vakalopoulou,Jose Dolz

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Deep neural networks, Deep neural, neural networks remain, networks remain vulnerable, undermining uncertainty calibration

备注:

点击查看摘要

Abstract:Deep neural networks remain vulnerable to adversarial perturbations, which can distort not only predictions but also confidence scores, undermining uncertainty calibration. While existing certification methods focus on preserving the predicted category, providing guarantees on how calibration behaves under adversarial attacks remains overlooked. In this work, we introduce CalCErt, a simple and efficient post-hoc strategy that certifies bin-wise confidence calibration for any pretrained differentiable classifier. Our approach combines empirical calibration estimates, statistical concentration bounds, and local Lipschitz estimates of the confidence function to derive data-dependent upper bounds on worst-case miscalibration within an $ell_2$-ball of radius R. We evaluate CalCErt across 11 medical image classification tasks and multiple adversarial perturbations, demonstrating substantially higher certified coverage than baseline strategies while maintaining competitive tightness. Our code is available at this https URL.

57. 【2610.03141】Behavior Pack Optimization for Video MLLM Post-Training

链接:https://arxiv.org/abs/2610.03141

作者:Zhaolu Kang,Shiyu Liu,Tailong Luo,Wei Zhang,Yingjie He,Lei Wei,Guansu Wang,Liang He,Siheng Wang,Guangyuan Dong,Jiaqi Su,Shuang Chen,Haoyu Ji,Qishi Zhan,Kaiyue Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:target object barely, question answering benchmarks, large language models, climbing video question, video question answering

备注: NeurIPS 2026 poster

点击查看摘要

Abstract:Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence the question asks for. We trace this to the unit of post-training: rewards are computed on a single response to the original clip, so the model is never asked to behave consistently across views. We propose Behavior Pack Optimization (BPO), which replaces the single response with a behavior pack of outputs across counterfactual views chosen by question type, scored jointly. The pack reward asks for stability when the intervention is irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains. To keep this objective stable at small pack sizes, BPO uses an anchor-relative advantage: the response on the original view serves as a per-prompt reference instead of a group mean over mixed views. On TempCompass, MVBench, and NExT-QA, BPO improves the macro accuracy of Qwen2.5-VL-7B-Instruct by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint. The gains transfer to Video-MME, LongVideoBench, and to LLaVA-Video-7B; ablations confirm they follow the view sets, not the rollout count. We hope this pack-level perspective offers a useful starting point for the video MLLM and multimodal post-training community as the field moves toward evidence-grounded video reasoning.

58. 【2610.03123】Foresight: planning future perception in streaming VLMs without retraining

链接:https://arxiv.org/abs/2610.03123

作者:Ashok Prasad Neupane,Dipan Bartaula,Ankit Belbase,Saugat Adhikari,Samip Ghimire,Saroj Poudel,Binod Bhattarai,Danda Pani Paudel

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Existing streaming vision-language, computational pathways remain, pathways remain fixed, streaming vision-language models, Existing streaming

备注:

点击查看摘要

Abstract:Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.

59. 【2610.03120】In-Distribution Forcing for Long Video Generation at Test Time

链接:https://arxiv.org/abs/2610.03120

作者:Jeongwoo Shin,Youngyoon Choi,Sangwoo Jo,Hyunmog Kim,Sungjoon Choi,Joonseok Lee,Jaewoong Choi,Jaemoo Choi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:motion dynamics decay, Modern autoregressive, remains challenging due, generating long videos, diffusion models excel

备注: Preprint

点击查看摘要

Abstract:Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the training horizon: nothing constrains the construction of KV entries during rollout, giving rise to the KV-provenance problem where cached entries themselves become out-of-distribution (OOD). To address this, we propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations. Its key mechanism, self-caching, prevents OOD KV entries at their source. Each chunk is cached without attending to prior KV entry, keeping the rolling window exactly in-distribution. Consequently, ID-Forcing seamlessly extends short-horizon models to minute-scale video generation. Extensive evaluations show that our method remains competitive on standard video generation benchmark while substantially outperforming prior work in mitigating drifting, as validated by both our drift metrics and a user study.

60. 【2610.03105】A Benchmark for Spatially Grounded Gesture Generation

链接:https://arxiv.org/abs/2610.03105

作者:Anna Deichler,Rishabh Dabral,Fethiye Irmak Dogan,Anindita Ghosh,Jonas Beskow

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Human-Computer Interaction (cs.HC)

关键词:shared space interweaves, space interweaves verbal, Communication in shared, gestures anchor language, non-verbal signals

备注: 13 pages, 9 figures. Benchmark of the Referential Gesture Challenge at the HSI Workshop, ECCV 2026. Data and video: [this https URL](https://huggingface.co/datasets/hsi-workshop/referential-gesture-challenge)

点击查看摘要

Abstract:Communication in shared space interweaves verbal and non-verbal signals, and pointing gestures anchor language to the environment: "put the cup on that one" is uninterpretable without the gesture that fixes the referent. Yet no common framework exists for evaluating whether generated gestures indicate their intended referent; distributional metrics reward a gesture aimed at the wrong object as long as it looks natural. We introduce a benchmark for spatially grounded gesture generation, comprising ~2K pointing-annotated clips from naturalistic VR dialogue with ground-truth 3D referents, a task in which systems must decide when, how and where to point within conversational speech, and a protocol that separates temporal alignment, spatial grounding and perceived naturalness. We also provide a flow-matching baseline, MM-Conv-Flow. Evaluating it alongside an independent retrieval-based system and captured human motion, we find that geometric grounding can exceed that of human pointing without any gain in perceived naturalness, showing that referential gesture quality must be measured along separate dimensions.

61. 【2610.03099】Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

链接:https://arxiv.org/abs/2610.03099

作者:Jinghan Zhao,Yiman Hu,Liang Wu,Jian Xu,Bo Zheng

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:assessing marketing strategies, consumers evaluating products, merchants assessing marketing, marketing strategies, information-dense and frequently

备注:

点击查看摘要

Abstract:E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. Cross- video reasoning requires models to locate fine-grained evidence among many redundant frames and integrate visual details, speech, and on-screen text. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. To address the sparse credit assignment of reinforcement learning, we develop an offline trajectory rectification mechanism that identifies reasoning errors and missing multimodal evidence in RL-generated trajectories. The corrected trajectories provide supervised fine-tuning signals that reduce biases learned during RL. This mechanism supports a rectified bootstrapping pipeline in which initial RL exposes reasoning bottlenecks, supervised fine-tuning corrects them, and a final RL stage further improves the policy. AdSeek achieves 74.30 percent accuracy on the AdsCVR test split, outperforming its Qwen3-VL-8B-Instruct backbone by 27.90 percentage points. It also generalizes to the open- domain CrossVid benchmark, demonstrating effective active evidence gathering.

62. 【2610.03084】NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models

链接:https://arxiv.org/abs/2610.03084

作者:Omar Elfatairy,Maria A. Bravo,Jessica Bader,Zeynep Akata

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:benchmarks largely overlook, satisfy negated constraints, non-red cup, requested content, largely overlook

备注: *Equal contribution

点击查看摘要

Abstract:Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.

63. 【2610.03068】Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition

链接:https://arxiv.org/abs/2610.03068

作者:Fangzheng Wu,Brian Summa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Understanding compositional failures, diffusion requires identifying, Understanding compositional, diffusion requires, Compositional Stress Index

备注:

点击查看摘要

Abstract:Understanding compositional failures in text-to-image diffusion requires identifying both where stress is detectable and how intervention changes the output. We study these questions through a controlled anchor--stress protocol that jointly evaluates text-encoder diagnostics and denoiser interventions. We introduce a text-only Compositional Stress Index (CSI), which separates common from rare compositions across SD1.5, SDXL, and the SD3 text path and provides an upstream diagnostic coordinate. A matched six-prompt localization study links intervention location to distinct outcomes: residual-minimizing embedding adapters improve representation fit, while downstream cross-attention intervention increases color hit rate (CHR) by 0.0272. Across SD1.5 and SDXL denoiser blocks, the largest positive signed diagnostic-accessibility mean occurs at the deep encoder, whereas selective boost has its largest positive mean CHR response at decoder blocks. Selective subtraction and broad ablation reveal further modality- and architecture-dependent responses, including a substantial CHR decrease when SDXL decoder cross-attention is broadly ablated. We find a diagnosis-control dissociation under our controlled attribute-object composition setting: compositional defects are diagnosable before denoising, but the representation coordinate that exposes risk is not necessarily the coordinate or modality that improves generation.

64. 【2610.03051】BeeWhere: Segmenting Bumble Bee Colonies to Quantify Behavioral Effects

链接:https://arxiv.org/abs/2610.03051

作者:Roberta Hunt,August Easton-Calabria,James Crall

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:crop pollination globally, colony-level behavior remains, behavior remains difficult, occluded nest environments, Social bees

备注: Preprint. Accepted to ECCV 2026 Computer Vision for Ecology Workshop Proceedings. Proceedings DOI pending

点击查看摘要

Abstract:Social bees are important pollinators that support biodiversity and crop pollination globally and serve as important model systems for collective behavior, but scalable measurement of individual- and colony-level behavior remains difficult in dense, occluded nest environments. Existing monitoring workflows use fiducial tags (e.g., ArUco) to preserve individual identity, yet tag-based tracking can fail when markers are obscured and provide limited information about body extent, spatial context, and untagged individuals. We present BeeWhere, an AI-assisted annotation and analysis workflow that combines ArUco detections with deep-learnt instance segmentations to quantify bumble bee behavior from high-resolution colony images and videos. Using bumble bee (Bombus impatiens) microcolonies as a test case, we annotate 483 frames containing 8,443 bee instances. We additionally annotate pollen balls, nest structures, and chamber boundaries, and train YOLO instance segmentation models for downstream behavioral analysis. Instance segmentations enable quantification of important behavioral metrics based on body contours, including nearest-neighbor distance, proximity to nest structures, spatial occupancy within the nest, and detection counts over time. We apply the BeeWhere models to tag-based tracking in an exploratory validation study assessing the behavioral impacts of neonicotinoid pesticide exposure. BeeWhere increased detection rates compared to tag-based tracking, particularly when bees were partially obscured or under challenging imaging conditions, and also captured treatment-associated changes in bee spatial organization not captured using tag-based tracking alone. These results suggest that instance segmentation can complement fiducial-marker tracking by recovering behaviorally meaningful signals under challenging colony conditions.

65. 【2610.03047】Parasitic Co-Denoising: Unlocking 3D Human Motion Generation in a Frozen Video Diffusion Model

链接:https://arxiv.org/abs/2610.03047

作者:Yunjiao Zhou,Junlang Qian,Lihua Xie,Jianfei Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:synthesize realistic human, diffusion models synthesize, models synthesize realistic, realistic human motion, synthesize realistic

备注:

点击查看摘要

Abstract:Despite never being supervised on explicit 3D motion, large-scale text-to-video diffusion models synthesize realistic human motion in their generated videos. We ask whether this implicit knowledge can be turned into explicit 3D motion generation, without training a separate motion model. Probing a frozen Wan2.1 reveals that a recoverable motion signal is present in its intermediate states across the entire denoising schedule, not confined to the clean output. Motivated by this, we introduce parasitic co-denoising, a paradigm in which motion is decoded from the host model along its denoising schedule rather than produced by an independent generator. We instantiate it as the Parasitic Motion Decoder (PMD), an efficient flow-matching decoder that shares the host's noise schedule and reads its intermediate features through a $\sigma$-adaptive multi-layer fusion, leaving the host unmodified. Drawing its coverage from the host rather than from motion data, PMD leads dedicated motion generators on text-motion alignment at a small fraction of their trainable parameters, while producing paired video and motion in a single pass that motion-only baselines cannot match.

66. 【2610.03036】WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

链接:https://arxiv.org/abs/2610.03036

作者:Jiangang Han

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:vision-based web agent, challenge evaluates agents, vision-based web, evaluates agents end, WebRetriever Challenge

备注: 10 pages, 4 figures, 7 tables. Technical report of the 2nd-place solution in the WebRetriever Challenge 2026

点击查看摘要

Abstract:We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models.

67. 【2610.03034】Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling

链接:https://arxiv.org/abs/2610.03034

作者:Ella Kemperman,Luca Ambrogioni

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Diffusion models rely, solvers requiring time-discretization, numerical solvers requiring, requiring time-discretization, models rely

备注: Submitted to ICLR 2027

点击查看摘要

Abstract:Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization important. We adapt proportional-integral (PI) step-size control to diffusion, using our diffusion noise-normalised error estimator. Unlike existing adaptive methods in diffusion that respond only to the current error, the PI solver also incorporates the previous error, yielding smoother step adaptation. We further show that these per-sample trajectories exhibit shared structure and can be aggregated into a fixed schedule that retains much of the benefit of adaptive sampling. We evaluate both approaches on natural-image and language datasets, in terms of quality, measured by FID at a matched number of neural network evaluations (NFE), comparing them with widely used stochastic solvers and schedules. For images, our fixed discretization outperforms the commonly used EDM schedule in terms of sample quality when used with the stochastic Heun sampler, and with the EDM-churn sampler at low NFE. Additionally, our PI adaptive solver obtains better FID than most stochastic and adaptive baselines, although it does not beat the EDM-churn sampler at low NFE. Moreover, we find our solver outperforms both the EDM and the entropy schedule on language diffusion at low-to-medium NFE in terms of perplexity, with the drawback of lower token entropy. Lastly, we find that the benefit of per-sample adaptivity is problem-dependent. It is highly beneficial in 1D toy examples, while only marginal for image and language data, where the average schedule sometimes even outperforms the PI-adaptive solver. Code is available at this https URL

68. 【2610.03031】CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments

链接:https://arxiv.org/abs/2610.03031

作者:Feiyang Chen,Jincheng Hu,Yiduo Chen,Jihao Li,Yue Liang,Bingzhao Gao,Yanjun Huang,Yuanjian Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:quadruped robots remains, robots remains underexplored, crowded indoor environments, real crowded indoor, human-scene occlusion disrupts

备注: 8 pages, 4 figures. Submitted to IEEE International Conference on Robotics and Automation (ICRA) 2027

点击查看摘要

Abstract:Monocular semantic scene completion (SSC) for quadruped robots remains underexplored in real crowded indoor environments, where human-scene occlusion disrupts static geometry and human occupancy predictions are often incomplete or spatially misplaced. We present CrowdOcc, an RGB-D dataset and monocular SSC framework for this setting. CrowdOcc contains 25.1K frames from 11 indoor scenes, with semantic occupancy annotations constructed through static dynamic decoupling. Our framework combines: (i) Normal Guided Scene Geometry Fusion (NGSGF) to complement depth-aware lifting with surface-normal cues for occlusion robust geometry; and (ii) Human-Centric Sparse Interaction (HCSI) to selectively model human-human and local human scene relations in 3D. Our method achieves state-of-the-art SSC performance on CrowdOcc's scene-disjoint test set, reaching 15.80 IoU, 11.40 mIoU, and 46.23 Human IoU, demonstrating generalization to unseen indoor scenes.

69. 【2610.03022】ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

链接:https://arxiv.org/abs/2610.03022

作者:Sihan Ren,Gaozheng Li,Yuanshang Quan,Yiming Qin,Fuyi Yang,Chang Liu,Lan Xu,Minye Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:existing methods remain, methods remain largely, remain largely confined, assume pre-segmented inputs, Simultaneous Sign Language

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.

70. 【2610.03016】From Expression to Reaction: Role-aware Visual Transfer and Stimulus-guided Reasoning for Interlocutor Emotion Recognition

链接:https://arxiv.org/abs/2610.03016

作者:Wei Wang,Zhaowu Li,Jianjie Luo,Fu Lee Wang,Lap-Kei Lee,Zhenguo Yang

类目:Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV)

关键词:Stimulus-guided Boundary Reasoning, Role-aware Visual Transfer, interlocutor emotion recognition, Role-aware Stimulus-guided, framework for interlocutor

备注: Technical report of the second-place solution in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026

点击查看摘要

Abstract:In this paper, we propose a Role-aware Stimulus-guided (RASG) framework for interlocutor emotion recognition, which predicts listener emotions from listener-only videos and speaker-only audios. RASG consists of Role-aware Visual Transfer (RVT) and Stimulus-guided Boundary Reasoning (SBR) modules, which address supervision mismatch due to the lack of labeled listener data and ambiguity among visually similar listener reactions whose interpretation depends on speaker context, respectively. More specifically, RVT selects speaker samples whose facial expressions support their emotion labels. It then filters listener tracks and uses reliable pseudo-labels to train a listener-centric visual expert. SBR uses a two-class language reasoner only when the visual model is uncertain. It treats speaker audio and text as context rather than direct emotion evidence to distinguish similar listener reactions. Experiments conducted on MER-Cross dataset shows that RASG achieves 76.25\% on MER-Cross and improves the performance of the baseline over 17\%. Our team ranks second in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026.

71. 【2610.03015】OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

链接:https://arxiv.org/abs/2610.03015

作者:Runtong Wu,Fei Teng,Di Wen,Guoqiang Zhao,Kunyu Peng,Kailun Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:Vision Foundation Models, Vision Foundation, mobile embodied agents, Foundation Models, detection is essential

备注:

点击查看摘要

Abstract:Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes geometry and visual information differently, making object-relevant cues hard to model, localize, and preserve. We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. To resolve geometric mismatch, the ERP-Ray Geometry Adapter (ERGA-Ray) models spherical viewing rays and periodic spatial structure. To localize evidence in scene-wide context, the Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence and converts it into a structured geometric action. To recover local cues lost under fixed token budgets, the Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution for heading estimation. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. With target-specific geometry adaptation, VARC retains 95--98% of the same-configuration mAP, indicating reusable object-level 3D reasoning across sensing configurations. The source code will be made publicly available at this https URL.

72. 【2610.03013】RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

链接:https://arxiv.org/abs/2610.03013

作者:Hakjin Lee,Junghoon Seo,Jaehoon Sim

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Category-level object pose, pose estimation predicts, pose estimation, Category-level object, predicts the rotation

备注: Project page: [this https URL](https://yopo-series.github.io/RYOPO-project-page/)

点击查看摘要

Abstract:Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RGB-D methods rely on external instance segmentation and crop-based pose estimation, introducing separate stages and object-dependent processing costs that hinder real-time inference. To bring accurate pose estimation into real time, we present \ours{}, an end-to-end trainable query-based RGB-D set predictor. It jointly detects and segments objects and estimates their \mbox{9-DoF} poses without explicit CAD-derived shape priors or a separately trained instance segmentor. Shared image and scene encoding avoids repeated per-object crop encoding. A query-conditioned geometry pathway associates observed 3D points and RGB features with object queries and incorporates shared scene context. Object-centric refinement uses the resulting point descriptors to update an explicit pose state through pose-conditioned cross-attention and recurrent residual corrections. On NOCS, \ours{} substantially improves on published RGB-D joint detection and pose estimation results. It achieves competitive performance compared with two-stage methods under all-object evaluation on REAL275 and HouseCat6D, while enabling real-time full-frame pose estimation at $31.8$ FPS on an RTX~A6000. Project page: this https URL.

73. 【2610.03012】Rethinking Fixed Temporal Grids: Frequency-Disentangled Motion Generation

链接:https://arxiv.org/abs/2610.03012

作者:Yunjiao Zhou,Junlang Qian,Gen Li,Xinying Guo,Lihua Xie,Jianfei Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:fixed time window, uniform temporal grid, time window, fixed time, motion generation methods

备注:

点击查看摘要

Abstract:Most human motion generation methods encode motion as tokens on a uniform temporal grid, where every token spans the same fixed time window. Human motion, however, is temporally heterogeneous: slowly evolving global trajectories coexist with rapid transient events such as foot contacts and joint impulses. Forcing such multi-scale dynamics onto tokens of identical temporal resolution entangles motion frequencies, leaving slow regions redundant while smoothing out the rapid details that distinguish realistic motion. We propose \textbf{FreqMo}, a scale-adaptive motion representation that decomposes motion into wavelet frequency bands, separating dynamics across temporal scales while preserving temporal localization and exact reconstruction. Unified Frequency Residual Quantization (UFRQ) then encodes all bands within a single shared codebook, compressing the token sequence threefold and enabling stable single-stage generation. Experiments show FreqMo attains SOTA fidelity with substantially improved high-frequency preservation, and the same decomposition transfers to continuous diffusion backbones.

74. 【2610.03002】Recursive Self-Improvement in Unified Multimodal Models

链接:https://arxiv.org/abs/2610.03002

作者:Huijuan Wang,Chufan Shi,Cheng Yang,Yaokang Wu,Taylor Berg-Kirkpatrick,Xuezhe Ma

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Unified multimodal models, Unified multimodal, Unified, training, model

备注:

点击查看摘要

Abstract:Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.

75. 【2610.02974】From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation

链接:https://arxiv.org/abs/2610.02974

作者:Simon Schwaiger,David Seyser,Alessandro Scherl,Zlatan Ajanović,Wilfried Wöber,Gerald Steinbauer-Wagner

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:robot platform, purpose-trained models, Image-based traversability estimation, estimation is inherently, inherently dependent

备注:

点击查看摘要

Abstract:Image-based traversability estimation is inherently dependent on the robot platform, deployment domain, and mission preferences, which limits the applicability of purpose-trained models. To facilitate domain adaptation, this work aims to reduce the number of required annotations in the target domain using sample-efficient preference learning. Our method represents traversability through von Mises-Fisher mixture prototypes in a frozen vision-language feature space. Relative natural-language rules provide a commonsense prior, while sparse relative image annotations adapt the prototype directions and utilities to a target domain through computationally and sample-efficient fine-tuning. Experiments on WayFAST demonstrate accuracy competitive with end-to-end trained estimators while enabling sample-efficient image-based adaptation. Qualitative experiments further demonstrate the language prior's zero shot applicability and the fine-tuned estimator's improved dense prediction on semantic maps. Evaluation is complemented via semantic interpretation of learned prototypes by dissecting semantically close natural language prompts. Code and trained estimators available at this https URL

76. 【2610.02967】Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

链接:https://arxiv.org/abs/2610.02967

作者:Yuanhao Ban,I-Hung Hsu,Anastasios Angelopoulos,Wei-Lin Chiang,Ion Stoica,Cho-Jui Hsieh

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:remarkable visual quality, achieved remarkable visual, post-training remains challenging, single reward signal, visual quality

备注:

点击查看摘要

Abstract:Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (this https URL), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.

77. 【2610.02959】rraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

链接:https://arxiv.org/abs/2610.02959

作者:Shuai Fu,Jing Gu,Jian Zhou,Zicheng Duan,Gengze Zhou,Qi Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:made substantial progress, progress in photorealism, Recent, text-image alignment, aesthetics

备注: Accepted by NeurIPS 2026 (ED Track)

点击查看摘要

Abstract:Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationships. Such failures are not well captured by existing fidelity, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations spanning object-, interaction-, and scene-level failures, and employs a multi-stage evaluation framework to identify and quantify them. Given an image, TerraVis first uses an MLLM to assess its eligibility for evaluation, then detects violations across 18 taxonomy-defined types and classifies them as minor or major to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis achieves the strongest correlation with human judgments of world consistency among existing metrics. Our benchmark results further show that models that achieve strong performance on conventional metrics can still exhibit substantial world-consistency failures. These findings highlight world consistency as a complementary evaluation dimension and demonstrate that TerraVis enables systematic quantification, diagnosis, and comparison of such failures. Our code is publicly available at this https URL.

78. 【2610.02946】When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

链接:https://arxiv.org/abs/2610.02946

作者:Jihwan Hong,Woohyeon Park,Jaeik Kim,Jaeyoung Do

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video Object Segmentation, long temporal horizons, real-world applications, Video Object, increasingly important

备注: NeurIPS 2026 ED

点击查看摘要

Abstract:Video Object Segmentation (VOS) in complex and long videos is increasingly important for real-world applications, where target objects often appear only intermittently within long temporal horizons. However, existing benchmarks largely focus on temporally salient objects that remain visible for most of the video. To address this gap, we introduce FaVOS (A Benchmark for Video Object Segmentation with Fractional Temporal Visibility), a benchmark designed to evaluate VOS methods under low temporal visibility. We show that, in this regime, the standard JF metric can collapse VOS evaluation into absence classification, because empty predictions receive high rewards on target-absent frames. Consequently, even a trivial empty-mask predictor can outperform strong models such as SAM 3, revealing a fundamental mismatch between current metrics and practical VOS performance. To mitigate this issue, we propose Volumetric JF, which evaluates mask sequences as spatio-temporal volumes and reduces the dominance of target-absence rewards while preserving sensitivity to segmentation quality and temporal structure. Project page: this https URL.

79. 【2610.02943】Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy

链接:https://arxiv.org/abs/2610.02943

作者:Kaushik Bhargav Sivangi,Fani Deligianni

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Human Pose Estimation, information from RGB, combines complementary information, raise privacy risks, Human Pose

备注:

点击查看摘要

Abstract:Multimodal 3D Human Pose Estimation (3D HPE) combines complementary information from RGB, LiDAR, and mmWave radar, but models trained on correlated observations from the same individuals, raise privacy risks overlooked by record level analysis. We present a unified framework for multimodal 3D HPE that couples kinematics-induced sensor fusion with subject level privacy auditing and private training. First, our multimodal model aligns modality specific joint representation, injects skeletal structure and adaptively aggregates complementary sensor evidence for accurate pose prediction. Second, we formulate a black-box subject membership inference attack for 3D HPE, complemented by an empirical pointwise maximal leakage analysis, which characterizes how individual attack score outcomes change inference about the membership outcome. Third, we instantiate user-level differential privacy via Action Temporal Stratification, a population weighted within-subject sampling strategy that enforces action and temporal coverage. We evaluate our framework on the MM-Fi dataset across three diverse experimental protocols. Source-code will be released upon acceptance.

80. 【2610.02914】Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation

链接:https://arxiv.org/abs/2610.02914

作者:Yunseung Ok(1),Hyunsoo Kim(2),Minseo Kim(1),Suhyun Kim(1) ((1) Kyung Hee University, (2) The University of Texas at Austin)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generate minute-long videos, user-provided images, generate minute-long, Custom Forcing, Autoregressive video

备注: 31 pages. Project page: [this https URL](https://gustn9609.github.io/custom-forcing/)

点击查看摘要

Abstract:Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5--28.5 times faster than these long-video baselines.

81. 【2610.02903】ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling

链接:https://arxiv.org/abs/2610.02903

作者:Hailun Xu,Kanchan Sarkar

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:current VITOK progress, jointly preserves global, preserves global recognition, VITOK progress, single multi-teacher distillation

备注:

点击查看摘要

Abstract:We study how to consolidate the current VITOK progress into a single multi-teacher distillation recipe that jointly preserves global recognition and dense semantics. Our starting point is an AM-RADIO-style student distilled from SigLIP2 and DINOv3-L, where SigLIP2 supplies strong global semantics and DINOv3-L supplies stronger dense features. The central empirical issue is that the same recipe does not optimize all objectives equally well: changes that improve ImageNet-1K kNN accuracy can still degrade ADE20K segmentation. We summarize a progression of modifications that make this trade-off more explicit and more manageable: split adaptor heads for CLS and patch tokens, asymmetric cosine/MSE losses, initialization from a DINOv3-L checkpoint, teacher reweighting, masked image modeling (MIM), and PHI-S feature balancing. The resulting model reaches 83.2 patch kNN and 85.2 CLS kNN, slightly surpassing the DINOv3-L teacher on ImageNet-1K kNN classification, while PHI-S restores ADE20K performance from 46.5/58.1 to 48.5/61.0 mIoU/mAcc, matching the teacher on this dense benchmark. We also summarize negative results: scaling distillation from ImageNet-1K to ImageNet22K does not consistently help, and naively adding extra teachers such as SAM3 or HOG features introduces interference. Rather than claiming a final recipe, this paper distills the current project state into a compact empirical story and a concrete set of lessons for future iterations.

82. 【2610.02887】Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking

链接:https://arxiv.org/abs/2610.02887

作者:Haoyang Luo,Linwei Tao,Jie Gui,Xinghao Chen,Chang Xu,Jianyuan Guo,Minjing Dong

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, ensure reliable deployment, Large Language

备注: Accepted by NeurIPS 2026

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model's sensitivity to non-causal correlations, establishing its ability to capture MLLM's limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50% with comparable performance.

83. 【2610.02880】Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models

链接:https://arxiv.org/abs/2610.02880

作者:Qingtao Xia,Siyao Cheng,Jiahua Bao,Jiaxing Du,Jie Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Retrieval-augmented document question, document question answering, question answering assumes, Retrieval-augmented document, vision-language model

备注: 5 pages, 5 figures, 3 tables. Submitted to ICASSP 2027

点击查看摘要

Abstract:Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this assumption often fails, leaving a retrieval-reading gap: evidence found but not used. A paired protocol isolates this gap by comparing answers from the retrieved page images alone with answers from the same images plus their extracted text. On FoveDoc-Bench, our benchmark with traceable evidence, retrieval finds nearly every evidence page, yet adding CPU-OCR text raises strict accuracy by 13 to 16 points. An exact text layer roughly doubles the gain, which appears across six VLMs from three families and, within one family, narrows with scale without closing. The reader can read this evidence but cannot find it: crops of it recover most of the text gain, boxes around it on the page none. The same protocol identifies two boundaries. Extracted text helps on textual evidence but is neutral or harmful on charts and figures. Its advantage shrinks as retrieval degrades, and unrelated text of the same form adds nothing detectable. Extracted text is an amplifier of retrieval that works, not a substitute for retrieval that does not. Our code is available at this https URL.

84. 【2610.02876】Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

链接:https://arxiv.org/abs/2610.02876

作者:Jinchang Zhang,Guoyu Lu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large language model, multimodal large language, state, large language, spatial

备注:

点击查看摘要

Abstract:A multimodal large language model that correctly reports a spatial fact does not necessarily use that fact in subsequent reasoning. To study this distinction, we introduce \textsc{SpaceConflict}, a benchmark of 23{,}196 inputs for the construction and use of spatial state. Under a unified Supported/Contradictory/Unknown judgment interface, it covers local fact binding (L1), relational composition (L2), cross-observation consistency (L3), and state judgment under transformation (L4). Posing a direct-state query, a full-transformation query, and an explicit-initial-state query on the same world reveals an availability--utilization gap: models recover the initial state from visual evidence yet fail when that state must drive a transformation. For Qwen3.5-9B, 50 of 100 sequences with a correctly recovered initial state fail the full transformation, and supplying the state explicitly repairs all 50; the gap narrows with scale but does not close. We therefore propose Operational State Supervision (OSS), which supervises task-relevant spatial states and their transformation trajectories and aligns shared facts across contexts. OSS improves paired accuracy on matched judgments most on L3 and L4, the levels that depend on organizing and using state. Evaluating multimodal spatial reasoning thus requires asking not only whether a model can see and state a spatial fact, but whether that fact becomes a usable state in subsequent computation.

85. 【2610.02840】PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

链接:https://arxiv.org/abs/2610.02840

作者:Chunghyun Park,Beomjun Kim,Seungcheol Park,Heeseung Kwon,Yashu Shukla,Seunghoon Sim,Jinwoo Shin,Minsu Cho

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:dynamics guide accurate, internal world dynamics, world dynamics guide, World Action Model, learned internal world

备注: Preprint. Project page: [this https URL](https://chrockey.github.io/PointWAM)

点击查看摘要

Abstract:World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.

86. 【2610.02832】FastOPD: On-Policy Distillation for Lightweight VLA Deployment

链接:https://arxiv.org/abs/2610.02832

作者:Yoojin Oh,Jeongsol Kim,Yeonwoo Seo,Jangho Park,Seonghyun Jin,Sunwoo Park,Youngmin Kim,Youngjun Jun,Kyumin Choi,Jong Chul Ye

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:scaling incurs high, incurs high computational, high computational costs, deployment increasingly challenging, enhance manipulation performance

备注: Project page: [this https URL](https://fastopd.github.io/)

点击查看摘要

Abstract:Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of $\pi_{0.5}$ with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.

87. 【2610.02825】rrainForge: Physics-Grounded road geometry Editing for Counterfactual Autonomous Driving

链接:https://arxiv.org/abs/2610.02825

作者:Yang Chen,Yicheng Zhu,zhenning Li,Tao Li,Zilin Bian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:onboard cameras observe, surface conditions, wet or icy, icy pavement, vehicle motion

备注:

点击查看摘要

Abstract:Road geometry (e.g., crests, sags, and speed humps) and surface conditions (e.g., wet or icy pavement) affect how vehicles move, what drivers and onboard cameras observe, and how much clearance remains between vehicles. Editing these properties in a driving scene therefore requires corresponding changes in vehicle motion. Capturing these differences in a driving video requires a road edit to propagate to vehicle motion, camera viewpoint, and the clearance between vehicles. We present TerrainForge, a framework for generating road geometry-focused counterfactuals from reconstructed multi-vehicle driving episodes. A unified road model connects scene deformation with four-wheel vehicle dynamics, allowing crests, sags, speed humps, and friction changes to propagate through vehicle motion, camera viewpoint, and inter-vehicle clearance. Vehicle dynamics are evaluated against CarSim, and prescribed road geometry is verified in reconstructed Waymo scenes. Across 18 episodes, leaving surrounding vehicles on their recorded trajectories instead of recomputing their responses produces median peak differences in predicted ego-lead distance of 1.52 m for crests and 1.41 m for sags. We further simulate the ego response to 15,758 road edits across 983 braking episodes, pairing each edit with its safety outcomes relative to an unedited replay. These pairs train a first-stage screening surrogate that takes the original driving context and candidate road-edit parameters as input and predicts the resulting change in the ego's terminal gap. On held-out scenes, this prediction achieves 22-40% lower mean absolute error than predicting no change, so candidates can be screened cheaply before the full multi-vehicle rollout.

88. 【2610.02799】FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

链接:https://arxiv.org/abs/2610.02799

作者:Wenya Su,Kai Luo,Di Wen,Ruiping Liu,Yufan Chen,Junwei Zheng,Kunyu Peng,Kailun Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO); Image and Video Processing (eess.IV)

关键词:strong radial distortion, cameras give mobile, give mobile robots, practitioners routinely reuse, routinely reuse fail

备注: Source code will be available at [this https URL](https://github.com/Su-wenya/FUSEye)

点击查看摘要

Abstract:Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes much of the gap but requires abundant fisheye labels and compute. We present FUSEye, a training-light framework that turns a frozen-backbone COCO-pretrained extra-large YOLO26 detector (YOLO26-x) into a fisheye detector. FUSEye adds roughly 227k new parameters while updating the inserted modules and the pretrained detection head. It addresses the transfer gap at three causally linked levels. At the input level, overlapping grid view generation and box remapping (GridViews) enlarge compressed boundary regions. At the feature level, zero-initialized residual adapters (Z-Adapters) correct distortion-induced feature misalignment. At the decision level, learned cross-projection agreement fusion (AgreeFusion) promotes low-confidence detections only when they are supported by consistent evidence across multiple views. On the WoodScape surround-view fisheye benchmark, FUSEye raises YOLO26-x from 0.148 to 0.266 mAP50 and retains 84.3% fully fine-tuned accuracy. Moreover, randomly using only 25% of the labeled training images, FUSEye achieves 0.2597 mAP50, retaining 97.6% of its full-label performance. FUSEye also consistently improves YOLOv8-11 detectors, showing that the recipe is architecture-agnostic. Source code will be available at this https URL.

89. 【2610.02779】RAC: Trajectory-aware Reuse and Adaptive Correction for Efficient Autoregressive Video Generation

链接:https://arxiv.org/abs/2610.02779

作者:Jiaxing Song,Weiqi Yan,You Huang,Mingte Qiu,Huazhong Liu,Xiaofeng Zhu,Yunshan Zhong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:training-free framework, framework for efficient, video generation, generation, TRAC

备注: Preprint under review

点击查看摘要

Abstract:In this paper, we present trajectory-aware reuse and adaptive correction (TRAC), a training-free framework for efficient autoregressive (AR) video generation. Existing acceleration methods mainly target single-trajectory generation with bidirectional attention. AR video generation, by contrast, sequentially couples chunk-level denoising trajectories. Consequently, approximation errors accumulate and propagate through the generation process. TRAC addresses this challenge with three components, including robust cumulative scheduling (RCS), autoregressive trajectory-aware guidance scheduling (ATGS), and spectral structure correction (SSC). RCS selects cache reuse schedules by cumulative rollout error and cross-chunk/prompt variation. ATGS coordinates CFG refreshes along the global AR trajectory. SSC restores low-frequency structure of the first chunk to correct long-term structural loss. Experiments on SkyReels-V2 and FramePack-F1 show that, compared with existing methods, TRAC achieves both the highest inference efficiency and the best generation quality for AR video generation.

90. 【2610.02755】FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter

链接:https://arxiv.org/abs/2610.02755

作者:Yuqian Chen,R. Jarrett Rushmore,Guikun Chen,Fan Zhang,Edward Yeterian,Nikos Makris,Yogesh Rathi,Lauren J. O'Donnell

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:critical brain region, remains incompletely characterized, superficial white matter, abundant short-range association, short-range association fibers

备注: 22 pages, 3 figures

点击查看摘要

Abstract:The superficial white matter (SWM), a critical brain region for cognition across the lifespan and brain disease, contains abundant short-range association fibers whose organization remains incompletely characterized, in part because the short trajectories and highly variable cortical folding make correspondence across individuals challenging. Anatomically corresponding connections may vary in spatial location across individuals and therefore may not be adequately defined by geometric proximity alone. We introduce FiberGeoText (FGT), a vision-language model (VLM) for organizing short-range superficial white matter (SWM) streamlines reconstructed from ultra-high-resolution diffusion MRI into population-level clusters. FGT jointly represents three complementary properties of each streamline: its three-dimensional trajectory, its cortical anatomical context, and its shape. Cortical endpoint information from multiple parcellation schemes is expressed as text and encoded using a pretrained large language model (LLM), enabling heterogeneous anatomical descriptions to contribute to a common continuous representation. We evaluated FGT on acquired submillimeter 0.76 mm diffusion MRI data. Compared with state-of-the-art (SOTA) methods, FGT produced substantially greater cortical parcel coherence, within-cluster shape consistency, cluster-size consistency, and cross-subject correspondence. The trained model also generalizes well to unseen subjects with an average of 96.7% of the 5,000 learned clusters recovered, and high consistency of cluster structure between training and testing data. Together, these findings demonstrate that integrating geometric, anatomical, and shape information by learning multimodal deep embeddings with a VLM model enables robust learning of population-consistent SWM organization despite interindividual anatomical variability.

91. 【2610.02753】Correcting Guided Diffusion Trajectories with Spectral Alignment

链接:https://arxiv.org/abs/2610.02753

作者:Gihoon Kim,Taesup Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:image generation hinges, visual fidelity, hinges on fine-grained, fine-grained differences, differences in condition

备注:

点击查看摘要

Abstract:The practical success of conditional image generation hinges on fine-grained differences in condition alignment and visual fidelity. Classifier-free guidance (CFG) is central to this success, but its lack of an explicit criterion makes it difficult to assess whether the guided trajectory is progressing as intended. To address this gap, we show that spectral alignment provides a principled criterion for understanding guidance behavior and improving guided diffusion sampling through adaptive correction. Our analysis identifies the spectra of intermediate states as an indicator of consistency with the expected spectral evolution of the forward process. Based on this observation, we introduce Spectral Correction Guidance, a method that corrects deviations from an analytic reference spectrum during sampling. The proposed method is training-free and applicable across diffusion backbones and conditional generation tasks without modifying the underlying model. Experiments demonstrate consistent gains in preference-based metrics over baseline guidance methods in text-to-image generation and improved generation quality over CFG on ImageNet. These improvements persist across a range of guidance scales and with fewer denoising steps. Our analyses and ablations provide insight into guidance behavior and how the proposed method affects generation quality.

92. 【2610.02726】SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

链接:https://arxiv.org/abs/2610.02726

作者:Xi Ye,Yuzhu Wang,Xiaoyang Liu,Jiayi Wang,Yangyang Xu,Ruyu Wang,Wenlin Chen,Duo Su,Jun Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multi-view world models, world models generate, models generate realistic, generate realistic videos, fixed camera rigs

备注:

点击查看摘要

Abstract:Flow-matching-based multi-view world models generate realistic videos, but are commonly restricted to fixed camera rigs. Extending them to continuously varying camera poses requires paired pose--video observations with dense pose coverage, which are costly to acquire. We introduce \emph{SymRegFlow}, a symmetry-regularized flow-matching framework for multi-view-consistent video generation across continuous viewpoints without ground-truth novel-view RGB supervision. For each target pose, SymRegFlow geometrically warps source views into noisy anchors and combines masked dual-anchor supervision with cross-anchor denoising-output consistency to mitigate anchor-specific errors. Under an affine Gaussian surrogate, we prove that suitable consistency regularization recovers the clean-reference optimum at fixed noise levels, strictly outperforming single- and merged-anchor baselines. Experiments on Cosmos-Drive-Dreams and nuScenes demonstrate high-quality, multi-view-consistent autonomous-driving video generation: on nuScenes, SymRegFlow achieves the lowest FVD and FVMD among the evaluated baselines, reducing FVD by over 31\% relative to the best baseline, and source-conditioned inference also attains the best FID and instance preservation.

93. 【2610.02718】Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis

链接:https://arxiv.org/abs/2610.02718

作者:Peilin Yang,Xiaoyu Liu,Jian Sun,Qinghua Tao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:CLIP exhibit strong, strong semantic generalization, exhibit strong semantic, fine-grained visual perception, Vision-language models

备注:

点击查看摘要

Abstract:Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.

94. 【2610.02697】GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

链接:https://arxiv.org/abs/2610.02697

作者:Yixuan Jiang,Wentong Li,An Liu,Zihao Xin,Fulin Tang,Cong Leng,Yang Gao,Jian Cheng

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:egocentric RGB observations, systems increasingly adopt, map egocentric RGB, increasingly adopt streaming, adopt streaming Video-LLM

备注:

点击查看摘要

Abstract:Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inference-time price: depth sensors, 3D encoders, or per-step perception tool calls. We propose GeoScaffold, a geometric supervision framework that pays this price once, at training time, by internalizing geometry into the policy itself. It first learns a compact depth tokenizer on depth maps from the training trajectories and freezes it. It then fine-tunes the policy with a handful of learnable geometry query tokens, training their hidden states to reconstruct navigation-critical geometry such as depth, connectivity, and traversability. This supervision turns the query states into compact geometric latents for action decoding, and through the shared weights also internalizes geometry into the backbone's own representations. Like a scaffold, the tokenizer, target generators, and reconstruction heads are discarded after training, leaving the backbone and action interface unchanged. Extensive experiments show that GeoScaffold consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.

95. 【2610.02666】CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation

链接:https://arxiv.org/abs/2610.02666

作者:Jin Hyun,Jung Gyu Min,Gyuhyun Jung,Youngjoo Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models map visual, map visual observations, diffusion-based action expert, continuous robot actions, low-bit post-training quantization

备注: Accepted at ACCV 2026. 22 pages, including references and supplementary material

点击查看摘要

Abstract:Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration scales can be mismatched with activation ranges that vary with denoising progress and intended motion. We propose CHASE-VLA, a chunk-aware PTQ method that exploits a VLA-specific signal readily available from the policy: the generated action chunk, including its unexecuted future suffix. Rather than relying only on static scale matching for AE layers, CHASE-VLA combines the previously generated chunk as causal action context with denoising step group information to adapt AE activation scales. This enables W4A4 quantization of both MLP and attention projections in the repeated AE without modifying the pretrained policy. On LIBERO, CHASE-VLA achieves 97.3% average success rate on $\pi_{0.5}$ when both MLP and attention projections in the AE are quantized to W4A4, restoring FP16-level performance. CHASE-VLA also reduces the weight storage of the quantized AE linear layers by 73.4% and their single-chunk memory traffic by 70.9% and 71.2% on $\pi_{0.5}$ and GR00T N1.6, respectively, with a predictor overhead of at most 1.26% of the saved storage.

96. 【2610.02660】SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching

链接:https://arxiv.org/abs/2610.02660

作者:Zhendong Mi,Pu Zhao,Ziyu Hu,Xiaodong Yu,Yanzhi Wang,Grace Li Zhang,Shaoyi Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:repeated Transformer evaluations, enable high-quality interactive, high-quality interactive environment, repeated Transformer, Diffusion-based world models

备注:

点击查看摘要

Abstract:Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical structure of diffusion features largely unexplored. In this work, we reveal that world-model features exhibit highly stable singular subspaces across nearby denoising steps, while their singular values follow predictable evolution patterns. Building on this observation, we propose SpectralCache, a training-free spectral caching framework that reuses stable singular subspaces and estimates only low-dimensional singular values through linear extrapolation. We further exploit the spectral consistency between neighboring full-computation features to skip selected expensive backbone evaluations via singular value scaling. Extensive experiments on representative world models demonstrate that SpectralCache consistently improves inference efficiency while preserving generation quality. On HunyuanWorld-Voyager-13B, SpectralCache achieves 5.22x acceleration while maintaining a WorldScore of 65.90 for static scenes, substantially outperforming existing training-free caching methods in inference efficiency.

97. 【2610.02647】Capturing Dynamics: The 4D Facial Expression Intensity Dataset

链接:https://arxiv.org/abs/2610.02647

作者:Zesheng Wang,Alexandre Bruckert,Pierre Lebreton,Patrick Le Callet,Yante Li,Guoying Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:facial expression intensity, facial expression, expression intensity, expression intensity play, estimating facial expression

备注:

点击查看摘要

Abstract:The estimation and analysis of facial expression intensity play a crucial role in affective communication and human-computer interaction. Previous research has primarily focused on detecting and estimating facial expression intensity from frame-level 2D representations. However, this limitation restricts a comprehensive understanding of real-world facial expressions, as they are inherently 3D and temporally continuous. This paper investigates the perception of facial expression intensity by introducing the 4D Facial Expression Intensity Dataset (4DFEID). We employ a parametric face model and compile a total of 2,869 mesh sequences with controlled geometric variations, generating 4D data instances with diverse peak intensities and identity attributes. Using a Likert scale, we collect more than 90,000 subjective intensity perception ratings via a crowdsourcing platform. We explore various architectures and aggregation methods to establish baselines for episode intensity estimation on the new dataset, revealing that spatial-temporal graph models consistently outperform traditional frame-aggregation methods. In contrast to existing datasets that rely on 2D static imagery, the proposed 4D-FEID dataset provides the community with a unique and vital resource for investigating the perception of facial expression intensity through the use of dynamic 3D stimuli. By offering high-fidelity, spatio-temporally coherent facial data, 4D-FEID establishes a new foundation for research into more nuanced and naturalistic expression analysis, thereby addressing a gap in the current landscape of affective computing and human-computer interaction studies. The dataset is available at link.

98. 【2610.02626】Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

链接:https://arxiv.org/abs/2610.02626

作者:Shenglan Li,Zhendong Mi,Hengyi Zhu,Jingwu Luo,Chun Kit Chan,Geng Yuan,Yanzhi Wang,Pu Zhao,Shaoyi Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:improve robotic manipulation, increasingly incorporate intermediate, existing approaches primarily, approaches primarily reason, models increasingly incorporate

备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, however, introduces substantial computational overhead. We propose IG-VLA, a VLA reasoning framework that enables models to imagine the future and internalize the gist. Our Latent Spatiotemporal Reasoning learns to imagine task-relevant future scene evolution directly in visual representation space, guiding action prediction without costly pixel-level video generation. To further reduce inference overhead, we introduce Scene Gist Memory, which internalizes reasoning-derived scene-behavior associations into a compact Scene Gist Token, preserving the benefits of future reasoning while bypassing explicit future imagination at inference. Extensive experiments on LIBERO, LIBERO-Plus, and VLABench demonstrate the effectiveness and efficiency of IG-VLA. On the LIBERO-Plus Language suite, both the reasoning and gist policies outperform the strongest baseline by nearly 6% in success rate. The gist policy also achieves up to 6.38x speedup over baselines, reducing inference latency from 1081ms to 169.5ms per action chunk on a single NVIDIA A6000 GPU. These results demonstrate that future spatiotemporal reasoning can be effectively internalized for efficient VLA deployment.

99. 【2610.02611】Scale-Recursive Rectified Flows for Few-Step Precipitation Ensembles

链接:https://arxiv.org/abs/2610.02611

作者:Shunya Nagashima,Takumi Bannai

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Fine-resolution precipitation estimates, precipitation estimates support, estimates support flood, support flood risk, flood risk assessment

备注:

点击查看摘要

Abstract:Fine-resolution precipitation estimates support flood risk assessment and water management, but coarse satellite products cannot resolve rainfall within each grid cell. Generative models address this ambiguity by producing ensembles of plausible high-resolution rainfall fields. Among these models, rectified flows generate samples by iteratively transforming random noise into rainfall fields. Reducing the number of sampling steps accelerates generation but can make ensemble members too similar, understating uncertainty. We propose a scale-recursive rectified flow that generates broad patterns before local details and guides sampling-step allocation by comparing ensemble variability with prediction error across spatial scales. Validation scores and rainfall power spectra constrain the allocation to avoid excessive amplification. In satellite-to-radar downscaling over the contiguous United States, our analysis identified broad rainfall patterns as the main source of insufficient ensemble variability under reduced sampling budgets. Allocating more steps to the coarse flow improved probabilistic accuracy and rain detection across training seeds at fixed architecture and computational cost. The proposed model also achieved better probabilistic accuracy with shorter sampling time than a nonrecursive flow using more steps.

100. 【2610.02597】GRAFT: Growing Agglomerative Foundation Models via Continual Teacher Distillation

链接:https://arxiv.org/abs/2610.02597

作者:Zhenghao Zhao,Chi Zhang,Qingshuang Chen,Yelin Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:knowledge remains distributed, develop complementary capabilities, develop complementary, pretraining objectives, distributed across separate

备注:

点击查看摘要

Abstract:Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowledge remains distributed across separate, specialized models. Multi-teacher knowledge distillation offers a path toward consolidating these capabilities into a single agglomerative backbone, but existing approaches assume a fixed set of teachers, and incorporating a new teacher requires repeating expensive joint distillation over the entire teacher set. We introduce GRAFT, a continual multi-teacher distillation framework that enables a unified backbone to progressively acquire capabilities from an open-ended sequence of foundation models. When a new teacher arrives, GRAFT treats the previously distilled model as a teacher for preserving learned capabilities, while the current student jointly learns from both the previous model and the incoming teacher. Furthermore, to reconcile the incompatible representation geometries of heterogeneous teachers, we introduce Teacher Specific Readout Tokens, which grant each teacher an independent read-out of the shared encoder, together with Geometry Agnostic Relational Loss that aligns a vision-language teacher by matching image-text similarity structures rather than raw feature values. We provide GRAFT model, which is a single, continually extensible backbone that unifies five domains, including image understanding, 2D dense prediction, 3D human pose estimation, 3D vision, and vision-language, delivering strong performance across all of them while acquiring each new capability at the cost of a single distillation rather than a full re-distillation.

101. 【2610.02580】Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces

链接:https://arxiv.org/abs/2610.02580

作者:Yuxing Wang,Yizhou Wang,Anqi Li,Shuo Wang,Sameer Satish Pusegaonkar,Haoquan Liang,Jiajun Li,Shenxin Jiang,Jianhe Yuan,Shangru Li,Tongwei Dai,Zihao Chen,David C. Anastasiu,Sujit Biswas,Xunlei Wu,Zheng Tang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:indoor smart spaces, Smart Spaces, simultaneously provide large-scale, Physical AI Smart, indoor smart

备注: Accepted at NeurIPS 2026, Evaluations Datasets Track (poster)

点击查看摘要

Abstract:Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hospitals, retail venues, and similar settings, together with automatic annotations for multi-camera identities, 2D bounding boxes, 3D bounding boxes, camera calibration, and depth where available. The benchmark spans Isaac Sim synthetic generation, Cosmos Transfer appearance augmentation, and real-world Sim2Real evaluation. For the real-world target, we include two warehouse deployments with time-synchronized streams, automatic VGGT-based calibration, and a 3D labeling interface that projects world-frame 3D boxes into each view for cross-camera verification. We describe the dataset scope, annotation and calibration schema, generation workflow, benchmark protocols, and official evaluation system, which standardizes submission format, and leaderboard reporting. A central contribution is a 3D instantiation of Higher Order Tracking Accuracy (HOTA), extending the usual 2D box-based tracking evaluation to 3D locations and 3D boxes. We further report empirical baselines from the AI City Challenge leaderboards, showing how methods evolve from person-only 3D location tracking to multi-class 3D box tracking under realistic smart-space constraints. The release is available at this https URL.

102. 【2610.02567】DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering

链接:https://arxiv.org/abs/2610.02567

作者:Karthik Mohan Kumar,Damian Andrysiak,Pedro Antonio Pena,Kunal Tyagi,Rama Harihara

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Machine Learning (cs.LG)

关键词:outputs carry large, carry large variance, desired target depends, target depends heavily, generate high-fidelity images

备注: 5 pages, 3 figures, 2 tables. Accepted to SIGGRAPH Asia 2026 Technical Communications

点击查看摘要

Abstract:Diffusion transformers (DiTs) generate high-fidelity images from text and image conditions, but their outputs carry large variance and their faithfulness to a desired target depends heavily on how the condition is supplied. We present DAGS, a lightweight, attention-free, disentangled appearance and geometry conditioning scheme that steers a frozen image DiT to produce high-fidelity, highly faithful, and independently controllable renders. Two small convolutional encoders compute conditioning features once per frame and inject them as a learned, per-layer, element-wise residual into the image tokens, avoiding the quadratic cost of stacking conditions through attention. Because control and temporal handling live outside the frozen backbone, we retain its vast pretrained prior and eliminate backbone-overfitting risk. We further add a small recurrent lighting stabilizer and a training-free temporal guidance term that, coupled with our conditioning, elevate a per-frame image model into a streaming renderer. DAGS produces controllable, high-quality renders at a fraction of the compute of path tracing; it is not real-time, trading compute for controllability and quality. On a matched 1-spp + G-buffer input, per-frame DAGS reconstructs +8.6 dB / +10.1 dB PSNR over the real-time denoiser Intel OIDN and the diffusion renderer RGB-X while being 2.5-8x more temporally stable perceptually (temporal-LPIPS flicker).

103. 【2610.02561】Oracle headroom without signal: null-calibrated evaluation of candidate selection for thermal heart rate estimation

链接:https://arxiv.org/abs/2610.02561

作者:Mohammad Rakibur Rahman,Nhi Nguyen,Sasan Sharifipour,Le Nguyen,Manuel Lage Cañellas,Miguel Bordallo López,Constantino Álvarez Casado

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Camera-based physiological monitoring, produce multiple estimates, Camera-based physiological, extraction methods, facial regions

备注: 5 pages, 3 figures

点击查看摘要

Abstract:Camera-based physiological monitoring can produce multiple estimates from several facial regions, extraction methods, and processing settings. Signal quality indices aim to select reliable estimates without a physiological reference, and their potential is often assessed with an oracle that selects the estimate closest to the reference in each window. This retrospective selection can reward chance agreement. We model the effect with order statistics. For K independent candidates unrelated to the reference, the expected oracle error decreases approximately as 1/K. We analyze thermal heart rate estimation on 96 iBVP recordings with 168 candidates per 10 s window. The oracle achieves a mean absolute error of 0.91 bpm, compared with 10.74 bpm for the best fixed configuration, 18.03 bpm for the best quality index, and 8.61 bpm for a constant predictor. With K = 24, a forehead signal from another recording matches the correct one, with 4.62 against 4.61 bpm. Oracle evaluations should report candidate count, valid coverage, and matched null controls.

104. 【2610.02521】Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

链接:https://arxiv.org/abs/2610.02521

作者:Ying Yang,Guiyu Zhang,Lianghua Huang,Chang Nie,Chenyang Si,Haofan Wang,Shaoshuai Shi,Li Jiang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:shown strong potential, predicting future observations, future observations conditioned, long-video world models, shown strong

备注: 32 pages. Project page: [this https URL](https://spatial-memory-intelligence.github.io/)

点击查看摘要

Abstract:Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.

105. 【2610.02513】From Fragments to Global Maps: Learning Vectorized Map Aggregation with Large Language Models

链接:https://arxiv.org/abs/2610.02513

作者:Ziwei Li,Yi-Tang Chen,Xiaoqi Wang,Wenbin He,Han-Wei Shen,Liu Ren

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:structured road information, provide structured road, maps provide structured, Large-scale vectorized, essential for perception

备注:

点击查看摘要

Abstract:Large-scale vectorized HD maps provide structured road information that is essential for perception, localization, and planning in autonomous driving. Constructing such maps requires aggregating noisy, fragmented, and overlapping local predictions collected along a vehicle trajectory into a coherent global map. Existing aggregation methods typically rely on hand-crafted rules for fragment association and refinement. However, a fixed set of thresholds cannot effectively handle variations in road structures and prediction errors, often requiring detector-specific tuning or manual adjustment. To address this limitation, we propose MapMergeLLM, a data-driven framework that formulates vectorized map aggregation as conditional sequence generation with a large language model. Given serialized local vectorized maps, our model directly predicts the aggregated global map polylines. To reduce dependence on any particular upstream detector, we train the model on synthetic local maps generated from clean vector maps using corruptions that simulate representative prediction errors. We further introduce a coordinate tokenizer with geometry-aware pretraining to precisely represent map coordinates. In addition, we propose a line-level association loss that explicitly supervises correspondences between local observations of the same map element. Experiments on Argoverse2 and nuScenes using multiple recent upstream detectors demonstrate that MapMergeLLM substantially outperforms heuristic and optimization-based aggregation baselines without detector-specific retraining.

106. 【2610.02508】World Action Modeling with Progressive Visual Planning

链接:https://arxiv.org/abs/2610.02508

作者:Fei Zhang,Zhaochong An,Duncan Frost,Yikai Wang,Pengfei Liu,Ya Zhang,Michal Drozdzal,Amir Bar

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:observation and instruction, promising paradigm, paradigm for robotic, initial observation, future visual dynamics

备注: Project Page: [this https URL](https://sii-ferenas.github.io/ProWAM-page)

点击查看摘要

Abstract:World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in this https URL.

107. 【2610.02507】MeshQuery: Agentic Seam Planning for UV Parametrization

链接:https://arxiv.org/abs/2610.02507

作者:Marco Schouten,Arthur Roullier,Elie Michel,Ruben Wiersma,Axel Paris,Tamy Boubekeur

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:training-free agentic approach, production-grade quad meshes, training-free agentic, agentic approach, approach to automatic

备注:

点击查看摘要

Abstract:We present MeshQuery, a training-free agentic approach to automatic UV unwrapping of production-grade quad meshes. A Vision-Language Model (VLM) plans artist-aligned seams using a set of edge-selection tools, conditioned on domain-specific UV-unwrapping knowledge expressed in natural language and refined with a feedback loop. We design a queryable mesh representation together with a domain-specific language (DSL) that enables the agent to retrieve mesh information on demand, express a seam plan as a compact program of edge-selection operators over topological, geometric, and semantic mesh attributes, and iteratively refine it from UV quality feedback. On Adobe Substance 3D and Toys4K meshes, MeshQuery produces 2.9x/4.29x fewer charts and 1.63x/1.7x shorter seams than the strongest baseline, and professional artists prefer its results in 80.9% of comparisons. Ultimately, decoupling high-level intent planning from low-level edge selection and compact mesh representation lets MeshQuery run on different backend VLMs and scale to meshes an order of magnitude larger than autoregressive seam prediction

108. 【2610.02494】DeepStratNet: A Context-Aware Coordinate Regression Framework for Seismic Horizon Tracking under Sparse Labels

链接:https://arxiv.org/abs/2610.02494

作者:Aniq Ahmad,Musham Ahmad Malik,Ahmad Mustafa,Heather Bedle

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Automatic horizon tracking, Automatic horizon, foundational task, Automatic, horizon

备注:

点击查看摘要

Abstract:Automatic horizon tracking is a foundational task in 3D seismic interpretation. Most existing deep learning approaches formulate it as dense semantic segmentation, typically using U-Net-based architectures. The model produces a probability map over all pixels that must be post-processed to extract precise horizon coordinates, while horizon picks in time/depth must be converted into dense masks for training. Unpicked seismic traces are consequently treated as background, which can hinder convergence, and both pre- and post-processing can introduce errors into the final interpretation. Moreover, 2D segmentation models do not inherently capture inter-slice context, while 3D models are often computationally prohibitive. We instead formulate horizon tracking as a bounded coordinate regression problem, where the model directly predicts the time/depth coordinate of the target horizon at each lateral position. We propose a lightweight regression head compatible with any pretrained vision backbone, coupled with an LSTM module to model inter-slice context and produce a continuous horizon surface across the volume. A combination of L1 and L2 losses supervises predictions at valid horizon picks, while a geology-informed regularization enforces lateral continuity between successive traces. Under controlled experimental conditions, we evaluate four pretrained vision backbones under both segmentation and regression configurations on a seismic volume from New Zealand. The proposed approach consistently outperforms its segmentation counterparts quantitatively, using metrics including RMSE and PCC, and qualitatively, while also demonstrating greater robustness to increasing sparsity of training picks. Finally, we show that prediction variation across successive traces captures local variations in geological complexity, providing an automated quality control measure for downstream seismic interpretation.

109. 【2610.02468】Windfoil: Closed-Form Coverage for Real-Time and Differentiable Vector Graphics

链接:https://arxiv.org/abs/2610.02468

作者:Matt DesLauriers

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:differentiable vector graphics, quadratic Bézier contours, box-filtered winding number, present Windfoil, closed form

备注:

点击查看摘要

Abstract:We present Windfoil, a GPU-friendly algorithm that treats rasterisation and differentiable vector graphics as two sides of the same problem by evaluating the box-filtered winding number of quadratic Bézier contours in closed form. We implement this in WebGPU, allowing it to run across a range of environments, including a web browser on a consumer laptop, and apply the system to real-time 2D rendering, high-resolution rasterisation for print media, and a differentiable renderer. We compare our renderer against Skia, a production-grade engine, and Slug, a popular GPU rasterisation algorithm for games and real-time applications, measuring fidelity to a reference box-filtered coverage. Our renderer matches the reference more closely than either, at performance comparable to Slug. We also compare our optimiser against DiffVG and Bézier Splatting, where it reaches equivalent or better reconstruction quality at a fraction of the per-step cost, scaling to tens of thousands of shapes at interactive rates.

110. 【2610.02451】A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

链接:https://arxiv.org/abs/2610.02451

作者:Duowen Chen,Yuchen Sun,Zhiqi Li,Yuxuan Liao,Sinan Wang,Bart van Bloemen Waanders,Bo Zhu

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Effective wildfire monitoring, requires relating visual, relating visual evidence, Effective wildfire, monitoring requires relating

备注:

点击查看摘要

Abstract:Effective wildfire monitoring requires relating visual evidence to physical fire dynamics, yet real videos with synchronized physical annotations are scarce and high-fidelity 3D simulation is costly. We present a simulation-grounded vision-language model (VLM) framework that automatically converts 2D wildfire simulations into labeled video episodes. A fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues; controllable video generation supplies richer appearance. The proxies are intermediate representations rather than finely rendered final scenes. Generated videos and simulator labels form reusable multimodal memory for a training-free multi-agent VLM system that retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports. On held-out generated episodes, video memory achieves 51.5% exact four-tag accuracy, compared with 22.6% for direct VLM querying and 16-17% for text-only memory; the complete system achieves 77.3% accuracy on six simulator-derived report fields. Component ablations, cross-generator tests, and three real-UAV evaluations assess retrieval, reporting, generator changes, and observable monitoring tasks. The framework connects automatic simulation-to-proxy conversion with memory-based VLM reasoning under scarce real-world physical annotations.

111. 【2610.02421】An AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images

链接:https://arxiv.org/abs/2610.02421

作者:Mahmoud Raslan,Nada Omar,Omar Khaled,Tarek Waleed,Mohamed Hazem,Rania Mounir,Solwan Elsamanoudy,Ahmed Mourad,Noura Adel,Muhammad Rushdi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:patterned follicular miniaturization, increased single-hair follicular, single-hair follicular units, altered hair-shaft diameter, Androgenetic alopecia

备注:

点击查看摘要

Abstract:Androgenetic alopecia (AGA) is characterized by patterned follicular miniaturization, increased single-hair follicular units, and altered hair-shaft diameter. We present an automated quantitative scalp-analysis and clinical decision-support framework combining FU localization, ordinal visible-shaft counting, calibrated shaft-width estimation, regional aggregation, and an interpretable rule layer. The clinical cohort comprised 243 patients (127 AGA, 116 non-AGA), while the computer-vision experiments used 160 expert-annotated patients, 2,400 trichoscopic images, and approximately 158,000 FU annotations. Under patientdisjoint evaluation, YOLOv8m achieved test mAP@0.5=0.920 and recall=0.860; EfficientNet-B5 with a support-map channel achieved 87.0% expert-box count accuracy (macro F1=0.85). A separate 500-image set was processed end-to-end with detector-generated boxes, yielding MAE of 6.56 for follicle detection and 16.59 for follicle classification relative to human-expert annotations. The system is intended to assist, rather than replace, dermatologist interpretation.

112. 【2610.02388】Octrees as an Explicit 3D Language

链接:https://arxiv.org/abs/2610.02388

作者:Ran Dan,Si-Tong Wei,Pengfei Xiong,Wei Zhang,Yadong Mu,Peng-Shuai Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:latent codebook indices, removes spatial structure, large language models, model observes, latent codebook

备注: Project Page: [this https URL](https://plurato.github.io/OctLLM-page/) Code: [this https URL](https://github.com/octree-nn/octllm)

点击查看摘要

Abstract:Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by $17.4\%$ and raising render-grounded captioning by $28.7$ points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.

113. 【2610.02382】FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering

链接:https://arxiv.org/abs/2610.02382

作者:Zhongpai Gao,Benjamin Planche,Meng Zheng,Anwesa Choudhuri,Terrence Chen,Ziyan Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image-trained Gaussian proxies, Gaussian proxies typically, medical volume rendering, proxies typically bake, Transfer functions

备注:

点击查看摘要

Abstract:Transfer functions (TFs) control color and visibility in medical volume rendering, but image-trained Gaussian proxies typically bake one transfer function into their appearance. We present FactorSplat, a per-scene N-dimensional Gaussian splatting (N-DGS) proxy that accepts region-specific intensity-to-RGBA curves at inference. A local lookup applies the authored color and opacity change, while a shared functional encoder and low-rank per-Gaussian factors learn the residual appearance response. Geometry and directional appearance remain shared across presets, with visibility control and TF-aware pruning preserving the ability to hide and reveal structures. On seven CT and MR scans, FactorSplat improves mean PSNR and changed-region error over region-aware VEG across validation, interpolation, unseen composition, and out-of-distribution (OOD) edits. Across these four splits, seven-scan mean PSNR gains over VEG range from 1.10 to 1.52 dB. One checkpoint per scan supports unseen edits without retraining. At $1600^2$, the cached fast renderer averages 524 FPS with 1.17 ms TF switches. Project page: this https URL.

114. 【2610.02375】EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision

链接:https://arxiv.org/abs/2610.02375

作者:Ruiyang Hao,Zhi Qin Tan,Yulan He,Owen Addison,Yunpeng Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Dento-maxillofacial cone-beam, dozens of tooth-specific, spatial findings, Dento-maxillofacial, document image findings

备注:

点击查看摘要

Abstract:Dento-maxillofacial cone-beam CT (CBCT) reports may contain dozens of tooth-specific, anatomical, and spatial findings from a single 3D scan. Learning to generate such reports from limited clinical data is challenging because routine reports may not exhaustively document image findings, and a non-mention may reflect either absence or non-reporting. We present EviDent-CBCT, an evidence-bottlenecked framework designed for this incomplete supervision. An anatomy-aware network maps each CBCT scan to a discrete record of tooth-level, global, and tooth-IAC evidence. A dental-logic consistency projection reconciles incompatible evidence before a deterministic renderer and an image-blind local language model generate the report using only this record. For tooth-level evidence, reliability-aware training uses eligible non-mentions as reduced-weight negatives, while unreported global and tooth-IAC labels remain unknown. A metal-sensitive input channel preserves intensity cues from dental materials. Across three validation runs, EviDent-CBCT achieves $0.666\pm0.006$ merged evidence set-F1 and $0.402\pm0.003$ RadFact-Lite-Dental logical-F1, versus $0.371\pm0.018$ for the strongest controlled direct baseline. In the ODIN 2026 challenge, it ranked second in automated evaluation and third in blinded clinical Arena comparison on the hidden test set. These results support the discrete evidence record as an effective and auditable interface for CBCT report generation.

115. 【2610.02364】Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift

链接:https://arxiv.org/abs/2610.02364

作者:Ruben Dario Florez-Zela

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Explainability is increasingly, intelligent vehicles, poorly understood, increasingly required, models in intelligent

备注: Accepted at the 2026 IEEE International Conference on Vehicular Electronics and Safety (ICVES 2026). 6 pages, 4 figures

点击查看摘要

Abstract:Explainability is increasingly required for perception models in intelligent vehicles, yet whether explanations remain faithful under driving domain shift is still poorly understood. This work audits post-hoc explanations of a fixed YOLOv8s pedestrian detector across PIE and JAAD using ROI-based D-Deletion, frozen confidence terciles, rank-based tests, bootstrap intervals, and Holm correction. The audit shows that deletion-based faithfulness is strongly coupled to detection strength at explanation time, with Spearman correlations between 0.70 and 0.82 for D-RISE, making naive confidence-stratified comparisons unreliable. After controlling for detection strength within fixed f0 bins, D-RISE faithfulness remains domain-dependent in the central f0 range, with PIE showing higher D-Deletion than JAAD and Holm-adjusted significance. A non-perturbative EigenCAM baseline is less faithful than D-RISE but also exhibits score coupling, suggesting that the effect is not specific to D-RISE and is related to the deletion-based evaluation setup. These results motivate confidence-controlled XAI audits for safety-critical perception under domain shift.

116. 【2610.02343】SCOPE-4D: Endoscopic 4D Geometry Foundation Models

链接:https://arxiv.org/abs/2610.02343

作者:Chaoyi Zhou,Zhongpai Gao,Anwesa Choudhuri,Meng Zheng,Benjamin Planche,Run Wang,Terrence Chen,Siyu Huang,Ziyan Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:learning reliable endoscopic, scarce geometric annotations, robotic assistance, faces two challenges, endoscopic geometry faces

备注: Project page: [this https URL](https://chaoyizh.github.io/SCOPE-4D-page/)

点击查看摘要

Abstract:Geometric understanding supports endoscopic navigation and robotic assistance, but learning reliable endoscopic geometry faces two challenges: scarce geometric annotations and ambiguity between camera motion and tissue deformation. We present SCOPE-4D, an endoscopic 4D geometry foundation model that jointly predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. Our curation and annotation pipeline constructs SCOPE-5K, a collection of approximately 5,000 clips spanning real and synthetic gastrointestinal endoscopy and laparoscopy. The collection provides rich geometric supervision and includes newly collected phantom and real-colonoscopy evaluation sets. Geometric supervised fine-tuning on SCOPE-5K learns endoscopic priors that improve camera and depth estimation. Common--Residual Motion (CRM) further constrains local deformation relative to common tissue movement. Together with geometric supervision, CRM and trajectory supervision further improve camera and depth estimation over geometric fine-tuning alone while enabling dense 3D tissue tracking. Evaluations on public and newly collected benchmarks demonstrate strong in-domain and out-of-domain geometry, superior 3D tracking, and more stable long-sequence colon reconstruction. A blinded user study further supports the perceived reconstruction quality on real clinical video. Together, these results demonstrate the value of large-scale endoscopic supervision and motion constraints for joint geometry estimation and tissue tracking.

117. 【2610.02323】World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

链接:https://arxiv.org/abs/2610.02323

作者:Jie He,Wei Li,Junwen Tong,Rui Shao,Wei-Shi Zheng,Liqiang Nie

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:task-agnostic isotropic Gaussian, isotropic Gaussian source, policies generate action, generate action chunks, isotropic Gaussian

备注: 25 pages, 10 figures. Project page: [this https URL](https://github.com/JiuTian-VL/ProAct-page)

点击查看摘要

Abstract:Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As this source is conditioned on neither recent execution nor predicted future evolution, (i) it discards the local continuity established by recently executed motion. (ii) Even when predictive world representations are introduced, they often only condition the transport dynamics rather than determine where generation starts, how far it may deviate, or along which action directions it may expand. Building on this observation, we introduce ProAct, a world-calibrated proposal-to-action framework that makes the generative source itself predictable. (i) To preserve motion continuity, a lightweight Proposal Expert converts recent actions into a scene-aware hypothesis via one motion-anchored endpoint flow-matching step, initializing generation near the demonstrated action manifold. (ii) To jointly capture intended scene evolution and proposal-future compatibility, a prospective World Expert treats the hypothesis as a soft motion prior while predicting the task-consistent latent future. (iii) From this compatibility, the model calibrates a proposal-centered anisotropic source, where a bounded per-step extent controls the allowed deviation and a trace-normalized low-rank geometry under a condition-number budget allocates refinement over coupled translation, rotation, and gripper directions. Compared with $\pi_{0.5}$, ProAct improves performance across simulation and real-world tasks while reducing denoising steps by 50%, inference latency by up to 25.8%, and increasing throughput by up to 34.8%.

118. 【2610.02322】SCION: Scene Composition with Instanced Neural Primitives

链接:https://arxiv.org/abs/2610.02322

作者:William Koch,Amogh Joshi,Cyrus Vachha,Cheng Zheng,Felix Heide

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:tree leaves recur, Real-world scenes, blades of grass, natural environments, tree leaves

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hier- archical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this represen- tation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts. Project webpage: this https URL

119. 【2610.02320】DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

链接:https://arxiv.org/abs/2610.02320

作者:A. Said Gurbuz,Ahmed Nassar,Sunghwan Hong,Marc Pollefeys,Peter W. J. Staar

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:visually similar controls, similar controls compete, reliably ground action, ground action targets, complex desktop scenes

备注: 37 pages, 15 figures, 12 tables. Project page: [this https URL](https://saidgurbuz.github.io/deskforge/)

点击查看摘要

Abstract:Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: this https URL

120. 【2610.02298】EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

链接:https://arxiv.org/abs/2610.02298

作者:Ruihan Yu,Yu-Ju Tsai,Muyao Niu,Runyi Li,Lian Fu,Hanqing Liu,Zheng-Hui Huang,Yonghao Yu,Sho Kuno,Ming-Hsuan Yang,Kaipeng Zhang,Zhixiang Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:asset is built, editing, long sequence, implement the requested, single edit

备注: Project page: [this https URL](https://alaya-lab.github.io/EditHero/) , Code: [this https URL](https://github.com/AlayaLab/EditHero)

点击查看摘要

Abstract:3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for both geometry and texture. A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand. We use EditHero to compare 2 opposite approaches to 3D editing. Non-agentic methods operate top down, regenerating the object from a learned 3D representation and inferring what to keep. In contrast, LLM/VLM agents operate bottom up, editing through code that inspects the mesh and rewrites only the parts required by instructions. The non-agentic methods often miss the requested change and disturb regions that should stay fixed. Most LLMs follow instructions more closely, and all of them preserve the unedited parts better, but each of their edits takes minutes. We will release the engine and the edit sequences to support research on reliable iterative 3D editing.

121. 【2610.01939】Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

链接:https://arxiv.org/abs/2610.01939

作者:Ruiyang Si,Jianxin Bi,Shunyu Yang,Rui Ni,Wenbo Huang,Qiang Wang,Shulong Jiang,Duomin Wang,Xiuyu Li,Haiwen Feng,Zhen Dong,Daquan Zhou

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Vision language model, repeated model invocations, redundant observations incur, observations incur substantial, Vision language

备注:

点击查看摘要

Abstract:Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.

122. 【2608.23927】GlanceWAM: Sparse Test-Time Imagination for World-Action Models

链接:https://arxiv.org/abs/2608.23927

作者:Linhan Wang,Zijian An,Mingyuan Zhang,Chen Dai,Yi Xu,Can Cui,Jiayan Wang,Zichong Yang,Yinlin Chen,Lifeng Zhou,Chang-Tien Lu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:provide rich physical, rich physical priors, abandoning test-time visual, imagination sacrifices task, sacrifices task success

备注: Add real-robot experiments

点击查看摘要

Abstract:Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency $24\times$ relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than $\pi_{0.5}$ without any robot-data pretraining. Code is available at this https URL.

123. 【2610.03414】Iterating Consistency Models: Stability, Error Bounds and Noise Schedules

链接:https://arxiv.org/abs/2610.03414

作者:Alessio Spagnoletti,Abdul-Lateef Haji-Ali,Andrés Almansa,Alain Oliviero Durmus,Eric Moulines,Marcelo Pereyra

类目:Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:generating high-quality samples, Consistency models, leading approach, approach for generating, generating high-quality

备注: 27 pages, 6 figures

点击查看摘要

Abstract:Consistency models (CMs) have become a leading approach for generating high-quality samples in few steps. However, adding steps can improve or degrade sample quality in ways that are highly sensitive to the schedule and that existing theory does not fully explain. To provide accuracy guarantees and guide CM sampler design, we analyze multistep CM sampling as a composition of noising and approximate denoising operators. Under explicit, verifiable stability assumptions, we derive a non-asymptotic error bound that separates contraction of the initialization error from accumulation of approximation error. The bound assigns distinct roles to the schedule: large early noise levels drive contraction, while small late noise levels control the residual bias. As a corollary, we obtain explicit constants for strongly log-concave and semi-log-concave targets. We further establish a complementary guarantee whose assumptions, one-step accuracy and stability, can be estimated for a given trained model. Experiments show that the contraction and approximation profiles entering our bounds can be reliably measured and closely match the predicted functional forms. Together, these results provide a meaningful convergence theory for multi-step CMs and a practical route to sampler design.

124. 【2610.03290】Wrong Organ, Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening

链接:https://arxiv.org/abs/2610.03290

作者:Christiaan M. Geldenhuys,Joshua M. Jansen van Vüren,Véronique Suttels,Trevor Brokowski,Ablo P. Wachinou,Mary-Anne Hartley,Rensu P. Theart,Grant Theron,Thomas R. Niesler

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Lung ultrasound, attractive for tuberculosis, screening at primary-care, primary-care level, percentage points

备注: 10 pages, 3 figures, 4 tables. Accepted at SATNAC 2026, Drakensberg, South Africa, 11-14 October 2026

点击查看摘要

Abstract:Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on that high-resource ultrasound domain carries representations that remain usable in the low-resource one. Only the encoder varies, across seventeen encoders spanning three architecture families. Among them, a latent-predictive video encoder pretrained on generic video (V-JEPA2-L) and its echocardiography counterpart (EchoJEPA-L) differ in pretraining corpus alone. The choice among these encoders does not resolve the classification, the whole family spanning 2.50 percentage points against a measurement resolution of 2.71. What moves the task instead is feature conditioning. Standardising the features between the encoder and the classifier improves all seventeen encoders by a mean of +1.23 percentage points at $p=1.5\times10^{-5}$. On the held-out test set every encoder selected on the development folds stands above the baseline system by up to +2.57 percentage points of area under the receiver operating characteristic curve (AUROC), and specificity at 90% sensitivity reaches 79.3% against 60.3%. The contrast specified in advance, EchoJEPA-L against V-JEPA2-L, measures -0.16 percentage points at $p=0.926$. We therefore find no evidence that shared ultrasonic physics alone makes echocardiography a more productive pretraining corpus than generic video, and any advantage, if present, is smaller than this cohort can resolve. The video encoders receive replicated still images, however, so whether this absence of an effect reflects the pretraining domain or a video encoder applied to static frames cannot be separated. The limiting factor is the labelled cohort rather than the encoder.

125. 【2610.02675】One Photon, Many Worlds: Posteriors and Predictions with Single-Photon Cameras

链接:https://arxiv.org/abs/2610.02675

作者:Haejoon Lee(Carnegie Mellon University),Mohit Gupta(University of Wisconsin-Madison),Vijayakumar Bhagavatula(Carnegie Mellon University),Aswin C. Sankaranarayanan(Carnegie Mellon University)

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Single-photon avalanche diode, conventional cameras due, Single-photon avalanche, cameras operate fundamentally, operate fundamentally differently

备注: 18 pages, 18 figures

点击查看摘要

Abstract:Single-photon avalanche diode (SPAD) cameras operate fundamentally differently from conventional cameras due to their photon-counting nature. Each frame produces a binary image: pixels report zero if no photons arrived during exposure, and one if one or more photons arrived. Reconstructing a scene or inferring its properties from a single binary frame is difficult because many different images could produce the same measurement; thus, the inverse problem is fundamentally one-to-many. As we gather more binary measurements, the inherent uncertainty associated with the inverse problem and any associated inference diminishes. With sufficient photon counts, photon noise becomes negligible relative to the signal mean, enabling near-deterministic scene recovery and inference. This work characterizes the transition from stochastic to near-deterministic scene understanding as photon budget increases, analyzing how the stochasticity in photon arrival affects downstream inference tasks. Technically, we develop a conditional generative framework based on a Hypergeometric frame-thinning process for accumulated binary SPAD measurements. Generative models capture the one-to-many nature of photon-starved inverse problems, enabling empirical characterization of how this ambiguity diminishes with increasing measurements and its impact on downstream tasks like character recognition, QR code decoding, and facial analysis.

126. 【2610.02270】Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models

链接:https://arxiv.org/abs/2610.02270

作者:Xinye Yang,Zhusi Zhong,Scott Collins,Grayson Baird,Xuyu Wang,Zhicheng Jiao

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:prompting strategy, benchmark construction, factors in isolation, Medical vision-language model, studies treat

备注: 6 pages, 4 figures, 5 tables. Accepted version of a workshop paper presented orally at the 2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE). Code and per-case MIMIC-CXR results: [this https URL](https://github.com/Yangxinyee/cxr-vlm-routing)

点击查看摘要

Abstract:Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs (CheXagent, MedGemma-4B, and MedGemma-27B) are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal" predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.

127. 【2610.02269】Confidence-Gated Cloud-Edge Cascade Triage via Variational Risk Minimization for Medical Imaging

链接:https://arxiv.org/abs/2610.02269

作者:Xinye Yang,Zhusi Zhong,Scott Collins,Michael Bernstein,Grayson Baird,Terrence Healey,Michael Atalay,Mahesh Jayaraman,Xuyu Wang,Zhicheng Jiao

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Emergency chest X-ray, structural modality gap, require image-text inputs, multimodal foundation models, foundation models require

备注: 14 pages, 6 figures, 14 tables. Accepted manuscript of the article published in Smart Health 41 (2026) 100689. Presented as an oral at IEEE/ACM CHASE 2026. Code: [this https URL](https://github.com/Yangxinyee/vrm-edge-triage)

点击查看摘要

Abstract:Emergency chest X-ray (CXR) triage has a structural modality gap: reports arrive after triage decisions, yet multimodal foundation models require image-text inputs. We present Variational Risk Minimization (VRM), a distillation framework that treats LVLM-generated report variants as Monte Carlo samples of latent clinical interpretations. Rather than distilling from a single teacher target, VRM learns from a variationally marginalized teacher distribution, enabling uncertainty-aware supervision under missing-modality constraints. Under matched encoder families, VRM outperforms direct fine-tuning baselines and improves calibration with strong recovery from hallucinated supervision. Marginalized supervision reduces report-selection instability. In our compact edge-student instantiation, a confidence-gated cascade reaches AUC 0.941 at 103ms average latency with 20.3% cloud escalation, yielding an explicit reliability-latency operating point for cloud-edge clinical workflows.

128. 【2610.02265】Event-guided Neural Video Compression

链接:https://arxiv.org/abs/2610.02265

作者:Jiyun Kong,Jungwoo Kim,Enes Eray Demirtas,Touradj Ebrahimi,Jong-Seok Lee

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Neural video codecs, Neural Video Codec, video codecs derive, Event-guided Neural Video, codecs derive motion

备注: 28 pages. 21 figures

点击查看摘要

Abstract:Neural video codecs derive motion and temporal contexts mainly from RGB frames, leaving room for cross-modal guidance from complementary temporal observations. Event streams can provide such observations by recording brightness changes between frames. In this work, we propose an Event-guided Neural Video Codec (ENVC) that uses events shared by the encoder and decoder to improve RGB compression efficiency. For motion coding, ENVC forms an event-guided motion prior and codes the remaining motion residual. For frame coding, an event-conditioned predictor supplies multi-scale features for gated temporal context refinement. To support training and evaluation on standard video datasets, we synthesize paired RGB-event data and assess its predictive utility through comparisons with real events. Across six benchmarks, ENVC achieves average BD-rate savings of 39.13% using PSNR-RGB and 67.63% using LPIPS relative to DCMVC. Further analyses show that our gains persist on large-motion sequences and that ENVC effectively learns to integrate event information. These results demonstrate the potential of events as a complementary modality for reducing the RGB coding rate. Our model and code are available at this https URL.

129. 【2610.02263】ZAGNet: Zone-Aware Graph Aggregation Network for Patient-Level Lung Ultrasound Diagnosis

链接:https://arxiv.org/abs/2610.02263

作者:Li Chen,Shubham Patil,Rashid Al Mukaddim,Jochen Kruecker,Balasundar Raju,Alvin Chen

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:clinical examinations frequently, examinations frequently involve, frequently involve variable, incomplete scanning protocols, requires integrating findings

备注: The 2026 IEEE International Ultrasonics Symposium (IUS)

点击查看摘要

Abstract:Patient-level lung ultrasound (LUS) diagnosis requires integrating findings acquired across multiple anatomical zones, yet clinical examinations frequently involve variable and incomplete scanning protocols with missing zones. Existing diagnostic AI methods primarily analyze individual frames or video loops, relying on heuristic aggregation strategies such as max or mean pooling that ignore inter-zone relationships for patient-level inference. This paper presents ZAGNet, a Zone- Aware Graph Neural Network that represents temporally tracked pathology findings as graph nodes connected by anatomical zone adjacency. A graph transformer network propagates contextual information across neighboring lung regions, while a virtual global node aggregates graph-level features to predict patientlevel consolidation and pleural effusion using only patient-level supervision. ZAGNet accommodates missing zones by computing on a graph structure without fixed input format or size. We evaluate ZAGNet on a multicenter dataset of 714 subjects (20,256 LUS video loops) with exams varying from 4 to 16 zones across anterior, posterior, and lateral thoracic regions. For consolidation diagnosis, ZAGNet achieved an AUC of 0.803 compared to 0.677 (max pooling) and 0.674 (mean pooling). For pleural effusion, AUC increased to 0.893 from 0.804 (max pooling) and 0.815 (mean pooling). These represent improvements of up to 19% and 11% for consolidation and pleural effusion respectively. The results demonstrate that graph-based inter-zone reasoning provides an effective and clinically consistent framework for automated patient-level LUS assessment.

130. 【2610.02247】oward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots

链接:https://arxiv.org/abs/2610.02247

作者:Nam H. Le,Douglas Blackiston,Michael Levin,Josh Bongard

类目:Quantitative Methods (q-bio.QM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Robotics (cs.RO)

关键词:Artificial intelligence increasingly, letting people accomplish, intelligence increasingly serves, people accomplish sophisticated, accomplish sophisticated tasks

备注:

点击查看摘要

Abstract:Artificial intelligence increasingly serves as a natural-language interface to complex technical systems, letting people accomplish sophisticated tasks by describing what they want rather than specifying how to do it. Extending this interface to living systems is harder: unlike code or images, a biological intervention has no closed-form linguistic meaning, and the paired language-intervention-outcome data needed to learn such a mapping is expensive to collect, since each example requires its own wet-lab experiment. One way around this is to treat an existing archive of interventions and their already-observed outcomes as a fixed, offline dataset, and use a vision-language model to judge, without any new experiments, whether an archived outcome matches a natural-language description. But whether that judgment is reliable enough to train a language-to-intervention mapping on -- without new experiments and without human validation -- has remained unclear. Here we show that a natural-language interface for a living organism -- a xenobot, a synthetic multicellular construct with no nervous system -- can be learned entirely offline this way, using a vision-language model's own judgment as the sole training reward: an instruction is mapped to the intervention already on record as producing the described behavior. This mapping generalizes to entirely new instructions, evaluated against archive data withheld from training (80.0% held-out accuracy vs a $66.7% chance baseline, matching a network trained directly on ground-truth labels).