本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新1249篇论文,其中:

  • 自然语言处理185篇
  • 信息检索21篇
  • 计算机视觉215篇

自然语言处理

1. 【2610.02206】KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

链接:https://arxiv.org/abs/2610.02206

作者:Pengfei Li,Naufal Suryanto,Sicheng Zhang,Muzammal Naseer

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)

关键词:translate analysts' intent, LLMs are increasingly, increasingly applied, expected to translate, translate analysts'

备注: Accepted at NeurIPS 2026 Evaluations and Datasets Track. Project page: [this https URL](https://risys-lab.github.io/KaliBench/) | Github: [this https URL](https://github.com/RISys-Lab/KaliBench)

点击查看摘要

Abstract:LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

2. 【2610.02202】ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

链接:https://arxiv.org/abs/2610.02202

作者:Sohyeon Kim,Yoonho Lee,Bo Liu,Dayoon Ko,Rulin Shao,Seungone Kim,Graham Neubig,Pang Wei Koh,Aakanksha Chowdhery,Akari Asai,Omar Khattab,Yejin Choi,Gunhee Kim,Chelsea Finn

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:great scientists great, makes great scientists, great scientists, scientists great, makes great

备注: 57 pages

点击查看摘要

Abstract:What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

3. 【2610.02193】Hierarchical Continuous Diffusion Language Models

链接:https://arxiv.org/abs/2610.02193

作者:Hui Ren,Zihan Li,Chang Liu,Huidong Liu,Alexander Schwing

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:global constraint satisfaction, tasks demanding bidirectional, diffusion language models, Continuous diffusion language, language models offer

备注:

点击查看摘要

Abstract:Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: this https URL.

4. 【2610.02173】Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

链接:https://arxiv.org/abs/2610.02173

作者:Areeb Ahmad,Pratinav Seth,Vinay Kumar Sankarapu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:adjust and compensate, Ablate, Ablate a component, language model, mechanism remains unclear

备注:

点击查看摘要

Abstract:Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $\lambda$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(\lambda)=\mathrm{own}_r+\gamma_r\lambda$. The slope $\gamma_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $\gamma_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.

5. 【2610.02163】AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

链接:https://arxiv.org/abs/2610.02163

作者:Xuan Zhang,Longtao Zheng,Cunxiao Du,Bo An,Xin Dong

类目:Computation and Language (cs.CL)

关键词:solve repository-level software, repository-level software engineering, software engineering tasks, agents solve repository-level, code inspection

备注:

点击查看摘要

Abstract:Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.

6. 【2610.02150】From Knowledge Access to Source Learning: Developing Source-Specific Competence

链接:https://arxiv.org/abs/2610.02150

作者:Lucheng Fu,Kejing Xia,Yiyang Wang,Yiqiao Jin,Jinjin He,Xiyuan Yang,Haoxin Liu,Ye Yu,Haibo Jin,Yijia Xiao,Wenke Lee,B. Aditya Prakash,Haohan Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language model, agents increasingly rely, Large language, source, persistent external sources

备注: Website: [this https URL](https://sourcelearn.github.io/) Code: [this https URL](https://github.com/luchengfu6/SourceLearn)

点击查看摘要

Abstract:Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.

7. 【2610.02142】Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

链接:https://arxiv.org/abs/2610.02142

作者:Juan S. Santillana

类目:Computation and Language (cs.CL)

关键词:Keyword-matching benchmarks, benchmarks can credit, credit small models, Spanish security language, Keyword-matching

备注: 24 pages, 12 tables, preprint

点击查看摘要

Abstract:Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on |tool_call|), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.

Comments:
24 pages, 12 tables, preprint

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2610.02142 [cs.CL]

(or
arXiv:2610.02142v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.02142

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
8. 【2610.02140】Finetuning with Sampling: SFT Learns Better Than You Think

链接:https://arxiv.org/abs/2610.02140

作者:Aayush Karan,Sitan Chen,Yilun Du

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:predominantly employs supervised, employs supervised finetuning, predominantly employs, employs supervised, Introducing new capabilities

备注:

点击查看摘要

Abstract:Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

9. 【2610.02122】Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

链接:https://arxiv.org/abs/2610.02122

作者:Gabriel Tomitsuka,Arman Raayatsanati,Emma Xing,Duke Gand,Joseph J Ma

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)

关键词:performing statistical analyses, Real-world enterprise data, workflows require reasoning, Real-world enterprise, performing statistical

备注: 41 pages, 4 figures, 18 tables. Code: [this https URL](https://github.com/TextQLLabs/Argo-Bench) . Data: [this https URL](https://huggingface.co/datasets/textql/Argo-Bench) . Website: [this https URL](https://argo-bench.com)

点击查看摘要

Abstract:Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

10. 【2610.02117】Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

链接:https://arxiv.org/abs/2610.02117

作者:Sophia Sirko-Galouchenko,Monika Wysoczanska,Andrei Bursuc,Nicolas Thome,Spyros Gidaris

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:improving language-model reasoning, frozen or EMA, EMA version, recently emerged, improving language-model

备注:

点击查看摘要

Abstract:On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: this https URL

11. 【2610.02116】A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

链接:https://arxiv.org/abs/2610.02116

作者:Javier Diaz Esteban-Herreros,David Muñoz-Valero,Raquel Martínez-España,Jose M. Juarez,Juan Moreno-Garcia

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:comparative explainability framework, Explainable Artificial Intelligence, presented to audit, comparative explainability, explainability framework

备注: 18 pages, 6 figures

点击查看摘要

Abstract:A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.

12. 【2610.02092】Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

链接:https://arxiv.org/abs/2610.02092

作者:Zilin Du,Bowen Yang,Boyang Albert Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:training large language, large language models, Training-data Selection offers, critical for training, training large

备注:

点击查看摘要

Abstract:Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.

13. 【2610.02076】LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

链接:https://arxiv.org/abs/2610.02076

作者:Yinheng Li,Justin Wagle

类目:Computation and Language (cs.CL)

关键词:enabling software systems, categorical probability distributions, return categorical probability, generating free-form text, enabling software

备注:

点击查看摘要

Abstract:Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.

14. 【2610.02040】ypological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages

链接:https://arxiv.org/abs/2610.02040

作者:Nadine El-Naggar,Tatsuki Kuribayashi,Ted Briscoe

类目:Computation and Language (cs.CL)

关键词:attested natural languages, SOV, typological commonality, thousands of attested, attested natural

备注: EMNLP 2026 Main Conference

点击查看摘要

Abstract:Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.

15. 【2610.02039】CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

链接:https://arxiv.org/abs/2610.02039

作者:Yafei Zhang,Songshuo Lu,Sicong Liao,Zhi Chen,Yaohua Tang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language model, Recent years, language model, years have witnessed, witnessed the rapid

备注: 28 pages, 11 figures, 5 tables

点击查看摘要

Abstract:Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.

16. 【2610.02022】Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

链接:https://arxiv.org/abs/2610.02022

作者:Noy Sternlicht,Simra Shahid,Peter Jansen,Daniel S. Weld,Pao Siangliulue,Tom Hope

类目:Computation and Language (cs.CL)

关键词:judgment is increasingly, increasingly delegated, large language models, language models, novelty

备注:

点击查看摘要

Abstract:Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper's originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods.

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2610.02022 [cs.CL]

(or
arXiv:2610.02022v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.02022

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
17. 【2610.02019】Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

链接:https://arxiv.org/abs/2610.02019

作者:Guangyu Yang,Jingbiao Mei,Mingsheng Sun,Jinghong Chen,Yingtong Bu,Pengda Qin,Da Chen,Bill Byrne

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:video-based social media, increased users' exposure, video safety detection, automated video safety, safety detection

备注:

点击查看摘要

Abstract:The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at this https URL .

18. 【2610.02002】Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

链接:https://arxiv.org/abs/2610.02002

作者:Ahmad Yehia,Aly O. Abdelkareem,Islam Ahmed,Hesham Omran,Khaled Alashmouny,Christian Claudel,Abduallah Mohamed

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Model, Large Language, Language Model, LLM, Mem

备注: 15 pages, 4 figures

点击查看摘要

Abstract:Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at this https URL.

19. 【2610.02001】Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

链接:https://arxiv.org/abs/2610.02001

作者:Hao Wang,Ting Huang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Small open-weight models, tool prefill overflows, tool demonstrations loop, rarely complete real, Small open-weight

备注: 44 pages, 9 figures. Code, benchmark protocol, scoring code, and all 288 per-cell results: [this https URL](https://github.com/Mingbird/Mingbird-agent)

点击查看摘要

Abstract:Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $\tau^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.

20. 【2610.01984】Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

链接:https://arxiv.org/abs/2610.01984

作者:Hyunsik Kim,Youngmoon Jung

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:UBE, UBE matches BBPE, encoding, BBPE, English

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.

21. 【2610.01963】Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

链接:https://arxiv.org/abs/2610.01963

作者:Manar Aljohani,Brandon Ho,Kenneth McKinley,Dennis Ren,Xuan Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:influence acuity assignment, high-stakes prioritization task, improperly influence acuity, Emergency Severity Index, Emergency department

备注:

点击查看摘要

Abstract:Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.

22. 【2610.01947】Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry

链接:https://arxiv.org/abs/2610.01947

作者:Xinjian Zhao,Yaoyao Xu,Xuemin Chen,Xiaozhuang Song,Tianshu Yu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Large language models, multistep problem solving, language models offer, Large language, problem solving

备注:

点击查看摘要

Abstract:Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.

23. 【2610.01938】A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

链接:https://arxiv.org/abs/2610.01938

作者:Zhangshu Joshua Jiang,Zina Ibrahim,James T. Teo

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Exam-style accuracy, large language models, establish whether large, problem representation, clinical LLM benchmarks

备注: 13 pages, 1 table. Structured narrative review

点击查看摘要

Abstract:Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.

Comments:
13 pages, 1 table. Structured narrative review

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.01938 [cs.CL]

(or
arXiv:2610.01938v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.01938

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Zhangshu Joshua Jiang [view email] [v1]
Thu, 1 Oct 2026 16:07:12 UTC (26 KB)

24. 【2610.01921】Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

链接:https://arxiv.org/abs/2610.01921

作者:Lucas Bandarkar,Clark Peng,Ahmed Haj Ahmed,Aditi Khandelwal,Nanyun Peng

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:multilingual encoder training, varying multilingual tokenization, Cross-lingual contrastive learning, encoder training, core component

备注:

点击查看摘要

Abstract:Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.

25. 【2610.01917】MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

链接:https://arxiv.org/abs/2610.01917

作者:Yingcheng Liu,Tianyi Jiang,Yujuan Ding,jiangbo Ai,Xun Jiang,Guoqing Wang,Wei Ye,Yi Bin

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:repeated image operations, Latent, visual, reasoning equips vision, Latent visual

备注:

点击查看摘要

Abstract:Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.

26. 【2610.01889】Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

链接:https://arxiv.org/abs/2610.01889

作者:Yohan Chatelain(1),Pablo de Oliveira Castro(2) ((1) Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada, (2) Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France)

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:low-precision transformer inference, low-precision transformer, transformer inference, stochastic rounding, MLP

备注: 35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at [this https URL](https://github.com/big-data-lab-team/fuzzy-llm) and archived on Zenodo at [this https URL](https://doi.org/10.5281/zenodo.23066028)

点击查看摘要

Abstract:Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.

Comments:
35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at this https URL and archived on Zenodo at this https URL

Subjects:

Machine Learning (cs.LG); Computation and Language (cs.CL)

ACMclasses:
G.1.0; I.2.6; I.2.7

Cite as:
arXiv:2610.01889 [cs.LG]

(or
arXiv:2610.01889v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2610.01889

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
27. 【2610.01873】Where LLMs Fail with Visualization DSLs

链接:https://arxiv.org/abs/2610.01873

作者:Chang Han,Andrew McNutt,Katherine Isaacs

类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)

关键词:visualization domain-specific languages, necessarily easy, longer apply, domain-specific languages, role of authoring

备注: VIS 2026 VISxGenAI, 6 pages, 3 figures

点击查看摘要

Abstract:As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.

28. 【2610.01847】Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning

链接:https://arxiv.org/abs/2610.01847

作者:Zichen Xie,Mrigank Pawagi,Lize Shao,Yang Hu,Wenxi Wang

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:guiding alignment training, Model specifications define, define how large, large language models, inference-time behavior

备注:

点击查看摘要

Abstract:Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at this https URL.

29. 【2610.01846】Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

链接:https://arxiv.org/abs/2610.01846

作者:Serli Kopar,Alkis Koudounas,Roshan P. Rane,Sam Gijsen,Paula A. Perez-Toro,Kerstin Ritter

类目:ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)

关键词:Speech-based Alzheimer disease, Speech-based Alzheimer, assessments increasingly rely, pretrained self-supervised learning, Alzheimer disease

备注:

点击查看摘要

Abstract:Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.

30. 【2610.01828】he Asymptotics of Language Model Alignment with Memory

链接:https://arxiv.org/abs/2610.01828

作者:Haricharan Balasundaram,V. Arvind Rameshwar

类目:Computation and Language (cs.CL); Information Theory (cs.IT)

关键词:alignment broadly aims, Language model, higher expected reward, broadly aims, aims to perturb

备注:

点击查看摘要

Abstract:Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

31. 【2610.01821】Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

链接:https://arxiv.org/abs/2610.01821

作者:Tido Specht,Elias Benedict Krey,Nils Neukirch,Nils Strodthoff

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Understanding information processing, internal token representations, Understanding information, large language models, requires dissecting

备注: 24 pages, 13 figures. Code: [this https URL](https://anonymous.4open.science/r/NLMCD-NLP-C5E7)

点击查看摘要

Abstract:Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.

32. 【2610.01801】A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

链接:https://arxiv.org/abs/2610.01801

作者:Sahil Kadadekar

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:scored by cosine, cosine similarity, known-safe responses, reference, responses

备注: Accepted at the NeurIPS 2026 Workshop on Foundations of Language Model Security (FLMSec). 15 pages, 3 figures, 11 tables. Code, results, and a verifier are in the ancillary files

点击查看摘要

Abstract:Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.

33. 【2610.01785】VETO: Video Efficient Token Optimization for Vision Language Models

链接:https://arxiv.org/abs/2610.01785

作者:Gueter Josmy Faure,Hao Ping Wang,Min-Hung Chen,Winston H. Hsu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Processing long videos, Processing long, making long-form inference, inference prohibitively expensive, Vision-Language Models

备注:

点击查看摘要

Abstract:Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.

34. 【2610.01767】A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering

链接:https://arxiv.org/abs/2610.01767

作者:Gianluca Bonifazi,Christopher Buratti,Michele Marchetti,Federica Parlapiano,Giulia Quaglieri,Davide Traini,Domenico Ursino,Luca Virgili

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:multi-hop Question Answering, Question Answering, Large Language Models, Retrieval-Augmented Generation, balance retrieval quality

备注:

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.

35. 【2610.01702】ask-Oriented Rank Adaptation for Continual Learning in Text Classification

链接:https://arxiv.org/abs/2610.01702

作者:Rey Sanchez Lopez,Eduardo Morales Manzanares,Hugo Jair Escalante

类目:Computation and Language (cs.CL)

关键词:Continual learning, critical challenges, catastrophic forgetting, faces two critical, forgetting and negative

备注: Preprint submitted to CIARP2026

点击查看摘要

Abstract:Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.

36. 【2610.01696】Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information

链接:https://arxiv.org/abs/2610.01696

作者:Tian Lan,Xiaoqing Cheng,Han Zhang,Jiang Li

类目:Computation and Language (cs.CL)

关键词:Large language models, motivating extensive research, Large language, reproduce social stereotypes, motivating extensive

备注: 15 pages, 0 figures

点击查看摘要

Abstract:Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at this https URL.

37. 【2610.01688】Compound interpretation is based on analogy

链接:https://arxiv.org/abs/2610.01688

作者:Tian Shen,Harald Baayen

类目:Computation and Language (cs.CL)

关键词:remains a central, central question, CAM, computational models, CAOSS model

备注:

点击查看摘要

Abstract:How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.

38. 【2610.01634】Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

链接:https://arxiv.org/abs/2610.01634

作者:Ahmad Samuel Gali(1),Shamsuddeen Hassan Muhammad(2 and 3) ((1) University of Lagos, (2) Bayero University Kano, (3) Imperial College London)

类目:Computation and Language (cs.CL)

关键词:avoid lexical ambiguity, widely spoken tonal, spoken tonal language, Natural Language Processing, lexical ambiguity

备注: 7 pages, 3 figures, 3 tables. Code and outputs: [this https URL](https://github.com/lazy-monster/yo-byt5)

点击查看摘要

Abstract:Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.

39. 【2610.01627】What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

链接:https://arxiv.org/abs/2610.01627

作者:Peng Cui,Qiaoyuan Zheng,Rudolf Debelak,Mrinmaya Sachan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:differing ability, fundamental properties, meaningfully discriminate, Difficulty, Item Response Theory

备注:

点击查看摘要

Abstract:Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

40. 【2610.01616】Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

链接:https://arxiv.org/abs/2610.01616

作者:Laura van Weesep,Riccardo Tedoldi,Jens Sjölund,Hossein Azizpour,Susanne Winiwarter,Ola Engkvist,Jon Paul Janet,Samuel Genheden,Juan Viguera Diez

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB); Quantitative Methods (q-bio.QM)

关键词:molecular property prediction, property prediction requires, including reliable metadata, reliable metadata annotation, data readiness

备注: Accepted to the AIDaR workshop at NeurIPS

点击查看摘要

Abstract:The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's $\sim$2 million bioassays is critically sparse, 36\% lacking an assay format, 89\% a BioAssay type, and 99.9\% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.

41. 【2610.01592】Which LLM to pick? Online Active Model Selection for Large Language Models

链接:https://arxiv.org/abs/2610.01592

作者:Alessandro Turrin,Patrik Okanovic,Torsten Hoefler,Nezihe Merve Gürel

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:ONLINE LLM PICKER, approximate real performance, Large Language Models, LLM PICKER, ONLINE LLM

备注:

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online settings. Given an arbitrary stream of queries and a limited annotation budget, ONLINE LLM PICKER selects the most informative prompts for annotation to identify the best LLM among candidate models. Across multiple tasks including 10 datasets, for over 130 language models, we show that ONLINE LLM PICKER saves annotation cost by up to 71.67% while reliably identifying the best or near-best model for the stream. We also show that using the returned model for sequential generation on unannotated prompts across the stream reduces regret by up to a factor of 2.51x, indicating that ONLINE LLM PICKER can identify the best or near-best model well before processing all streaming prompts.

42. 【2610.01560】AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

链接:https://arxiv.org/abs/2610.01560

作者:Yuxiang Wang,Kunyu Feng,Yuancheng Wang,Zihang Liu,Shengbo Cai,Qinke Ni,Wan Lin,Tao Feng,Yingda shen,Ming-Hao Hsu,Zhixian Zhao,Liqiang Zhang,Teddy Sun,Steve Yves,Zhizheng Wu

类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)

关键词:fast response jointly, response jointly shape, intelligence and fast, shape the quality, quality of interaction

备注:

点击查看摘要

Abstract:Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.

43. 【2610.01554】QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

链接:https://arxiv.org/abs/2610.01554

作者:Ivan Ilin,Peter Richtárik

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:prunes large language, scoring weights independently, large language models, prunes large, dot products

备注: 81 pages, including appendices

点击查看摘要

Abstract:Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.

44. 【2610.01553】From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search

链接:https://arxiv.org/abs/2610.01553

作者:Nikolai Zenovkin,Sebastian Björkqvist

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:search requires processing, Patent search requires, routinely exceeding tens, requires processing documents, processing documents routinely

备注: Accepted for publication at the ECML PKDD 2026 conference (Applied Data Science track)

点击查看摘要

Abstract:Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.

45. 【2610.01514】How the Audit Rule Shapes Faithful Factor Explanations in LLMs

链接:https://arxiv.org/abs/2610.01514

作者:Taolin Zhang,Hanyu Wang,Jiuheng Wan,Tingyuan Hu,Chengyu Wang

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, input factors influenced, influenced their outputs, language models

备注:

点击查看摘要

Abstract:Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.

46. 【2610.01511】GAW-PO: Preference Optimization with Gradient-Aligned Token Weights

链接:https://arxiv.org/abs/2610.01511

作者:Andreea Dutulescu,Stefan Ruseti,Mihai Masala,Traian Rebedea,Mihai Dascalu

类目:Computation and Language (cs.CL)

关键词:autoregressive language models, Direct Preference Optimization, apply preference supervision, Direct Preference, autoregressive language

备注:

点击查看摘要

Abstract:Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter $\beta$ decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.

47. 【2610.01508】OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents

链接:https://arxiv.org/abs/2610.01508

作者:Taolin Zhang,Jiuheng Wan,Hanyu Wang,Tingyuan Hu,Chengyu Wang

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:access external services, request explicitly requires, user request explicitly, LLM agents, private user data

备注:

点击查看摘要

Abstract:LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.

48. 【2610.01493】No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

链接:https://arxiv.org/abs/2610.01493

作者:Lewis Mitchell

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Theory (cs.IT); Data Analysis, Statistics and Probability (physics.data-an); Machine Learning (stat.ML)

关键词:output diversity narrows, Iterative fine-tuning, progressively lost, narrows as rare, rare patterns

备注: 17 pages, 8 figures, NeurIPS 2026

点击查看摘要

Abstract:Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $h_k$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric ($p 0.23$), whereas $h_k$-filtering yields $+42\%$ unique trigrams, $+30\%$ vocabulary, and $-19\%$ repetition (all $p 0.001$). We validate $h_k$ as a cross-domain entropy proxy ($\beta = 0.924$, $R^2 = 0.746$) and collapse detector ($\rho = +0.454$, $p 0.0001$) across 4~domains, 2~temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.

49. 【2610.01492】Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization

链接:https://arxiv.org/abs/2610.01492

作者:Jeeyoung Yun,Seohwan Yun,Sungwoong Kim

类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:speech language models, Neural speech codecs, codecs increasingly serve, Neural speech, language models

备注:

点击查看摘要

Abstract:Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.

50. 【2610.01491】Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

链接:https://arxiv.org/abs/2610.01491

作者:Chengguang Gan,Zimeng He,Yoshihiro Tsujii,Ken-ichiro Kobayashi,Hiroki Itoh,Kotaro Funakoshi

类目:Computation and Language (cs.CL)

关键词:large language models, large language, language model evaluators, important application, application of large

备注: 13 pages, 1 figure, 10 tables. Accepted as a poster at the NeurIPS 2026 Workshop "Who Verifies the Agents? Toward Reliable Agent Development"

点击查看摘要

Abstract:Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.

51. 【2610.01490】he Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents

链接:https://arxiv.org/abs/2610.01490

作者:David Fraile Navarro

类目:Computation and Language (cs.CL)

关键词:Claude Opus, striking dissociation-like state, always-on personal agent, user on Discord, Paul

备注: 10 pages, 5 figures

点击查看摘要

Abstract:In February 2026, an always-on personal agent (``Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated ``heartbeat'' checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to ``Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emph{represented} from \emph{enacted} identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to ``I.''

Comments:
10 pages, 5 figures

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2610.01490 [cs.CL]

(or
arXiv:2610.01490v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.01490

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: David Fraile Navarro MD PhD [view email] [v1]
Thu, 1 Oct 2026 11:30:39 UTC (1,099 KB)

52. 【2610.01471】When Does a Second Model Help? Cross-Model Review in LLM Verification

链接:https://arxiv.org/abs/2610.01471

作者:Tae-Eun Song

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)

关键词:Large language models, Large language, generate code, review, Large

备注: 15 pages, 2 figures, 6 tables. Follow-up to [arXiv:2603.12123](https://arxiv.org/abs/2603.12123) and [arXiv:2603.21454](https://arxiv.org/abs/2603.21454)

点击查看摘要

Abstract:Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CCR plus one cross-model review matches more planted errors than two CCR reviews (56.7% vs. 42.7%; Holm-adjusted p=.006), but not significantly more than two reviews by the top-tier cross-model reviewer, so model difference and reviewer capability are not separated. A lightweight cross-model reviewer scores no higher than same-model review. Withholding requirements from the reviewer raises F1 for the two lower tiers but not the top tier, in untested point estimates whose pattern depends on how failed sessions are scored. Before analysis we audited all session records, excluding one baseline run of uncertain provenance and 14 failed calls; results with all sessions are also reported. A partial check on public detector outputs from another benchmark neither replicates nor contradicts the main comparison. Records, artifacts, and scripts are available from the author on request.

53. 【2610.01434】MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

链接:https://arxiv.org/abs/2610.01434

作者:Xudong Wang,Hao Wu,Haozhe Hu,Peiran Yin,Xinghao Chen,Yunpu Ma,Wei Zhang,Xiaoyu Shen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Multimodal large language, large language models, incur substantial inference, substantial inference costs, processing long visual-textual

备注:

点击查看摘要

Abstract:Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at this https URL.

54. 【2610.01428】Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

链接:https://arxiv.org/abs/2610.01428

作者:Nagham Omar,Mahmoud Jabarin,Maya Rozenshtein,Rom Himelstein,Avi Mendelson,Amit LeVi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:semantically stable outputs, large language models, large language, ability to produce, semantically stable

备注: Accepted at the TAE (Trust-AI-Eval) Workshop: Can We Trust AI Evaluation?, NeurIPS 2026

点击查看摘要

Abstract:Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.

55. 【2610.01427】SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic

链接:https://arxiv.org/abs/2610.01427

作者:Ben Sapirstein,Roy Mattar,Guy Mor-Lan,Ahlam Mohamed,Letizia Cerqueglini,Morris Alper

类目:Computation and Language (cs.CL)

关键词:Levantine Arabic, millions of people, creating a pressing, speech-language technologies, spoken by tens

备注: Accepted to ArabicNLP 2026. Project page: [this https URL](https://shams-nlp.github.io/)

点击查看摘要

Abstract:Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotated Multi-dialect Speech), a benchmark comprising 1,300 utterances drawn from open audio corpora, balanced across five LA varieties (Urban and Rural Palestinian, and Urban Jordanian, Lebanese, and Syrian). Each utterance is represented across four aligned tiers: audio, unvocalized orthography, diacritized text, and phonetic transcription. This structure supports evaluation of various downstream tasks such as diacritization, grapheme-to-phoneme conversion, automatic speech recognition, and audio-to-phoneme, grounded in audio and stratified by variety. We benchmark open and proprietary models across these tasks to demonstrate the utility of this benchmark for measuring progress across LA. We release SHAMS at this https URL .

56. 【2610.01393】LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction

链接:https://arxiv.org/abs/2610.01393

作者:Nouha Hayouni,Sheeba Samuel,Alsayed Algergawy

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Constructing typed, interdisciplinary knowledge domains, essential for enabling, enabling interoperability, interoperability across heterogeneous

备注:

点击查看摘要

Abstract:Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for ontology network construction that automates the discovery and generation of both intra-domain and inter-domain relationships. Our approach combines domain-adapted DistilBERT embeddings for dense contextual representation, clustering-based pre-filtering to reduce the candidate search space, and GPT-4o-driven relationship generation via iterative prompt engineering to produce semantically rich, interpretable links. Applied to ReproduceMeON - a network of 33 ontologies spanning machine learning, microscopy, computational science, and experimental workflow - the pipeline reduces approximately 800k raw concept pairs to 95k high-quality candidates. Human expert validation of 429 generated relationships by two independent annotators yields an overall precision of 80.19% (91.49% on high-certainty annotations) and an F1 of 0.890, with substantial inter-annotator agreement. Comparative experiments against five similarity-based baselines, including Sentence-BERT, show a substantial performance gap (best baseline F1 = 0.581), while an ablation study demonstrates that similarity-based methods alone fail to discriminate valid from invalid relationships (AUC approx 0.5) on the filtered candidate set. These findings highlight the necessity of LLM-based reasoning over concept roles and domain semantics for accurate relationship construction.

57. 【2610.01382】Gacha Decoding: Eliciting Diverse Generations Through Instruction Following

链接:https://arxiv.org/abs/2610.01382

作者:Scott Geng,Yufei Zhang,Joseph Lee,Jerry Li,Marjan Ghazvininejad,Pang Wei Koh

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:introduce Gacha Decoding, eliciting diverse language, Gacha Decoding, diverse language model, Gacha Decoding significantly

备注:

点击查看摘要

Abstract:We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.

58. 【2610.01378】Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

链接:https://arxiv.org/abs/2610.01378

作者:Sidi Chang,Peiying Zhu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Attributing model behavior, Attributing model, data requires knowing, knowing what produced, synthetic training data

备注: Accepted to the Third NeurIPS Workshop on Attributing Model Behavior at Scale: Data Attribution and Provenance. 4 pages, 0 figures, 1 table. An aggregate reproducibility package is available from the authors on request!

点击查看摘要

Abstract:Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.

59. 【2610.01353】Does AI-Generated Scientific Text Follow Human Argumentation Patterns? A CARS-Based Comparison of Research Article Introductions

链接:https://arxiv.org/abs/2610.01353

作者:Abdelrahman Sadallah,Narjes Sheikh Asadi,Lonneke van der Plas

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, human writing, helping write, write up research

备注:

点击查看摘要

Abstract:Large language models are moving from helping write up research to helping do it, which makes it important to know how the scientific text they produce differs from human writing. Work on this question has stayed mostly at the surface, using lexical and stylistic cues that light paraphrasing erases. We look instead at rhetorical structure, the sequence of argumentative moves through which a text makes its case. We study research-article introductions under Swales' CARS model, and compare original introductions from published linguistics articles with generated counterparts of the same papers. We find that human-written introductions are more flexible in which moves they use and in what order, while the generated ones are more uniform. Giving the models the CARS definitions makes them more rigid.

60. 【2610.01345】ARCCS: An Automated Regulatory Compliance Checking System

链接:https://arxiv.org/abs/2610.01345

作者:Giorgos Filandrianos,José Menezes,Chrysoula Zerva,Alessandro Gianola

类目:Computation and Language (cs.CL)

关键词:requires interpreting dense, interpreting dense legal, dense legal text, target document satisfies, Regulatory compliance checking

备注: This is the extended version of a paper accepted to EMNLP 2026 (System Demonstrations)

点击查看摘要

Abstract:Regulatory compliance checking - deciding whether a target document satisfies the obligations of a regulation - requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications. This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. We evaluate ARCCS in two complementary settings. First, in a GDPR policy-document evaluation, LLM-based judges find its decisions and justifications legally and evidentially consistent in up to 96.67% of the assessed cases. Second, on an EU public-procurement benchmark comprising more than 1,200 individual rule checks, the system attains 98.8% accuracy in violation detection. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.

61. 【2610.01324】Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes

链接:https://arxiv.org/abs/2610.01324

作者:Maryam Shahbaz Ali,Laura B. Strachan,Caitlin Sherman,Mark Kovler,Eleanor Mackey,Syed Muhammad Anwar

类目:Computation and Language (cs.CL); Emerging Technologies (cs.ET)

关键词:Patient-specific clinical question, question answering requires, answering requires locating, clinical question answering, heterogeneous longitudinal clinical

备注:

点击查看摘要

Abstract:Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.

62. 【2610.01316】What Wins a Vote? Formatting, Length, and Lexical Diversity in the French Compar:IA LLM Arena

链接:https://arxiv.org/abs/2610.01316

作者:Simonas Zilinskas,Maayeesha Farzana,Christophe Benavent

类目:Computation and Language (cs.CL)

关键词:LLM arenas turn, arenas turn pairwise, turn pairwise human, pairwise human preferences, LLM arenas

备注:

点击查看摘要

Abstract:LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 137,113 battles across 116 models, and the joint estimates use the 127,092 battles with all required measurements. For each battle, we reconstruct the response visible when the user voted. We then compare the raw ranking with rankings adjusted for formatting, length, readability, vocabulary variety, and sentence structure. Presentation is associated with winning, but length, bold text, and lists tend to occur together, making their individual contributions hard to separate. Across the measured features, two associations change least across specifications: bold usage (+11.0% win odds per standard deviation in the joint model) and moving-average type-token ratio (MATTR), a measure of vocabulary variety that is less sensitive to answer length (+16.8%). The bold association is substantially smaller in observed multi-turn conversations, whereas the MATTR association changes little; because users choose whether to continue, this difference is descriptive rather than causal. The full adjustment moves 36 of 116 models by at least ten ranks. Yet comparisons with external benchmarks do not show that adjusted rankings better measure capability. We therefore recommend publishing raw and adjusted rankings side by side as a transparent sensitivity analysis.

63. 【2610.01306】DAYJOB: A Benchmark for Long-Horizon Professional Work

链接:https://arxiv.org/abs/2610.01306

作者:Stephanie Finley,Liudas Panavas,Thomas Mikkelson,Cam Hinton,Stacey Ganss,Bradley Monton,Emily Kendall,Michelle Spradlin,Lydia Bye,Michael O'Brien,Lauren Ylvisaker,Derek Ray,Suhaas Garre,Sushant Mehta,Edwin Chen

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:request premise holds, documents matter, Professional work, work often starts, request that leaves

备注: 11 pages, 4 figures, 3 tables. An earlier version was accepted to the 2nd Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks (AABA4ET) at NeurIPS 2026. Evaluation harness: [this https URL](https://github.com/surge-ai/dayjob)

点击查看摘要

Abstract:Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.

64. 【2610.01278】SCOPE-AD: Sequential cost-aware ordinal-belief planning with energy-based models for diagnostic agents

链接:https://arxiv.org/abs/2610.01278

作者:Ziwen Yu,Ivan Koychev,Elizabeth Coulthard,Ting Zhou,Bolin Chen,Dian Hong,Zinuo You,Yujiao Wang,Anthony Mulholland,Qiang Liu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Alzheimer disease, requires sequential evidence, patient burden, diagnosis requires sequential, heterogeneous test costs

备注: 5 pages,2 figures

点击查看摘要

Abstract:Alzheimer's disease (AD) diagnosis requires sequential evidence acquisition under heterogeneous test costs and patient burden. Fixed-modality predictors do not jointly decide which test to acquire or when the available evidence is sufficient for diagnosis. We propose SCOPE-AD (Sequential Cost-Aware Ordinal-Belief Planning with Energy-Based Models for Diagnostic Agents) for cost-aware classification of cognitively normal (CN), mild cognitive impairment (MCI), and AD cases. A mask-aware ordinal model represents uncertainty along the ordered CN--MCI--AD continuum. Retrospective training records provide sampled Bellman targets for an energy-based teacher, whose action distributions are distilled into a Qwen policy. At deployment, the agent selects acquisition or diagnosis actions under availability and budget constraints without access to unacquired values. After each acquisition, the evidence and ordinal belief are updated before the next decision. On ADNI, SCOPE-AD achieves 77.70\% Macro-F1 at an average acquisition cost of \$50.46, exceeding the strongest evaluated baseline by 9.34 percentage points. Full-modality evaluation raises Macro-F1 by only 1.89 points while increasing acquisition cost by 116.7 times. These results support selective acquisition for cost-effective diagnosis.

65. 【2610.01275】Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models

链接:https://arxiv.org/abs/2610.01275

作者:Mojtaba Nafez,James Henderson

类目:Computation and Language (cs.CL)

关键词:Uniform-state diffusion models, Uniform-state diffusion, advantage over masked, Uniform-state, denoising step

备注: 38 pages, 8 figures

点击查看摘要

Abstract:Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173--270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3--11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.

66. 【2610.01257】Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems

链接:https://arxiv.org/abs/2610.01257

作者:Yiqiao Jin,Yiyang Wang,Lucheng Fu,Bing He,Siheng Xiong,Yijia Xiao,B. Aditya Prakash,Josiah Hester,Srijan Kumar,James Evans,Jindong Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:scientific literature co-evolve, literature co-evolve, Scientific progress emerges, Scientific progress, scientific literature

备注: [this https URL](https://ahren09.github.io/ScienceUtopia/)

点击查看摘要

Abstract:Scientific progress emerges from a longitudinal ecosystem in which researchers, institutions, funding agencies, collaboration networks, and the scientific literature co-evolve. As AI becomes increasingly involved throughout the scientific research cycle, understanding these interconnected and evolving processes becomes increasingly important. We introduce SciUtopia, a persistent, closed-loop LLM-agent simulation framework for studying academic research ecosystems. SciUtopia models interconnected scientific processes such as research-direction choice, collaboration, submission, peer review, resubmission, citation, funding, and researcher attrition, while maintaining evolving states across simulated years. Its configurable institutional mechanisms and information channels provide a controlled testbed for matched counterfactual experiments and targeted interventions. Across 61 simulation worlds, SciUtopia simulates over 40,000 researchers from 8,000 institutions, producing around 400,000 publication decisions and 1.2 million LLM-generated peer reviews. Using these longitudinal simulations, we find that rejection-driven resubmission substantially amplifies reviewer burden beyond population growth alone, cautious exploration balances citation impact with career success and long-term topic diversity, and resource inequality can emerge even without detectable cumulative advantage from narrowly winning early funding. Code is available at this https URL.

67. 【2610.01249】Revision-Aware Independent Agent Graphs for Dynamic Reasoning

链接:https://arxiv.org/abs/2610.01249

作者:Yan Luo,Selim-Antoine Lali,Jeremy Moebel,Iliass Khoutaibi,Ahmadou Aidara,Mengyu Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:propagates relevant updates, Conventional reasoning protocols, preserves unaffected work, agent propagates relevant, reasoning protocols present

备注:

点击查看摘要

Abstract:Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dynamic task routing}, in which an event stream revises task bindings and a system must select the document version valid at each query time before solving it. To study this problem, we repurpose six widely used benchmarks: MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval into 31{,}119 dynamic episodes comprising 373{,}428 temporally categorized queries. This setting exposes a central trade-off: recomputing after every event wastes work, whereas unguarded reuse returns stale conclusions. We introduce the Revision-Aware Independent Agent Graph (RIAG), a bounded multi-agent policy that separates deterministic temporal resolution from task reasoning. RIAG caches solutions by immutable document identity, starts each fresh task with two unexposed attempts, and conditionally invokes audit and repair, using at most four calls per document version. On this collection, homogeneous RIAG achieves 54.24\% joint routing-and-answer accuracy at 0.62 calls/query, compared with 32.22\% at 18.00 calls/query for the strongest comparison method; heterogeneous RIAG reaches 49.78\% at 0.63 calls/query.

68. 【2610.01244】Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration

链接:https://arxiv.org/abs/2610.01244

作者:Herun Wan,Jiaying Wu,Minnan Luo,Zihan Ma,Fanxiao Li,Nancy F. Chen,Min-Yen Kan

类目:Computation and Language (cs.CL)

关键词:Multi-agent systems, Multi-agent, correct answer, state, correct

备注:

点击查看摘要

Abstract:Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introduce OffQuery, which separately evaluates evidence verification (T1), shared-state reconstruction (T2), and task resolution (T3) in two representative high-stakes settings: healthcare and disaster response. Across GPT, Gemini, and Qwen models, standard collaboration shows much stronger task performance than state reliability. Averaged over 21 model--setting combinations, task resolution reaches 64.7%, while evidence verification and state reconstruction reach only 14.3% and 43.1%. We trace this gap to selective information use: current queries often bypass corrupted facts, which become consequential when later tasks require them. We further introduce ReGround, which resolves conflicting evidence, verifies shared facts, reconstructs a trusted state, and reasons over that state. Across seven models from three families, ReGround improves all three capabilities in every evaluated setting, with average relative gains of 309.0%, 82.9%, and 17.6% on T1, T2, and T3. Reliable collaboration therefore requires both a correct decision and a reliable shared state for future reasoning.

69. 【2610.01241】Evaluating the Robustness of Japanese LLMs to IME-Related and Typographical Errors

链接:https://arxiv.org/abs/2610.01241

作者:Ryota Mibayashi,Hiroaki Ohshima

类目:Computation and Language (cs.CL)

关键词:Large language models, language processing tasks, natural language processing, achieved strong performance, Large language

备注:

点击查看摘要

Abstract:Large language models (LLMs) have achieved strong performance across various natural language processing tasks. However, their robustness to typographical errors remains underexplored, particularly in Japanese, where text input involves multiple writing systems and IME-based conversion. In this study, we evaluate the robustness of Japanese LLMs against realistic Japanese-specific typos. We introduce five typo categories: Character Transposition, Character Replacement, Homophone Conversion, Japanese IME Conversion, and Full-Width Conversion. These perturbations are applied to three Japanese benchmark datasets (JMMLU, JCommonsenseQA, and JamC-QA), and eleven Japanese and multilingual LLMs are evaluated. The results show that Character Transposition and Character Replacement typos consistently reduce accuracy across benchmarks, whereas IME Conversion, Full-Width Conversion, and Homophone Conversion have relatively limited impact. These findings reveal that current Japanese LLMs remain vulnerable to realistic Japanese typing errors, particularly those that substantially distort the original input, highlighting the importance of robustness evaluation in practical input environments.

70. 【2610.01235】Harness Annealing: Learning to Act with Less External Control

链接:https://arxiv.org/abs/2610.01235

作者:Yingxuan Yang,Huacan Chai,Ying Wen

类目:Computation and Language (cs.CL)

关键词:Language agents rely, organize workflows, track state, verify answers, Language agents

备注:

点击查看摘要

Abstract:Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories can improve task performance while leaving these decisions dependent on runtime intervention. We ask whether harness-supported experience can also teach the model to make these decisions, allowing the division of control to change as the model learns. We call this objective harness internalization: learning to assume specified control responsibilities while retaining task performance after the corresponding support is withdrawn. We introduce HARNESS ANNEALING TRAINING (HAT), which combines explicit control supervision with a curriculum over teacher trajectories collected under progressively weaker harnesses. Experiments with 9B and 35B models on SWE-QA and SWE-QA-Pro evaluate every checkpoint under four deployment harnesses. Selected annealed checkpoints operating with tools alone achieve scores close to those of their respective starting checkpoints deployed with the full harness. The benefits vary with model scale and deployment configuration, and further annealing does not uniformly improve performance. These findings suggest that harness-supported experience can help reduce the runtime control required by a trained agent.

71. 【2610.01234】ASCRIBE: Atomic and Significance-Based Reasoning for Thai Clinical SOAP Note Generation

链接:https://arxiv.org/abs/2610.01234

作者:Tarm Kalavantavanich,Teerawut Ponarchar,Pattaramanee Arsomngern,Jenta Wonglertsakul,Watcharakorn Chuthong,Chiraphat Boonnag,Knot Pipatsrisawat,Titipat Achakulvisut

类目:Computation and Language (cs.CL)

关键词:Automatic SOAP note, generate unsupported content, Automatic SOAP, SOAP note generation, omit clinically important

备注:

点击查看摘要

Abstract:Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physician-inspired reasoning framework that ascribes a clinical-significance level to each extracted atomic fact in the conversation before summarization, making a general-purpose LLM a more reliable scribe. We also release ThaiClinicBench, the first de-identified Thai clinical summarization benchmark of real encounters, together with a synthetic training corpus derived from real clinical notes. As a prompt, ASCRIBE outperforms chain-of-thought prompting on GPT-5.4 and Gemini 3.1 Pro across the physician-aligned LLM-judge metrics and improves on standard prompting by up to 10.3 points on the completeness LLM-judge metric. As a GRPO reward, it enables a Gemma-4-E4B model trained solely on synthetic data to match Gemini 3.1 Pro in factual precision and surpass it in completeness. Code and data can be found at this https URL.

72. 【2610.01218】AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

链接:https://arxiv.org/abs/2610.01218

作者:Giulio Zeloni,Enrico Lo Conte,Salvatore Rionero,Giuseppe Santoro,Alessandro Rastelli,Fabio Sorrentino

类目:Computation and Language (cs.CL)

关键词:Enterprises adopting retrieval-augmented, adopting retrieval-augmented generation, Enterprises adopting, recurring operational decision, retrieval-augmented generation

备注: 14 pages, 1 figure, 4 tables. Submitted version (pre-review). Accepted at NFMCP 2026, ECML PKDD 2026 Workshops

点击查看摘要

Abstract:Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining deterministic checks, local guardrails, and structured LLM evaluation; a stratified beta-binomial gate that quantifies regression risk probabilistically; and a mandatory meta-evaluation protocol to validate the LLM judge before it influences decisions. Since engagement data is proprietary, we evaluate the judge layer on RAGBench, a public benchmark of 100k annotated RAG traces across 12 datasets. On identical stratified test samples (N=1200 per judge), a low-cost judge (gpt-4.1-nano) detects non-adherent answers barely above chance (AUROC 0.603 [0.570, 0.634]), despite producing flawless protocol output, while gpt-4o reaches 0.783 [0.756, 0.807] -- yet its per-domain performance still ranges from 0.62 to 0.88. A fixed-seed gate study spanning regression, no change, and improvement quantifies unsafe promotion, false-alarm cost, and improvement throughput. Under regression, the decision-grade profile reduces unsafe promotion to 22.2%-35.1%, against 29.3%-41.8% for a naive gate. These results support the design choices that judge quality must be measured per engagement and that point estimates alone are not a release decision.

73. 【2610.01184】ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning

链接:https://arxiv.org/abs/2610.01184

作者:Bingchen Pei,Lichong Chen,Bingxi Zhao,Ziang Wu,Sirui Wang,Min Zhang,Yanhao Chen,Qingxu Liu,Qiang Gao,Chang-Tien Lu,Bo Gao

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:risks exposing sensitive, exposing sensitive content, multimodal models offer, models offer strong, offer strong numerical

备注: 24 pages, 10 figures

点击查看摘要

Abstract:Remote multimodal models offer strong numerical reasoning capabilities over charts and speech, but sending private inputs risks exposing sensitive content. Text-only sanitization cannot directly satisfy fixed media interfaces, while identity anonymization leaves the underlying task content exposed. We introduce ReCast, an agentic plug-in framework that replaces source-specific content while preserving task-relevant relations and the required input modality. ReCast locally converts inputs into a shared textual evidence-query record, jointly rewrites entities and topics with a distilled 4B model, and substitutes values through a locally invertible, role-aware numerical map. A reconstruction agent generates and validates the required media from the protected record. The remote solver returns a program whose protected operands are restored locally before execution. On 4,000 held-out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of unprotected remote accuracy, while a model-based audit flags source-content leakage in 7.95% of solver-bound requests. It outperforms all evaluated local baselines, preserving the benefit of remote reasoning while reducing source-content exposure under existing media interfaces.

74. 【2610.01177】mporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models

链接:https://arxiv.org/abs/2610.01177

作者:Darpan Aswal,Céline Hudelot

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:extends Integrated Gradients, Layer Integrated Gradients, Integrated Gradients, Diffusion Layer Integrated, diffusion language models

备注:

点击查看摘要

Abstract:This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.

75. 【2610.01170】HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix

链接:https://arxiv.org/abs/2610.01170

作者:Zirui He,Haiyan Zhao,Jingyu Hu,Yinghao Wu,Chenxi Yuan,Yingcong Li,Yandong Bai,Mengnan Du

类目:Computation and Language (cs.CL)

关键词:eliminate behavioral errors, model, language models, model behavior, eliminate behavioral

备注: 31 pages, 18 figures, 7 tables

点击查看摘要

Abstract:Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.

76. 【2610.01165】Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

链接:https://arxiv.org/abs/2610.01165

作者:Shengye Tao,Yinzhu Cheng,Haihua Xie

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:language models, model representations, probe the internal, analyses examine, representations and computations

备注: 24 pages, 12 figures

点击查看摘要

Abstract:Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-forced contexts across five released trajectories and 11 model-domain combinations. We find that block-bypass responses retain recognizable depth ordering while their magnitudes redistribute: nearby checkpoints preserve stronger rank correspondence than distant ones, and large changes concentrate at positions that recur across text samples and transfer across evaluation domains. Controlled experiments further show that changes in the natural bypass effect cannot be reduced to a single downstream sensitivity: in replicated Pythia runs, local missing-update magnitude grows while the pooled matched downstream response decreases, whereas OLMo-2 7B exhibits a different balance. These matched responses also depend on perturbation strength and direction, without identifying targeted compensation. Together, our results show that longitudinal layer sensitivity is structured but not static, and that single-checkpoint intervention responses should be interpreted in the context of how the underlying perturbation pathway evolves during training.

77. 【2610.01161】My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

链接:https://arxiv.org/abs/2610.01161

作者:Yihua Zhu,Qianying Liu,Weixu Qiao,Xuan Ren,Weiwei Xu,Wenbo Li,Wei Wang,Ruijia Chen,Xinmiao Luan,Yin Luo,Hao Huang,Xiang Zheng,Hidetoshi Shimodaira

类目:Computation and Language (cs.CL)

关键词:Agentic reinforcement learning, Agentic reinforcement, training large language, credit-assignment problems, language model agents

备注: Preprint

点击查看摘要

Abstract:Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.

78. 【2610.01150】BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

链接:https://arxiv.org/abs/2610.01150

作者:Hasin Almas Sifat

类目:Computation and Language (cs.CL)

关键词:natural language processing, hostile Bangla text, Bangla natural language, remains an important, important challenge

备注: 5 pages, 3 figures, 1 table. Dataset Version 1.0 available on Zenodo: [https://doi.org/10.5281/zenodo.23074319](https://doi.org/10.5281/zenodo.23074319)

点击查看摘要

Abstract:Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class. The resource was constructed using a corpus-grounded synthetic procedure incorporating regional variation in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, vocabulary, and Bengali-script spelling conventions while preserving the underlying hostile or abusive meaning. Descriptive analysis shows broadly comparable sentence-length distributions but partially distinct lexical spaces across the four classes. Pairwise Jaccard vocabulary similarity ranges from 0.37 to 0.56. The primary task is four-class regional dialect identification rather than binary abusive-text detection. The dataset is publicly available through Zenodo under a Creative Commons Attribution 4.0 license. The current version is intended as a research and prototyping corpus rather than a native-speaker-validated gold-standard linguistic resource. Keywords: Bangla, Bengali, dialect identification, regional dialect, abusive language, low-resource NLP, Chattagram, Sylhet, Barishal, dataset

79. 【2610.01139】Do Multilingual Encoders Produce Language-Consistent Semantic IDs?

链接:https://arxiv.org/abs/2610.01139

作者:Abhinav Bohra,Anuj Bohra

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:compress item embeddings, Semantic IDs, discrete code sequences, compress item, generative retrieval

备注: 7 pages, 8 tables. Accepted as a short paper at WiNLP 2026, co-located with EMNLP 2026

点击查看摘要

Abstract:Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.

80. 【2610.01127】Counting and Min-Cost Encoding for Tokenization in Large Language Models

链接:https://arxiv.org/abs/2610.01127

作者:Shuming Shi,Xiang Zhang,Hao Yu,Wenbo Fei,Changjian Wang,Zhan Wang,Guoqing Pang,Guangye Yu,Quan Lu,Ning Jiang

类目:Computation and Language (cs.CL)

关键词:Mainstream large language, Mainstream large, BPE, token, text

备注:

点击查看摘要

Abstract:Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.

81. 【2610.01118】Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives

链接:https://arxiv.org/abs/2610.01118

作者:Zhiyun Shi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:long-term conversational assistant, long-term conversational, conversational assistant, assistant must recall, thousand LLM calls

备注: 17 pages, 4 figures

点击查看摘要

Abstract:A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand LLM calls per memory bank and up to several thousand context tokens per query. We argue that association is a learnable relevance: the pointwise mutual information of memories under how human lives unfold. We introduce Madeleine, which learns amortized association: offline, an LLM life simulator writes simulated lives, whose cue-trigger pairs teach a query encoder a residual association on top of frozen similarity; online, it calls no LLM and plugs into any vector memory by replacing only the query encoder. On LoCoMo-Plus under the official protocol, Madeleine (I) reaches 66.6 when plugged into HyperMem, the highest among all systems evaluated under this protocol; (II) used alone, reaches the score of HyperMem as released (52.4 vs. 52.9) with zero LLM calls and about 1/21 of its answer context; and (III) lifts T-Mem by 26.2 points, significantly outperforms the same untrained backbone inside both systems, and leaves ordinary QA intact on the 4B backbone.

82. 【2610.01108】AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

链接:https://arxiv.org/abs/2610.01108

作者:Sumin Lee,Sukmin Cho,Suengjae Lim,Youngjin Kwon

类目:Computation and Language (cs.CL)

关键词:repeatedly reproduce code, reproduce code, earlier attempts, tokens by copying, copying continuations

备注:

点击查看摘要

Abstract:Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37$\times$ at batch size 1 and 4.76$\times$ at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.

83. 【2610.01082】Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models

链接:https://arxiv.org/abs/2610.01082

作者:Grzegorz Kulik,Mikołaj Pokrywka,Adam Jatowt,Wojciech Kusa

类目:Computation and Language (cs.CL)

关键词:machine translation remains, translation remains challenging, remains challenging due, Dialectal machine translation, strong linguistic variation

备注: Accepted at EMNLP 2026 Findings

点击查看摘要

Abstract:Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testset, alongside established BOUQuET and FLORES benchmarks. Results show our rule-based system is consistently strongest on SiLTT and BOUQuET datasets and that TranslateGemma fine-tuned on a curated dataset improves over strong neural baselines but does not surpass the rule-based system in dialectal settings. We release SiLTT and our best neural model to support further research.

84. 【2610.01066】Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability

链接:https://arxiv.org/abs/2610.01066

作者:Yu Mao,Lei Yu,Zining Zhu,Yusheng Zheng,Haohang Li,Freda Shi,Yutong Yin,Zhaoran Wang,Jingcheng Niu

类目:Computation and Language (cs.CL)

关键词:random-reward reinforcement learning, propose random-reward reinforcement, reinforcement learning, addressing a decade-long, propose random-reward

备注:

点击查看摘要

Abstract:We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.

85. 【2610.01064】JoinGR: Learning to Traverse Join Graphs for Table Retrieval

链接:https://arxiv.org/abs/2610.01064

作者:Sandipan De,Abhijit Chakraborty,Sambaran Bandyopadhyay,Vivek Gupta

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB); Information Retrieval (cs.IR)

关键词:Retrieving, JOINGR, realistic databases, tables, table

备注: 12 pages, 6 figures, 5 pages

点击查看摘要

Abstract:Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, but this ignores a key source of evidence: some required tables are not mentioned in the question and become identifiable only through their join relationships to already relevant tables. We introduce JOINGR, a join-aware table retrieval method that treats the database join graph as the retrieval space. Columns are represented as graph nodes, while intra-table and foreign-key relationships are represented as typed edges. Given a question, JOINGR selects semantically similar anchor tables, traverses join edges with a query-conditioned scorer, and aggregates the resulting edge deposits into table scores. The scorer is a lightweight MLP on top of frozen query, node, and edge embeddings, trained with a pairwise margin loss over gold tables. On BIRD and Spider datasets, JOINGR is competitive with the strongest retrieval baselines. On BEAVER, a challenging enterprise benchmark with multi-hop table requirements, JOINGR substantially improves recall over dense retrieval and re-ranking baselines. Cross-domain experiments show that the learned scorer transfers across benchmarks, indicating that the method captures reusable joingraph traversal behavior.

86. 【2610.01054】Capturing In-Context Learning Dynamics with Task Operators

链接:https://arxiv.org/abs/2610.01054

作者:Guangzhi Xiong,Zhenghao He,Bohan Liu,Sanchit Sinha,Wenqian Ye,Aidong Zhang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:In-context learning, enables language models, language models, models to perform, ICL

备注: NeurIPS 2026

点击查看摘要

Abstract:In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at this https URL.

87. 【2610.01046】Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

链接:https://arxiv.org/abs/2610.01046

作者:Rocker D'Antonio,Thomas Benton Townsend,Dimitrios Michael Manias

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Collaboration depends, shared context, context persists, depends on shared, persists across people

备注: 17 pages, 2 figures. Accepted for publication in the 2026 IEEE 12th International Conference on Collaboration and Internet Computing (CIC)

点击查看摘要

Abstract:Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicated. This work audits sentence-specificity scoring artifacts on technical documentation and tests whether scores applied only after generation help choose among fixed LLM-generated revisions. Across Wikipedia and three technical-documentation corpora, the fixed general-domain predictor SpeciTeller and the pinned post-publication author-repository implementation of Ko et al.'s target-adapted predictor produce different corpus orders and same-sentence rank agreement from -0.066 to 0.510. Strict filtering and token-length adjustment change these patterns without reconciling them. In the Gemma set, SpeciTeller ranking raises direction-valid selection from 71.7% to 83.3% (+11.7 points; 95% source-case bootstrap interval +1.7 to +21.7); in the GPT-OSS-120B set, SpeciTeller ranking raises direction-valid selection from 51.7% to 56.7% (+5.0 points; 95% source-case bootstrap interval -6.7 to +16.7), and every primary single-score GPT-OSS-120B interval includes zero. These findings tie score interpretation and decision value to the predictor and candidate set.

88. 【2610.01045】Empty Commitments: When Agents Promise What Their Runtime Cannot Deliver

链接:https://arxiv.org/abs/2610.01045

作者:Jiaqi Tang,Lan Wei,Bingyu Shen,Boyang Li

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:remind you tomorrow, user writes, Abstract, define empty commitments, promise

备注: 4 pages, 3 tables

点击查看摘要

Abstract:A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness follows from the agent's configuration alone; no later trajectory is needed. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and a response-level outcome taxonomy. We then describe a measurement protocol: follow-up requests run in five setups that add one persistence affordance at a time, with the environment either left implicit or stated.

89. 【2610.01042】Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

链接:https://arxiv.org/abs/2610.01042

作者:Shixuan Li,Wei Yang,Peiyu Zhang,Anzhe Cheng,Heng Ping,Paul Bogdan

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Multi-agent communication aims, Multi-agent communication, Multi-agent, communication, favorable agent architecture

备注:

点击查看摘要

Abstract:Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.

90. 【2610.01027】LawCompass: Navigating from Legal QA to Multi-Agent Deep Research with Grounded Evidence

链接:https://arxiv.org/abs/2610.01027

作者:Xiaoxia Cheng,Linnan Wang,Jiahao Ma,Zhichuan Ye,Xuemei Zhou,Chuanyu Tong,Bo Jiang,Qing Zhu

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Language Models, Large Language, significantly democratized access, Recent advances

备注:

点击查看摘要

Abstract:Recent advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have significantly democratized access to legal information. Nevertheless, most existing legal assistants remain confined to multi-turn conversational QA, failing to support complex legal tasks that require systematic evidence retrieval, multi-step reasoning, and report-level synthesis. In this paper, we present LawCompass, an evidence-grounded legal assistant that navigates the transition from standard Legal QA to multi-agent deep research. LawCompass provides three task-oriented functions: Legal QA, which delivers precise, evidence-backed answers to legal questions; Professional Retrieval, which enables structured exploration of statutes and judicial cases via query rewriting; and Deep Research, which employs a multi-agent workflow to decompose complex legal tasks and synthesize comprehensive research reports. Crucially, LawCompass maintains explicit citation links across all modules, empowering users to directly verify system outputs against original legal sources. Evaluation results demonstrate that LawCompass provides a practical and scalable paradigm for transforming conversational AI into trustworthy and evidence-grounded legal research assistance.

91. 【2610.01026】It Takes Workflows to Evolve Better Workflows

链接:https://arxiv.org/abs/2610.01026

作者:Xuehang Guo,Haoyu Wang,Haifeng Chen,Yangyi Chen,Zhenhailong Wang,Qingyun Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Tackling complex real-world, Tackling complex, single large language, coordinate specialized agents, large language model

备注:

点击查看摘要

Abstract:Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to $+7.41\%$, with co-evolving ($+5.03\%$) more roles gaining more than optimizing one of them alone ($+2.83\%$). Our project page: this https URL.

92. 【2610.01023】Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

链接:https://arxiv.org/abs/2610.01023

作者:Junyu Guo,Shangding Gu,Ming Jin,Javad Lavaei

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:omit required behavior, return plausible patches, Coding agents, required behavior, return plausible

备注:

点击查看摘要

Abstract:Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.

93. 【2610.01017】Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

链接:https://arxiv.org/abs/2610.01017

作者:Xuehang Guo,Haoyu Wang,Shengyu Chen,Zach Chen,Wei Cheng,Qingyun Wang,Haifeng Chen

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language models, Large language, increasingly construct multi-agent, assign specialist agents, language models

备注:

点击查看摘要

Abstract:Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to $+11.97\%$, achieving $+9.64\%$ with in-flow optimization. Our project page: this https URL.

94. 【2610.01016】Scaling and Distilling Text Embeddings for Better Diffusibility

链接:https://arxiv.org/abs/2610.01016

作者:Zekai Zhang,Yunjie Tian,Yanjin He,Xiaoyan Zhang,Dongdi Zhao,Qing Qu,Di Fu

类目:Computation and Language (cs.CL)

关键词:Diffusion language models, offer a promising, Diffusion language, language models, Gen. PPL

备注: 28 pages, 12 figures. Code is available at [this https URL](https://github.com/la0ka1/diffusing-scaled-text-embeddings)

点击查看摘要

Abstract:Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.

95. 【2610.00997】Distilling Directional Verification

链接:https://arxiv.org/abs/2610.00997

作者:Jungseob Lee,Sugyeong Eo,Seongtae Hong,Seungyoon Lee,Chanjun Park,Jaehyung Seo,Heuiseok Lim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:large language models, Knowledge distillation aims, factual knowledge, language models, smaller models

备注: 29 pages, 7 figures, 31 tables

点击查看摘要

Abstract:Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at this https URL.

96. 【2610.00983】he Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization

链接:https://arxiv.org/abs/2610.00983

作者:Chao Li,Shigeng Wang,Anbang Yao

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:reconstruction loss, Post-training quantization, learning-based PTQ, transformer blocks, PTQ methods

备注: Project page: [this https URL](https://github.com/IntelChina-AI/RMSE)

点击查看摘要

Abstract:Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss. A common practice is to use mean squared error (MSE) as the reconstruction loss function, yet its induced optimization behavior remains largely unexplored. In this work, we take a holistic view of sequential quantization and systematically investigate how optimization evolves from the first quantization stage to the last, aiming for a deep understanding of optimization in learning-based PTQ schemes. Through extensive empirical studies spanning representative learning-based PTQ methods, LLM families, model scales, architectures, quantization settings and various tasks, we consistently uncover Optimization Imbalance: reconstruction loss magnitudes vary dramatically across stages, accompanied by highly uneven gradient magnitudes and parameter updates under MSE. We term the cross-stage range of loss magnitudes the reconstruction loss scale, and reveal that MSE translates the unexpectedly large reconstruction loss scale into highly uneven gradient magnitudes, which in turn lead to uneven optimization strength across quantization stages. This finding suggests a general principle for improving learning-based PTQ: optimization strength across stages should be decoupled from the reconstruction loss scale. Theoretically, we show that root mean squared error (RMSE) variants defined at the sample, channel, token, and element levels naturally realize this principle through implicit gradient normalization, outperforming MSE significantly as a drop-in replacement.

97. 【2610.00969】A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models

链接:https://arxiv.org/abs/2610.00969

作者:Yingzhu Zhao,Vlad Pandelea,Han Yuan,Bo Hu,Wuqiong Luo,Li Zhang,Zheng Ma

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, including earnings call, earnings call transcripts, Large language, language models

备注:

点击查看摘要

Abstract:Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generating standalone claims, users increasingly prefer grounded analyses that pair claims with verifiable citations from source documents to enable independent validation. However, evaluating such analytical claims typically requires extensive expert annotation, which is costly and difficult to scale, and real-world financial analysis commonly involves long context-question-answer triplets, further increasing task complexity. To address these challenges and benchmark the current landscape of grounded analysis by LLMs, we propose a numeric evidence evaluation method that enables groundedness assessment without reliance on expert annotation. We also introduce an automated dataset construction pipeline and construct ECTs-100 from the top 100 constituents of the SP 500 to support benchmark of both groundedness and correctness. In addition, we examine conscious incompetence, a practical failure mode in financial analysis in which LLMs must detect when available evidence is insufficient and refrain from producing unsupported hallucinations. Empirical results show that LLMs perform well in groundedness but face notable limitations in correctness, with informational insufficiency presenting an additional challenge.

98. 【2610.00964】RPTune: Learned Context Curation for LLM Catalog Search

链接:https://arxiv.org/abs/2610.00964

作者:Chuxuan Hu,Hejie Cui,Norman Huang,Shubham Kumar Bharti,Wang-Chiew Tan,Sercan Ö. Arık

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:full-catalog prompting offers, multi-stage retrieval designed, retrieval designed primarily, small merchant businesses, full-catalog prompting

备注: 23 pages, 9 figures, 4 tables

点击查看摘要

Abstract:For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average.

Comments:
23 pages, 9 figures, 4 tables

Subjects:

Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2610.00964 [cs.IR]

(or
arXiv:2610.00964v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.00964

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
99. 【2610.00958】Role-aware Heuristic Episodic Attention for Conversational LLMs

链接:https://arxiv.org/abs/2610.00958

作者:Wanyang Hong,Zhaoning Zhang,Yi Chen,Libo Zhang,Baihui Liu,Linbo Qiao,Zhiliang Tian,Dongsheng Li

类目:Computation and Language (cs.CL)

关键词:multi-turn conversations grow, Large language models, Large language, conversations grow, lose track

备注:

点击查看摘要

Abstract:Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91$\times$. Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.

100. 【2610.00954】Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles

链接:https://arxiv.org/abs/2610.00954

作者:Alexei N. Skurikhin,Emily M. Taylor,Nathan A. DeBardeleben

类目:Computation and Language (cs.CL)

关键词:large language models, scalar leaderboard accuracy, language models, small language models, extend beyond scalar

备注: 8 pages, 9 figures, Presented at ACM CAIS 2026 Workshop RLEval: Methods and Reinforcement Learning Environments for Evaluating AI Agents. Resubmission of permitted appeal, Ticket #MOD-104177

点击查看摘要

Abstract:As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.

101. 【2610.00940】ReHoPER: Receding-Horizon Planning for Enhanced Reasoning

链接:https://arxiv.org/abs/2610.00940

作者:Saeed Ahmadnia,Cornelia Caragea

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:improves large language, large language models', answering intermediate questions, language models' reasoning, zero-shot method

备注:

点击查看摘要

Abstract:We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.

102. 【2610.00928】Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations

链接:https://arxiv.org/abs/2610.00928

作者:Jungwon Park,Changin Choi,Jimyeong Kim,Nojun Kwak,Wonjong Rhee

类目:Computation and Language (cs.CL)

关键词:diverse downstream tasks, large language models, efficient task adaptation, central challenge, large language

备注: Accepted by AACL-IJCNLP 2026 Main

点击查看摘要

Abstract:As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-injection approaches. However, these lines of work have largely evolved within individual paradigms, leaving their cross-paradigm relationships and trade-offs underexplored, especially for recently emerging embedding-based adaptations. This survey presents a unified framework that categorizes task adaptation methods by where and how task information is encoded: model weights, input prompts, or injected task embeddings. We provide a comprehensive taxonomy that integrates these paradigms, analyze their key strengths and limitations to explain how different adaptation paradigms have evolved, clarify relationships across paradigms, and highlight open problems for future research.

103. 【2610.00910】he Geometry of Contextual Relations: Language Models Address Facts by Order of Mention

链接:https://arxiv.org/abs/2610.00910

作者:Yufa Zhou

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:emph, fact, Alice eats, related within propositions, objects are related

备注: Code: [this https URL](https://github.com/MasterZhou1/order-of-mention)

点击查看摘要

Abstract:Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes when the question switches from what Alice eats to what Bob eats. Averaged over many lists, this change is a steering vector, which we call the \emph{ordinal vector}. It points to a fact by its \emph{order of mention}, the order in which the facts were stated in the context. We find that LLMs represent the fact a question asks about by its order of mention, not by the name the question contains. We state this as the \textit{ordinal addressing hypothesis}: each order of mention has a \emph{fact address} in the model's state, shared by all contexts, and a question moves the state to the fact address of the fact it asks about, while the context supplies what that fact says. Across Qwen, Gemma, and Llama, fact addresses are (1) \emph{ordered by mention}: query states are organized by the order of facts, not of names, even when one fact has multiple subjects; (2) \emph{steerable}: added to a question about the first fact of a new list, the ordinal vector makes the model answer with the second fact of that list; (3) \emph{low-rank}: they span a low-rank subspace in which the first-mentioned fact is the easiest to reach, surprisingly similar to human recall; and (4) \emph{emergent}: they are shared in late-middle layers, hold from 1.5B to 32B parameters, and form early in pretraining. Language models reach a stated fact by where it was mentioned, deepening our understanding of LLM reasoning.

104. 【2610.00906】ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

链接:https://arxiv.org/abs/2610.00906

作者:Sungho Park,Wonjoong Kim,Jue Zhang,Wook-Shin Han,Pengfei Gao,Chanyoung Park,Yongqiang Yao,Rao Fu,Elsie Nallipogu,Qingwei Lin,Victor Rühle

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Software Engineering (cs.SE)

关键词:improve LLM agents, substantially improve LLM, improve LLM, LLM agents, tool interfaces

备注: 37 pages, 16 figures. Project website and code: [this https URL](https://autosaddler-projectpage.github.io/activesaddler/)

点击查看摘要

Abstract:Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.

105. 【2610.00883】DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text

链接:https://arxiv.org/abs/2610.00883

作者:Mohamed Mady,Yupei Li,Johannes Reschke,Björn W. Schuller

类目:Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:Robust detection, conditions is challenging, domains and generators, input surface, perform well in-domain

备注: Accepted at AACL-IJCNLP 2026 (main conference). 9 pages plus appendix. Code and checkpoint: [this https URL](https://github.com/MohamedMady19/deberta-conpara) , [this https URL](https://huggingface.co/mohamedmady/deberta-conpara)

点击查看摘要

Abstract:Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.

106. 【2610.00864】Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

链接:https://arxiv.org/abs/2610.00864

作者:Jiawei Fan,Sifeng Wang,Yuqing Hou,Anbang Yao

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Robotic Foundation Models, Foundation Models, Robotic Foundation, aiming to overcome, generation in Robotic

备注: Project page: [this https URL](https://github.com/IntelChina-AI/K-MF)

点击查看摘要

Abstract:In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at this https URL.

107. 【2610.00852】Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

链接:https://arxiv.org/abs/2610.00852

作者:Abner Hernandez,Tomás Arias Vergara,Andreas Maier,Paula Andrea Pérez-Toro

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:Structured phonological representations, phonological representations provide, generic speech embeddings, Structured phonological, adult

备注: Submitted for review at ICASSP 2027

点击查看摘要

Abstract:Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) across two initialization strategies (Adult PhonoQ and scratch). Generalization is evaluated against manual child-speech annotations. On 1,352 consonant targets from 58 typically developing children, child-speech adaptation improves voicing recognition across all supervision conditions, from 0.922 macro-F1 for Adult PhonoQ to 0.972--0.987 after adaptation. Manner is more sensitive to alignment supervision: Adult+Child MFA reaches 0.804 and 0.796, compared to approximately 0.70 under Adult MFA supervision. Place remains comparatively strong across systems (0.871--0.902), although per-class performance varies substantially. The velar-fronting contrast is preserved across all seven model variants. Longitudinal UltraPhonix analysis further reveals speaker-specific velar and post-alveolar changes that are largely preserved across models and broadly consistent with reported clinical progress.

108. 【2610.00850】AuraForge: Scaling Security Supervision for Training Coding Agents

链接:https://arxiv.org/abs/2610.00850

作者:Danqing Wang,Songwen Zhao,Harsh Sharma,Jierui Wang,Andre Vicente Duarte,Ivan Bercovich,Lei Li

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:generate complex software, complex software applications, single prompt, Coding agents, generate complex

备注:

点击查看摘要

Abstract:Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a critical risk: functional correctness alone does not guarantee a secure implementation. Despite growing attention to code security, training safer coding agents remains challenging because reliable security supervision is difficult to obtain at scale from real-world repositories. We introduce AuraForge to synthesize and validate executable security tests for training secure coding agents. Our approach combines attack-oriented test synthesis, language-extensible task construction, and safeguards against reward hacking. Using AuraForge, we construct AuraGym, a multi-language and multi-CWE executable training gym: 679 executable feature-implementation tasks from 344 real-world repositories across Python, JavaScript, and TypeScript, covering 177 CWE categories. On the subset with human-written security tests, AuraForge produces about 3 times as many test cases on average and reduces the false-positive rate by 83.23%, allowing alternative secure implementations to receive correct supervision. Training Qwen3.5-4B with synthesized security tests gains larger improvements than human-written security tests (average 19.7 FuncPass and 6.2 SecPass vs. 14.9 FuncPass and 4.4 SecPass) on three languages. These results demonstrate that AuraForge provides more diverse and reliable security supervision to train secure coding agents.

109. 【2610.00840】Contextual trajectory and incremental contextual displacement: Towards using LLMs to understand dynamic, utterance-specific meaning construction

链接:https://arxiv.org/abs/2610.00840

作者:Grayson Wycliffe Storer,Julia Witte Zimmerman

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Transformer-based large language, large language models, RoBERTa represent text, Transformer-based large, language models

备注:

点击查看摘要

Abstract:Transformer-based large language models (LLMs) such as RoBERTa represent text using contextual word embeddings (CWEs), which alter the embeddings associated with each token based on surrounding context. We construct token-wise incremental trajectories by repeatedly recomputing a token's CWE as successive words are added to a sentence, yielding a representation of how contextualized embeddings evolve as the utterance unfolds. We evaluate this approach using garden-path sentences as a test case with characteristic features. Token-wise trajectories reproduce known features of garden-path processing, including disruption around the critical region, and reliably distinguish garden-path sentences from matched disambiguated controls. We introduce several metrics for quantifying representational displacement across contextual increments and show that trajectory information can be highly predictive of sentence type. We find that ambiguity-related information is recoverable not only from the sentence-level CLS representation but also from ordinary vocabulary tokens, suggesting that utterance-level information is distributed across multiple representational scales. In exploratory analyses, we find qualitatively similar trajectory structures in other ambiguity- and misdirection-related linguistic phenomena. Together, these results establish token-wise incremental trajectories as a promising framework for studying utterance-specific meaning construction using LLMs.

110. 【2610.00833】VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence

链接:https://arxiv.org/abs/2610.00833

作者:Sachin Gupta

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Fluent LLM explanations, Fluent LLM, LLM explanations, structured system, Sonnet

备注: 16 pages, 5 figures, 7 tables. Accepted at Grounding Language Models: Learning Faithfully and Efficiently (GroundLM 2026), co-located with EMNLP 2026. Code: [this https URL](https://github.com/sachinkg12/yukti/releases/tag/veritygate-groundlm-2026-v1.0.0) Supplementary artifact: [this https URL](https://doi.org/10.5281/zenodo.22668710)

点击查看摘要

Abstract:Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify every fact in the prose. At r=0 and r=1, we test 900 instances per setting (450 grounded-ungrounded pairs) with GPT-4o-mini, Llama-3.3-70B, and Claude Sonnet 4.6. Under this schema-level contract and before repair, 80.3% of mini claims and 47.9% of Sonnet claims fail. These are verifier rejection rates, not prose-hallucination rates. One repair pass raises claim survival from 19.7% to 28.0% for mini and from 52.1% to 54.3% for Sonnet. Verified claims per example change by +0.14 for mini, -0.71 for Llama, and -0.47 for Sonnet, so survival and output volume must be reported together. A second Sonnet pass gives no clear gain. At r=1, Gate 4 covers 97.0%, 98.7%, and 100% of failing claims for mini, Llama, and Sonnet. Small human studies support the rules but show gaps between schema checks and correct prose. A domain-specific GPT-4o judge test shows an order effect, so it is only a usefulness check. We release the code and data.

111. 【2610.00827】Verbalized and Internal Probabilities Are Coupled in Large Language Models

链接:https://arxiv.org/abs/2610.00827

作者:Sinead Williamson,Jiaxuan Li,Nick Foti,Russ Webb,Masha Fedzechkina

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, Large language, language models carry, training data, place on generating

备注:

点击查看摘要

Abstract:Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that internal probabilities track relative frequencies in the training data, and that verbalized probabilities track explicit probabilistic assertions in the training data. However, we do not know whether these two readouts are aligned, except when frequencies and probabilistic assertions in the training data happen to align. This limits our understanding of when we can use verbalized uncertainties as a proxy for either training data frequencies, or a model's internal distribution. We resolve this gap by systematically exploring how LLMs probability readouts are impacted by training and in-context data, via intervening on the underlying uncertainty sources in the data. We find that both internal and verbalized probability readouts are impacted by both distributional and asserted uncertainty in the training data. Further, we find that verbalized and internal probabilities are aligned beyond what would be expected by independently tracking the same uncertainty sources, suggesting that verbalized probabilities can be used to probe a model's internal distribution.

112. 【2610.00817】abJoinBench: A Benchmark for Joinable Table Discovery

链接:https://arxiv.org/abs/2610.00817

作者:Sandipan De,Jin Wang,Vivek Gupta

类目:Databases (cs.DB); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:enabling downstream tasks, join discovery methods, Join discovery aims, large data repositories, Join discovery

备注: 13 pages, 8 Tables, 1 Figure

点击查看摘要

Abstract:Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.

113. 【2610.00809】Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics

链接:https://arxiv.org/abs/2610.00809

作者:Zhixi Zhu,Kristina Gligoric

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:Vision-Language Models, costs accumulate quickly, processing a typical, accumulate quickly, requires millions

备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single $2\times8$ image grid built via simple shot-transition detection approaches full-video understanding ($\kappa$ within~.05), at $\sim 15\%$ of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.

114. 【2610.00797】Sapien: A Stateful Policy Engine for Autonomous AI Agents

链接:https://arxiv.org/abs/2610.00797

作者:Corinn Tiffany,Wen Zhang,Eugene Bagdasarian,Lillian Tsai

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:security defenses prevent, Contextual security defenses, taking rogue actions, security defenses, defenses prevent

备注:

点击查看摘要

Abstract:Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).

115. 【2610.00795】Can large language models unlock discrete data in ophthalmic diagnostic reports?

链接:https://arxiv.org/abs/2610.00795

作者:Umair A. Zaidi,An-Lun Wu,Wei-Chun Lin,Thomas S. Hwang,Michelle R. Hribar

类目:Computation and Language (cs.CL)

关键词:RNFL Single Exam, large language model, OCT Glaucoma Overview, OCT Thickness Map, Single Exam

备注: 10 pages, 5 figures. Presented at the Association for Research in Vision and Ophthalmology (ARVO) Annual Meeting, Denver, Colorado, May 4, 2026

点击查看摘要

Abstract:Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer Single Exam, and OCT Thickness Map; n = 5 each) were processed using two GPT-4o-assisted pipelines and compared with a reconciled manual ground truth. Schema-Constrained used Structured Output mode with a predefined JSON Schema; Prompt-Only used a detailed instruction prompt followed by Python conversion to JSON. Outcomes were value accuracy, formatting accuracy, and extraction time. Results: Schema-Constrained value accuracy was 100.00% for Visual Field and RNFL Single Exam, 97.45% for Glaucoma Overview, and 98.00% for Thickness Map; Prompt-Only achieved 100.00% across all four report types. Formatting accuracy was 100.00% for Schema-Constrained across all report types and 100.00% for Prompt-Only except RNFL Single Exam (90.14%). Mean extraction time was 56.51 s per report for manual review versus 5.04 s for Schema-Constrained and 4.70 s for Prompt-Only, an approximately 92% reduction. Conclusions: In this small proof-of-concept dataset, general-purpose LLM-assisted pipelines extracted structured data from ophthalmic diagnostic PDFs with high accuracy and substantially reduced processing time. Prompt-Only achieved the highest value accuracy, while Schema-Constrained produced schema-compliant output with 100% formatting accuracy. These complementary strengths support further evaluation of hybrid, validation-aware workflows for research and clinical data abstraction.

116. 【2610.00779】Effective Synthetic Data Curation Requires Group-Level Signals

链接:https://arxiv.org/abs/2610.00779

作者:Cathy Jiao,Chenyan Xiong

类目:Computation and Language (cs.CL)

关键词:long-horizon task execution, strengthen advanced capabilities, essential to LLM, LLM training, Synthetic data

备注:

点击查看摘要

Abstract:Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.

117. 【2610.00767】Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints

链接:https://arxiv.org/abs/2610.00767

作者:Peter Nutter,Dani Roytburg,Clément Dumas,Jinghua Ou,Shi Feng

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:alignment research, critical to alignment, model, pre-training shape, post-trained model

备注: 78 pages. Code: [this https URL](https://github.com/peternutter/grafting-beliefs)

点击查看摘要

Abstract:Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.

118. 【2610.00724】Reason in Style: Discovering and Controlling Style in Language Models

链接:https://arxiv.org/abs/2610.00724

作者:Ioana Marinescu,Eric Karl Oermann,Kyunghyun Cho

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:difficult to identify, Language models learn, making stylistic variation, models learn content, styles

备注: 38 pages, 12 figures

点击查看摘要

Abstract:Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@$k$ over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.

119. 【2610.00717】Sequential Functional Structured Tucker Compression for Large Language Model Attentions

链接:https://arxiv.org/abs/2610.00717

作者:Jiangfeng Chen,Xinyu Wang,Tianshuo Yan,Hanwei Wu,Xiao-Wen Chang,Yang Zhang,Lei Ding

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)

关键词:independent matrix approximation, representation shift introduced, Post-training compression, formulated as independent, independent matrix

备注:

点击查看摘要

Abstract:Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.

120. 【2610.00707】Initialization Improves LLM-Driven Discovery

链接:https://arxiv.org/abs/2610.00707

作者:Mansi Sakarvadia,Marco Ciccone,Colin Raffel

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, Language Models, optimize an objective, iteratively optimize

备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.

121. 【2610.00694】How Divergence Becomes Decision Flips in Compressed Language Models

链接:https://arxiv.org/abs/2610.00694

作者:Beatriz Almeida Felicio

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:Compression reports summarize, dense model outputs, language model moved, Compression reports, total variation

备注: Preprint

点击查看摘要

Abstract:Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed copies of 19 open models on five corpora and nine mechanically unrelated perturbation families, the rate at which the arg-max token changes (the \emph{flip rate}) tracks total variation at a ratio with median $1.05$, with no fitted constant. KL converts into flips only through its square root and a factor that varies fourfold across models and corpora, because KL averages over tokens before the root is taken; first-order statistics averaged per token, such as Hellinger distance, avoid this, but reports rarely give them. As a result, of two compressors reported on different models and corpora whose flip rates differ by at least $10%$, KL assigns the smaller divergence to the one that changes more decisions in $11%$ of cases, total variation in $1%$. Two pre-registered tests mark the limits: on a held-out code corpus the ratio held for all eight models while three predictions about KL each failed for half of them or more, and on three new models with real kernels it stayed in its band for 37 of 38 checkpoints but fell below one on code for two models. In vLLM speculative decoding, total variation measured under teacher forcing predicts greedy draft acceptance with a mean relative error of $1.1$--$2.4%$, without the task-specific calibration that KL needs.

122. 【2610.00689】owards Robust Numerical Claim Verification

链接:https://arxiv.org/abs/2610.00689

作者:Peter Røysland Aarnes,Vinay Setty

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, sharply degrade accuracy, claim verification, remain brittle

备注: Accepted to AACL-IJCNLP 2026 Findings

点击查看摘要

Abstract:Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B$\unicode{x2013}$8B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.

123. 【2610.00679】Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow

链接:https://arxiv.org/abs/2610.00679

作者:Polina Tsvilodub,Andreas Waldis,Linlu Qiu,Tal Linzen,Michael Franke

类目:Computation and Language (cs.CL)

关键词:normatively correct solution, Language models, Bayesian, correct solution, hidden variables

备注: under review, 27 pages, 25 figures

点击查看摘要

Abstract:Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us $\textit{why}$ tuning on a $\textit{Bayesian}$ or an $\textit{oracle}$ (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.

124. 【2610.00673】Closing the Loop: Practical Training Recipes for Looped Language Models

链接:https://arxiv.org/abs/2610.00673

作者:Andrei Marchenko,Viacheslav Bezrukov,Oleg Kashurin,Inessa Fedorova,Dmitry Bocharov,Yuliana Shakhvalieva,Maria Tikhonova,Valerii Ternovskii

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:increase effective depth, large-scale recipes require, Looped language models, models increase effective, language models increase

备注:

点击查看摘要

Abstract:Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-stage schedules. (2) Under controlled comparisons, our 1.4B LoopLM outperforms a parameter-matched dense model trained on the same data and token budget on all 12 evaluated benchmarks, including +14 points on GSM8K, +10 on MATH, and +22 on DROP. At matched inference compute, it approaches a 3.9B dense model on mathematical reasoning and reading comprehension while using only 36\% as many parameters. (3) We introduce a minimal recipe for converting pretrained dense models into looped ones: a single learned input-mixing scalar and a smoothed exit loss, with no step-specific parameters. Applied to Qwen3-1.7B-Base, Looped Qwen improves over an identically continued dense baseline on every evaluated benchmark across two data regimes, with statistically clear gains on GSM8K, MATH, and MMLU-Pro on the curated mixture. Together, these results make looped language models substantially cheaper to train from scratch and practical to introduce into existing pretrained checkpoints, while isolating the gains due to recurrence itself.

125. 【2610.00664】PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation

链接:https://arxiv.org/abs/2610.00664

作者:Rashid Azraf Jahin,Saadman Sajid,Khan Raiyan Ibne Reza,Sumaiya Tabassum Nimi

类目:Computation and Language (cs.CL)

关键词:Bengali secondary education, secondary education lacks, education lacks curriculum-grounded, lacks curriculum-grounded benchmarks, physics problems demand

备注: 6 pages, 3 figures, 5 tables. Accepted at 11th IEEE Asia-Pacific Conference on Computer Science and Data Engineering (IEEE CSDE 2026)

点击查看摘要

Abstract:Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.

126. 【2610.00656】Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference

链接:https://arxiv.org/abs/2610.00656

作者:Jiangang Chen

类目:Computation and Language (cs.CL)

关键词:large language model, language model computes, remains difficult, autoregressive inference, training probes

备注: 15 pages, 4 figures, 7 tables. An earlier version was publicly released on Zenodo (DOI: [https://doi.org/10.5281/zenodo.23068698](https://doi.org/10.5281/zenodo.23068698) )

点击查看摘要

Abstract:Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept anchors, constructed without labeled concept examples, outcome labels, gradient fitting, or activation-space optimization, producing a structured per-step concept-coordinate signal. Across code generation and grade-school mathematical reasoning, this signal exhibits a robust association with predictive uncertainty: the association survives problem-identity and token-position controls and is not attributable to a single token type, is not explained by a simple correct/incorrect mixture on GSM8K, and is not reproduced by matched random anchors; it is markedly weaker or direction-inconsistent in K-means and PCA projections. Two structures emerge: a recurring uncertainty-linked activity signal whose functional geometry is task-conditioned (distinct activity-entropy shapes on HumanEval, MBPP, and GSM8K), and an execution-specific trajectory identity with strong local inertia but weak re-instantiation invariance--under completion-only elastic alignment, corruption at k=32 (approximately a median quarter of the completion) on the matched re-execution subset still retrieves the archived episode at 62.0%, while a fresh execution retrieves it only 11.7-16.0% of the time. Finally, a matched audit finds no evidence that the scalar concept-activity signal used here supplies a stable correctness coordinate under the tested protocol; we therefore treat correctness as externally supplied. Telemetry adds 0.7-1.6% per-token decode overhead for the 161-anchor code implementation, with unchanged generated tokens.

127. 【2610.00650】Self-Evolving Coding Rules for AI Coding Agents

链接:https://arxiv.org/abs/2610.00650

作者:Zhengyuan Jiang,Reachal Wang,Yuepeng Hu,Yupu Wang,Yuqi Jia,Neil Zhenqiang Gong

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:underlying coding rules, coding rules, agents is highly, highly dependent, coding agents

备注: Accepted by NeurIPS 2026

点击查看摘要

Abstract:The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, we propose RuleEvolve, a self-evolving framework for coding rules. RuleEvolve maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleEvolve outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).

128. 【2610.00621】Mixture of Decoders for Diverse Dialog Response Generation

链接:https://arxiv.org/abs/2610.00621

作者:Wenchao Du

类目:Computation and Language (cs.CL)

关键词:established machine learning, machine learning technique, learning large sets, long established machine, multi-modal data

备注: preprints

点击查看摘要

Abstract:Mixture modeling is a long established machine learning technique for learning large sets of multi-modal data. While it is known that sequence-to-sequence models for dialog response generation suffer from the problem of low diversity, we hypothesize that it is because sequence-to-sequence models tend to learn a degenerate uni-modal distribution of responses. We then propose to incorporate a mixture of decoders into sequence-to-sequence models and try to make each decoder learn specialized topics in order to improve the diversity of generated responses. Our model is developed under the framework of conditional variational autoencoder (CVAE). We evaluate our approach on an open domain chat corpus and show improvement over strong baselines in quantitative measures and human evaluation.

129. 【2610.00610】Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

链接:https://arxiv.org/abs/2610.00610

作者:Xuan Zhong Feng,Geoffrey Martin,Hexin Dong,Yifan Peng

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Explainable suicide-risk assessment, Explainable Suicide Risk, suicide-risk assessment requires, identify supporting language, estimate risk severity

备注: Accepted at IEEE BigData 2026

点击查看摘要

Abstract:Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three tasks: risk-level classification, evidence phrase extraction, and multi-label factor identification. Our approach adapts Qwen2.5-Instruct models using quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. We jointly train across all three tasks for risk classification, jointly train on Tasks~1a and 1b for evidence extraction, and adapt Task~2 separately for factor identification. We also tailor aggregation to each output: we average risk-level probabilities from the 32B and 72B models, combine evidence phrases through cross-fold consensus, and calibrate factor-specific decisions through rate matching based on out-of-fold operating points. On the official leaderboard, the final system achieved a composite score of 0.7738, with 0.8089 on Task~1 and 0.6919 on Task~2. Across the evaluated configurations, three-task training performed best for Task~1a, joint training on Tasks~1a and 1b performed best for Task~1b, and task-specific training performed best for Task~2. Probability averaging further improved Task~1a when component models had complementary errors. These findings highlight the value of tailoring both training objectives and aggregation strategies to the output structure of each task within a unified language-model framework.

130. 【2610.00609】Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

链接:https://arxiv.org/abs/2610.00609

作者:Katrina Drozdov,Oliver Chen,Langston Nashold,Rayan Krishnan

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:time-consuming legal workflow, Legal research, core and time-consuming, Legal Research Bench, Legal

备注:

点击查看摘要

Abstract:Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbf{Legal Research Bench} (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness with web search, case-law search, page parsing, and retrieval tools. We score agent responses through all-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4.8, is fully correct on 42.9\% of questions. Performance also varies substantially by task setting: all-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities. Across models, more turns, tool calls, and inference cost do not predict higher accuracy.

131. 【2610.00606】Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities

链接:https://arxiv.org/abs/2610.00606

作者:Dayeon Ki,Ruochen Zhang,Silviu Cucerzan,Ryen W. White,Ning Gao

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, knowledge-intensive information seeking, information seeking tasks, Models increasingly serve, Large Language

备注: 43 pages, 6 figures

点击查看摘要

Abstract:Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.

132. 【2610.00574】Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

链接:https://arxiv.org/abs/2610.00574

作者:Tong Zheng,Skylar Zhai,Zhan Cheng,TianMing Sha,Youling Huang,Shuo Zhou,Shaotong Qi,Jingcheng Liang,Xuwei Ding,Pengcheng Xu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Multi-reward reinforcement learning, trains large language, large language models, satisfy multiple behavioral, Multi-reward reinforcement

备注: 24 pages, 8 figures

点击查看摘要

Abstract:Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at this https URL.

133. 【2610.00568】Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

链接:https://arxiv.org/abs/2610.00568

作者:Pardis Sadat Zahraei,Janvijay Singh,Gokhan Tur,Dilek Hakkani-Tur

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, Large language, key properties, capability and faithfulness, capability

备注: Accepted at COLM 2026

点击查看摘要

Abstract:Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.

134. 【2610.00562】Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning

链接:https://arxiv.org/abs/2610.00562

作者:Taye Akinrele,Noorbakhsh Amiri Golilarz,Subash Neupane,Sudip Mittal,Shahram Rahimi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:extended patient histories, requires large language, clinical reasoning requires, large language models, patient histories

备注:

点击查看摘要

Abstract:Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.

135. 【2610.00540】Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models

链接:https://arxiv.org/abs/2610.00540

作者:Zhanyu Chen,Jaap Jumelet

类目:Computation and Language (cs.CL)

关键词:models vary sharply, language resource availability, vary sharply, resource availability, Claims

备注: EMNLP Main 2026

点击查看摘要

Abstract:Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.

136. 【2610.00526】Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning

链接:https://arxiv.org/abs/2610.00526

作者:Gunmay Jhingran

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:In-context learning, recover few-shot behavior, single synthetic demonstration, recent theory shows, word-level bijections

备注: Accepted to the NeurIPS 2026 Workshop on Linguistic Principles for Foundation Models (LP4FM). 5 pages

点击查看摘要

Abstract:In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections. We ask a linguistic version of this question: which linguistic operations can be amortized out of the prompt? We train a 2.6M-parameter network that reads the geometry of a few-shot support set (centroid, principal subspace, spectrum, computed once and cached) and produces an input-conditioned additive update to the query's residual stream at a mid-depth layer of a frozen GPT-2-large/XL. Across eight inflectional directions and one lexical relation, under a canonical split that bars inverted-pair leakage between directions, three regimes emerge. On forward inflection, where 10-shot ICL is strong (0.67-0.89) and extracted task vectors collapse (=0.06), the transform matches ICL at strictly zero-shot per-query cost. On lemmatization directions, which frozen GPT-2 can execute but 10 demonstrations systematically fail to convey (ICL 0.13-0.48 at 1.5B), the transform is not capped by ICL at all: it reaches 0.78-0.92, up to +72 points over ICL (past to present: 0.85 vs. 0.13). On arbitrary pairings (antonymy) every amortizer plateaus near half of ICL at every scale, capacity, and seed tested. Controls show the support manifold acts as a causally necessary task fingerprint: wrong-task manifolds collapse accuracy to =0.06, query-only variants cannot disambiguate tasks sharing an input space, and leave-one-task-out transfer is zero. Productive rules amortize into latent task representations, sometimes better than prompting can convey them; memorized pairings do not.

137. 【2610.00492】EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

链接:https://arxiv.org/abs/2610.00492

作者:Jiayi Geng,Zhengxuan Wu,Kevin S. Chen,Seungone Kim,Joseph Janssen,Zora Zhiruo Wang,Bhupalee Kalita,Runtian Gao,Aaron Ho,Andrew Oakleigh Nelson,Olexandr Isayev,Francisco Villaescusa-Navarro,Ching-Yao Lai,Howard Chen,Graham Neubig

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Isaac Newton discovered, Isaac Newton, analyzing observed data, Moon orbit, planetary patterns

备注:

点击查看摘要

Abstract:When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.

138. 【2610.00447】Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

链接:https://arxiv.org/abs/2610.00447

作者:Anson Y. Lam,Shuqing Li,Michael R. Lyu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Multimedia (cs.MM)

关键词:scene stays fixed, generated scene stays, stays fixed, generated scene, scene stays

备注: 26 pages, 6 figures

点击查看摘要

Abstract:Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment evaluators plus one perceptual-quality control, then test four targeted scene degradations. Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators, with prompt-bootstrap lower bounds above 1 for 11/19. Rankings are more stable than scores, yet 18/19 evaluators change their point-estimate winner under some configuration. Pairwise protocol margin envelopes show which comparisons keep their direction across the tested settings. Selected pairs have opposite pointwise intervals, but no reversal survives simultaneous inference over the full search. Thus the observed winner changes are descriptive, not confirmed changes in generator superiority. Sensitivity remains separate: no evaluator, even the prompt-free control, exceeds 67% tie-adjusted directional discrimination on layout scrambling, which is diagnostic rather than human-validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) with protocol-dependent comparisons and selection-aware uncertainty.

139. 【2610.00408】UniBuc at SemEval-2024 Task 2: Tailored Prompting with Solar for Clinical NLI

链接:https://arxiv.org/abs/2610.00408

作者:Marius Micluta-Campeanu,Claudiu Creanga,Ana-Maria Bucur,Ana Sabina Uban,Liviu P. Dinu

类目:Computation and Language (cs.CL)

关键词:Safe Biomedical Natural, Biomedical Natural Language, Natural Language Inference, Safe Biomedical, Clinical Trials

备注:

点击查看摘要

Abstract:This paper describes the approach of the UniBuc team in tackling the SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials. We used SOLAR Instruct, without any fine-tuning, while focusing on input manipulation and tailored prompting. By customizing prompts for individual CTR sections, in both zero-shot and few-shots settings, we managed to achieve a consistency score of 0.72, ranking 14th in the leaderboard. Our thorough error analysis revealed that our model has a tendency to take shortcuts and rely on simple heuristics, especially when dealing with semantic-preserving changes.

140. 【2610.00406】LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian

链接:https://arxiv.org/abs/2610.00406

作者:Claudiu Creanga,Liviu P. Dinu

类目:Computation and Language (cs.CL)

关键词:Evaluating Retrieval-Augmented Generation, Evaluating Retrieval-Augmented, Retrieval-Augmented Generation, Low-Resource Languages, metrics fall short

备注:

点击查看摘要

Abstract:Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation models (Gemini 2.5 and Gemini 3). We introduce AdminRo-Eval, a curated dataset of Romanian administrative documents annotated by native speakers, to serve as a ground truth for benchmarking automated evaluators. We compare three evaluation methodologies - direct scoring, comparative ranking, and granular decomposition - across metrics for Faithfulness, Answer Relevance, and Context Relevance. Our findings reveal that evaluation strategies must be metric-specific: granular decomposition achieves the highest human alignment for Faithfulness (96% with Gemini 2.5 Pro), while comparative ranking outperforms in Answer Relevance (90%). Furthermore, we demonstrate that while lightweight models struggle with complex reasoning in LRLs, the Gemini 2.5 Pro architecture establishes a robust, transferable baseline for automated Romanian RAG evaluation.

141. 【2610.00402】Dissonant ballerinas and crafty carrots: a comparative multi-modal analysis of Italian brain rot

链接:https://arxiv.org/abs/2610.00402

作者:Anca Dinu,Andra-Maria Florescu,Marius Micluta-Campeanu,Stefana-Arina Tabusca,Claudiu Creanga,Andreiana Mihail

类目:Computation and Language (cs.CL)

关键词:Italian Brain rot, brain rot memes, brain rot, Romanian brain rot, multi-modal brain rot

备注:

点击查看摘要

Abstract:This paper presents a comparative multi-modal analysis of Italian and Romanian brain rot memes, investigating the factors that contribute to its appeal and the linguistic and cultural distinctions between the two versions. To conduct this analysis, we introduce a multi-modal brain rot dataset named CRIB (Collection of Romanian and Italian Brain rot), a manually curated collection of 240 TikTok videos stratified by language (Italian, Romanian) and popularity, on which we examine textual, acoustic, and visual features. Our findings indicate that popularity is not significantly correlated with textual elements like sentiment, absurdity, or rhyme, or acoustic elements such as vocal features or sentiment of the sound. Instead, in Romanian language, video-level dynamics, specifically faster cutting speeds and a more rapid overall pace, are strong predictors of a video's success. The cross-linguistic analysis reveals significant differences. Italian brain rot is textually more negative, exhibits higher perplexity, and uses more rhyme, while its sound is characterized by higher melodic range and loudness. Romanian audio is spectrally brighter with more erratic pitch variations.

142. 【2610.00385】FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training

链接:https://arxiv.org/abs/2610.00385

作者:Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Tianshu Fu,Daren Zha,Jun Xiao

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:rank cached trajectories, format feedback, response length, distinct objectives, rank cached

备注: 35 pages, 6 figures

点击查看摘要

Abstract:Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches $0.6476\!\pm\!0.0139$ over eight seeds (median 0.6481; paired 95% interval $[+0.079,+0.122]$) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.

143. 【2610.00374】Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

链接:https://arxiv.org/abs/2610.00374

作者:Yuxin Yue,Yingchen Zhang,Ruqing Zhang,Jiafeng Guo,Maarten de Rijke,Xueqi Cheng

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Graphics (cs.GR)

关键词:encode quantitative claims, Analytical charts, faithfully grounded, grounded in supporting, evidence

备注:

点击查看摘要

Abstract:Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.

144. 【2610.00373】When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls

链接:https://arxiv.org/abs/2610.00373

作者:Juli Huang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Attention-head ablation, measuring the resulting, resulting change, common method, method for inferring

备注: Code available in the accompanying repository. 2 figures

点击查看摘要

Abstract:Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.

145. 【2610.00369】A Shared Taste for Model-Written Text: The Generator-by-Selector Matrices of "AI-AI Bias" Show No Detectable Own-Model Premium

链接:https://arxiv.org/abs/2610.00369

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, language models choosing, model, selector prefer text, language models

备注: 10 pages, 4 figures, 3 tables. Reanalysis of publicly available generator-by-selector matrices

点击查看摘要

Abstract:Laurito et al. (PNAS 2025) showed that large language models choosing between two descriptions of the same product, paper or film prefer the description written by a language model over the one written by a person, by a wide margin over what human judges do. Their design crosses five generators with the same five models as selectors, which permits a second question the paper does not headline: does a selector prefer text from its own model beyond what the generator and selector main effects predict? We rebuild the three 5x5 matrices from the per-item counts in the authors' public repository (21,828 valid trials; every cell matches the published value) and fit a two-way fixed-effects model with an own-model term gamma, tested by the exact permutation test over the 120 relabellings of the selectors. The premium is +0.013 on products (exact one-sided p = 0.24), -0.010 on paper abstracts (p = 0.74), +0.054 on films (p = 0.07) and +0.019 pooled (p = 0.14; 95% interval -0.008 to 0.046). The same-vendor term for the GPT-3.5 and GPT-4 pair is negative in all three datasets. Position bias moves single cells by up to 0.42 share points in either direction, and the own-model contrast is unchanged once order-driven items are removed. The design would have detected a premium of 0.05 with 82% (products), 88% (papers), 42% (films) and 97% (pooled) power; the minimum detectable effect at 80% power is 0.034 pooled. The absence is informative down to about 0.04 share points and silent below that. The 4x4 matrix of Tan et al. (ACL 2024) gives gamma = +0.148 at the smallest p its 24 relabellings allow, with a same-family term of the same size. The main result of Laurito et al. stands: models share a taste for model-written text, with GPT-4's descriptions chosen 77% to 95% of the time by every selector on products. What these data do not show is a model recognising and favouring its own prose.

146. 【2610.00333】LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation

链接:https://arxiv.org/abs/2610.00333

作者:Jaeyun Shin,Hangeol Chang,Jong Chul Ye

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Multimodal on-policy distillation, strong reasoning capabilities, on-policy distillation, grounding, visual

备注:

点击查看摘要

Abstract:Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribution entangles its visual grounding signal with its own language prior, preventing the grounding information from being transferred independently. Conversely, increasing the strength of visual supervision can improve perception but may overemphasize visual evidence and degrade language reasoning. To address this trade-off, we introduce LEGO-OPD, which selectively composes factors from a Language Expert and a Grounding expert into One teacher distribution for multimodal OPD. Under a generalized Bayesian formulation, the language expert provides a prior over candidate tokens, while the grounding expert contributes a visual likelihood that updates this prior, rather than transferring its complete predictive distribution. This factorized composition allows language reasoning and visual grounding to be controlled independently. We further introduce adaptive calibration to determine how strongly the visual likelihood should update the language prior at each decoding prefix. Specifically, LEGO-OPD uses the grounding expert's image-induced prediction shift as a prefix-dependent reference, preventing both insufficient and excessive visual supervision. Experiments with Qwen3 models show that LEGO-OPD consistently outperforms the evaluated single- and multi-teacher OPD baselines on both multimodal and text-only reasoning tasks. Moreover, it improves the initial student's visual perception while preserving text-only reasoning.

147. 【2610.00328】ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair

链接:https://arxiv.org/abs/2610.00328

作者:Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Yina Sa,Daren Zha,Jun Xiao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Structured tool calls, Structured tool, execution contract, tool calls, calls often fail

备注: 29 pages, 8 figures

点击查看摘要

Abstract:Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify a contract-constrained group-relative objective for patch, retry, and abstention decisions while keeping canonical targets and semantic labels outside the online state until trace freeze. Under identical verifier information, ContractRL attains 0.9362 semantic success with 34.4 generated tokens, compared with 0.9076 and 44.9 tokens for Patch-SFT and 0.9148 and 137.2 tokens for full regeneration over 192 cases per seed and five seeds. Policy optimization improves semantic success from 0.9186 for supervised ContractRL to 0.9375. A separate three-seed paired evaluation against Patch-SFT yields a semantic difference of +0.0396 (95% CI $[+0.0137,+0.0662], p=0.0039$). Feedback, action-mask, budget, and schema-shift analyses connect these gains to localized correction, while adversarial and multi-turn evaluations characterize the remaining failure modes.

148. 【2610.00327】Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing

链接:https://arxiv.org/abs/2610.00327

作者:Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Yina Sa,Daren Zha,Jun Xiao

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:critical association unaudited, Tool-using agents, association unaudited, logs while leaving, leaving a critical

备注: 35 pages, 8 figures

点击查看摘要

Abstract:Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being transplanted across claims, actions, runs, or source versions. We introduce a claim-anchored execution contract that jointly binds the emitted claim, its exact source span, the ordered execution prefix that produced it, and the source version and access state observed by that execution. Each receipt contains an emission anchor that deterministically locates the claim inside a committed answer or claim-bearing action, together with source identifiers, offsets, hashes, quotes, and a domain-separated execution commitment. A deterministic integrity verifier reconstructs these bindings before semantic or task labels are joined. We separate this integrity plane from a pluggable support plane, so structural validity is not used as a proxy for entailment. The contract exposes seven independently testable properties: claim-emission binding, source binding, ordered-execution binding, oracle separation, persisted-object replay, execution-rerun consistency, and version/access binding. Across 1,280 cross-object attacks, the joint contract detects 1,275 substitutions (0.9961). Removing a targeted property reduces its attack-detection rate to 0.0156-0.0625. On an independently adjudicated 384-pair split, the conflict-aware support guard reaches F1 0.8865 and false acceptance 0.0729; on unseen failure families, these rates are 0.8679 and 0.0938.

149. 【2610.00321】CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

链接:https://arxiv.org/abs/2610.00321

作者:Jungseob Lee,Sugyeong Eo

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:accelerates large language, future tokens cheaply, decoding accelerates large, Speculative decoding accelerates, large language model

备注: 28 pages, 7 figures, 17 tables

点击查看摘要

Abstract:Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs the verification time it adds. The width therefore adapts to each deployment from a latency measurement, without sweeping over widths. We evaluate CAST across five domains on three GPU generations and two model families. At its predicted width, CAST is faster than the standard chain in all eight settings, by up to 43%. We also find that the best width depends strongly on the deployment. Where verification cost jumps at a kernel boundary, a 128-token tree is only 2% faster than the standard chain, whereas the tree at the predicted width is 20% faster. Furthermore, we prove that CAST leaves the target output distribution unchanged under both greedy and sampled decoding. Code is available at this https URL.

150. 【2610.00320】Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

链接:https://arxiv.org/abs/2610.00320

作者:Jungseob Lee,Dongyub Jude Lee,Sugyeong Eo,Seongtae Hong,Seungyoon Lee,Heuiseok Lim

类目:Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)

关键词:adapts aligned large, aligned large language, large language models, Fine-tuning adapts aligned, downstream tasks

备注: 24 pages, 7 figures, 21 tables. Jungseob Lee and Dongyub Jude Lee contributed equally

点击查看摘要

Abstract:Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful examples, refusal remains near zero on all six checkpoints, with recovery transitions above the frozen boundary. In a second study, removing the update's top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, ordinary training changes weaken this repair and an attacker who spreads the update defeats it. A spectral detector calibrated on benign Llama fine-tunes misses most repair failures on that checkpoint. Localized freezing can nevertheless help preserve refusal when a few harmful examples enter training data unintentionally. These results show that an attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at this https URL.

151. 【2610.00316】DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents

链接:https://arxiv.org/abs/2610.00316

作者:Puneet Mathur,Nedim Lipka,Zeyu Jin,Dinesh Manocha

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:agents enable low-latency, Voice agents enable, documents remains underexplored, external documents remains, Voice agents

备注: Under submission at EACL 2027

点击查看摘要

Abstract:Voice agents enable low-latency, natural interaction, yet their ability to faithfully ground responses in external documents remains underexplored. We introduce DuplexSpeechBench-Document Grounding (DSB-DG), a benchmark for evaluating document grounding in voice agents across five professional domains. DSB-DG targets three failure modes: Context Saturation, which measures grounding under increasing document length; Grounding Decay, which measures retention of document facts across multi-turn dialogue; and Proactive Grounding, which evaluates whether context re-injection mitigates conversational drift. The benchmark contains 1,636 adversarially verified QA pairs from 50 documents covering five professional domains, and supports fully automatic evaluation of grounding accuracy, hallucination, and response latency. Across systems spanning cascaded, proprietary full-duplex and real-time, and open-weight speech2speech architectures, we find substantial differences in effective grounding capacity. While cascaded pipeline (ASR-LLM-TTS) achieves the highest grounding accuracy, Gemini-Live and GPT-Realtime closely trail behind. Open-weight systems exhibit distinct failure modes, most notably an abrupt context-capacity collapse and multi-turn grounding decay. More broadly, grounding fidelity degrades with context and conversational load, and failures frequently manifest as unsupported generations rather than abstention. We show that contextual grounding as a key unresolved challenge for reliable full-duplex voice agents.

152. 【2610.00309】okenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs

链接:https://arxiv.org/abs/2610.00309

作者:Mohamed Shaaban,Mohamed Elmahallawy

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language models, personally identifiable information, disclose personally identifiable, Large language, privacy-critical domains

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to memorize and disclose personally identifiable information (PII) poses serious security and compliance risks. Existing defenses typically force a trade-off between model utility, privacy protection, and access to fine-tuned private knowledge. We propose LoRA-Oriented Control via Keyed Entry Tokens (Locket), a practical framework that embeds fine-grained, policy-driven access control directly into LLM generation. Locket trains a set of lightweight LoRA (Low-Rank Adaptation) adapters, each encoding a distinct access policy (e.g., full reveal, partial redaction via PII masking, or reveal under a specified differential privacy level). A compact gating module is trained to associate a learned keyed entry token with exactly one LoRA adapter via sequence-level hard routing; the presence of a valid token acts as an authorization key that unlocks corresponding private knowledge, while an invalid or absent token triggers a privacy-preserving adapter that redacts or sanitizes sensitive content. This design ensures Locket remains fully compatible with off-the-shelf LLMs, supporting scalable deployment while satisfying regulatory and privacy requirements. We evaluate Locket across multiple datasets (Enron, ECHR, Yelp) and a diverse set of state-of-the-art LLMs, including Qwen3 (1.7B and 8B), Meta's Llama-3.2 (1B and 3B), and Google's Gemma-2-2B. Our extensive experiments demonstrate that, when the correct token is provided, Locket preserves perplexity comparable to fine-tuning on raw data (without any defense). Conversely, when the token is missing or invalid, it substantially reduces PII leakage while maintaining utility and perplexity on par with strong baseline defenses.

153. 【2610.00296】Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

链接:https://arxiv.org/abs/2610.00296

作者:Yunfan Zhou,Ye Zhu,Zhihai Wang,Jianguo Yao,Haibing Guan,Xijun Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:LLM training, Token-level certainty, training and inference, certainty, Token-level

备注:

点击查看摘要

Abstract:Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through controlled empirical evaluations across models and tasks. We distinguish two prediction targets: identifying questions a model is more likely to answer correctly and distinguishing correct from incorrect responses to the same question. In our experiments, certainty is generally better at identifying questions a model is likely to answer correctly than at distinguishing correct from incorrect responses to the same question. Certainty also varies systematically across token types and positions within words, reflecting local properties of words and text form. Information about question difficulty appears early in generation, while the weaker information about answer correctness is more concentrated near the end. These findings show that the information certainty provides for decisions depends on the prediction target, the model, the certainty metric, and which token positions in the response are included in aggregation. We further demonstrate the practical value of these findings for test-time compute. We allocate the number of responses using certainty early in generation and weight answer votes using certainty near the end of each response. Compared with a fixed-sampling majority-voting baseline, this approach increases overall accuracy from 78.71\% to 79.54\% while reducing generated-token cost by 82.4\%.

154. 【2610.00262】Signed Lexical Confidence for Risk-Calibrated Intent Routing

链接:https://arxiv.org/abs/2610.00262

作者:Yezhou Cheng,Zehua Yang,Bojun Lin

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:deferring uncertain requests, Selective intent routing, Selective intent, uncertain requests, assistant to act

备注:

点击查看摘要

Abstract:Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests. Standard confidence scores primarily reflect the base model's representation, leaving an opportunity to incorporate complementary evidence without changing its decisions. We introduce a signed lexical gate that combines a sentence classifier's logit margin with a sparse lexical model's support for the classifier's predicted intent. By assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent, the gate retains more information than either unsigned lexical confidence or a hard agreement rule. An independent binomial calibration stage selects an operating threshold for a specified risk target. Across ten runs on BANKING77, CLINC150, and HWU64, the proposed score reduces area under the risk-coverage curve by 15.8%, 15.1%, and 11.8% relative to a learned semantic-only gate. At a nominal 5% error target, it increases accepted coverage by 1.83 and 5.14 percentage points on BANKING77 and HWU64, while CLINC150 is already near full coverage. At a stricter 2% target, the simultaneous binomial procedure yields a nonempty policy in all 30 dataset-run combinations at the available calibration budgets. Matched controls show that the proposed feature improves average error ranking over the tested unsigned lexical-confidence feature, with dataset-dependent gains over binary agreement. The resulting two-feature gate provides a compact, interpretable confidence enhancement for risk-calibrated intent routing while preserving the base classifier's predictions.

155. 【2610.00253】System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer

链接:https://arxiv.org/abs/2610.00253

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:deployed language models, visibility summarise, language models, models into per-system, deployed language

备注: 30 pages, 5 figures, 9 tables. Appendix D documents corrections to an earlier manuscript

点击查看摘要

Abstract:Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. We test whether such a profile describes the system on one corpus of 6,475 stored responses (6,324 analysable) collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation and category-ownership queries. The collection harness cut many answers short: 83.1% of Gemini 3 Flash answers in category ownership end mid-sentence under a 1,024-token output cap. With every answer cut to its first 800 characters, a character n-gram classifier cross-validated by prompt attributes one response to GPT-5.2, Gemini 3 Flash, Gemini 3 Flash with search, Grok or Perplexity sonar-pro with 97.84% accuracy (5,028 responses, 383 prompts, majority class 31.5%, 30 split seeds). Length alone falls to the majority rate, 24 formatting statistics reach 95.79%, and masking brand names and capitalised tokens leaves 97.72%. Held-out query conditions keep 97.43% weighted by size and 88.0% unweighted; in a retrieval-grounded arm that changes the harness, no Grok answer is attributed to Grok (0/120). Aggregated into 50 model-by-domain-by-condition units, twelve behavioural features separate four systems at 66.53% under grouped cross-validation, against a label-permutation null with mean 33.71% and 95th percentile 46.0%. Across domains the aggregate profile fails: a forest trained on category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume (8.41 against 0.94 brands per response in gifts, 3.01 against 3.91 in category ownership), while single responses transfer at 89.92% balanced accuracy. The surface form of an answer carries the system across the query domains tested; aggregated brand behaviour does not, and the uncrossed design cannot separate the system from the domain or the harness.

156. 【2610.00238】CAVE-Mem: Boundary-Aware Experience Validation for Memory Search

链接:https://arxiv.org/abs/2610.00238

作者:Xinyu Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:agents increasingly rely, memory agents increasingly, large personal, agents increasingly, increasingly rely

备注:

点击查看摘要

Abstract:Long-term memory agents increasingly rely on it- erative search and reusable experience to answer questions over large personal, factual, or narrative histories. However, current experience-memory systems largely optimize relevance: they re- trieve past search lessons that appear similar to the current state and inject them into the prompt. A relevant experience can still be harmful when the memory substrate, question intent, answer granularity, or evidence boundary changes. We propose CAVE- Mem, a training-free framework that represents experience as a typed intervention operator with applicability, boundary, and utility conditions. CAVE-Mem first obtains a base memory-search answer, then allows an operator to change it only if the oper- ator matches the current substrate, answer contract, evidence boundary, and cross-fitted utility; otherwise the system abstains. Experiments across long-term conversational memory, multi-hop question answering, and long-document narrative reasoning show consistent gains over relevance-only experience reuse.

157. 【2610.00233】Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole

链接:https://arxiv.org/abs/2610.00233

作者:Cris Huynh

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)

关键词:constrained signalling channel, informed adversary shares, signalling channel, shares the audience, constrained signalling

备注: 11 pages, 3 figures

点击查看摘要

Abstract:When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool, the two differ on only 2,748 items --- lying exactly where the prior salience-to-Bayes coordinate is undefined. Where defined, robustness is achieved by moving from Bayesian discrimination entirely to salience. We show this by introducing an adversary to a forced-choice task (abstracted from Deception: Murder in Hong Kong). The adversary knows the target, observes the signal, and argues for the strongest wrong answer using a persuasion budget, $\beta$. As $\beta$ grows, the optimal signal shifts from the posterior-maximizing option to the margin-maximizing one; at $\beta = 0$, the game reproduces the original oracle model with a listener temperature of $\tau = 1$. This effect is real: 18.2 percent of the pool has an optimum that shifts under a finite budget, and each item's critical budget is exact. This coincidence structurally limits empirical evaluation. Two adversary framings change the chosen option of seven language models on 30 to 77 of 108 items against an exact no-effect rate. Yet, no measurement can determine whether this movement is toward the adversary-aware optimum or toward salience, because the two options are identical. This is a structural limit, not a null result. The diagnostic check is cheap: before evaluating adversary-awareness, verify whether the robust target coincides with a heuristic target on the evaluation items.

158. 【2610.00223】A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model

链接:https://arxiv.org/abs/2610.00223

作者:Imad Lakim,Ebtesam Almazrouei,Ibrahim Abu Alhaol,Merouane Debbah,Julien Launay

类目:Computation and Language (cs.CL)

关键词:grow more ubiquitous, larger language models, language models grow, larger language, models grow

备注: 11 pages, 3 figures, 2 tables. Published in Proceedings of BigScience Episode #5 -- Workshop on Challenges Perspectives in Creating Large Language Models (ACL 2022)

点击查看摘要

Abstract:As ever larger language models grow more ubiquitous, it is crucial to consider their environmental impact. Characterised by extreme size and resource use, recent generations of models have been criticised for their voracious appetite for compute, and thus significant carbon footprint. Although reporting of carbon impact has grown more common in machine learning papers, this reporting is usually limited to compute resources used strictly for training. In this work, we propose a holistic assessment of the footprint of an extreme-scale language model, Noor. Noor is an ongoing project aiming to develop the largest multi-task Arabic language models -- with up to 13B parameters -- leveraging zero-shot generalisation to enable a wide range of downstream tasks via natural language instructions. We assess the total carbon bill of the entire project: starting with data collection and storage costs, including research and development budgets, pretraining costs, future serving estimates, and other exogenous costs necessary for this international cooperation. Notably, we find that inference costs and exogenous factors can have a significant impact on total budget. Finally, we discuss pathways to reduce the carbon footprint of extreme-scale models.

159. 【2610.00205】From Web(logs) to Web(AI): Questions, Platforms, and Methods across Twenty Editions of ICWSM

链接:https://arxiv.org/abs/2610.00205

作者:Koustuv Saha,Eshwar Chandrasekharan

类目:ocial and Information Networks (cs.SI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)

关键词:examined social life, social life online, twenty editions, ICWSM community, examined social

备注:

点击查看摘要

Abstract:Over twenty editions, the ICWSM community has examined social life online as platforms, interactions, and research methods have changed. What can this body of research tell us at this critical juncture, as AI increasingly reshapes how people communicate online? We analyzed 2,139 indexed contributions from 2007 to 2026, distinguishing topics identified through nonnegative matrix factorization from problem framings captured through explicit textual cues. We find that platform mentions shift from blogs toward Twitter and, more recently, Reddit. Online community research maintains a similar topic share (10.5% to 10.0%), but governance cues within it increase from 4.0% to 34.3%. Harm-related cues also increase after restricting abstracts to a fixed length. Our review also traces advances in sampling, measurement, and causal and experimental methods. We discuss how AI-mediated interactions complicate these questions and provide a reporting checklist to support research across changing platforms.

160. 【2610.00202】When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

链接:https://arxiv.org/abs/2610.00202

作者:Esther Xin

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:RLVR training data, invent plausible wrong, pipelines build RLVR, build RLVR training, plausible wrong answers

备注: 9 pages, 2 figures,4 tables;Code and data [this https URL](https://github.com/ethxin0011/rlvr_authenticity_audit)

点击查看摘要

Abstract:Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principle learn the second instead of the first. We audit that possibility in GooseReason-0.7M. First we ask whether the asymmetry is visible at all: a classifier reading only five surface statistics (never the meaning) reaches AUROC 0.562 over 315,499 options, barely above chance. The aggregate hides something, though. Code sits at 0.416, below chance, and manual inspection explains why: code distractors turn out to be single-operator mutations of the gold answer rather than freely written alternatives, so the two classes are nearly identical by construction. Detecting a signal is not the same as showing a model uses it, so we then run an intervention. We build a paraphrase-matched control corpus, hold training-set size identical across arms, and train two policies under one fixed budget. The exploitation gap does not favour the unmodified-data arm: 0.021 against 0.027 for the control. Under our budget, in other words, a detectable artifact went unexploited. We think that dissociation, along with the domain-specific construction finding, is worth knowing for anyone curating corpora of this kind, and we release the audit as a mostly CPU-only protocol.

161. 【2610.00197】Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor

链接:https://arxiv.org/abs/2610.00197

作者:Sam Larson

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:conversational humor, training language models, reward, investigate automated rewards, language models

备注: 11 pages, 3 figures, 4 tables

点击查看摘要

Abstract:We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails further validation. An audience model's predicted laughter is instead vulnerable to laughter cues in either speaker's messages. Normalizing these cues across speakers blocks the covered attacks, although unmatched expressions remain exploitable. Three reinforcement-learning runs evaluate training with successive reward revisions. The final run improves the combined evaluation score by 0.0903 and reduces zero-score sessions by 40%, but its humor-specific improvement remains below our preregistered target. These findings illustrate a broader challenge for automated reward design: countermeasures must block exploitable shortcuts while preserving the behavior the reward was intended to encourage.

162. 【2610.00170】On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries

链接:https://arxiv.org/abs/2610.00170

作者:Hyojung Han

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:stable identifier leaving, study commercial intent, commercial intent inference, user device, work began

备注: 36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): [this https URL](https://github.com/hyojunguy/ondevice-intent-evidence)

点击查看摘要

Abstract:We study commercial intent inference that runs entirely on the user's device, under three constraints frozen before the work began: the downloaded payload under 3 MiB, Tier-0 inference under 20 ms at p95, and no raw text, content embedding, or stable identifier leaving the device. Under them we build a retrieval path over a 6,020-leaf commercial taxonomy: a static embedding table distilled from a Korean sentence transformer, quantized to 4 bits, no inference runtime. Our main result is where that constraint costs accuracy. On real Korean commerce text labelled by others (22,900 AI-Hub shopping reviews), mid-category top-5 on real product names is 75.0% against an 18.4% permutation baseline, but splits on one observable: a query containing some leaf name as a substring scores 83.5%, one containing none 45.2%. A generic 196.6x larger teacher seemed to localize the gap (+20.1 pp without an anchor, +0.1 with). That null was two effects cancelling: the same teacher fine-tuned on the student's own contrastive pairs reaches 0.8586 and beats the pure-encoder student by +10.6 pp with an anchor and +20.9 pp without. The cost is not uniform, but it is not free anywhere; where the anchor is absent, task adaptation buys the teacher nothing, so what the constrained encoder lacks there is capacity. The expensive regime is detectable on-device from the ranker's own score margin: declining the least confident fifth lifts the rest to 0.8296. A second axis we first reported, a manufacturer model code, does not survive source-category fixed effects (-4.0 pp, p=0.51); the anchor does (+13.0 pp). Payload is 2,942,652 bytes, all three library links measured. Tier-0 p95 is 4.431 and 3.670 ms on two iPhones (A14, A16) and 5.080 ms on a budget Android tablet (Snapdragon 695), all slower than three server CPUs on the same code. Taxonomy supervision is mostly synthetic Korean utterances.

Comments:
36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): this https URL

Subjects:

Information Retrieval (cs.IR); Computation and Language (cs.CL)

Cite as:
arXiv:2610.00170 [cs.IR]

(or
arXiv:2610.00170v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.00170

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Hyo Jung Han [view email] [v1]
Wed, 16 Sep 2026 00:34:01 UTC (74 KB)

163. 【2610.00164】Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery

链接:https://arxiv.org/abs/2610.00164

作者:Hyojung Han(ThakiCloud)

类目:Computation and Language (cs.CL)

关键词:Korean dialect corpora, Korean dialect, dialect corpora, underlying data, synthetic data recovers

备注: 22 pages, 6 figures, 6 tables

点击查看摘要

Abstract:Korean dialect corpora are available but not redistributable: weights may be released, while reproducing training and evaluation from the underlying data cannot be. We ask how much of that supervision synthetic data recovers, and whether that recovery can be measured independently of the synthesis pipeline. We contribute KoDialectBench, 1,000 items across five regions on three axes, released as identifier hashes and scoring code so users reconstruct the items from their own licensed copy. Recovery is strongly axis-dependent: our best synthetic arm reaches 91.2% of the real-data gain on region identification but 63.7% on comprehension. On generation the answer depends on the metric: the deployed marker lexicon reports 92.3% on dialectness and 119.3% on region match, the latter exceeding the real-data reference, whereas reference-based generation reaches 72.4%. We find the marker metrics' scoring inventory is entirely contained in the inventory our transformation rules can emit. We test the effect of construction access directly with an exact-form construction-disjoint arm that withholds 20% of marker types from the rules. At exactly matched training size (8,600 examples) it reduces dialectness recovery from 91.8% to 8.1% and region-match recovery from 101.9% to 25.6% on the held-out marker inventory, while the three pipeline-independent measurements do not fall at all. A complementary evaluator sweep defines metric-construction coverage (MCC) and finds measured dialectness recovery increasing monotonically as overlap rises from MCC=0 to MCC=1. Shared construction and evaluation inventories can therefore substantially inflate estimates of synthetic-data recovery.

164. 【2610.00154】How Robust Are Neural Audio Codecs for African Speech? A Multi-Task Benchmark and the Limits of Perceptual Quality

链接:https://arxiv.org/abs/2610.00154

作者:Chibuzor Okocha,Christan Earl Grant

类目:ound (cs.SD); Computation and Language (cs.CL)

关键词:speech remains under-evaluated, enable low-bitrate speech, low-bitrate speech compression, audio codecs enable, codecs enable low-bitrate

备注: Accepted to IEEE Speech Language Technology

点击查看摘要

Abstract:Neural audio codecs enable low-bitrate speech compression and tokenization, yet their robustness on accented and multilingual speech remains under-evaluated. We benchmark seven open-source neural codecs (DAC, EnCodec, FocalCodec, LanguageCodec, SemantiCodec, UniCodec, WavTokenizer) on three African speech datasets afrinames, afrispeech dialog, afrispeech multilingual, reporting signal-level quality (NISQA, UTMOS, ViSQOL, STOI, F0-RMSE) alongside two downstream tasks: automatic speech recognition (ASR) and speaker verification (ASV). Our analysis yields four findings. First, signal metrics differ sharply in downstream validity: reference-based structural/intelligibility measures (ViSQOL, STOI) and prosodic error (F0-RMSE) track ASR and ASV degradation far more reliably than neural mean-opinion-score predictors (NISQA, UTMOS). Second, intelligibility and speaker-identity preservation diverge sharply across architectures, and the apparent identity ranking itself depends on the ASV backend. Third, degradation is strongly domain-dependent and largest for conversational dialog. Fourth, the resulting degradation is partly recoverable: parameter-efficient \emph{codec} adaptation (LoRA, $\sim$1--3\% of parameters) on roughly 35 hours of African speech reduces the compression-induced word-error-rate gap, recovering recognition toward the uncompressed baseline (developed in companion work). These results motivate task-aware, domain-representative, and adaptation-aware evaluation of speech codecs as a prerequisite for inclusive deployment.

165. 【2610.00136】Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

链接:https://arxiv.org/abs/2610.00136

作者:Parker Hummel(Minot State University),Ryne Skabo(Minot State University),Muhammad Abusaqer(Minot State University)

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:presents a reproducible, educational study, image classification, paper presents, SMS Spam Collection

备注: Presented at the 58th Midwest Instruction and Computing Symposium (MICS 2026), Eau Claire, WI, March 27 to 28, 2026. 14 pages, 7 figures, 4 tables

点击查看摘要

Abstract:This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $\epsilon$ = 0.15 and 1.72% at $\epsilon$ = 0.30; under PGD it fell to 32.47% and 0.41%, and a bit-depth-reduction defense recovered only part of the loss. In the second experiment, DistilBERT fine-tuned on the SMS Spam Collection reached 98.75% accuracy and a 94.96% F1-score, but a controlled sequence of pre-defined perturbations (character substitutions, whitespace noise, and a benign suffix) produced only modest probability shifts in most displayed examples and no flip from spam to ham. Adversarial vulnerability is strongly modality-dependent: the MNIST experiment is a clear evasion demonstration, whereas the text experiment is a controlled robustness evaluation. Robustness must be tested empirically rather than inferred from clean accuracy.

166. 【2610.00097】DramaAgent: Agentic Storytelling Video Generation

链接:https://arxiv.org/abs/2610.00097

作者:Ting Huang,Biao Wu,Ronghao Chen,Zeyu Zhang,Tengfei Cheng,Qizhen Lan,Huacan Wang,Hao Tang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:audio remains challenging, aligned audio remains, Recent diffusion, producing coherent long-form, substantially improved

备注:

点击查看摘要

Abstract:Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross-scene continuity, and audio-visual mismatch over extended sequences. We propose DramaAgent, a hierarchical, agentic, and model-agnostic framework for long-form text-to-video-and-audio generation. Rather than improving the underlying video backbone itself, DramaAgent introduces an upper-level control layer that decomposes generation into story planning, persistent character conditioning, scene-wise synthesis, and reflection-guided targeted repair. The framework maintains reusable story and character states across scenes, diagnoses failures such as identity drift, missing scene semantics, temporal discontinuity, and cross-modal mismatch, and repairs problematic clips in a stage-specific manner. Experiments across multiple video generation backbones show that DramaAgent improves long-horizon coherence, character consistency, narrative fidelity, and scene-level audio-visual consistency over direct generation and strong baselines. These results suggest that hierarchical agentic control is a practical direction for controllable long-form audiovisual generation. Code: this https URL. Website: this https URL.

167. 【2610.00096】FACET at WMT 2026 Automated Translation Quality Evaluation Task

链接:https://arxiv.org/abs/2610.00096

作者:Ahrii Kim,Chanjun Park,Seong-heum Kim

类目:Computation and Language (cs.CL)

关键词:machine translation require, Automated Translation Quality, error type requires, Translation Quality Evaluation, Quality Evaluation Task

备注: Accepted at WMT 2026 (shared task system paper)

点击查看摘要

Abstract:Different error types in machine translation require different evidence. Whether meaning is preserved can be judged only against the source, while whether the target is well-formed, or whether it names one entity consistently, can be judged from the target alone. We present FACET, our reference-free submission to the WMT26 Automated Translation Quality Evaluation Task, which decomposes evaluation into Fluency, Accuracy, and Consistency passes and gives each pass only the context its error type requires. A single fixed model is prompted three times, and the merged error spans yield the three task outputs, error spans, quality scores, and error-free labels, with no trained components. We also submit FACET-C, which omits the Consistency pass. Without gold labels, we characterize the predictions of FACET. Its system rankings place post-edited human translation first, and the Consistency pass changes about a tenth of segment scores while leaving the ranking nearly unchanged.

168. 【2610.00092】BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL

链接:https://arxiv.org/abs/2610.00092

作者:Chen Shen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)

关键词:Data agents, context window, fit database schema, model context window, agents over structured

备注:

点击查看摘要

Abstract:Data agents over structured sources must fit database schema into the model's context window. Large catalogs can span many databases and thousands of columns, so cost constraints may require choosing between table coverage and serialization detail well before the context window is full. We introduce BudgetSchemaBench, an execution-grounded diagnostic for this setting. Its construction derives relevance labels mechanically from gold SQL, without human- or LLM-authored ground truth. Using a pooled 80-database catalog, we sweep four schema-context budgets and compare three representations while keeping each retriever's table ranking fixed. A source-namespace check rejects queries that obtain the correct result from the wrong database. The evaluation covers three conditions: end-to-end retrieval; frozen-gold, in which the required tables are guaranteed; and a probe that removes those tables. For the primary solver with raw serialization, raising the budget from 2.5% to 50% of the catalog improves execution accuracy on 1,279 held-out questions by 18 percentage points under lexical retrieval but only 3 under dense retrieval; the dense retriever already finds most required tables at the smallest budget. When the required tables are removed, 94.6% of correct predictions name one of them exactly, consistent with reconstruction of absent schema from parametric knowledge. For the two main solvers in the frozen-gold condition, the three representations differ by at most 2 percentage points, and the widest paired 95% confidence interval bounds the difference within +/-4 points. We observe the same qualitative patterns with one reasoning model from a different family. When retrieval is coverage-limited, execution accuracy is more sensitive to the schema budget than to the tested serializations. The diagnostic and the code used to construct and evaluate it are publicly available.

169. 【2610.00087】Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights

链接:https://arxiv.org/abs/2610.00087

作者:Jeongmin Lee

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:expanded AI-based text, natural language processing, AI-based text classification, legal, legal text classification

备注: 22 pages, 3 figures. Published in Artificial Intelligence and Law

点击查看摘要

Abstract:The advancement of natural language processing (NLP) has expanded AI-based text classification in the legal domain. However, accurately classifying legal documents remains challenging due to the complexity of legal texts and subtle differences between legal categories. This study evaluates legal text classification models ranging from traditional machine learning techniques to large language models (LLMs) using ten categories of Korean sexual offense precedents. The results show that fine-tuning small-scale models such as KLUE-BERT on legal data outperforms general-purpose models such as GPT-3.5 and GPT-4.0, as well as traditional machine learning models. KLUE-BERT achieved the highest accuracy of 99.3%, indicating that domain adaptation and fine-tuning can be more important than model size for legal document classification. We further employ explainable AI (XAI) techniques to analyze model predictions and misclassification cases. XAI analysis identifies linguistic features influencing model decisions and limitations in capturing subtle textual cues. Using KICS data, which closely resembles real-world legal case records, we further evaluate the model's generalization capabilities and find that it struggles to interpret implicit contextual cues. These findings highlight the importance of both performance and interpretability in legal AI and demonstrate how XAI can improve transparency in legal text classification. AI-assisted tools can support legal professionals in tasks including document classification, legal information retrieval, and case assessment.

170. 【2610.00084】Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

链接:https://arxiv.org/abs/2610.00084

作者:Timothy Kassis

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Detailed profession-specific system, Detailed profession-specific, system prompts raise, profession-specific system prompts, evaluate Scientific Agents

备注: 46 pages (11 pages main text, references, 33-page appendix); 10 figures, 29 tables. Evaluated corpus: [this https URL](https://github.com/K-Dense-AI/scientific-agents) (commit 48dedd2); evaluation code and item-level records are not released

点击查看摘要

Abstract:Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific this http URL profiles, with Gemini 3.8 Flash via OpenRouter in the Pi agent harness. We compare matched profiles with four controls: a minimal baseline ("You are a helpful assistant"), the profile's opening role sentence, a generic scientific rigor guide, and a profile from an unrelated domain. Across nine text-based science benchmarks (4,531 sampled questions, 100 matched profiles), 4,488 items completed all five conditions after API-error retries, scored with automated, rule-based grading. The average profile-baseline accuracy difference is -0.6 percentage points (95% bootstrap interval [-1.5, +0.2] across fixed tasks), and no benchmark shows a statistically clear improvement. Matched profiles produced 1.5-2.3 times as many output tokens and cost 2.2-4.5 times more per successful call. On 60 tool-using BioMysteryBench bioinformatics problems (three runs each for baseline and profile), mean solve rates were 46.7% with the profile and 56.7% at baseline, a difference of -10.0 percentage points (95% interval [-16.7, -3.3]) driven by more frequent token- and time-limit stops under the profile. Longer prompts had one unexpected operational advantage: on SuperGPQA, frequent provider API drops left the short baseline with a correct first-pass answer on only 54.0% of items, against 71.6% with the profile. Generic and mismatched prompts were about as reliable, so this gain comes from prompt length or formatting rather than domain expertise. For the tested model and tasks, loading full profession profiles by default does not improve accuracy and costs considerably more; whether selective retrieval of profile sections or open-ended scientific tasks would change this remains to be tested.

171. 【2610.00070】Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation

链接:https://arxiv.org/abs/2610.00070

作者:Antonela Tommasel,Markus Schedl

类目:Computation and Language (cs.CL)

关键词:Large Language Models, study Large Language, Large Language, Researchers increasingly, including social-cognitive constructs

备注:

点击查看摘要

Abstract:Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring, framing effects, and confirmation bias. Such approaches offer alternatives to overt bias probes, particularly when direct questioning may obscure bias or when model behaviour appears normatively acceptable. However, adapting human bias constructs to LLMs introduces an inferential gap. Psychological instruments were developed to study human cognition and social behaviour, whereas LLM evaluations rely on probabilities, text completions, rankings, or simulated decisions. This paper critiques human-centered bias evaluation in LLMs. We show how this gap arises from mismatches pertaining to human-derived constructs, human-model differences, and evaluation contexts, which can blur distinct interpretations of model bias. We then introduce a framework providing an analytical lens for relating these elements to warranted interpretations, with attention to target constructs, operationalizations, scope of inference, and limits of human analogy.

172. 【2610.00063】Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

链接:https://arxiv.org/abs/2610.00063

作者:Pakorn Nathong,Kunat Pipatanakul

类目:Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Instance-billed serverless platforms, serverless platforms charge, Instance-billed serverless, direct serving cost, platforms charge

备注: 6 pages, technical report

点击查看摘要

Abstract:Instance-billed serverless platforms charge for CPU and memory over the lifetime of a warm instance, making idle inference state a direct serving cost. We present billing-aware neural text-to-speech (TTS) serving on serverless CPUs, optimizing CPU-seconds and GB-seconds rather than throughput or latency alone. Conventional runtimes are poorly suited to this setting: per-request parallelism causes CPU contention under concurrency, while warm instances retain gigabytes of billable inference and page-cache state. We address these costs with request-sized concurrent inference, which bounds per-request CPU parallelism, and a reclaimable instance lifecycle, which releases inference state and page-cache memory after idle periods while retaining the server process and compile cache. On Kokoro-82M, our system achieves 2.71 audio-seconds per CPU-second versus 0.89 with ONNX Runtime defaults and reduces cost per audio-hour from $0.0631 with PyTorch to $0.0153, a 4.1x reduction. Idle billed memory falls from 8.7 GB to 1.33 GB, while restoration reaches first audio in 2.2 s versus 7.7 s for a PyTorch cold start. Under bursty traffic, lifecycle reclamation is essential for translating inference efficiency into lower serverless cost.

Comments:
6 pages, technical report

Subjects:

Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

Cite as:
arXiv:2610.00063 [cs.DC]

(or
arXiv:2610.00063v1 [cs.DC] for this version)

https://doi.org/10.48550/arXiv.2610.00063

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
173. 【2610.00054】he First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

链接:https://arxiv.org/abs/2610.00054

作者:Gnaneswar Villuri,Hashmath Shaik,Alex Doboli

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Reading an LLM, evaluation harnesses produce, likelihood-scoring evaluation harnesses, LLM judge verdict, harnesses produce

备注:

点击查看摘要

Abstract:Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama-3.1-8B and Phi-3.5-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422, 95% CI [+0.365, +0.467]). The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one. A second, smaller failure occurs even when the judge does lead with a verdict token, since it sometimes opens with one letter and reasons its way to the other, on 0 to 5.5% of pairs at a rate uncorrelated with compliance. We recommend reporting the rate at which a judge leads with a verdict token, which costs one forward pass and no labels, alongside any position-bias figure.

174. 【2610.00052】Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-49 Outputs

链接:https://arxiv.org/abs/2610.00052

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:distinct random integers, integers from 1-49, evaluate six language-model, random integers, English prompt variants

备注: 6 pages, 1 figure. Data, code, and collector at [this http URL](http://github.com/Rankfor/rankfor-open) (research/lotto-models)

点击查看摘要

Abstract:We evaluate six language-model configurations on requests for six distinct random integers from 1-49. Across 1,200 attempted calls using four English prompt variants, 1,184 responses yielded valid tickets. Effective diversity of number frequencies ranged from 9.9 to 18.0, compared with simulated fifth-percentile thresholds of 46.6-46.7 under independent uniform six-of-49 sampling at the corresponding sample sizes. Systems produced 8-93 distinct unordered tickets, and their modal tickets accounted for 22.5-68.0% of valid responses. Two archived Polish Lotto samples provided a physical-lottery comparison, with effective diversities of 41.1 and 41.4 at smaller sample sizes. These results demonstrate substantial concentration under the tested deployment settings. They do not identify its mechanism or establish performance under other prompts, temperatures, or tool configurations.

175. 【2610.00047】Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

链接:https://arxiv.org/abs/2610.00047

作者:Khawaja Murad ul Hassan,Mehran Ebrahimi

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:process reward model, Diversity collapse, motivated inference-time interventions, inference-time interventions built, PRM-Pruned Fragment Grafting

备注: 24 pages, 4 figures, 22 tables

点击查看摘要

Abstract:Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells. A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation. A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.

176. 【2610.00045】SCM-based Fairness and Faithful Explainability for Legal Document Classification

链接:https://arxiv.org/abs/2610.00045

作者:Yasmina El Kacemi,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag

类目:Computation and Language (cs.CL)

关键词:legal decision support, Transformer models, decision support, raising concerns, legal decision

备注:

点击查看摘要

Abstract:Transformer models such as LegalBERT are increasingly used in legal decision support, raising concerns about both fairness and the transparency of model explanations. These properties are usually evaluated separately, leaving open whether a debiasing intervention that changes fairness also changes how faithfully explanations reflect model reasoning. This study investigates that relationship on the ECtHR alleged-violations corpus from LexGLUE. It compares a LegalBERT baseline with a fairness-regularized variant that penalizes stereotypical warmth and competence representations during fine-tuning. The evaluation covers predictive performance, demographic fairness, and SHAP explanation faithfulness across five random seeds. At the performance-optimal regularization strength, the intervention does not reduce demographic disparity. This null result holds across two fairness definitions and a conventional word-pair control on the gender axis. Classification performance is largely unchanged. However, the intervention consistently degrades explanation sufficiency across all five seeds and three thresholds. A shuffled-pair control reproduces this degradation while leaving performance and fairness unchanged, indicating that the effect arises from contrastive representational regularization rather than specifically from the warmth and competence structure. The results demonstrate a dissociation between fairness and explanation faithfulness: changes in explanation behavior do not necessarily indicate changes in fairness, and fairness must therefore be evaluated directly.

177. 【2610.00035】Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System

链接:https://arxiv.org/abs/2610.00035

作者:Bente Hinkenhuis,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:support meaningful intervention, demographic information introduces, Predicting student performance, educational interaction data, data requires models

备注:

点击查看摘要

Abstract:Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of unfair predictions. This study investigates a multi-objective framework that combines reinforcement learning-based multiple instance learning (RL-MIL), adversarial debiasing, and preference-conditioned hypernetworks for student-at-risk prediction. MIL represents each student as a bag of weakly labeled interactions, while an RL agent selects informative instances for downstream classification. Two hypernetwork variants are evaluated to determine whether a user-defined preference scalar can continuously control the trade-off between predictive performance and Equalized Odds. The underlying RL-MIL baseline achieves strong classification performance, but both hypernetwork extensions exhibit mode collapse: changing the preference weight produces little systematic movement along the intended fairness-performance frontier. The failure is associated with objective dominance, weak gradient propagation through the conditioning mechanism, and interactions between dynamically generated parameters. The results show that fairness objectives can be incorporated into an interpretable RL-MIL pipeline, but preference conditioning alone does not guarantee controllable multi-objective behavior. Robust fair RL-MIL therefore requires explicit mechanisms for gradient balancing, objective separation, and stability analysis.

178. 【2610.00026】High-Value Synthetic Supervision for Parameter-Efficient Adaptation of a Compact Japanese Speech Model

链接:https://arxiv.org/abs/2610.00026

作者:Sidi Chang,Peiying Zhu

类目:ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Private domain speech, Private domain, collect and redistribute, task-specific supervision, difficult to collect

备注: Submitted to On-Device Intelligence: Foundation Models under Real-World Constraints (NeurIPS 2026 workshop). 4 pages, 0 figures, 1 table

点击查看摘要

Abstract:Private domain speech is difficult to collect and redistribute, while compact models need task-specific supervision. We study an auditable synthetic pipeline that maps Japanese care handoffs directly to six-field structured notes. Using 182 synthetic training and development clips, we adapt a 1.47B audio model by full fine-tuning and rank-16 LoRA. On a 39-clip scenario-seed-disjoint synthetic test, an unadapted model obtains a model-judged factuality-recall score of 0.0500, full tuning 0.8664, and LoRA 0.8461. LoRA reaches 97.7% of the full-tuning aggregate as a descriptive ratio while project telemetry reports 12.4M trainable parameters, about 0.85% of the backbone. Both adaptations show large paired gains over the same base; the full-versus-LoRA interval crosses zero, and differing optimization settings preclude an equivalence claim. This is a parameter-efficient capability-acquisition result, not a device-performance result: latency, memory, energy, and real-time factor were not measured. All evaluation speech and targets are synthetic, references are model-proposed, and the judge is uncalibrated. The evidence shows that a compact model can acquire a narrow audio-to-structure transformation from a few hundred provenance-linked synthetic examples; it does not establish clinical validity, real-speech transfer, or superiority to a clean cloud system.

179. 【2610.00024】Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

链接:https://arxiv.org/abs/2610.00024

作者:Genpei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:universal negative finding, vision-language model architectures, mid-layer interpretability, vision-language model, universal negative

备注: 13 pages, 4 figures

点击查看摘要

Abstract:Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p 1e-4). Despite the null, the errors separate operationally into three failure modes -- Perception Failure, Encoded-but-Disconnected, Prior-Override -- learnable above 60% on all three architectures, and the architecture's prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.

180. 【2610.00018】What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA

链接:https://arxiv.org/abs/2610.00018

作者:Jiameng Zhang,Hongqiu Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:pipelines increasingly pass, Role-specialized QA pipelines, stronger support assessment, increasingly pass rationales, message actually buys

备注: 14 pages, 6 figures, 9 tables. Preprint

点击查看摘要

Abstract:Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. We introduce a message-intervention diagnostic that fixes the evidence and candidate answer while varying only the rationale passed across the reasoner-to-verifier boundary. On 400 MuSiQue, HotpotQA, and 2WikiMultiHopQA examples with DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy over no rationale, while corrupted rationales strongly alter support judgments. Under a blind verifier prompt, harmless paraphrases shift support by only 0--2.5%, whereas corrupted rationales shift support by 10--22%; an explicit rationale-checking prompt amplifies the same pattern to 34--55%. Final answers move less (2--30%), and only 2.9--35.3% of corrupted support flips co-occur with answer changes. Human audits show why this matters: 16/42 valid corruptions are corruption-overtrust cases, and blind humans reject or mark unclear 9/10 audited corrupted rationales that the model accepts. Cross-model and task-boundary checks show when the channel is active, amplified, inert, or folded into the task label. Rationale sharing should be evaluated as a verification-message mechanism, not merely as a route to higher answer accuracy.

181. 【2610.00009】FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law

链接:https://arxiv.org/abs/2610.00009

作者:Athanasios Zeris

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Signal Processing (eess.SP)

关键词:achieves large gains, standard dot-product attention, achieves large, Frequency-collapse attention, dot product

备注: 9 pages, 1 figure, 2 tables

点击查看摘要

Abstract:Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about filter properties -- DC suppression, Nyquist suppression, bandwidth, centre frequency, and multi-scale coverage -- using a controlled ablation on character-level language modelling (TinyShakespeare, 6-layer GPT). Our main findings are: (1) DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), confirming that oscillatory bandpass structure is essential, not just any low-dimensional spectral summary; (2) the optimal single-scale bandwidth is sigma ~= 2 bins centred at paragraph scale (~70 tokens), giving a clean gain of Delta = +1.15 nats over BASE-DOT; (3) admissible filters (zero-mean, Mexican Hat DOG m = 2) outperform non-admissible Gaussians at the same scale and provide partial protection against bilateral FFT leakage; (4) bilateral FFT leakage scales monotonically with spectral coverage -- narrowband filters (gap +4) are clean, wideband filters (gap +2) are leaky; and (5) causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context), motivating word-level experiments in the companion MorletQK paper [Zeris, 2026f]. Together, findings (1)-(5) characterise FourierQK as effective in bidirectional attention settings (encoder-style, e.g. BERT), where full-sequence context is available at both training and inference time; autoregressive generation requires a causal spectral variant such as MorletQK [Zeris, 2026f] (decoder-style, e.g. GPT). Code available at: this https URL

182. 【2610.00007】On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence

链接:https://arxiv.org/abs/2610.00007

作者:Vinay Kumar Chaganti

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:increasingly wanted on-device, Named-entity recognition, data kept local, wanted on-device, increasingly wanted

备注: 7 pages, 5 figures, 12 tables. Code and per-span records reproduce all reported numbers offline

点击查看摘要

Abstract:Named-entity recognition (NER) is increasingly wanted on-device (no API, low latency, data kept local). The practitioner's question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. We answer these jointly. We place nine systems across three paradigms and 13 M to 8 B parameters: a classical tagger (spaCy), bidirectional-encoder specialists (GLiNER, 166 to 460 M), and generative LLMs run locally (Qwen3-0.6B/1.7B/4B-Instruct, DeepSeek-R1-1.5B/8B), on three datasets of differing character, and report accuracy plus two axes the literature omits: latency and output validity. Because our corpus (RSS-News) had no gold, we built silver gold from a cross-family LLM judge panel, then measured its fidelity against benchmark gold and a full human re-validation of the corpus (strict F1 0.95, an upper bound since the human gold was silver-seeded); gold provenance flips the paradigm ranking, moving from LLM-authored silver to human gold raises every encoder and lowers every generative model. On accuracy alone a 4 B instruct LLM is competitive (it leads on clean newswire), so the encoder's case is deployability: it matches or slightly trails at one-ninth to one-twenty-fourth the size, at millisecond-to-second latency, with zero malformed output, while the smallest generative models emit up to 27% invalid output on long inputs, a failure fixed by scale, not output budget. We then characterize GLiNER's per-span confidence: it ranks correctness well (AUROC 0.76 to 0.86) but is overconfident (ECE 0.24 to 0.47, halved by temperature scaling); thresholding gives a small honest out-of-sample F1 gain; an all-local small-to-large cascade gives a modest, corpus-dependent gain over cost-matched random routing; and confidence tracks correctness but not novelty. Every number recomputes offline from per-span records.

183. 【2608.22124】LLM assisted writing deserves empirical evaluation

链接:https://arxiv.org/abs/2608.22124

作者:Xuan Zhong Feng,Yi Lin,Yiye Zhang,Chunhua Weng,Yifan Peng

类目:Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:LLM-assisted writing, detection problem, questions about clarity, Health Informatics papers, raises questions

备注:

点击查看摘要

Abstract:LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, integrity, equity, and evaluation. An analysis of 69,209 Health Informatics papers links it to more focused presentation, broader citation practices, and more globally distributed authorship. These patterns do not prove better science, but they support evaluating manuscripts by scholarly quality and accountability rather than by tool use.

184. 【2607.18270】rajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation

链接:https://arxiv.org/abs/2607.18270

作者:Kyunghoon Jeon,Youmin Ko,Woohwan Jung,Hyunjoon Kim

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Electronic Health Records, Electronic Health, Health Records, clinical risk remains, heterogeneous external knowledge

备注: Accepted to KDD 2026

点击查看摘要

Abstract:While Electronic Health Records (EHRs) offer a wealth of clinical data, effectively augmenting a patient's records with heterogeneous external knowledge to predict the patient's clinical risk remains a significant challenge. Existing methods fail to capture disease severity, treatment responses, and nuanced clinical progression, due to data sparsity and the underutilization of unstructured clinical notes. To address these challenges, we propose TRACER (a trajectory-aware and clinically grounded prediction framework) that (1) constructs a medical knowledge graph enriched with severity information from medical literature, (2) retrieves clinically relevant, severity-weighted paths of a patient's progression from the knowledge graph, (3) extracts clinically relevant events from unstructured clinical notes, and (4) augments patient context with similar peer cases. Experiments on the MIMIC-III and MIMIC-IV datasets demonstrate large gains over state-of-the-art baselines, with up to 28.5% increase in Macro F1 score for the mortality prediction task, and 19.7% increase for the readmission prediction task.

185. 【2610.01450】Code-Switching Spoken Language Identification as Multi-Label Set Prediction

链接:https://arxiv.org/abs/2610.01450

作者:Shunsuke Mitsumori,Matthew Wiesner,Shigeo Morishima,Shinji Watanabe

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:CS-aware LID, massive speech corpora, curate massive speech, LID, calling for CS-aware

备注: Accepted at IEEE SLT 2026

点击查看摘要

Abstract:Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.

信息检索

1. 【2610.02202】ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

链接:https://arxiv.org/abs/2610.02202

作者:Sohyeon Kim,Yoonho Lee,Bo Liu,Dayoon Ko,Rulin Shao,Seungone Kim,Graham Neubig,Pang Wei Koh,Aakanksha Chowdhery,Akari Asai,Omar Khattab,Yejin Choi,Gunhee Kim,Chelsea Finn

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:great scientists great, makes great scientists, great scientists, scientists great, makes great

备注: 57 pages

点击查看摘要

Abstract:What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

2. 【2610.02057】Optimizing Effective Training Time for Large-Scale Recommendation Systems

链接:https://arxiv.org/abs/2610.02057

作者:Mingming Ding,Ruilin Chen,Yuzhen Huang,Hang Qi,Menglu Yu,San Tan,Damian Reeves,Boris Sarana,Kevin Tang,Satendra Gera,Gagan Jain,Sahil Shah,Vishwa Karia,Fuzail Khan,Yashasvi Makin,Edward Z. Yang,Oguz Ulgen,Jia Chen Ren,Laith Sakka,Mayank Garg,Meet Vadakkanchery,Aici Lin,Wei Sun,Mengjiao Zhou,Shuai Yang,Junqing Zhou,Max Leung,Apoorv Purwar,Musharaf Sultan,John Bocharov,Zhenyu Tang,Vivek Trehan

类目:Information Retrieval (cs.IR)

关键词:silently consumes accelerator, consumes accelerator capacity, overhead silently consumes, Lifecycle overhead silently, large-scale recommendation training

备注:

点击查看摘要

Abstract:Lifecycle overhead silently consumes accelerator capacity across large-scale recommendation training fleets. Our largest recommendation workloads process tens of billions train- ing examples per day on thousands of GPUs. Before this work, only 50-60% of their end-to-end wall time advanced training on new data. We present a fleet-scale study of this lifecycle overhead and a set of optimizations spanning the full training stack. We use Effective Training Time (ETT%) as an operational framework to instrument lost time, localize it to independently owned infrastructure components, and expose work repeated across job restarts. This analysis guides optimizations like communication elimination and pipeline overlap during trainer initialization; dynamic-shape handling, autotuning pruning, and reusable Py- Torch 2 compilation caches; asynchronous checkpointing; stan- dalone model publishing; and reductions in recovery cost. We evaluate the optimizations on representative models and measure their impacts in our training fleet. ETT% improves on every benchmark, by 15.5% on average, and reaches 85% on our largest workload. Fleet-wide ETT% rose from about 80% to above 90% after deployment.

3. 【2610.01767】A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering

链接:https://arxiv.org/abs/2610.01767

作者:Gianluca Bonifazi,Christopher Buratti,Michele Marchetti,Federica Parlapiano,Giulia Quaglieri,Davide Traini,Domenico Ursino,Luca Virgili

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:multi-hop Question Answering, Question Answering, Large Language Models, Retrieval-Augmented Generation, balance retrieval quality

备注:

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.

4. 【2610.01705】AgentWebRec: Compact Evidence Fusion over the Agent Web for Personalized Recommendation

链接:https://arxiv.org/abs/2610.01705

作者:Haoran Qiang,Guannan Liu,Liang Zhang,Junjie Wu

类目:Information Retrieval (cs.IR)

关键词:user knowledge locally, maintaining richer user, richer user knowledge, LLM-based personal agents, LLM-based personal

备注:

点击查看摘要

Abstract:LLM-based personal agents are emerging as persistent carriers of user semantics and intermediaries between users and recommendation platforms, maintaining richer user knowledge locally. As agents interact with one another, the conventional \textit{User--Platform} relation evolves into a \textit{User--Agent Web--Platform} information pathway, enabling distributed user-side information to complement item-side information. This new pathway, however, defies conventional recommendation: evidence is scattered across mutually opaque agents and reachable only through bounded queries, only a small portion of it is relevant to the current recommendation decision, and the responses returned by different agents are semantically heterogeneous. We therefore recast recommendation over the agent web as a \emph{task-time evidence acquisition and fusion} problem under a finite evidence budget by deciding what to ask and what to keep, rather than learning from aggregated data. We propose AgentWebRec, a user-agent-oriented framework that progressively acquires and fuses distributed evidence for each user-item decision while keeping underlying agent memories local. It grounds each decision in platform-provided item semantics and task-relevant evidence from the target user agent's private memory, and conditionally queries neighboring user agents for complementary preference patterns when local evidence is insufficient. Experiments on four InstructRec datasets show that AgentWebRec consistently outperforms baseline recommenders, and ablations verify that the evidence layers contribute complementary gains.

5. 【2610.01553】From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search

链接:https://arxiv.org/abs/2610.01553

作者:Nikolai Zenovkin,Sebastian Björkqvist

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:search requires processing, Patent search requires, routinely exceeding tens, requires processing documents, processing documents routinely

备注: Accepted for publication at the ECML PKDD 2026 conference (Applied Data Science track)

点击查看摘要

Abstract:Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.

6. 【2610.01533】Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)

链接:https://arxiv.org/abs/2610.01533

作者:Aleksei Medvedev,Alejandro Ariza-Casabona,Steven Derby,Gonzalo Fiz Pontiveros,Xinyang Shao,Florian Spiess

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:recommendation treats SID, generative recommendation treats, Existing work, treats SID construction, quantised latent space

备注:

点击查看摘要

Abstract:Existing work on Semantic IDs (SIDs) for generative recommendation treats SID construction as a representation learning problem: encode items into a quantised latent space and read off codes. We argue this view is incidental. SID construction is, at heart, a recursive clustering problem, and once stated this way the natural object to cluster is a graph whose nodes carry semantic content and whose edges carry collaborative signal; SID assignment becomes a hierarchical graph partition. This reframing yields a unified framework, Graph-Informed Semantic IDs (GrIS), that subsumes prior approaches rather than displacing them. RQ-VAE and RQ-KMeans are recovered as the special case where the graph is empty, exposing content-only quantisation as one corner of a larger design space along two so-far-collapsed axes: graph construction and recursive partition algorithm. We explore two contrasting instantiations: RecDMoN, which performs hierarchical assignment via differentiable graph pooling, and RQ-GAE, which extends RQ-VAE with graph-aware item representations and a graph reconstruction objective. On multiple real-world datasets, GrIS consistently improves over CF-aware SOTA, with gains of up to +52\% Hit@10. Because graph construction and partition are explicit, separately configurable components, improvements on either axis can be combined and evaluated systematically.

7. 【2610.01463】Learning to structure data from user-generated thematic corpora

链接:https://arxiv.org/abs/2610.01463

作者:Elishay Avram,Oren Glickman,Elad Yom-Tov

类目:Information Retrieval (cs.IR)

关键词:unstructured text describing, text describing data, social media data, social media, unstructured text

备注:

点击查看摘要

Abstract:Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.

8. 【2610.01270】Not All Is Lost: Repairing Lossy User Preference States of Personalization Encoders

链接:https://arxiv.org/abs/2610.01270

作者:Parthiv Chatterjee,Dhiraj Golhar,Ummesalma Diwan,Sourish Dasgupta,Manjunath Joshi,Tanmoy Chakraborty

类目:Machine Learning (cs.LG); Information Retrieval (cs.IR)

关键词:Personalization encoders compress, encoders compress evolving, evolving interaction histories, compress evolving interaction, Personalization encoders

备注: Accepted to NeurIPS 2026. Author-prepared archival version with expanded discussion and interpretation. 59 pages, including references and appendices

点击查看摘要

Abstract:Personalization encoders compress evolving interaction histories into preference states used to rank items or condition text generation. A task head operating only on this state can miss useful evidence that remains in the frozen encoder's cached representations for individual timesteps. We study this recoverability gap and propose REPAIR, which compares cached representations with the current preference state in a compact learned coordinate space. It resolves corrective evidence over extended history, recent interactions, and localized bursts. It then selects which patterns at which timesteps contribute and adds their aggregate correction to the state before the task head. Encoder-host repair reuses representations from the existing forward computation without re-encoding the history. Across MovieLens, PENS, MIND, and Amazon Reviews 2023, training only REPAIR improves MRR and nDCG@10 for all twelve representative recommendation hosts while both encoder and task head remain frozen. Head-only finetuning of the same hosts yields smaller gains. For example, Mamba4Rec on MovieLens gains 3.96 MRR points, compared with 0.19 from head-only finetuning. Rank and temporal diagnostics support a compact, host-dependent corrective structure. In personalized generation, IMPerSumm improves the two reported weighted PerSEval variants, which assess responsiveness to user preference, by up to 25.23%. These results support post-compression state correction and distinguish the availability of preference evidence from its downstream use.

9. 【2610.01139】Do Multilingual Encoders Produce Language-Consistent Semantic IDs?

链接:https://arxiv.org/abs/2610.01139

作者:Abhinav Bohra,Anuj Bohra

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:compress item embeddings, Semantic IDs, discrete code sequences, compress item, generative retrieval

备注: 7 pages, 8 tables. Accepted as a short paper at WiNLP 2026, co-located with EMNLP 2026

点击查看摘要

Abstract:Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.

10. 【2610.01118】Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives

链接:https://arxiv.org/abs/2610.01118

作者:Zhiyun Shi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:long-term conversational assistant, long-term conversational, conversational assistant, assistant must recall, thousand LLM calls

备注: 17 pages, 4 figures

点击查看摘要

Abstract:A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand LLM calls per memory bank and up to several thousand context tokens per query. We argue that association is a learnable relevance: the pointwise mutual information of memories under how human lives unfold. We introduce Madeleine, which learns amortized association: offline, an LLM life simulator writes simulated lives, whose cue-trigger pairs teach a query encoder a residual association on top of frozen similarity; online, it calls no LLM and plugs into any vector memory by replacing only the query encoder. On LoCoMo-Plus under the official protocol, Madeleine (I) reaches 66.6 when plugged into HyperMem, the highest among all systems evaluated under this protocol; (II) used alone, reaches the score of HyperMem as released (52.4 vs. 52.9) with zero LLM calls and about 1/21 of its answer context; and (III) lifts T-Mem by 26.2 points, significantly outperforms the same untrained backbone inside both systems, and leaves ordinary QA intact on the 4B backbone.

11. 【2610.01064】JoinGR: Learning to Traverse Join Graphs for Table Retrieval

链接:https://arxiv.org/abs/2610.01064

作者:Sandipan De,Abhijit Chakraborty,Sambaran Bandyopadhyay,Vivek Gupta

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB); Information Retrieval (cs.IR)

关键词:Retrieving, JOINGR, realistic databases, tables, table

备注: 12 pages, 6 figures, 5 pages

点击查看摘要

Abstract:Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, but this ignores a key source of evidence: some required tables are not mentioned in the question and become identifiable only through their join relationships to already relevant tables. We introduce JOINGR, a join-aware table retrieval method that treats the database join graph as the retrieval space. Columns are represented as graph nodes, while intra-table and foreign-key relationships are represented as typed edges. Given a question, JOINGR selects semantically similar anchor tables, traverses join edges with a query-conditioned scorer, and aggregates the resulting edge deposits into table scores. The scorer is a lightweight MLP on top of frozen query, node, and edge embeddings, trained with a pairwise margin loss over gold tables. On BIRD and Spider datasets, JOINGR is competitive with the strongest retrieval baselines. On BEAVER, a challenging enterprise benchmark with multi-hop table requirements, JOINGR substantially improves recall over dense retrieval and re-ranking baselines. Cross-domain experiments show that the learned scorer transfers across benchmarks, indicating that the method captures reusable joingraph traversal behavior.

12. 【2610.00971】he Other Half of Workflow Portability: Evidence-Backed HPC Site Profiles with Agentic Discovery

链接:https://arxiv.org/abs/2610.00971

作者:Md Saiful Islam,Douglas Thain

类目:Distributed, Parallel, and Cluster Computing (cs.DC); Information Retrieval (cs.IR)

关键词:trial and error, HPC site, developed and tested, rarely succeeds, amount of trial

备注: Accepted to the 21st Workshop on Workflows in Support of Large-Scale Science (WORKS 2026), held with SC26, Chicago, IL, USA. 8 pages, 7 figures, 3 tables

点击查看摘要

Abstract:Moving a workflow developed and tested at one HPC site to another rarely succeeds without some amount of trial and error. Package managers rebuild software environments, containers ship whole filesystems, and workflow specifications such as backpacks package a workflow with its software, data, and resource requirements. These approaches address one half of workflow portability: what a workflow needs. But none describes how a given HPC site must be used, and that missing half is why even a portable workflow requires manual adjustment at each new site. That gap includes the site's resource shape, storage configuration, network permissions, and operating policies. This information may be explicit in the batch system, hidden in the prose of documentation, or buried deep within a router's configuration, making it difficult for an automated deployment tool to turn site knowledge into useful deployment decisions. We propose the HPC site profile, a structured, evidence-backed document that makes this knowledge actionable. We automatically construct it in three steps that mirror where the information lives: measuring the login node, extracting typed fields from documentation with a bounded language-model agent, and submitting pilot jobs for eligible unresolved fields. Every field is verified against its evidence or discarded, so a rule, not the model, decides what enters the profile. The profile then preflights a workflow into an execution plan or an early, explainable failure. We build profiles at Purdue Anvil, TACC Stampede3, and Notre Dame CRC and present a case study of preflighting a real workflow.

13. 【2610.00964】RPTune: Learned Context Curation for LLM Catalog Search

链接:https://arxiv.org/abs/2610.00964

作者:Chuxuan Hu,Hejie Cui,Norman Huang,Shubham Kumar Bharti,Wang-Chiew Tan,Sercan Ö. Arık

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:full-catalog prompting offers, multi-stage retrieval designed, retrieval designed primarily, small merchant businesses, full-catalog prompting

备注: 23 pages, 9 figures, 4 tables

点击查看摘要

Abstract:For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average.

Comments:
23 pages, 9 figures, 4 tables

Subjects:

Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2610.00964 [cs.IR]

(or
arXiv:2610.00964v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.00964

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
14. 【2610.00923】CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

链接:https://arxiv.org/abs/2610.00923

作者:Hyojeong Yun,Jueun Kim,Wook-Shin Han

类目:Information Retrieval (cs.IR)

关键词:Multimodal RAG retrieves, RAG retrieves text, Multimodal RAG, RAG retrieves, retrieves text

备注: 26 pages, 10 figures, project page: [this https URL](https://canopy-project-page.github.io)

点击查看摘要

Abstract:Multimodal RAG retrieves text, tables, images, and videos, but choosing a retrieval granularity does not determine how much context to retain within each item. Coarse units include irrelevant content, while uniformly fine selection can remove context needed to interpret the evidence. Existing compressors address this trade-off with modality-specific mechanisms, leaving open a shared procedure for adapting the retained extent region by region across heterogeneous items. We introduce CANOPY (Canonical Projection over Hierarchy), a framework for adaptive-granularity post-retrieval evidence compression. CANOPY represents retrieved items as hierarchies and uses a node encoder fine-tuned on gold evidence to score regions against the query. Parent-relative refinement compares these scores to select multiple regions at different granularities without LLM calls for node-level pruning. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added. Across five QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. Ablations indicate that additional retrieval drives the main accuracy gains on multi-hop QA. In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2-27.7% relative to the same iterative pipeline without compression, with comparable answer accuracy.

15. 【2610.00817】abJoinBench: A Benchmark for Joinable Table Discovery

链接:https://arxiv.org/abs/2610.00817

作者:Sandipan De,Jin Wang,Vivek Gupta

类目:Databases (cs.DB); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:enabling downstream tasks, join discovery methods, Join discovery aims, large data repositories, Join discovery

备注: 13 pages, 8 Tables, 1 Figure

点击查看摘要

Abstract:Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.

16. 【2610.00791】Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AI

链接:https://arxiv.org/abs/2610.00791

作者:Terry Dorsey,Kevin Huggins

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Enterprise Representation Simplification, Enterprise Representation, Enterprise Representation Complexity, organizational boundaries, shaped by applications

备注:

点击查看摘要

Abstract:Enterprise information is represented through artifacts shaped by applications, projects, technologies, organizational boundaries, and local requirements. These structures accumulate over time, creating representational complexity that must be maintained by the enterprise and interpreted by information consumers and AI systems. This paper introduces Enterprise Representation Simplification (ERS) as reducing unnecessary representational complexity while preserving required information within a defined scope, and Enterprise Representation Complexity (ERC), a representation-neutral model for comparing complexity across representation states. ERC characterizes representational extent through four dimensions: Representation Objects, Interactions, Behaviors, and Supporting Sources. Objects, Interactions, and Behaviors form dependent categories, while Supporting Sources characterize representation exposure. ERC is defined at representation and task levels, enabling comparison and distinguishing architectural simplification from retrieval optimization. The paper develops two consequences of ERS. First, representational structures create lifecycle obligations for maintenance, governance, dependencies, change, enhancement, and operation. An economic model distinguishes recurring global representation cost, recurring task-level cost, and one-time transformation cost, enabling evaluation over a defined time horizon. Second, reductions in task-level ERC reduce the representational extent an AI system must identify, relate, and interpret. Text-to-SQL research provides evidence that reduced schema and reasoning complexity can improve reasoning accuracy. ERC is not a universal complexity, performance, or cost metric. It provides measurable architectural variables for comparing representational alternatives, transformation effects, economic outcomes, and AI reasoning performance.

Subjects:

Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

Cite as:
arXiv:2610.00791 [cs.AI]

(or
arXiv:2610.00791v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2610.00791

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Terry Dorsey [view email] [v1]
Wed, 30 Sep 2026 22:31:40 UTC (477 KB)

17. 【2610.00587】Comparison of Common Crawl News GDELT

链接:https://arxiv.org/abs/2610.00587

作者:Ameir El Ouadi,David Beskow

类目:Information Retrieval (cs.IR)

关键词:large language models, natural language processing, knowledge graphs, technical efforts, important for natural

备注:

点击查看摘要

Abstract:The corpus of worldwide news is important for natural language processing, knowledge graphs, large language models, and other technical efforts. Additionally, this corpus is important for understanding the people, places, organizations, and events that interact in real-time every day. This paper compares two news datasets used for these tasks today, namely the Global Database of Events, Language, and Tone (GDELT) and Common Crawl News. Our research highlights the strengths and limitations of each dataset, analyzing their content and coverage. Notably, while GDELT relies on broadcasts, prints, and web news from across the globe, Common Crawl focuses on news sites from around the world gathered through web crawling. Our analysis revealed considerable differences in where the two datasets gather their news sources.

18. 【2610.00369】A Shared Taste for Model-Written Text: The Generator-by-Selector Matrices of "AI-AI Bias" Show No Detectable Own-Model Premium

链接:https://arxiv.org/abs/2610.00369

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, language models choosing, model, selector prefer text, language models

备注: 10 pages, 4 figures, 3 tables. Reanalysis of publicly available generator-by-selector matrices

点击查看摘要

Abstract:Laurito et al. (PNAS 2025) showed that large language models choosing between two descriptions of the same product, paper or film prefer the description written by a language model over the one written by a person, by a wide margin over what human judges do. Their design crosses five generators with the same five models as selectors, which permits a second question the paper does not headline: does a selector prefer text from its own model beyond what the generator and selector main effects predict? We rebuild the three 5x5 matrices from the per-item counts in the authors' public repository (21,828 valid trials; every cell matches the published value) and fit a two-way fixed-effects model with an own-model term gamma, tested by the exact permutation test over the 120 relabellings of the selectors. The premium is +0.013 on products (exact one-sided p = 0.24), -0.010 on paper abstracts (p = 0.74), +0.054 on films (p = 0.07) and +0.019 pooled (p = 0.14; 95% interval -0.008 to 0.046). The same-vendor term for the GPT-3.5 and GPT-4 pair is negative in all three datasets. Position bias moves single cells by up to 0.42 share points in either direction, and the own-model contrast is unchanged once order-driven items are removed. The design would have detected a premium of 0.05 with 82% (products), 88% (papers), 42% (films) and 97% (pooled) power; the minimum detectable effect at 80% power is 0.034 pooled. The absence is informative down to about 0.04 share points and silent below that. The 4x4 matrix of Tan et al. (ACL 2024) gives gamma = +0.148 at the smallest p its 24 relabellings allow, with a same-family term of the same size. The main result of Laurito et al. stands: models share a taste for model-written text, with GPT-4's descriptions chosen 77% to 95% of the time by every selector on products. What these data do not show is a model recognising and favouring its own prose.

19. 【2610.00253】System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer

链接:https://arxiv.org/abs/2610.00253

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:deployed language models, visibility summarise, language models, models into per-system, deployed language

备注: 30 pages, 5 figures, 9 tables. Appendix D documents corrections to an earlier manuscript

点击查看摘要

Abstract:Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. We test whether such a profile describes the system on one corpus of 6,475 stored responses (6,324 analysable) collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation and category-ownership queries. The collection harness cut many answers short: 83.1% of Gemini 3 Flash answers in category ownership end mid-sentence under a 1,024-token output cap. With every answer cut to its first 800 characters, a character n-gram classifier cross-validated by prompt attributes one response to GPT-5.2, Gemini 3 Flash, Gemini 3 Flash with search, Grok or Perplexity sonar-pro with 97.84% accuracy (5,028 responses, 383 prompts, majority class 31.5%, 30 split seeds). Length alone falls to the majority rate, 24 formatting statistics reach 95.79%, and masking brand names and capitalised tokens leaves 97.72%. Held-out query conditions keep 97.43% weighted by size and 88.0% unweighted; in a retrieval-grounded arm that changes the harness, no Grok answer is attributed to Grok (0/120). Aggregated into 50 model-by-domain-by-condition units, twelve behavioural features separate four systems at 66.53% under grouped cross-validation, against a label-permutation null with mean 33.71% and 95th percentile 46.0%. Across domains the aggregate profile fails: a forest trained on category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume (8.41 against 0.94 brands per response in gifts, 3.01 against 3.91 in category ownership), while single responses transfer at 89.92% balanced accuracy. The surface form of an answer carries the system across the query domains tested; aggregated brand behaviour does not, and the uncrossed design cannot separate the system from the domain or the harness.

20. 【2610.00170】On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries

链接:https://arxiv.org/abs/2610.00170

作者:Hyojung Han

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:stable identifier leaving, study commercial intent, commercial intent inference, user device, work began

备注: 36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): [this https URL](https://github.com/hyojunguy/ondevice-intent-evidence)

点击查看摘要

Abstract:We study commercial intent inference that runs entirely on the user's device, under three constraints frozen before the work began: the downloaded payload under 3 MiB, Tier-0 inference under 20 ms at p95, and no raw text, content embedding, or stable identifier leaving the device. Under them we build a retrieval path over a 6,020-leaf commercial taxonomy: a static embedding table distilled from a Korean sentence transformer, quantized to 4 bits, no inference runtime. Our main result is where that constraint costs accuracy. On real Korean commerce text labelled by others (22,900 AI-Hub shopping reviews), mid-category top-5 on real product names is 75.0% against an 18.4% permutation baseline, but splits on one observable: a query containing some leaf name as a substring scores 83.5%, one containing none 45.2%. A generic 196.6x larger teacher seemed to localize the gap (+20.1 pp without an anchor, +0.1 with). That null was two effects cancelling: the same teacher fine-tuned on the student's own contrastive pairs reaches 0.8586 and beats the pure-encoder student by +10.6 pp with an anchor and +20.9 pp without. The cost is not uniform, but it is not free anywhere; where the anchor is absent, task adaptation buys the teacher nothing, so what the constrained encoder lacks there is capacity. The expensive regime is detectable on-device from the ranker's own score margin: declining the least confident fifth lifts the rest to 0.8296. A second axis we first reported, a manufacturer model code, does not survive source-category fixed effects (-4.0 pp, p=0.51); the anchor does (+13.0 pp). Payload is 2,942,652 bytes, all three library links measured. Tier-0 p95 is 4.431 and 3.670 ms on two iPhones (A14, A16) and 5.080 ms on a budget Android tablet (Snapdragon 695), all slower than three server CPUs on the same code. Taxonomy supervision is mostly synthetic Korean utterances.

Comments:
36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): this https URL

Subjects:

Information Retrieval (cs.IR); Computation and Language (cs.CL)

Cite as:
arXiv:2610.00170 [cs.IR]

(or
arXiv:2610.00170v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.00170

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Hyo Jung Han [view email] [v1]
Wed, 16 Sep 2026 00:34:01 UTC (74 KB)

21. 【2610.00052】Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-49 Outputs

链接:https://arxiv.org/abs/2610.00052

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:distinct random integers, integers from 1-49, evaluate six language-model, random integers, English prompt variants

备注: 6 pages, 1 figure. Data, code, and collector at [this http URL](http://github.com/Rankfor/rankfor-open) (research/lotto-models)

点击查看摘要

Abstract:We evaluate six language-model configurations on requests for six distinct random integers from 1-49. Across 1,200 attempted calls using four English prompt variants, 1,184 responses yielded valid tickets. Effective diversity of number frequencies ranged from 9.9 to 18.0, compared with simulated fifth-percentile thresholds of 46.6-46.7 under independent uniform six-of-49 sampling at the corresponding sample sizes. Systems produced 8-93 distinct unordered tickets, and their modal tickets accounted for 22.5-68.0% of valid responses. Two archived Polish Lotto samples provided a physical-lottery comparison, with effective diversities of 41.1 and 41.4 at smaller sample sizes. These results demonstrate substantial concentration under the tested deployment settings. They do not identify its mechanism or establish performance under other prompts, temperatures, or tool configurations.

计算机视觉

1. 【2610.02210】Moore, Escher, Penrose: A Conformal Golden Braid

链接:https://arxiv.org/abs/2610.02210

作者:Sophia Feldman,Assaf Shocher

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:lithograph Print Gallery, print, dagger, Print Gallery, image

备注:

点击查看摘要

Abstract:I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved universe.'' So wrote M.C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map $z \mapsto z^\alpha$, $\alpha \in \mathbb{C}$. Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insufficient: the denoiser may "repair" the intended distortion or drift out of the prescribed geometry. We construct a generalized inverse $T^\dagger$ of the non-invertible image transformation $T$, adapted to its recursive constraint. In the idealized formulation, the Penrose identity $TT^\dagger T = T$ makes $TT^\dagger$ an idempotent projection onto geometrically admissible images. Yet denoising only the transformed image remains an out-of-distribution task, even with projection. We therefore braid denoising steps with $T$ and $T^\dagger$: source-space steps develop the untwisted scene, while transformed-space steps refine its appearance and connections in the final geometry. We generate Print Gallery-like compositions and explore further transformations. Rather than distorting a finished image, we let the scene and its distortion develop together.

2. 【2610.02208】Sphere Encoder 2

链接:https://arxiv.org/abs/2610.02208

作者:Kaiyu Yue,Sean McLeish,Ruchit Rawal,Brian Bartoldson,Menglin Jia,Tom Goldstein

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:high-dimensional latent sphere, decoding random points, Sphere Encoder, present Sphere Encoder, random points

备注: Code will be available at [this https URL](https://github.com/kaiyuyue/sphere2)

点击查看摘要

Abstract:Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{this https URL}{this http URL}.

3. 【2610.02207】One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

链接:https://arxiv.org/abs/2610.02207

作者:Ramazan Fazylov,Stamatis Lefkimmiatis,Ivan Laptev

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)

关键词:Gaussian avatars support, avatars support fast, support fast rendering, support fast, costly neural inference

备注:

点击查看摘要

Abstract:3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: this https URL

4. 【2610.02205】ROWBench: Do Video Models Render What the Program Specifies?

链接:https://arxiv.org/abs/2610.02205

作者:Zheng-Hui Huang,Guixu Lin,Yu-Ju Tsai,Jian-Kai Zhu,Fengbo Lan,Yu-Lun Liu,Yung-Yu Chuang,Kaipeng Zhang,Zhixiang Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Programmable world models, next-generation game engines, models separate executable, separate executable dynamics, world models separate

备注:

点击查看摘要

Abstract:Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.

5. 【2610.02203】Embedding Prediction Helps Image Generation

链接:https://arxiv.org/abs/2610.02203

作者:Sihan Xu,Ji Xie,Zilin Wang,Hui Shen,Stella X. Yu

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:class label, text prompt, prompt is embedded, diffusion transformers, Next-Embedding Predictive Autoregression

备注: Project page: [this https URL](https://sihanxu.me/nepa-dit)

点击查看摘要

Abstract:In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

6. 【2610.02201】SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

链接:https://arxiv.org/abs/2610.02201

作者:Tianjiao Yu,Xinzhuo Li,Yifan Shen,Ying Shen,Kiet A. Nguyen,Adheesh Sunil Juvekar,Ismini Lourentzou

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:synthesize local geometry, generation increasingly relies, predict active structure, increasingly relies, multi-stage pipelines

备注: Accepted at NeurIPS 2026. Project link: [this https URL](https://plan-lab.github.io/silsa)

点击查看摘要

Abstract:High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.

7. 【2610.02200】VISTA: A Visual Harness for Reasoning in an Interactive World

链接:https://arxiv.org/abs/2610.02200

作者:Qiushi Han,Keya Hu,Linlu Qiu,Cathy Wu,Kaiming He

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:possess strong reasoning, strong reasoning abilities, models possess strong, multimodal models possess, possess strong

备注: Tech report. An early version of this manuscript was in a blogpost published in Aug 5, 2026: [this https URL](https://vista-research.github.io/)

点击查看摘要

Abstract:We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.

8. 【2610.02197】HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

链接:https://arxiv.org/abs/2610.02197

作者:Tahira Kazimi,Shubhankar Borse,Munawar Hayat,Fatih Porikli,Pinar Yanardag

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:general-purpose world simulators, achieved remarkable visual, remarkable visual fidelity, world simulators, models have achieved

备注: Project page: [this https URL](https://hiphy-video.github.io/)

点击查看摘要

Abstract:Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.

9. 【2610.02196】InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

链接:https://arxiv.org/abs/2610.02196

作者:Zhuo Lin,Sirui Xu,Liuyu Bian,Yu-Xiong Wang,Liang-Yan Gui

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:study test-time evolution, study test-time, test-time evolution, evolution for humanoid, solving tasks

备注: Project page: [this https URL](https://sirui-xu.github.io/InterEvolve)

点击查看摘要

Abstract:We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

10. 【2610.02188】DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

链接:https://arxiv.org/abs/2610.02188

作者:Zhengming Yu,Junkun Yuan,Haotian Yang,Gordon Guocheng Qian,Yizhi Wang,Angtian Wang,Yiding Yang,Bo Liu,Xin Li,Wenping Wang,Chongyang Ma

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Distribution Matching Distillation, Distribution Matching, separately estimated target, recasts distribution matching, student evolving distribution

备注: 28 pages, 15 figures. Project page: [this https URL](https://yzmblog.github.io/projects/DMAD)

点击查看摘要

Abstract:Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at this https URL.

11. 【2610.02181】OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

链接:https://arxiv.org/abs/2610.02181

作者:Haibo Wang,Jiteng Mu,Jialu Li,Jingru Yi,Yuanjun Xiong,Jianming Zhang,Lifu Huang,Mingze Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Omni Large Language, Large Language Model, Omni Large, Large Language, transforms an Omni

备注:

点击查看摘要

Abstract:We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.

12. 【2610.02180】Generative Cinematographer: Composing Camera and Object Motion in 3D

链接:https://arxiv.org/abs/2610.02180

作者:Jiahan Zhang,Chaohao Yang,Namitha Guruprasad,Vivekjyoti Banerjee,Trong-Tung Nguyen,Alan Yuille,Anand Bhattad

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:sparse drag signals, controllable video generation, sparse drag, drag signals, motion

备注:

点击查看摘要

Abstract:Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.

13. 【2610.02162】World Observer: Joint Actor-Observer Generation for Persistent World Modeling

链接:https://arxiv.org/abs/2610.02162

作者:Hyunwook Choi,Dahyun Chung,Hyunsung Kim,Siyoon Jin,Jinhyeok Choi,Junyoung Seo,Seungryong Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:model continuously observe, continuously observe regions, world model continuously, continuously observe, actor current view

备注:

点击查看摘要

Abstract:How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor's view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. Since the observers are decoupled from the actor, they can be placed freely across the scene, extended to multiple locations for broader coverage, and driven by control signals to steer out-of-view evolution. To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while remaining competitive in visual fidelity, camera control, and 3D adherence.

14. 【2610.02160】4Director: Controlling Video World Models with Rigid 3D Geometry

链接:https://arxiv.org/abs/2610.02160

作者:Wei Cao,Hao Zhang,Vikram Voleti,Yuqun Wu,Mallikarjun B R,Shimon Vainer,Mark Boss,Yaoyao Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:professional video production, Precise control, essential for professional, Precise, video production

备注: 28 pages, 15 figures. Project page: [this https URL](https://stability-ai.github.io/4director/)

点击查看摘要

Abstract:Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

15. 【2610.02153】MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation

链接:https://arxiv.org/abs/2610.02153

作者:Yiwen Zhang,Haocheng Xi,Michael Tian-Yue Liu,Alexei A. Efros,Hadar Averbuch-Elor,Qianqian Wang,Haiwen Feng

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:finite context window, autoregressive video generation, generation is limited, Long-horizon autoregressive video, context window

备注: 27 pages. Project page: [this https URL](https://mosaichunk.github.io/)

点击查看摘要

Abstract:Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain access to such visual details, we introduce MosaiChunk, a spatio-temporal memory mechanism that composes a mosaic of selected historical key-value (KV) entries across space and time. Our approach is motivated by the observation that a frozen video generator can directly consume such non-contiguous historical KV and recover the corresponding visual content. We therefore keep the generator fixed and learn only a lightweight router that determines which historical sections to include in the mosaic under a fixed active-memory budget. We further introduce RememBench, a benchmark of long-horizon revisits with prompt-driven text-to-video (T2V) and camera-driven image-to-video (I2V) splits. Our experiments show that MosaiChunk consistently improves revisit consistency over both sliding-window inference and whole-chunk retrieval under matched memory budgets, across both T2V and I2V settings.

16. 【2610.02148】Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

链接:https://arxiv.org/abs/2610.02148

作者:Mohammed Irfan Kurpath,Jaseel Muhammad Kaithakkodan,Sahal Shaji Mullappilly,Ivan Laptev,Hisham Cholakkal

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:typically degrades text, existing omni-modal embedders, modalities typically degrades, omni-modal embedders compensate, text retrieval quality

备注: Findings of EMNLP 2026. 26 pages, 8 figures, 14 tables. Project page: [this https URL](https://omniembed.cvmbzuai.com)

点击查看摘要

Abstract:Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: this https URL

17. 【2610.02136】MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

链接:https://arxiv.org/abs/2610.02136

作者:Negin Kafee Hernashki,Soumick Chatterjee

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV); Medical Physics (physics.med-ph)

关键词:Unsupervised anomaly detection, Unsupervised anomaly, brain MRI, MRI are ranked, anomaly detection

备注:

点击查看摘要

Abstract:Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model's voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.

18. 【2610.02123】Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

链接:https://arxiv.org/abs/2610.02123

作者:Damiano Marsili,Raphi Kang,Aditya Mehta,Pietro Perona,Georgia Gkioxari

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:architectures scale model, scale model capacity, architectures scale, sparse computation, capacity through sparse

备注: Project page: [this https URL](https://glab-caltech.github.io/expertlens/)

点击查看摘要

Abstract:Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.

19. 【2610.02117】Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

链接:https://arxiv.org/abs/2610.02117

作者:Sophia Sirko-Galouchenko,Monika Wysoczanska,Andrei Bursuc,Nicolas Thome,Spyros Gidaris

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:improving language-model reasoning, frozen or EMA, EMA version, recently emerged, improving language-model

备注:

点击查看摘要

Abstract:On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: this https URL

20. 【2610.02114】Surface-volume self-supervised representation learning of brain MRI for genetic discovery

链接:https://arxiv.org/abs/2610.02114

作者:Tian Xia,Nuo Chen,Zihao Zhu,Huiwen Han,Ziqian Xie,Zhiwen Fan,Degui Zhi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Existing genome-wide association, genome-wide association studies, imaging provide predefined, brain imaging provide, Existing genome-wide

备注: 17 pages, 3 figures, 1 table, 2 supplementary tables

点击查看摘要

Abstract:Existing genome-wide association studies (GWAS) of brain imaging provide predefined or deep-learning-derived imaging phenotypes, yet these phenotypes come from either volumetric scans or cortical surface meshes, so each captures only part of the heritable variation in brain anatomy. Here we introduce MEVA (Mesh-Enhanced Volumetric Autoencoder), a self-supervised framework that encodes voxel-level image intensity together with cortical mesh geometry, including curvature and cortical thickness at each surface vertex, into one shared set of imaging features. Combining the mesh and volumetric inputs in MEVA yields modest performance gains in age and sex prediction over models that use either input alone. When these features serve as phenotypes for GWAS in the UK Biobank, they reveal more genome-wide significant loci than features learned from volumes alone or from meshes alone. These results suggest that adding cortical surface geometry to volumetric self-supervised learning captures additional heritable variation and so increases the number of loci detected.

21. 【2610.02091】GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

链接:https://arxiv.org/abs/2610.02091

作者:Yakun Zhu,Yi Bin,Yujuan Ding,Zheng Wang,Pengpeng Zeng,Duo Peng,Jingkuan Song,Heng Tao Shen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:vision-language models, progress in vision-language, geometry, spatial reasoning, CR-GEO

备注: 23 pages, 6 figures

点击查看摘要

Abstract:Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.

22. 【2610.02051】Learning from Failure: Leveraging Unreliable Predictions in Semi-Supervised Real-World Adverse Weather Removal

链接:https://arxiv.org/abs/2610.02051

作者:Cap Dang Xuan Kiet,Tat-Jen Cham

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:outdoor vision systems, image restoration aims, recover images degraded, weather-induced artifacts, vision systems

备注:

点击查看摘要

Abstract:Adverse weather image restoration aims to recover images degraded by rain, haze, snow, and other weather-induced artifacts, thereby improving the robustness of outdoor vision systems. Existing unified restoration models exhibit limited generalization to real-world scenes due to their reliance on synthetic supervision and insufficient semantic constraints. In this paper, we propose a novel student--teacher semi-supervised framework that addresses both challenges. Specifically, we introduce an unreliable database that preserves failed teacher predictions as informative negative samples for contrastive learning, while a reliable database stores high-quality teacher predictions as positive samples. By jointly exploiting reliable pseudo-ground truths and unreliable teacher outputs, the proposed framework learns to enhance desirable restoration characteristics while avoiding common failures. We further propose a phase spectrum-based semantic constraint that replaces computationally expensive text-based supervision with an efficient and naturally aligned semantic prior. An adaptive phase consistency loss is also designed to dynamically balance supervision between the degraded input and teacher pseudo-ground truths according to degradation severity. Extensive experiments on real-world benchmarks demonstrate that the proposed method consistently outperforms existing state-of-the-art approaches in restoration quality and perceptual fidelity while exhibiting stronger generalization to real-world adverse weather conditions.

23. 【2610.02045】Form and Void: Entangled Composition through an Autonomous AI Agent

链接:https://arxiv.org/abs/2610.02045

作者:Shiwen Wang,Jian Yang,Xu Wang,Xincan Wang,Weiming Dong

类目:Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)

关键词:layered semantic relationships, Positive and negative, fundamental principle, forms and layered, textbf

备注:

点击查看摘要

Abstract:Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.

24. 【2610.02044】DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization

链接:https://arxiv.org/abs/2610.02044

作者:Tao Wu,Alexandra Gomez-Villa,Senmao Li,Yaxing Wang,Joost van de Weijer,Kai Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, rectified flow-based, enabled high-fidelity, advances in rectified, asset generation

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently -- a limitation we formalize as Disentangled 3D Stylization(Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.

25. 【2610.02021】ask-Adaptive Grounded 3D-Programmers Using 2D VLMs

链接:https://arxiv.org/abs/2610.02021

作者:Arman Raayatsanati,Sombit Dey,Anna-Maria Halacheva,Jan-Nico Zaech,Luc Van Gool,Danda Pani Paudel

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Recent vision-language models, exhibit remarkable generalization, Recent vision-language, Canonical Coordinate Framing, training diversity

备注: 18 pages, 9 figures, 11 tables

点击查看摘要

Abstract:Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.

Comments:
18 pages, 9 figures, 11 tables

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.02021 [cs.CV]

(or
arXiv:2610.02021v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.02021

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
26. 【2610.02019】Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

链接:https://arxiv.org/abs/2610.02019

作者:Guangyu Yang,Jingbiao Mei,Mingsheng Sun,Jinghong Chen,Yingtong Bu,Pengda Qin,Da Chen,Bill Byrne

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:video-based social media, increased users' exposure, video safety detection, automated video safety, safety detection

备注:

点击查看摘要

Abstract:The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at this https URL .

27. 【2610.02010】Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking

链接:https://arxiv.org/abs/2610.02010

作者:Kirill Aistov,Khaled Abud,Irina Serzhenko,Egor Kovalev,Aleksey Yakushev,Aleksandr Akimenkov,Dmitry Obydenkov,Yury Markin,Sergey Lavrushkin,Dmitriy Vatolin,Anastasia Antsiferova

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)

关键词:open security question, Latent Frequency Masking, tracing AI-generated images, adaptive removal attacks, removal attacks remains

备注: This work has been accepted for publication at IEEE ICDM 2026 conference. The final published version will be available via IEEE Xplore

点击查看摘要

Abstract:Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion regeneration for improved image preservation. We provide a theoretical distortion bound relating the change between the reconstructed adversarial image and the masked latent-frequency perturbation. We evaluate the proposed attack against six diffusion watermarking methods on images generated from DiffusionDB and MS-COCO prompts. Latent Frequency Masking removes or substantially weakens several watermarks while preserving perceptual quality and achieving favorable runtime compared with existing attacks. These results identify latent-frequency manipulation as a practical attack surface and highlight the need to include such attacks in robustness evaluations of generative image watermarking.

28. 【2610.02000】Weather-Aware Domain Adaptation for Street-View Weather Recognition

链接:https://arxiv.org/abs/2610.02000

作者:Hossein Maghsoumi,George Atia,Yaser P. Fallah

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:dust remain challenging, Adverse conditions, Discriminative Domain Adaptation, dust remain, Adversarial Discriminative Domain

备注: 7 pages, 3 figures, 4 tables. Published in the 2026 IEEE Conference on Technologies for Sustainability (SusTech)

点击查看摘要

Abstract:Adverse conditions such as rain, snow, fog, and dust remain challenging for camera-based perception in autonomous driving. We study multi-class weather recognition from street-view images under domain shift, where most available training data come from non-street-view sources that differ markedly from real driving scenes. We propose Weather-Aware Adversarial Discriminative Domain Adaptation (WA-ADDA), which conditions the domain discriminator on predicted weather to promote features that are both domain-invariant and weather-sensitive. We also assemble a multi-dataset benchmark by unifying diverse non-street-view weather collections as sources and real street-view images as targets, and define a standardized evaluation protocol with macro accuracy as the primary metric. Across backbones (ResNet-50, EfficientNet, VGG, DenseNet), WA-ADDA consistently improves street-view performance and yields strong per-class recalls in challenging conditions while preserving clear-weather accuracy. These findings highlight the feasibility of domain-adapted weather recognition and the value of our benchmark for advancing robust, on-board perception.

29. 【2610.01999】From Reasoning Failures to Composable Video Spatial Intelligence

链接:https://arxiv.org/abs/2610.01999

作者:Pengzhan Sun,Junbin Xiao,Ramanathan Rajaraman,Shiu-hong Kao,Angela Yao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:underlying capabilities account, evaluate vision-language models, vision-language models, models across diverse, account for success

备注:

点击查看摘要

Abstract:Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9\% to 60.2\% on ReVSI and improves the SpatialClaw result from 62.8\% to 66.3\% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.

30. 【2610.01994】Comparing a gradient boosting algorithm to the GOES FDC for wildfire detection

链接:https://arxiv.org/abs/2610.01994

作者:Asaf Vanunu,Boaz Nadler,Arnon Karnieli

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:pose severe risks, Wildfires pose severe, human life, pose severe, severe risks

备注:

点击查看摘要

Abstract:Wildfires pose severe risks to human life, ecosystems, and property. This study presents a machine learning approach for wildfire detection from GOES ABI imagery. A CatBoost model was trained on a large dataset with thousands of ABI images and over 300,000 matching VIIRS fire detections. An evaluation on a separate dataset across five regions showed that the learned CatBoost model outperformed the operational GOES Fire Detection and Characterization (FDC) product. It achieved higher precision, recall, and F1 scores both within and outside the training area. The CatBoost model achieved F1 scores that were 0.16 to 0.38 higher than the GOES FDC in all regions. In addition, out of 51 historical fire events, the CatBoost detected 26 fires before both VIIRS and GOES FDC, compared to only six earlier detections by the GOES FDC. Importantly, the CatBoost model achieved accurate wildfire detection also during nighttime, whereas the GOES FDC obtained very low recall values, around 0.03. This study demonstrates that machine learning models may offer significant improvements over existing geostationary fire products, including higher accuracy, fewer false alarms, and earlier detection.

31. 【2610.01989】Continual Concept Erasure in Diffusion Models by Suppressing Cross-Edit Interference

链接:https://arxiv.org/abs/2610.01989

作者:Yongliang Wu,Haori Lu,Jinqi Luo,Wei Cao,Xingyu Zhu,Yaoyao Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:support content governance, Concept erasure removes, erasure removes copyright-protected, undesirable concepts, governance and compliance

备注: 24 pages. Project page: [this https URL](https://continual-erasure.cvmlgroup.web.illinois.edu/)

点击查看摘要

Abstract:Concept erasure removes copyright-protected, privacy-sensitive, or otherwise undesirable concepts from pretrained text-to-image diffusion models to support content governance and compliance. As erasure requests arrive over time, models must remove new targets without undoing prior erasures. Existing methods do not constrain interference across edits: residual perturbations outside the retain set interact and accumulate, degrading unrelated generations and sometimes collapsing previously erased targets into noise. We propose CEASE (Continual Erasure via Adaptive Subspace Editing), a training-free method that imposes two subspace constraints on a closed-form solver. CEASE adds the token representation of the shared replacement to the solver's invariance matrix and, when interference is detected, projects the current update onto the orthogonal complement of dominant output directions extracted from cumulative past updates. A closed-form decomposition attributes the accumulated interference to repeated activation of the shared replacement and overlap between successive update directions, showing that the two constraints suppress these respective sources. Across continual erasure of celebrities, artistic styles, and instances, CEASE achieves the most consistent erase-preserve trade-off, while existing methods either degrade general generation or insufficiently erase targets.

32. 【2610.01973】oken-Level Video Reinforcement Learning

链接:https://arxiv.org/abs/2610.01973

作者:Yifan Wang,Gordon Guocheng Qian,Yanyu Li,Anil Kag,Yun Fu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:entire sampled video, Reinforcement learning, Video Reinforcement Learning, generation usually assigns, entire sampled

备注:

点击查看摘要

Abstract:Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.

33. 【2610.01969】RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models

链接:https://arxiv.org/abs/2610.01969

作者:Yongliang Wu,Haori Lu,Yulun Wu,Jinqi Luo,Xingyu Zhu,Yaoyao Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:copyrighted style, recognizable character, preserving its ability, ability to generate, activation steering

备注: 20 pages. Project page: [this https URL](https://rasteer.cvmlgroup.web.illinois.edu/)

点击查看摘要

Abstract:Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-image diffusion model while preserving its ability to generate other content. Existing activation steering methods build an erasure direction mainly from the target concept and adjust model activations along it at inference time. However, target and retained concepts often overlap in the model's representation space, so this direction also contains shared components that retained concepts rely on. Steering directly along this direction can therefore suppress retained concepts and harm the generation of non-target content. To address this issue, we propose Retain-aware Activation Steering (RASteer), a training-free method. RASteer first builds a retain subspace from the concepts to preserve. Retain-Orthogonal Steering (ROS) then removes components aligned with this subspace from the erasure direction, making steering more specific to the target. Since fully removing the shared components can weaken erasure, we further introduce Overlap-Adaptive Calibration (OAC). At each layer and denoising step, OAC uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing target erasure and concept preservation. Experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show that RASteer matches or outperforms the activation steering and weight editing baselines we evaluate, achieving a better balance between erasure and preservation.

34. 【2610.01962】SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning

链接:https://arxiv.org/abs/2610.01962

作者:Si Qi Goh,Cap Dang Xuan Kiet,Tat-Jen Cham,Kwok-Yan Lam

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:associate visual identities, biographical information creates, personally identifiable information, ability of vision-language, identities with biographical

备注:

点击查看摘要

Abstract:The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.

35. 【2610.01956】EndoLive: Real-Time Style Transfer for Endoscopic Endonasal Skull Base Surgical Video

链接:https://arxiv.org/abs/2610.01956

作者:Griffin Hurt,Calvin Brinkman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:skull base surgery, endonasal skull base, requires significant practice, endoscopic endonasal skull, base surgery

备注:

点击查看摘要

Abstract:Complex surgical procedures around critical anatomy, such as the endoscopic endonasal skull base surgery, requires significant practice and training on the part of the surgeon before they are allowed to perform the operation on a live patient. This training in typically done in cadaveric specimens, due to them containing the same critical structures as a living human. However, cadavers are not a perfect 1-to-1 substitute for a living patient. The dead and preserved tissues of a cadaver are colored completely differently than a living human, and -- without complex and expensive pumping systems -- do not bleed in the same way. As a result, identifying the critical pieces of anatomy that make this procedure so complex can be quite different in a live case than in a surgeon's cadaveric practice. This paper presents EndoLive, a framework for real-time style transfer between cadaveric endoscopic video and living human endoscopic video. Our method combines the ConStructS GAN model for realistic style transfer for surgical applications, with the HyPER-GAN model that can learn complex translations and perform them in real-time. We train EndoLive on unpaired cadaveric and live images taken from an endoscope, and test the trained model with cadaveric video, on a variety of devices. Experimental results demonstrate that EndoLive can perform cadaveric-to-live translation at speeds well above the minimum necessary for real-time, while maintaining semantic consistency of critical anatomical structures. Our source code is available at this https URL.

36. 【2610.01944】Anti-Persona: Disrupting Unauthorized Identity Binding and Recognition in Personalized Vision--Language Models

链接:https://arxiv.org/abs/2610.01944

作者:Abhishek Basu,Fahad Shamshad,Karthik Nandakumar

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Few-shot personalization enables, enables large vision, personalization enables large, learn user-specific visual, user-specific visual concepts

备注: Code available at [this https URL](https://github.com/iabh1shekbasu/anti-persona)

点击查看摘要

Abstract:Few-shot personalization enables large vision--language models (LVLMs) to learn user-specific visual concepts for applications such as personalized retrieval and subject-aware querying. However, it also creates a privacy risk: an adversary can bind a target identity from a few reference images and subsequently detect that identity in new images through natural-language queries. We introduce Anti-Persona, an image-level defense against unauthorized identity binding and recognition in personalized LVLMs. Our key insight is that identity personalization relies on visual features shared across multiple reference images. We aggregate these features into an identity prototype and optimize visually subtle perturbations that disrupt prototype alignment in the vision-encoder space. Spatial smoothing and low-frequency preservation further promote visual fidelity and practical resilience to image compression. The resulting protection does not depend on a specific prompt and supports both proactive anti-personalization and reactive image protection. Experiments on two representative personalized LVLMs demonstrate protection rates of up to $95.0\%$ while preserving visual fidelity. The method remains stable across prompt variations and evaluated identity-query tasks, and improves black-box transfer under encoder mismatch.

37. 【2610.01942】Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

链接:https://arxiv.org/abs/2610.01942

作者:Efstathios Karypidis,Spyros Gidaris,Nikos Komodakis

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision Foundation Models, future scene understanding, world modeling, fundamental capability, capability for world

备注:

点击查看摘要

Abstract:Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of Vision Foundation Models (VFMs) yields semantically rich representations that support diverse future scene understanding tasks. However, existing approaches rely on two-stage pipelines, where VFM features are first compressed using fixed dimensionality reduction (e.g., PCA) or independently trained autoencoders, and a separate predictor is trained on top of the resulting frozen latent space. This decoupling between representation learning and temporal prediction, as well as approaches that apply predictors directly on raw VFM features, provides no guarantee that the latent space is structured for predictable dynamics. In this work, we propose Latent-Foresight, an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly shaping the representation to support temporal predictability. To enable stable joint optimization, we introduce several key design choices that prevent latent collapse and align reconstruction with generative objectives. Extensive experiments show that our approach learns more temporally coherent latent representations and consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons, while eliminating separate training stages, including during high-resolution adaptation. We provide the implementation code and model weights at this https URL

38. 【2610.01939】Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

链接:https://arxiv.org/abs/2610.01939

作者:Ruiyang Si,Jianxin Bi,Shunyu Yang,Rui Ni,Wenbo Huang,Qiang Wang,Shulong Jiang,Duomin Wang,Xiuyu Li,Haiwen Feng,Zhen Dong,Daquan Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision language model, repeated model invocations, redundant observations incur, observations incur substantial, Vision language

备注:

点击查看摘要

Abstract:Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.

39. 【2610.01927】CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction

链接:https://arxiv.org/abs/2610.01927

作者:Moyang Li,Zihan Zhu,Wei Zhang,Marc Pollefeys,Daniel Barath

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:recently shown remarkable, Feedforward foundation models, Feedforward foundation, shown remarkable, recently shown

备注: Authors contributed equally to this work. Author order is interchangeable

点击查看摘要

Abstract:Feedforward foundation models have recently shown remarkable 3D reconstruction capabilities. However, existing models exhibit large tracking drift in long-context streaming reconstruction due to error accumulation. In this paper, we revisit loop closure with streaming reconstruction foundation models to enable accurate, drift-free, kilometer-scale reconstruction. Specifically, our method detects loop candidates through global descriptor retrieval, and constructs loop-conditioned windows to estimate the relative poses between looped frames. Given the observation that our adopted streaming reconstruction backbone produces a globally consistent scale, we optimize all frame poses on the SE(3) manifold with sequential and loop closure constraints, avoiding the pose graph optimization on the Sim(3) or higher-dimensional SL(4) manifolds employed in prior works. Extensive experiments show that our method reduces drift and produces consistent geometry on kilometer-scale sequences, significantly outperforming the state of the art. Code is available at this https URL.

40. 【2610.01917】MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

链接:https://arxiv.org/abs/2610.01917

作者:Yingcheng Liu,Tianyi Jiang,Yujuan Ding,jiangbo Ai,Xun Jiang,Guoqing Wang,Wei Ye,Yi Bin

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:repeated image operations, Latent, visual, reasoning equips vision, Latent visual

备注:

点击查看摘要

Abstract:Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.

41. 【2610.01914】DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction

链接:https://arxiv.org/abs/2610.01914

作者:Junfeng Ni,Zirui Zhou,Yixin Chen,Yu Liu,Nan Jiang,Zhifei Yang,Song-Chun Zhu,Siyuan Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:heavy occlusions, aims to reconstruct, level of quality, quality under heavy, scene reconstruction aims

备注: SIGGRAPH Asia 2026 - Journal Track (TOG). Project page: [this https URL](https://decomvoxel.github.io/DecomVoxel-Webpage/)

点击查看摘要

Abstract:Decompositional scene reconstruction aims to reconstruct high-quality objects and background, yet existing methods still struggle with the level of quality under heavy occlusions. While generative priors offer a potential solution, 2D image-based priors often suffer from multi-view inconsistency due to a lack of 3D awareness. Conversely, 3D-native priors provide stronger structural inductive biases but frequently lead to spatial drift and misalignment within complex scenes. To address these issues, we propose DecomVoxel, formulating object completion as a guided in-situ denoising optimization that bridges 3D-native priors with neural scene reconstruction. Our framework introduces a reformulated epsilon-based distillation loss to ensure stable latent refinement, alongside adaptive spatial guidance that utilizes occupied and vacant anchors with temporal annealing to suppress generative hallucinations and mitigate spatial drift. Experiments on Replica and ScanNet++ show that DecomVoxel significantly outperforms state-of-the-art methods while faithfully preserving the original spatial layout, structural fidelity, and style-consistent texture. Our method pushes the boundary of decompositional reconstruction by delivering high-quality textured meshes with clean topology, geometry, and appearance, providing a robust solution for the decompositional reconstruction of complex real-world scenes. Code is available at this https URL.

42. 【2610.01905】MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens

链接:https://arxiv.org/abs/2610.01905

作者:Shen Zheng,Anurag Ghosh,Mani Ramanagopal,Srinivasa Narasimhan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scaling safe autonomous, safe autonomous driving, Online vectorized, map, requires accurate

备注:

点击查看摘要

Abstract:Online vectorized HD map construction is essential for scaling safe autonomous driving and requires accurate, real-time inference. Prior methods typically rely on dense bird's-eye-view (BEV) grids as the intermediate representation. We propose \textit{MapLightning}, which replaces the dense BEV grid with a compact set of 1D learnable map tokens. To construct map tokens from image features, we choose self-attention over vanilla cross-attention because it enables joint interactions and contextual aggregation among image and map tokens. Our transformer-based mapper concatenates map and image tokens, applies full self-attention, discards the image tokens, and retains the updated map tokens for decoding. This design offers three advantages. First, our representation is efficient, using fewer tokens, consuming less memory, and running faster. Second, the lightweight design allows the map decoder to use full rather than deformable cross-attention for better global context. Third, unlike BEV-based methods, our network does not use camera projection parameters, making it robust to camera-extrinsic perturbations. MapLightning uses up to 16.7$\times$ fewer intermediate tokens than dense BEV-based methods and achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse~2. Its lightweight variant surpasses MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse~2, while delivering 1.73$\times$ faster inference (40+ FPS) with 53\% less memory. We further show improvements on uncertainty-aware map construction and downstream trajectory prediction. Code and models will be released.

43. 【2610.01890】Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching

链接:https://arxiv.org/abs/2610.01890

作者:Victor Enescu,Assaad Zeghina,Matthieu Meignin,Nicolas Viltard,Cécile Mallet

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Deep generative networks, sophisticated textual prompts, recently achieved unprecedented, achieved unprecedented performance, Deep generative

备注:

点击查看摘要

Abstract:Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the rare occasions they do. In this paper, we investigate the potential of flow matching models for unsupervised domain adaptation of satellite radiometer images. Our main contribution is a novel unsupervised method that achieves precise domain alignment by leveraging parts of the deterministic ordinary differential equations in flow matching models, conditioned on different satellite instruments. A key strength of our approach is its ability to preserve essential information while adapting across any domains since the perturbations are in theory bijective. Extensive experiments conducted on the GPM-Core constellation show the benefit of our conditional domain adaptation, particularly in improving rain precipitation estimation from radiometer imagery.

44. 【2610.01884】Memory-Guided B-Roll Generation from User Video Collections

链接:https://arxiv.org/abs/2610.01884

作者:Cusuh Ham,Fabian Caba Heilbron,Josef Sivic,Bryan Russell

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:collection-grounded B-roll sequence, B-roll sequence generation, collection-grounded B-roll, introduce an approach, approach for collection-grounded

备注: Project page at [this https URL](https://cusuh.github.io/MemComposer)

点击查看摘要

Abstract:We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complements the user's primary footage (A-roll) while preserving the collection's characters, settings, objects, and style. This task is challenging as one must choose the visual evidence from hours of captured footage that should guide the generation of each shot in the sequence. We address this challenge with MemComposer, a three-stage system that turns raw footage into a structured memory with visual references (characters, settings, objects, and style) and uses it to plan, retrieve, and generate grounded B-roll sequences. First, in a one-time offline stage, MemComposer constructs an entity-centric memory from raw video. Second, it uses the memory and user directive to plan a grounded sequence and retrieve conditioning frames for each shot. Third, it iteratively generates and critiques the sequence to enforce identity, setting, and sequence-level consistency. We evaluate MemComposer in a user preference study along two dimensions: prompt adherence and visual alignment to the user's collection. Against an ungrounded text-to-video planner, MemComposer wins 60.0\% of prompt-adherence and 92.8\% of visual-alignment comparisons, showing the grounding benefit of collection memory and reference retrieval. Against retrieval-only sequences assembled from captured footage, MemComposer wins 94.5\% of prompt-adherence comparisons, showing the value of generating missing shots, while retrieval-only sequences are preferred for visual alignment in 58.2\% of comparisons.

45. 【2610.01876】EvenSplat: Coupled 2D-3D Decomposition for Gaussian Splatting under Exposure and Illumination Variation

链接:https://arxiv.org/abs/2610.01876

作者:Tongyu Wu,Jacob Edwards,Ziteng Cui,Caigui Jiang,Cheng Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:surface photographed, Gaussian Splatting treat, locally strong light, strong light sources, light sources leave

备注:

点击查看摘要

Abstract:A surface photographed under even light presents nearly the same appearance from every angle; the same surface under uneven light does not. Exposure changes between views, illumination varies within a single image, and locally strong light sources leave one region bright and its neighbor in shadow. Multi-view reconstruction methods such as 3D Gaussian Splatting treat these lighting artifacts as if they were properties of the scene, entangling capture-specific illumination with the geometry and color they recover. We present EvenSplat, a framework that separates the two. EvenSplat couples an image-space illumination decomposition with an illumination field carried by the Gaussians, so that the same explanation of the lighting is shared between the two-dimensional and three-dimensional views of the scene; a camera-response network and a local exposure-compensation module absorb the global and residual differences that remain across training images. Through extensive experiments across multiple datasets and diverse forms of uneven illumination (cross-view exposure, spatial illumination variation, and high-contrast lighting) on both real-world captured and simulated benchmarks, EvenSplat generally outperforms state-of-the-art methods, particularly under high-contrast illumination.

46. 【2610.01870】From Pixels to Policy: A Multi-Agent System for Intervention and Geo-Spatial Decision Support

链接:https://arxiv.org/abs/2610.01870

作者:Hosam Elgendy,Utkarsh Mall

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:interventions remains costly, evaluating proposed interventions, proposed interventions remains, implications for health, remains costly

备注:

点击查看摘要

Abstract:Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often impractical. Existing geospatial vision methods largely focus on monitoring urban indicators from aerial and street-view imagery, rather than proposing interventions and estimating their effects on such indicators. Moving beyond recognition, we introduce the problem of discovering interventions that improve target indicators for a given aerial or street-view image. We argue that a black-box indicator model, combined with a generative editing model, can serve as an implicit digital twin for testing intervention hypotheses. We present VIDA-Geo , a multi-agent system that explores this intervention space by coordinating segmentation, diffusion-based inpainting, and indicator scoring models to produce interventions that are both perceptually realistic and aligned with real-world policies. We evaluate our system on 8 indicators across aerial and street-view imagery, measuring changes in factors such as perceived safety and greenery. Our approach outperforms existing baselines in many cases, achieving up to 2X higher perceptual quality and policy alignment scores. Finally, our model provides users with multiple candidate interventions, supporting an expert city-planner-in-the-loop workflow.

47. 【2610.01863】LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction

链接:https://arxiv.org/abs/2610.01863

作者:Zhening Huang,Yueyan Li,Johnathan Chiu,Xiaoyang Lyu,Matt Zhou,Yuxin Yao,Joan Lasenby,Shangzhe Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO)

关键词:reconstructing real indoor, real indoor environments, scenes from RGB-D, RGB-D scans, reconstructing real

备注: Code: [this https URL](https://github.com/LiteReality/LiteReality-Agent/) Webpage: [this https URL](https://litereality.github.io/agent/)

点击查看摘要

Abstract:We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes from RGB-D scans. At its core, LiteReality-Agent formulates 3D reconstruction as a coding problem, in which a coding agent gathers evidence using specialised tools and iteratively edits a Python script, this http URL, which can be executed to produce a 3D digital twin of the room. With this formulation, we develop a robust observe-edit-verify harness that supports evidence gathering, measurement, verification, layout optimisation, simulation readiness, and quality control throughout the reconstruction process. LiteReality-Agent produces high-quality reconstructions suitable for simulation and downstream embodied AI tasks. Furthermore, as agent capabilities continue to improve rapidly, the system introduced by LiteReality-Agent remains a strong orchestration framework for future agents: it equips them with specialised tools, structured workflows, and robust verification mechanisms that substantially improve reconstruction quality and reliability. We demonstrate that LiteReality-Agent produces reconstructions that are more geometrically accurate, visually realistic, and simulation-compatible than those generated by recent frontier models, such as Astra and Fable. We therefore view LiteReality-Agent as a practical and important building block for robust real-to-sim systems. Both the source code and the data-capture application are publicly available. Code:this https URL

48. 【2610.01807】PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization

链接:https://arxiv.org/abs/2610.01807

作者:Ahmed Sharshar,Asif Hanif,Naveen Kumar Kummari,Mohammad Yaqub,Mohsen Guizan

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Reliable clinical deployment, Reliable clinical, deep medical image, shifts across scanners, acquisition protocols

备注: The paper is accepted in MICCAI 2026

点击查看摘要

Abstract:Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existing domain generalization (DG) methods often focus on style or intensity diversification, but they can still leave networks dependent on domain-specific texture correlations. Inspired by evidence that Fourier phase encodes semantic structure, we introduce PhaseAT, a phase-aware adversarial training framework for medical DG. PhaseAT forms phase-perturbed training views in the Fourier domain by iteratively updating a bounded phase perturbation while keeping the amplitude spectrum unchanged, thereby stressing spatial organization under matched appearance statistics. Perturbations are applied only to the luminance channel in YCbCr color space to avoid chromatic artifacts. Additionally, a simple phase-saliency mask concentrates updates on the most influential frequencies. The model is trained with a weighted combination of losses on clean and phase-perturbed samples, supporting both single-source and multi-source DG. We validate our method on two challenging medical datasets and demonstrate that PhaseAT achieves over 20% improvement in single-source domain generalization, outperforming several state-of-the-art DG methods. The code implementation is available at: this https URL.

49. 【2610.01794】Continuous Conditioning of VLAs with Augmenting EMG and Visual Task Descriptors

链接:https://arxiv.org/abs/2610.01794

作者:Edward W. Staley,Connor O. Pyles,Rahul Hingorani,Frank Camargo,Griffin Milsap,Jared Markowitz,Matthew S. Fifer,Michael Wolmetz

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:describing task information, models rely strongly, rely strongly, VLA, task information

备注: Presented at IROS WORLDS Workshop 2026. Four main pages double-column format plus references and appendices

点击查看摘要

Abstract:Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that other modalities in the state space may present opportunities for supplemental task conditioning, which may be particularly relevant in cluttered or otherwise ambiguous scenes. We introduce two tuned models to test this hypothesis: (1) an electrophysiology-conditioned VLA (EC-VLA) that incorporates 8-channel electromyography envelopes as continuous conditioning input concatenated to the proprioceptive vector, and (2) a visually-annotated VLA (VA-VLA) that incorporates visual segmentation annotations to the image inputs. On a cube-selection task evaluated across three participants, EC-VLA matches a language-prompted baseline in uncluttered, in-distribution conditions and substantially outperforms it in cluttered, out-of-distribution scenes. Similarly, VA-VLA shows modest improvements over a language-prompted baseline in in-distribution scenes with substantial improvement in cluttered, out-of-distribution trials. Together, these results provide strong evidence for the potential benefit of task-conditioning beyond language.

50. 【2610.01785】VETO: Video Efficient Token Optimization for Vision Language Models

链接:https://arxiv.org/abs/2610.01785

作者:Gueter Josmy Faure,Hao Ping Wang,Min-Hung Chen,Winston H. Hsu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Processing long videos, Processing long, making long-form inference, inference prohibitively expensive, Vision-Language Models

备注:

点击查看摘要

Abstract:Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.

51. 【2610.01778】GIFTBench: Diagnosing Generalization in Image Forgery Localization and Informing Model Design

链接:https://arxiv.org/abs/2610.01778

作者:Baoke Dou,Ziye Wang,Hao Wang,Guoqing Cai,Wende Tan,Chenyang Si,Liucheng Guo,Yueming Lyu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires assessing models, entangle multiple factors, cover limited manipulation, limited manipulation conditions, Reliable evaluation

备注:

点击查看摘要

Abstract:Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cover limited manipulation conditions or entangle multiple factors in cross-dataset evaluation. Consequently, aggregate performance provides an incomplete view of localization generalization. We introduce GIFTBench, a multi-axis benchmark of 115,013 manipulated images with pixel-level annotations spanning manipulation source, semantic target, editing operation, and composition complexity. GIFTBench supports axis-specific transfer analysis and evaluation on twelve external datasets. Its diagnostic studies reveal asymmetric cross-source transfer, recall-dominated failures, and heterogeneous degradation across semantic, operational, and compositional changes. Beyond diagnosis, the scale and diversity of GIFTBench provide a substantially broader training distribution than conventional IFL datasets. Training representative localizers on GIFTBench consistently improves their aggregate transfer to external datasets, showing that the benchmark serves not only as an evaluation tool but also as an effective training resource for cross-domain localization. Guided by the diagnostic findings, we further develop ForenScope, a detection and localization framework combining classification-adapted representations with multi-depth, multi-scale spatial features, learned layer fusion, and selective coarse-scale conditioning. Experiments show improved cross-dataset localization while retaining image-level detection capability. The GIFTBench dataset showcase page is available at this https URL.

52. 【2610.01766】VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding

链接:https://arxiv.org/abs/2610.01766

作者:Bingjun Luo,Yuhuan Fan,Jialin Guo,Siqi Li

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Video temporal grounding, aims to localize, localize events, Video temporal, temporal grounding aims

备注:

点击查看摘要

Abstract:Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically evolves agent harnesses for video temporal grounding. VideoEvolve uses a Cloze-Structured Harness Representation that preserves stage interfaces while leaving agent workflows and instructions open to evolution. Branch-Guided Harness Evolution preserves promising code branches for continued refinement, using execution feedback to guide local edits and validation to determine which improvements are carried forward. Experiments demonstrate improved grounding performance across multiple benchmarks. Component analyses identify instruction refinement as a consistent source of gains, while the benefits of evolved code vary across evaluation settings. Together, these results support automated harness evolution as an effective approach to improving video temporal grounding. Code is available at this https URL .

53. 【2610.01762】OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

链接:https://arxiv.org/abs/2610.01762

作者:Xiangyu Zeng,Yuandong Yang,Zhiqiu Zhang,Yuhan Zhu,Xinhao Li,Qingyi Si,Dingyu Yao,Changlian Ma,Haoran Chen,Xinyu Chen,Yansong Shi,Junhao Zhou,Yifei Li,Jun Zhang,Chuanyu Qin,Chenxu Yang,Xinlei Yu,Kun Ouyang,Yuchen Shao,Qianshan Wei,Changhai Zhou,Jun Gao,Jiaqi Wang,Limin Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:LLMs must retain, relevance to future, respond when sufficient, Streaming, Streaming video

备注: 29 pages, 12 figures, 20 tables. Project page: [this https URL](https://mcg-nju.github.io/OneStreamer)

点击查看摘要

Abstract:Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.

54. 【2610.01759】PhysDEM: Physics-Defined Energy-Matching Diffusion for Spatiotemporal Field Generation under Scarce Measurements

链接:https://arxiv.org/abs/2610.01759

作者:Zhenyu Liang,Yining Huang,Yubo Zhao,Jack C.P. Cheng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Generating and predicting, predicting spatiotemporal physical, insufficient to characterize, characterize a distribution, distribution over complete

备注:

点击查看摘要

Abstract:Generating and predicting spatiotemporal physical fields from scarce measurements is challenging, as observations are insufficient to characterize a distribution over complete fields. This limits conventional data-driven diffusion models that rely on full-field datasets. We introduce PhysDEM, a physics-defined diffusion framework that combines governing equations with spatially sparse observations to generate multiple plausible fields. First, we construct a Gibbs target by reweighting a measurement-conditioned Gaussian reference with PDE residual energy. Second, we derive an exact conditional-mean identity that reduces denoising to supervised learning of the standardized energy-induced mean correction. Third, a physics-displacement probability flow cancels Gaussian reference terms and enables amortized sampling with changing measurements through Gaussian conditioning, without retraining. Experiments on synthetic PDE systems and real-world-informed applications demonstrate that PhysDEM supports coherent field recovery and efficient sampling while maintaining stable diagnostics under tested noise levels, illustrating its practical value for field assessment. To our knowledge, PhysDEM is the first physics-defined diffusion model enabling amortized spatiotemporal field inference without preassembled full-field datasets.

55. 【2610.01758】GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

链接:https://arxiv.org/abs/2610.01758

作者:Jian Liu,Wei Sun,Zhenqi Dai,Hui Yang,Jian Xiao,Nicu Sebe,Na Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Category-level object pose, intra-class unknown objects, Category-level object, capable of generalizing, generalizing to intra-class

备注: Accepted by NeurIPS'26

点击查看摘要

Abstract:Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. Code and demos are released at this https URL.

56. 【2610.01754】Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

链接:https://arxiv.org/abs/2610.01754

作者:Mohd Ubaid Wani,Sara Atito,Josef Kittler,Muhammad Awais

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:temporally localize abnormal, localize abnormal events, aims to temporally, temporally localize, localize abnormal

备注: Published in Transactions on Machine Learning Research (TMLR), 2026. 39 pages

点击查看摘要

Abstract:Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.

57. 【2610.01750】FFBL-Coop: Association-Decoupled Cooperative 3D Multi-Object Tracking

链接:https://arxiv.org/abs/2610.01750

作者:Haoxin Wu,Xiaokai Bai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:maintaining consistent identities, integrate complementary observations, tracking must integrate, consistent identities, integrate complementary

备注: 9 pages (main content), 21 pages total including references and appendix; 11 figures; under review as a conference paper at ICLR 2027

点击查看摘要

Abstract:Cooperative 3D tracking must integrate complementary observations across agents and time while maintaining consistent identities. When evidence integration and identity inheritance share a matching decision, errors arising from cross-view appearance differences and spatial misalignment can compromise both feature fusion and track continuity. We propose FFBL-Coop, a fuse first, bind later framework that separates instance admission from identity management. Confidence-ranked Slot Admission (CSA) allocates cooperative queries to available ego slots using confidence and spatial proximity. Unified Representation Aggregation (URA) uses cooperative semantic features and aligned anchors to guide ego-feature retrieval, refining the augmented query bank within a shared transformer decoder. After refinement, Cooperative-Priority Identity Anchoring (CPIA) combines learned association with persistent mappings to establish accepted identity assignments across frames. A shared codebook reduces transmitted payload while retaining AP and AMOTA close to the uncompressed variant. FFBL-Coop achieves AMOTA/AP of 0.611/0.548 on V2X-Seq and 0.688/0.653 on Griffin-25M. Code will be released.

58. 【2610.01746】End-to-End Learning vs. Modular Architectures: Comparative Insights into Autonomous Driving Systems

链接:https://arxiv.org/abs/2610.01746

作者:Kartik B. Kapse

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:intelligent transportation research, prominent design paradigms, Autonomous driving systems, Modular Architectures, Autonomous driving

备注: 27 pages, 7 figures, 4 tables

点击查看摘要

Abstract:Autonomous driving systems have become a central focus of intelligent transportation research, with End-to-End Learning and Modular Architectures offering two prominent design paradigms for their implementation. E2E Learning uses deep learning algorithms to map raw sensory inputs directly to driving actuators, providing a streamlined and adaptable solution. while Modular Architectures employ a pipeline-based approach, dividing the system into distinct subsystems for perception, cognition, planning, and control. This paper presents a comprehensive comparative analysis of these paradigms, focusing on their strengths, limitations, and trade-offs to provide insights into their suitability for various autonomous driving applications. The study evaluates key factors such as interpretability, scalability, robustness, and real-world applicability. While End-to-End Learning emphasizes simplicity and adaptability in dynamic environments, it lacks transparency and is highly dependent on large datasets. Conversely, Modular Architectures offer superior interpretability and task-specific optimization, but face challenges related to integration complexity and scalability. To address these limitations, hybrid approaches that combine the strengths of both paradigms have emerged, offering a promising direction for overcoming these challenges. Beyond this comparative synthesis, following work proposes a Four-Dimensional Architecture Selection Framework, comprising twelve binary criteria across safety, operating environment, data/computational resources, and deployment context, and validate it against ten published autonomous driving systems, correctly recommending 7/10 deployed architectures. This work synthesizes existing literature to highlight key trade-offs between the paradigms and identifies hybrid architectures as a promising direction for future research.

59. 【2610.01744】3DROID: A Renderable 3D Gaussian Dataset with Measured Per-Scene Reliability

链接:https://arxiv.org/abs/2610.01744

作者:Wonguen Cho,Junhoo Lee,Nojun Kwak

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:models primarily reason, manipulation models primarily, physical world, observations while acting, models primarily

备注: 12 pages, 3 figures

点击查看摘要

Abstract:Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented robot data with geometric priors such as depth, point clouds, and 3D trajectories, while renderable 3D Gaussian representations provide another promising form of 3D supervision. However, 3DGS representation is designed mainly for photometric fidelity and may not preserve real-world metric scale, particularly when the supplied camera extrinsics are unreliable. We study the effect of extrinsic reliability and pose conditioning on feed-forward 3DGS, and propose a calibration-aware pipeline that anchors reconstructed scenes to the robot's metric workspace. Our experiments show that pose conditioning improves novel-view fidelity, while its geometric benefit depends on the reliability of the injected extrinsics. Using this pipeline, we present a renderable, metric-pose-anchored dataset with scene-level reliability information for robot manipulation research. Our dataset is available at this https URL

60. 【2610.01742】World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories

链接:https://arxiv.org/abs/2610.01742

作者:Jiahui Lei,Qianqian Wang,Trevor Darrell,Angjoo Kanazawa

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Equipping artificial agents, spatial intelligence requires, comprehensive generative prior, Equipping artificial, World Motion Models

备注: Accepted at NeurIPS 2026 (Spotlight). Url: [this https URL](https://jiahuilei.com/projects/wmm/)

点击查看摘要

Abstract:Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.

61. 【2610.01741】ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

链接:https://arxiv.org/abs/2610.01741

作者:Yijie Zhu,Rui Shao,Jie He,Wei Li,Bo Zhao,Yelin Wang,Xiaochen Yuan,Tao Tan,Miao Zhang,Xiaojiang Peng,Zitong Yu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:world dynamics forecasting, improve robotic manipulation, dynamics forecasting, aim to improve, manipulation via future

备注: Accepted to NeurIPS 2026. Project page: [this https URL](https://jiutian-vl.github.io/ATI-VLA-page/)

点击查看摘要

Abstract:Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.

62. 【2610.01723】Rethinking Memorization Mitigation in Diffusion Models: Reinforcing Text Conditioning

链接:https://arxiv.org/abs/2610.01723

作者:Hyungjun Joo,Sehwan Kim,Hyeonggeun Han,Sangwoo Hong,Jungwoo Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieved remarkable progress, closely reproducing individual, reproducing individual training, models have achieved, achieved remarkable

备注:

点击查看摘要

Abstract:Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.

63. 【2610.01710】CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

链接:https://arxiv.org/abs/2610.01710

作者:Dongwei Sun,Yujie Zhang,Bowen Yao,Pei Liu,Jing Yao,Xiangyong Cao

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual grounding localizes, Visual grounding, localizes an object, Visual, grounding

备注:

点击查看摘要

Abstract:Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at this https URL.

64. 【2610.01707】MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation

链接:https://arxiv.org/abs/2610.01707

作者:Liwei Liao,Yingkui Zhang,Qianqian Tong,Ronggang Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, aims to endow, enabling explicit, explicit and precise, Spatial Visual Distillation

备注:

点击查看摘要

Abstract:Mesh extraction from 3D Gaussian Splatting (3DGS) aims to endow 3D Gaussians with accurate geometric structures, enabling explicit and precise 3D occupancy. However, existing methods primarily focus on scene-level mesh extraction, making them unable to represent object-level occupancy and often resulting in non-watertight surfaces. To overcome these limitations, we propose \textbf{MEGA} (\underline{M}esh \underline{E}xtraction from \underline{GA}ussians), a ``segment-then-mesh'' framework for extracting object-level, watertight meshes from complex 3DGS scenes. At the core of MEGA are \textbf{Spatial Visual Distillation (SVD)} and a mask-guided neural surface reconstruction module. SVD treats the 3DGS model as a teacher, sampling diverse camera poses and rendering the corresponding views of each segmented object. These observations are then used to train a mesh reconstruction model through photometric supervision. Extensive experiments on several widely used benchmarks demonstrate that MEGA achieves state-of-the-art performance in recovering accurate object-level 3D occupancy. Moreover, MEGA enables complex physical interactions by combining high-quality object-level meshes for geometric occupancy with 3DGS representations for photorealistic rendering.

65. 【2610.01687】Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

链接:https://arxiv.org/abs/2610.01687

作者:Akshit Singh,Shyam Marjit,Wei Lin,Leonid Karlinsky,M. Jehanzeb Mirza

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:fixed computation path, sampling multiple responses, scaling often seeks, seeks better answers, multiple responses

备注:

点击查看摘要

Abstract:Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.

66. 【2610.01682】Beyond Leaderboard Scores: A Deployment-Focused Protocol for Interpretable Tracking Evaluation in Pedestrian-Centric Environments

链接:https://arxiv.org/abs/2610.01682

作者:Dominik Wojcikiewicz,Diego Paez-Granados

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Mobile robots operating, embedded computing budget, Mobile robots, remain spatially credible, Order Tracking Accuracy

备注: 8 pages, 7 figures; supplementary video provided as ancillary material. Submitted to IEEE Robotics and Automation Letters (RA-L)

点击查看摘要

Abstract:Mobile robots operating among pedestrians need trajectories that become available quickly, remain spatially credible through missed observations, preserve identity, and fit within an embedded computing budget. Aggregate tracking scores provide limited insight into when and how trajectories fail, while varying detector inputs can confound tracker and detector quality. We present a deployment-focused, tracker-only evaluation protocol that uses shared detections to isolate tracker behavior and directly evaluates initialization, detector-gap continuation, identity recovery, close-neighbor association, and load-dependent tracker-step runtime, while Higher Order Tracking Accuracy (HOTA) is retained as a complementary aggregate measure. We apply the protocol to the JackRabbot Dataset and Benchmark (JRDB) using six open-source trackers and our lightweight Pedestrian Reference Tracker (PedRefTrack), together with a GT-assisted variant that estimates the remaining tracker-side gap under idealized association and motion. Under fixed detections, the non-GT trackers span only 24.26%-29.67% HOTA yet exhibit markedly different capability profiles. After 1.0 s without detector support, no tracker without GT assistance maintains spatially correct, same-identity output in more than half of eligible cases, making missing-observation continuation the dominant limitation among the tested properties. Close-neighbor failures are smaller and increase mainly at the shortest separations. Tracker-step runtime on an NVIDIA Jetson Orin is heavy-tailed and load-sensitive, causing several trackers to fall below the 10 Hz real-time target in crowded frames. The protocol provides a reproducible way to characterize tracker behavior and deployment suitability in pedestrian-centric environments. Code and evaluation scripts are released at this https URL.

67. 【2610.01681】When Text-to-Image Helps Editing: The Effects of Conditioning During Denoising

链接:https://arxiv.org/abs/2610.01681

作者:Lidia Troeshestova,Alexander Ustyuzhanin,Sergey Kastryulin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:instruction-based image editing, standard editing pipelines, instruction-based image, pipelines keep source-image, editing

备注: Under review as a conference paper at ICLR 2027

点击查看摘要

Abstract:Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-image conditioning throughout denoising. We ask whether editing can benefit from T2I, and study how the effects of conditioning vary across edits and denoising stages. In pure editing, source attention declines for some edits over the sampling trajectory. This observation led us to task switching, which lets the model draw on its T2I capabilities. Across three unified editors and four benchmarks, switching to the T2I task for bounded intervals improves edit quality, while mean perceptual preservation remains close to pure editing across all three models. Unified editors therefore benefit from using both conditioning modes they are trained for, and the timing of the switch sets the balance between quality and preservation.

68. 【2610.01670】Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation

链接:https://arxiv.org/abs/2610.01670

作者:Yuan Huang,Zirui Song,Xiuying Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:large language models, Multimodal large language, instruction-based image editing, model training, language models

备注: 30 pages, 9 figures

点击查看摘要

Abstract:Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is verified to preserve the underlying editing quality. To address this challenge, we introduce EditJudgeBias, a counterfactual benchmark with verified quality preservation, comprising 1,196 real editing samples and 13 cues injected across four evaluation sites. We verify quality preservation for the requested edit using calibrated multimodal validators, controls, and human inspection. We then audit five MLLM judges along three complementary dimensions: invariance to quality-preserving cues, agreement with human judgments, and stability of pairwise preferences. Importantly, observed shifts are evaluated against each judge's own zero-dose and re-query noise floors rather than against zero. Experiments show that quality-preserving cues move every judge beyond its own noise. Fabricated majority opinions increase ratings, irrelevant visual elements cause larger shifts than whole-image manipulations, and swapping candidate order reverses up to 60.9% of pairwise decisions. Edit-region cues also tend to reduce human agreement. The three measures characterize judges differently, showing that robustness cannot be captured by a single metric.

69. 【2610.01661】DiVid: Diagnosing Dimension-Specific Diversity Collapse in Video Generation Models

链接:https://arxiv.org/abs/2610.01661

作者:Huanran Hu,Zihui Ren,Dingyi Yang,Zhinan Song,Guozheng Wu,Tiezheng Ge,Qin Jin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:produce highly similar, video generation, remarkable progress, limiting their usefulness, creative exploration

备注:

点击查看摘要

Abstract:Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rely on global scalar metrics, which obscure where diversity collapses in the spatiotemporal space of videos. We introduce DiVid, a dimension-level diagnostic framework that decomposes video generation diversity into six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera. Each dimension is measured through a reproducible computer-vision pipeline and analyzed alongside quality and instruction faithfulness to examine potential trade-offs. Systematic evaluation of representative video generation models reveals that diversity is highly dimension-specific: models with strong global diversity scores still collapse on specific factors, particularly Motion and Camera. These rankings persist after filtering unfaithful generations, indicating genuine capability differences rather than off-prompt outputs. Beyond measurement, controlled prompt interventions identify two fundamental bottlenecks: default mode convergence, where models fall back to dominant patterns under open-ended prompts; and realization gaps, where models fail to faithfully realize diverse, explicitly requested alternatives, particularly for temporal factors. The larger faithfulness losses for temporal factors highlight the difficulty of controlling motion and camera variation through text alone. DiVid thus shifts the study of diversity from measuring whether it exists to diagnosing where and why it collapses, and provides actionable directions for dimension-aware training objectives and control signals. The framework will be released to facilitate future research on diverse and controllable video generation.

70. 【2610.01640】Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

链接:https://arxiv.org/abs/2610.01640

作者:Xinye Zhao,Yunkai Dang,Yunchen Wu,Wenbin Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:backbone size, Vision-language models, face a fixed-budget, complex reasoning, language backbone

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.

71. 【2610.01637】Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering

链接:https://arxiv.org/abs/2610.01637

作者:Cong Phu Nguyen,Huy Tien Nguyen,Tung Le

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual Question Answering, Vietnamese VQA, recent decades, made significant progress, progress in understanding

备注:

点击查看摘要

Abstract:In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.

72. 【2610.01625】Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models

链接:https://arxiv.org/abs/2610.01625

作者:Wentao Yue,Qingyu Mao,Tianyou Lai,Ahmed M. Abdelmoniem,Qilei Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Federated parameter-efficient fine-tuning, parameter-efficient fine-tuning enables, fine-tuning enables distributed, adapt pretrained vision-language, sharing raw data

备注:

点击查看摘要

Abstract:Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from client-specific style, but they largely treat each domain as a class-agnostic transformation. We show that this abstraction is insufficient: the cross-domain displacement associated with a fixed domain varies across semantic classes, and only a subset of these class-domain residuals damages the image-text decision margin. We therefore propose Margin-Oriented Semantic-Appearance Interaction Correction (MOSAIC), which first constructs a decision-aware harmfulness score that measures whether a training-derived class-domain residual favors a competing text prototype over the true class. It then models fine-grained class-domain interactions with a low-rank residual adapter whose class factors and residual basis are globally shared while domain factors remain client-private. An image-conditioned gate further controls candidate-wise correction, and harmful-pair-aware reweighting prioritizes decision-relevant residuals during local optimization. Extensive experiments on Office31, OfficeHome, and DomainNet100 demonstrate that MOSAIC consistently improves macro-client top-1 accuracy across all evaluated domain-shift and joint domain-label-shift settings.

73. 【2610.01614】Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models

链接:https://arxiv.org/abs/2610.01614

作者:Xindi Yang,Baolu Li,Liam Lee,Zhenfei Yin,Songxin Zhang,Zhuoyang Song,Xu Jia,Jianfei Cai,Tien-Tsin Wong,Bingyi Jing,Mengyue Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Generative video world, world state, synthesize open-ended environments, world, Generative video

备注: Project page: [this https URL](https://madaoer.github.io/projects/oneira)

点击查看摘要

Abstract:Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: this https URL

74. 【2610.01605】Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

链接:https://arxiv.org/abs/2610.01605

作者:Yuzhou Wang,Emile Anand,Ijay Narang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Logic in Computer Science (cs.LO)

关键词:requires composing multiple, Reliable visual reasoning, reasoning requires composing, composing multiple visual, multiple visual observations

备注: 29 pages, 6 figures, 14 tables

点击查看摘要

Abstract:Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each defined by a Boolean combination of ten visual statements, across 1,000 generated scenes and 46 diverse labeled photographs, along with 1,000 object-identification questions over the same photographs. Our question families are deliberately constructed to challenge reasoning through misleading local cues and nested logical operations, and include symbolic and structured natural-language presentations. Across eight model configurations with thinking disabled or minimized, Boolean accuracy ranges from 48.52% to 50.57%, while the identification accuracy reaches at most 43.0%. A thinking-enabled GLM configuration achieves uneven gains while retaining substantial errors and inconsistencies. Hob-VL exposes these failures through executable reference answers and matched evaluations.

75. 【2610.01595】Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

链接:https://arxiv.org/abs/2610.01595

作者:Youngwoo Shin,Yusung Ro,Minseo Kim,Junmo Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Video Large Language, Language Models, Large Language, visual content evolves

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $\tau_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $\tau_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at this https URL.

76. 【2610.01590】wo Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize

链接:https://arxiv.org/abs/2610.01590

作者:Yuan Huang,Zihan Chen,Runbin Zhang,Hongwei Ding,Changzeng Fu,Shiqi Zhao

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:pre-trained vision transformer, vision transformer accumulate, accumulate storage linearly, transformer accumulate storage, Continual learners

备注: 21 pages, 12 figures

点击查看摘要

Abstract:Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.

77. 【2610.01589】PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

链接:https://arxiv.org/abs/2610.01589

作者:Akira-Miranda Adeyomi Adeniran-Lowe,Binod Singh,Lars Arnold Dethlefsen,Lazaros Nalpantidis,Theodora Kontogianni

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:world coordinate frame, consistent world coordinate, reconstructed scenes expressed, coordinate frame, globally reconstructed scenes

备注:

点击查看摘要

Abstract:Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial observations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label-free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmentation probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides supervision only during training. Inference operates directly on the partial observation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet$\rightarrow$ScanNet++ transfer surpasses fully fine-tuned Sonata ($53.93$ vs.\ $48.09$ mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer.

78. 【2610.01544】Revisiting Cross-Reconstruction for Generalizable Deepfake Detection

链接:https://arxiv.org/abs/2610.01544

作者:Bingjian Yang,Shilei Zhao,Zheng Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Existing image forgery, unseen manipulation methods, manipulation methods due, Existing image, image forgery detectors

备注:

点击查看摘要

Abstract:Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \textbf{artifact diversity}, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.

79. 【2610.01542】Synthetic training for long-tail haemorrhagic lesion segmentation in data-scarce settings

链接:https://arxiv.org/abs/2610.01542

作者:Yuan Cao,Sumeet Dash,Antonia Zachariadis,Stefanie Schreiber,Katja Neumann,Jose Bernal

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cortical superficial siderosis, cerebral small vessel, Cerebral microbleeds, cerebral small, small vessel disease

备注: Accepted: MICCAI 2026 SASHIMI workshop

点击查看摘要

Abstract:Cerebral microbleeds (CMBs) and cortical superficial siderosis (cSS) are imaging markers of cerebral small vessel disease, but their automated segmentation is limited by the scarcity of positive cases and voxel-level annotations. We propose a synthetic training framework for long-tail haemorrhagic lesion segmentation that requires no real lesion annotations for training and leverages radiological description of the lesions. Starting from anatomical brain parcellations, the framework applies spatial augmentation and voxel resampling, procedurally inserts cSS and CMB labels using clinical priors on lesion location and morphology, and synthesises images through randomised intensity assignment, blurring, and Rician noise simulation. Models were trained on dynamically generated image-label pairs and evaluated against manual delineations in 10 cSS cases and 13 CMB cases. The proposed configurations outperformed classical filter baselines. For cSS, the hypointensity constrained model achieved higher AUPRC and AUROC than the Frangi filter (AUPRC: 0.284 vs 0.083; AUROC: 0.907 vs 0.731). For CMBs, explicit synthesis of blood vessels as lesion mimics improved performance over the classical baseline (AUPRC: 0.538 vs 0.004; AUROC: 0.999 vs 0.968). These results support our proposal as a feasible strategy for data-scarce haemorrhagic lesion segmentation.

80. 【2610.01531】owards Reliable Vision-Language Models for Autonomous Driving

链接:https://arxiv.org/abs/2610.01531

作者:Manasa Mariam Mammen,Priyanka Mary Mammen,Zafer Kayatas,Stefan Wagner

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:scene understanding, driving reasoning, autonomous driving, explored in autonomous, mathrm

备注:

点击查看摘要

Abstract:Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ($\mathrm{V}{\scriptstyle \mathrm{EA}}$), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that $\mathrm{V}{\scriptstyle \mathrm{EA}}$ improves performance for some models and datasets, although the gains are not consistent across all settings.

81. 【2610.01517】SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing

链接:https://arxiv.org/abs/2610.01517

作者:Fa-Ting Hong,Peter Wonka

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Text-driven human motion, Text-driven human, learned preservation gate, source, preserving compatible source

备注: Under review

点击查看摘要

Abstract:Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20\% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.

82. 【2610.01512】VoxelSynth3D: Interpretable Volumetric Image-Domain Metal Artifact Reduction with a Paired Synthetic CLINIC-Metal Benchmark

链接:https://arxiv.org/abs/2610.01512

作者:Amritesh Banerjee,Abdul Basit,Renil Renji Joseph,Nouhaila Innan,Muhammad Shafique

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:adjacent soft-tissue interfaces, soft-tissue interfaces, postoperative musculoskeletal, musculoskeletal CT obscure, obscure bone-implant

备注: 7 pages, 7 figures. Accepted for publication at BHI 2026

点击查看摘要

Abstract:Metal artifacts in postoperative musculoskeletal CT obscure bone-implant and adjacent soft-tissue interfaces. Many metal artifact reduction (MAR) methods require unavailable raw projections or learned models that may shift across scanners and implants. We present VoxelSynth3D, a training-free 3D image-domain framework for reconstructed CT. The framework combines support masking, normalized tissue synthesis, deviation gating, and restricted edge refinement. Detected implant voxels are preserved in the output, while correction targets metal-induced artifacts in the surrounding tissue. We also construct Synthetic CLINIC-Metal, a controlled paired synthetic evaluation resource, from no-metal CTPelvic1K volumes with clean targets, metal/artifact masks, fixed seeds, and patient-level splits; 75 unpaired real metal cases receive qualitative/no-reference evaluation only. The operating point was fixed in a near-flat validation basin. With exact-mask oracle localization, all methods share a metal-excluded tissue ROI. On 40 held-out cases, VoxelSynth3D reduced RMSE from 801.48 to 786.18 HU (paired gain 15.30 HU, 95% CI 11.68-19.23), improving every case and exceeding the evaluated 3D Gaussian smoother by 13.58 HU. Clean-edge agreement decreased next to metal but exceeded input beyond 5 mm. Thus, VoxelSynth3D provides case-consistent within-distribution tissue-error reduction with a localized structural tradeoff. Spacing-aware sensitivity retained aggregate broad-region improvement and identified near-metal calibration as a target.

83. 【2610.01510】FedCKA: Representation-Guided Layer Personalization for Federated 3D Perception Across Driving Domains

链接:https://arxiv.org/abs/2610.01510

作者:Jolle Verhoog,Ali Burak Ünal,Holger Caesar

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:intelligent vehicles demands, vehicles demands, time of day, intelligent vehicles, remain dependable

备注: 8 pages, 3 figures. Submitted to IEEE ICRA 2027

点击查看摘要

Abstract:Robust perception in intelligent vehicles demands 3D object detectors that remain dependable under domain shifts, such as changes in time of day, location, or weather. However, due to costly annotation and rare shifts, some environments lack sufficient data to train a standalone detector. Federated learning offers a privacy-preserving framework for collaborative model training, enabling clients to benefit from shared learning across diverse environments. Yet, this framework traditionally relies on a single global consensus model, which struggles to perform across heterogeneous local data distributions. Local conditions are better captured by adapting a subset of the model, but many personalization approaches rely on predefined layer partitions or fixed personalization ratios, thereby limiting adaptation to client-specific divergence. To reduce this rigidity, we propose FedCKA, a Centered Kernel Alignment (CKA)-based strategy that dynamically handles the personalization-globalization trade-off. Specifically, FedCKA computes layer-wise feature similarities between local client models and the global consensus model during training. By converting layer-wise similarity scores into client-specific aggregation masks, FedCKA selectively shares representation-consistent layers. Evaluation on a unified multi-domain benchmark based on nuScenes shows that FedCKA outperforms established federated baselines, including FedBN, FedRep, and FedSelect, improving average NDS by 7 percentage points over the strongest baseline. The findings offer both a comparative benchmark and a promising direction for robust federated 3D perception across shifts in location, weather, and illumination. Code is available at this https URL.

84. 【2610.01499】VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

链接:https://arxiv.org/abs/2610.01499

作者:Yu Huang,Jungang Li,Zhiyuan Wang,Yonghua Hei,Song Dai,Jiayu Yang,Deyuan Liu,Xiang Zheng,Xiaoshuang Shi,Hao Cheng,Kaidi Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:natural language instructions, approaching cinematic standards, produce highly realistic, quality approaching cinematic, Recent video generation

备注:

点击查看摘要

Abstract:Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at this https URL.

85. 【2610.01496】SALD: Self-Referenced Advantage Learning for Diffusion Models

链接:https://arxiv.org/abs/2610.01496

作者:Aryan Das,Surjo Dey,Koushik Biswas,Swalpa Kumar Roy,Moloud Abdar,Arnab Bhattacharya,Vinay Kumar Verma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrationor feedback-augmented contexts, obtain informative training, Recent work, informative training signals, student learned parameters

备注:

点击查看摘要

Abstract:Recent work on language-model adaptation has shown that single models can obtain informative training signals by evaluating their behavior in demonstrationor feedback-augmented contexts, with the help of a teacher network, which is driven by the student's learned parameters. Inspired by this internal-reference principle, we investigate how diffusion models can identify self-referenced training signals without external demonstrations or teacher networks. We introduce SALD, a self-referenced training framework that evaluates each image-caption pair at two noise levels using the same model. The easier, lower-noise path is evaluated without gradient tracking to provide a reference, while the harder, higher-noise path provides the training gradient. Rather than directly distilling the easy-path prediction, SALD uses the difference between two path errors to adapt the hardpath objective. The proposed Advantage-Guided Diffusion (AGD) converts this relative error into a differentiable sample-level weight. Temporal Advantage Memory (TAM) accumulates relative difficulty across training and adapts the future gap between the two noise levels. Spectral Advantage Decomposition (SAD) further compares the residual power spectra of the two paths and constructs a differentiable, frequency-derived latent-element weight. All components share a single set of model parameters, requiring neither an external teacher network nor additional trainable parameters during training or inference, and no modification to the inference procedure. Experiments across multiple architectures and datasets demonstrate consistent improvements in generation quality, while component-wise ablations quantify the contributions of the proposed components.

86. 【2610.01480】FiVOS: A Fish Segmentation Algorithm Based on Interactive Video Object Segmentation and Filter Enhancement

链接:https://arxiv.org/abs/2610.01480

作者:Yuqing Duan,Song Zhang,Shili Zhao,Daoliang Li,Ran Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reducing economic losses, economic losses, continuous expansion, increasingly critical, reducing economic

备注:

点击查看摘要

Abstract:With the continuous expansion of aquaculture, precise and efficient monitoring of fish behavior has become increasingly critical for improving farming efficiency and reducing economic losses. In particular, with the ongoing enhancement of computational capabilities in deep learning models, vision-based fish segmentation methods are garnering growing attention. By analyzing video segmentation results, fish behavior can be effectively tracked, thereby providing reliable data support for the precise regulation of aquaculture environments. However, existing deep learning-based video segmentation methods for aquaculture scenarios often overlook the dynamic correlations between video frames. In contrast, Interactive Video Object Segmentation (IVOS) employs an interaction-propagation scheme to achieve high-precision segmentation while minimizing user effort, thereby enhancing monitoring efficiency. Yet, IVOS applications in aquaculture remain limited due to data scarcity, and are susceptible to error accumulation and mask loss over long sequence propagation due to high intra-class similarity. In response, this paper proposes an improved interactive video object segmentation method (FiVOS) and constructs two fish-specific datasets. FiVOS utilizes a mask block filter to enable early detection and correction of erroneous propagated mask blocks, enhancing filtering accuracy through a rule-based thresholding approach. Additionally, it serializes noise filters to further eliminate erroneous mask noise, thereby improving model robustness. Experimental results demonstrate that FiVOS achieves state-of-the-art (SOTA) performance in fish video segmentation tasks, providing robust technical support for fish behavior research.

87. 【2610.01477】ALFRED: Requirement-driven development of an open-source mobile manipulator for long-term plant monitoring

链接:https://arxiv.org/abs/2610.01477

作者:Ciarán Miceal Johnson,Christopher Quail,Garry Ellard,Alistair McConnell,Steve Tonneau,Fernando Auat Cheein

类目:Robotics (cs.RO); Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV)

关键词:Tracking seasonal change, Tracking seasonal, forests requires observing, plants repeatedly, seasonal change

备注: 36 pages, 19 figures

点击查看摘要

Abstract:Tracking seasonal change in crops and forests requires observing the same plants repeatedly. Ground robots can do this at close range, and a manipulator gives their sensors more viewpoints. Yet the robots behind long-term field datasets are rarely released with their design files, and how a robot's own structure limits arm reach and occludes its sensors is seldom compared between builds. We present ALFRED, an open-source mobile manipulator built from commercially available components. It carries a six-degree-of-freedom arm, LiDAR, RGB-D cameras, RTK GNSS and an IMU on an Ackermann-steered base, all mounted on a reconfigurable aluminium strut frame, and runs containerised ROS software. It was developed through four builds against six requirements for repeated outdoor deployment: durability, modularity, repairability, sensing reach, endurance and reproducibility. Model-based analysis of the last three builds shows the usable share of the arm's reachable poses rising from 34.0% to 60.0% and then 66.1%, and ray casting shows that only the final build keeps the frame-mounted LiDAR's horizontal view clear both forwards and backwards. ALFRED completed a year of monthly forest surveys (528 traversals) without missing a scheduled collection. This was despite battery degradation, reconfiguration for another researcher's study, and the parallel development of ALFRED 2.0 for autonomous crop-row operation, with each switch between builds taking about six hours. The deployment also showed that mechanical modularity is only as dependable as the robot description that tracks it.

88. 【2610.01452】Uncertainty-Guided Handshake: Efficient Human-in-the-Loop Refinement for Surgical-Grade Glioma Segmentation

链接:https://arxiv.org/abs/2610.01452

作者:Samuel Hart,Ahmad Yahya,Ahmed Karam Eldaly

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:medical image segmentation, image segmentation achieve, segmentation achieve high, safe clinical deployment, preclude safe clinical

备注: 12 pages

点击查看摘要

Abstract:While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastrophic failures that preclude safe clinical deployment, particularly in neuro-oncology. Interactive segmentation frameworks mitigate this by incorporating human oversight, but traditionally impose prohibitive cognitive and temporal workloads by requiring clinicians to manually search for errors. In this project, we present an efficient, Hybrid Structural-Aleatoric Human-in-the-Loop framework for glioma segmentation that bridges the gap between automated baseline performance and surgical-grade precision, achieving sub-2.0 mm HD95 on curated benchmarks while providing safety-net routing for structural failures across real-world clinical data. By extracting voxel-wise Test-Time Augmentation (TTA) uncertainty and applying hierarchical topological filtering, our method proactively isolates high-risk structural anomalies. We comprehensively evaluated our approach on a challenging out-of-distribution clinical stress-test cohort (N = 362). Operating under a simulated Human Oracle, the framework improved the Whole Tumor (WT) Dice score from 0.891 to 0.914 and reduced the 95th percentile Hausdorff Distance (HD95) from 5.82 mm to 4.76 mm. Critically for surgical safety, the system rescued severe boundary failures in the Tumor Core, reducing mean HD95 from 17.96 mm to 14.83 mm (improving absolute TC Dice to 0.356). These spatial rescues were achieved while demanding a median interactive workload of just 11.3% of the target volume. Acknowledging this as a simulated upper bound lacking real-world cognitive friction, the framework nevertheless demonstrates a highly Pareto-efficient pathway for safely deploying clinical AI.

89. 【2610.01438】he Impact of Processing Parameters on High-Accuracy Measurements in UAV Photogrammetry

链接:https://arxiv.org/abs/2610.01438

作者:Paweł Ćwiąkała,Edyta Puniach,Elżbieta Pastucha,Wojciech Gruszczyński

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Unmanned aerial vehicle, Bundle Block Adjustment, requiring high accuracy, Unmanned aerial, applications requiring high

备注:

点击查看摘要

Abstract:Unmanned aerial vehicle (UAV) photogrammetry is increasingly used in applications requiring high accuracy, such as determining ground surface changes caused by landslides, mining, or microrelief transformation. While acquisition strategies have been widely studied, the influence of the processing workflow-particularly Bundle Block Adjustment parameter settings-remains insufficiently explored. This study addresses this gap through a systematic, full-factorial evaluation of 768 processing variants applied to ten UAV datasets collected over 1.5 years in a 220 ha study area. Eight key parameters were analysed. The results show substantial variability in final 3D accuracy: the best performing variant achieved a root mean square error (RMSE) of 16 mm, whereas the weakest reached 303 mm. The most influential factors were the number of ground control points, the application of additional camera calibration corrections, and the use of the Post-Processing Kinematic GNSS method for determining camera projection center coordinates. The study also evaluates how workflow optimization affects the accuracy of displacement, tilt changes, and horizontal strain determination. While random displacement errors remained stable (RMSE of ~6-7 mm), systematic errors were significantly reduced by over half in all axes, with vertical median absolute error decreasing from 14 mm to 7 mm in the optimized configuration compared to the baseline previously used by the authors. This study provides the first large-scale, practice-oriented assessment of how processing parameter selection shapes the accuracy of both photogrammetric products and deformation indices determination. The results offer actionable guidance for developing more robust and repeatable UAV photogrammetry workflows tailored to high-precision monitoring.

90. 【2610.01434】MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

链接:https://arxiv.org/abs/2610.01434

作者:Xudong Wang,Hao Wu,Haozhe Hu,Peiran Yin,Xinghao Chen,Yunpu Ma,Wei Zhang,Xiaoyu Shen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Multimodal large language, large language models, incur substantial inference, substantial inference costs, processing long visual-textual

备注:

点击查看摘要

Abstract:Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at this https URL.

91. 【2610.01409】Localisation-Aware Uncertainty for Pretrained Object Detection

链接:https://arxiv.org/abs/2610.01409

作者:Charmaine Barker,Daniel Bethell,Simos Gerasimou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reliable uncertainty estimation, covariate shift, Reliable uncertainty, estimation is essential, essential for deploying

备注:

点击查看摘要

Abstract:Reliable uncertainty estimation is essential for deploying object detectors when distribution/covariate shift and adversarial attacks may occur. Existing approaches often require detector retraining, architectural modification, or repeated inference, which may be infeasible or incur significant overheads. We introduce a lightweight post-hoc evidential meta-model that learns when object localisations should be considered uncertain while keeping the base detector frozen. Our approach automatically identifies localisation-relevant features and uses saliency-guided modification to construct an increasingly challenging curriculum. Detection-level targets combine localisation error, modification level, and prediction instability to guide an evidential meta-model to estimate uncertainty for each predicted bounding box. Our approach requires no changes to the detector and preserves its original localisation outputs. Across adversarial attacks and evaluated strengths, GRACE improves TP-FP AUROC by 22% relative to the strongest comparator in some cases while maintaining in-distribution detection performance.

92. 【2610.01408】Smoother Flow Matching via Contrastive Trajectory Repulsion

链接:https://arxiv.org/abs/2610.01408

作者:Ziqi Jiang,Zhenqi He,Long Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Trajectory crossing remains, Flow Matching, previous works typically, works typically view, causing velocity averaging

备注: 18 pages, 5 figures

点击查看摘要

Abstract:Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the intermediate trajectories. In this paper, we introduce a new network learning perspective: crossing points inherently induce large local Lipschitz constants in the target velocity field, leading to two drawbacks. First, high Lipschitz constants correspond to high-frequency signals in the velocity field that neural networks struggle to fit due to spectral bias. Second, they also imply drastic velocity variations, leading to severe numerical integration errors in few-step inference. To alleviate this, we propose CoFlow, a framework that introduces the contrastive learning paradigm into FM to explicitly repel trajectories during training, thereby lowering the local Lipschitz constants of the velocity field. Specifically, we formulate CoFlow from a Stochastic Differential Equation (SDE) perspective by injecting a repulsive drift term. This drift actively guides the forward process of positive samples away from negative trajectories, effectively reducing the local Lipschitz constant. Furthermore, we derive an equivalent stochastic interpolant formulation from this SDE, providing a simple and tractable design space to control the influence of negative samples. Extensive experiments on ImageNet 256x256 demonstrate that CoFlow significantly reduces FID compared to standard FM in few-step inference (e.g., 20 steps), with no added training overhead. The code can be accessed at: this https URL

93. 【2610.01389】AiSearch: Interactive Multi-Modal Search with VLMs

链接:https://arxiv.org/abs/2610.01389

作者:Ali Koksal,Mei Chee Leong,Vicky Sintunata,Ching Ling Chin,Wee Teck Fong

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Modern retrieval systems, Modern retrieval, real time, Vision Language Models, Modern

备注: The demo paper with 1 page main paper, 7 pages supplementary material accepted and presented in ECCV 2026

点击查看摘要

Abstract:Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearch supports interactive search refinement through user feedback to tailor results to the user's intent, and allows visual benchmarking across multiple VLMs, enabling users to select the most suitable model for their task.

94. 【2610.01388】Supervising Sound Localization by In-the-wild Egomotion

链接:https://arxiv.org/abs/2610.01388

作者:Anna Min,Ziyang Chen,Hang Zhao,Andrew Owens

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)

关键词:supervisory signal, learning binaural sound, learning binaural, sound, sound localization

备注: CVPR 2025 Highlight (IEEE/CVF Conference on Computer Vision and Pattern Recognition)

点击查看摘要

Abstract:We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks

95. 【2610.01385】Is it Possible to Generate Irreversible PolyProtected Templates from Face Embeddings using System-Specific Keys?

链接:https://arxiv.org/abs/2610.01385

作者:Vedrana Krivokuća Hahn,Jérémy Maceiras,Sébastien Marcel

类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)

关键词:subjects' face embeddings, traditional subject-specific keys, face embeddings, PolyProtect biometric template, template protection method

备注: Submitted to TIFS journal on 12 May 2026 (under review). Consists of: 13 pages, 9 figures, 3 tables

点击查看摘要

Abstract:This work aims to answer the question of whether it is possible to generate irreversible protected templates when the PolyProtect biometric template protection method is applied to face embeddings using system-specific keys (i.e., the same C and E parameters, which define the transform, are applied to all subjects' face embeddings), instead of the traditional subject-specific keys (i.e., each subject has their own C and E parameters). This is important for determining whether we can perform de-duplication of face identities in the PolyProtected domain, which is not possible in the subject-specific key scenario due to the clash with PolyProtect's unlinkability property (i.e., one could generate multiple protected templates belonging to the same identity, using different C and E parameters, such that those templates cannot be linked to each other). We present experiments (reproducible using our open-source code) to prove that there exist at least three ways of systematically selecting system-specific keys that produce irreversible PolyProtected templates: (i) from pre-selected subject-specific keys, (ii) by applying a previously proposed key selection algorithm to random vectors, and (iii) by approximating a "good" C/E pair distribution from which system-specific keys can be constructed. Our findings thus point to the conclusion that it is, indeed, possible to safely operate PolyProtect in the system-specific key scenario without degrading the template protection potential. This opens up the possibility for identity de-duplication in the PolyProtected domain.

96. 【2610.01352】MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

链接:https://arxiv.org/abs/2610.01352

作者:Juekai Lin,Honglin Lin,Yuqian Yuan,Xiaolong Wu,Jie Cao,Liang Liang,Yunqi Cao,Yun Zhu,Wenqiao Zhang,Lijun Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains challenging due, post-training remains challenging, uneven data quality, inefficient supervision construction, large-scale reasoning supervision

备注:

点击查看摘要

Abstract:Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.

97. 【2610.01331】CLASP: Continual Low-rank Adapters for Spatially Placed Concepts from One Hypernetwork

链接:https://arxiv.org/abs/2610.01331

作者:Wojciech Gromski,Patryk Krukowski,Jan Miksa,Maciej Zieba,Przemysław Spurek

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires sequentially acquiring, models requires sequentially, requires sequentially, sequentially acquiring, diffusion models requires

备注: 31 pages. Code: [this https URL](https://github.com/genwro-ai/clasp) , project page: [this https URL](https://genwro-ai.github.io/clasp)

点击查看摘要

Abstract:Continual personalization of text-to-image diffusion models requires sequentially acquiring new concepts while retaining previously learned ones. However, existing methods either suffer from catastrophic forgetting or rely on storing additional concept-specific parameters and spatial components, causing their parameter footprint to grow with the concept stream. This limits their ability to scale to long sequences of personalization tasks. We propose a rehearsal-free approach that uses a single fixed-size hypernetwork to continually personalize a frozen diffusion model. Instead of expanding the model as new concepts are acquired, the hypernetwork dynamically produces the concept-specific adaptations required for personalization while preserving previously learned concepts. Our framework further integrates spatial control into the personalization process, allowing users to specify where a personalized concept should appear without introducing additional per-concept components. This formulation enables continual personalization with a parameter footprint that remains independent of the number of learned concepts, aside from compact concept representations. Experiments demonstrate strong retention of previously learned concepts and reliable spatial grounding, matching or improving upon existing methods while scaling effectively to long streams of personalization tasks.

98. 【2610.01314】ARROW: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild

链接:https://arxiv.org/abs/2610.01314

作者:Ilya Fradlin,Christian Schmidt,Jens Piekenbrinck,Karim Knaebel,Gonzalo Martin Garcia,Bastian Leibe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multiple video streams, Dynamic scenes, moving camera, multiple video, video streams

备注: Project page at: [this https URL](https://www.vision.rwth-aachen.de/arrow)

点击查看摘要

Abstract:Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks. Code and weights are publicly available.

99. 【2610.01302】STAGE: Subspace-Targeted Affine Generative Erasure for Text-to-3D Models

链接:https://arxiv.org/abs/2610.01302

作者:Karol Dziekan,Przemysław Spurek,Dawid Malarz

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unrelated inputs, behavior on unrelated, Concept erasure suppresses, stage, erasure suppresses

备注:

点击查看摘要

Abstract:Concept erasure suppresses a target concept while preserving behavior on unrelated inputs. Existing closed-form methods were designed for 2D image diffusion and assume a single generative pathway, so one edit must cover geometry and texture at once. Native 3D generators, which synthesize structured 3D representations directly rather than by lifting 2D samples, violate this assumption. We show that shape and object concepts must be erased in the structural stage of the pipeline and material concepts in the appearance stage. We therefore formulate erasure in native text-to-3D as a stage-aware editing problem and introduce STAGE, a training-free, closed-form framework. STAGE confines each edit to the low-dimensional subspace spanned by the differences between erase and anchor embeddings, and relaxes the norm-preserving (orthogonal) constraint of prior editors into a least-squares affine correction that maps target activations onto safe anchors subject to a penalty on the displacement of retained prompts. The correction applies to the structural stage, the appearance stage, or both. We find that the stage an edit must reach is determined by concept type. On TRELLIS, the standard open native 3D generator, across 15 shape, material, and object concepts, STAGE reaches 66.7 on a composite score that balances forgetting the target concept against preserving everything else, aggregating CLIP-based semantic and physical metrics, versus 53.2 for the strongest adapted baseline. Code: this https URL Project Page this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.01302 [cs.CV]

(or
arXiv:2610.01302v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.01302

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
100. 【2610.01291】ODDR: One-Step Deshadow Diffusion via Reward Guidance

链接:https://arxiv.org/abs/2610.01291

作者:Junseong Shin,Kijun Kim,Minseong Kim,Dongjin Kim,Tae Hyun Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:One-step Deshadow Diffusion, Recent advances, real-world paired, real-world paired supervision, significantly enhanced image

备注:

点击查看摘要

Abstract:Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.

101. 【2610.01286】Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

链接:https://arxiv.org/abs/2610.01286

作者:Xinhao Xiang,Weiyang Li,Zhijie Zheng,Abhijeet Rastogi,Jiawei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent depth foundation, achieve remarkable multi-view, Recent depth, remarkable multi-view depth, real-world dynamic environments

备注:

点击查看摘要

Abstract:Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encodes cross-frame matching, a property absent in depth-only models like DA3. We present Dyna3, a training-free framework that extends DA3 for 4D dynamic scene reconstruction without any fine-tuning. Our key insight is that DA3's cross-view features, though trained only for depth consistency, implicitly encode motion-discriminative signals when combined with best-match feature search across frames. Its static surfaces find consistent matches globally, while dynamic objects cannot. We further adopt vision-language models (VLM) to automatically generate scene-specific semantic prompts for SAM 3, enabling precise instance-level segmentation that distinguishes which objects move from what objects exist. For reconstruction, we decouple the scene into a cross-frame aligned static background and per-frame dynamic point clouds. Experiments on four datasets demonstrate that Dyna3 surpasses correspondence-trained methods with +5.5pp J-Mean over state-of-the-art VGGT4D on dynamic object segmentation, while achieving up to 13x faster pose estimation and 3x faster 4D reconstruction with 4 to 8x lower memory. Dyna3 could therefore enable much denser temporal sampling that prior methods cannot support.

102. 【2610.01283】ShelfChange3D: Object-Level 3D Change Detection for Retail Shelf Monitoring

链接:https://arxiv.org/abs/2610.01283

作者:Lingyi Zhou,Yunke Wang,Mengyu Zheng,Wenbo Wang,Zijian Wang,Chang Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reliable shelf monitoring, downstream robotic systems, Reliable shelf, retail automation, lack metric

备注: Our code will be available on our project website at [this https URL](https://zerone0011.github.io/ShelfChange3D/)

点击查看摘要

Abstract:Reliable shelf monitoring is an important capability for retail automation, yet existing out-of-stock detection methods mainly operate in image space and lack metric 3D localization for downstream robotic systems. We formulate shelf monitoring as object-level 3D change detection: given two RGB-D observations captured at different times, the goal is to identify changed products and localize each change with a 3D bounding box. To support this task, we introduce ShelfChange3D, comprising 145K synthetic and 5K real-world paired RGB-D observations with object-level 3D change annotations. We further propose ChangeBox, an end-to-end framework that jointly reasons over paired observations and predicts object-level 3D change boxes. To improve localization accuracy, we introduce a geometry-based refinement stage that exploits depth and gravity prior to estimate relative pose and refine predicted boxes. Experiments show that ChangeBox outperforms existing change detection baselines, with further gains from refinement and effective transfer from synthetic to real-world observations.

103. 【2610.01279】PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video

链接:https://arxiv.org/abs/2610.01279

作者:Junseong Shin,Hyeonsu Jo,Daehyun Kim,Tae Hyun Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:continuous sharp signal, Motion blur arises, finite exposure window, sharp signal, existing learning-based methods

备注:

点击查看摘要

Abstract:Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.

104. 【2610.01243】When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts

链接:https://arxiv.org/abs/2610.01243

作者:Huichan Seo

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:Vision-language models, increasingly act, decide what users, judge, Vision-language

备注: 25 pages including appendix. Code and project page: [this https URL](https://github.com/seochan99/JudgeActs) ; data: [this https URL](https://huggingface.co/datasets/seochan99/JudgeActs)

点击查看摘要

Abstract:Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random and falls short of a CLIP similarity baseline. It picks the first image shown in 49% of calls (chance: 28%), and reordering changes its choice on 60% of prompts. For this judge, agreement across orders is informative: decisions that survive reordering are much better than random, whereas agreement with a weaker second judge keeps the wrong ones. An 8B judge shows almost no position bias and outperforms CLIP, yet for it the same filter mostly discards good decisions. Agreement helps only when it targets the judge's failure mode, so filters must be re-audited whenever the judge changes. The 4B judge's slight rise in stereotype ratings is no longer detectable after aggregating across orders or with the larger judge.

105. 【2610.01233】Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization

链接:https://arxiv.org/abs/2610.01233

作者:Zhen Zhou,Zhiwei Ning,Puhua Jiang,Sheng Zhang,Yifei Tang,Jie Yang,Xintong Han,Wei Liu,Chunchao Guo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Flow matching, reinforcement learning, practice its reinforcement, largely adapted, Flow

备注:

点击查看摘要

Abstract:Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without explicitly specifying a target velocity field toward preferred samples. In 3D generation, constrained by pretrained model capabilities, rollout diversity, and reward-distribution complexity, directly applying these RL methods yields limited gains in geometric quality. We introduce a forward-process RL method \textbf{Dynamic Homing Optimization (DHO)}, which reformulates negative-trajectory optimization as positive-sample attraction-guided dynamic homing. Specifically, Minimum-Cost Attractive Matching (MAM) assigns each negative sample a distinct positive target, and Time-Aware Dynamic Correction (TDC) then redirects its trajectory toward the target using a remaining-time-aware corrective velocity. Building on asynchronous online DHO, we develop \textbf{Flow3D-Pro}, an image-to-3D geometry generation framework. Experiments show that DHO outperforms representative DPO-, GRPO-, and NFT-style objectives in 3D generation, while Flow3D-Pro produces higher-quality 3D geometry than existing mesh generation methods.

106. 【2610.01229】A Compact Explicit 4D Representation for Dynamic Scenes

链接:https://arxiv.org/abs/2610.01229

作者:Di Yang,Zhihao Li,Yanhai Xiong,Yufei Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:compact dynamic-scene representation, compact dynamic-scene, dynamic-scene representation, representation must retain, appearance needed

备注:

点击查看摘要

Abstract:A compact dynamic-scene representation must retain both the surfaces seen over time and the appearance needed to render them from new viewpoints. We present Sparc4D, a feed-forward autoencoder that encodes a monocular video with known cameras into a sparse 4D scene state. Static features are shared across the clip, while spatially anchored temporal slots compress time-varying features. A sparse decoder produces 2D Gaussian surfels, while stored source pixels preserve fine texture through geometric re-projection. The state includes one full source frame and dynamic-region pixels sampled every fourth frame, alongside learned features and sparse occupancy. For a 32-frame MultiCamVideo clip, it averages 0.95M 32-bit-equivalent values on random windows and 0.92M on the first-32 protocol. On first-32, Sparc4D reaches 21.70\,dB, compared with 20.40\,dB for MoVieS. On randomly placed windows, their PSNR scores are comparable. With stored texture disabled, temporal slots compress the time-varying feature state by a median $4.0\times$ and reduce the mean state from 1.04M to 0.42M values, with essentially unchanged target-view reconstruction quality. Without fine-tuning on real data, Sparc4D transfers to DyCheck and Neu3D, where stored texture improves LPIPS while slightly reducing PSNR.

107. 【2610.01215】AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

链接:https://arxiv.org/abs/2610.01215

作者:Cheng Yang,Yifan Wu,Yutao Huang,Zhaohua Zhang,Beiduo Chen,Muxi Chen,Chenchen Zhao,Hexuan Deng,Haolin Yang,Geyuan Zhu,Sa Zhu,Jianhuan Zhuo,Qiuyong Xiao,Jianhao Ruan,Yiran Peng,Jiayi Zhang,Tian Ye,Xinlei Yu,Tianwen Jiang,Jihong Zhang,Yuyu Luo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:support multi-step workflows, agents require high-quality, GUI agents require, software environments respond, require high-quality interaction

备注:

点击查看摘要

Abstract:GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.

108. 【2610.01210】EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes

链接:https://arxiv.org/abs/2610.01210

作者:Hongming Fu,Jingcheng Shi,Wenjia Wang,Binhua Zuo,Bo Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:occlusion make difficult, recovering hand motion, hand occlusion make, Egocentric video, camera motion

备注:

点击查看摘要

Abstract:Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.

109. 【2610.01206】Resolving Mixed Single-Photon LiDAR Returns for Foreground-View and Hidden Scene Reconstruction

链接:https://arxiv.org/abs/2610.01206

作者:Ziting Wen,Runrong Deng,Zili Zhang,Haitao Zheng,Yuecong Xu,Xiaoqiang Ren,Guodong Shi,Kemi Ding

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:create mixed LiDAR, mixed LiDAR returns, robotic inspection, screens and protective, protective covers

备注:

点击查看摘要

Abstract:Partially transmissive screens and protective covers are common in robotic inspection, but they create mixed LiDAR returns from both the foreground material and the scene behind it. Conventional peak-based LiDAR usually discards weak hidden returns, while single-photon LiDAR records time-resolved histograms that preserve attenuated and overlapping echoes. However, existing transient reconstruction methods typically fit a single scene representation to the measured waveform. Under occlusion, weak or nearby foreground--hidden echoes can form a broad peak or subtle shoulder. Because such waveforms can also be explained by a displaced single surface or a thick density distribution, accurate transient fitting does not necessarily imply correct geometry. We propose a state-aware framework for foreground-view and hidden scene reconstruction from occluded single-photon histograms. For each ray, we estimate local echo evidence, identifying no reliable surface evidence, single-return evidence, or two returns. The inferred echo state routes supervision for a two-head neural field: all rays constrain waveform reconstruction, while reliable anchors provide geometry localization. We also introduce a real paired single-photon LiDAR occlusion dataset with occluded and clean captures at fixed poses. Experiments on a real dataset show improved hidden scene depth and point-cloud accuracy over baselines. Our results demonstrate single-photon layered reconstruction as a practical route for 3D perception through partially transmissive occluders.

110. 【2610.01205】Semantic RGB--Depth Based Surgical Skill Assessment in Microscopic Stereo Videos

链接:https://arxiv.org/abs/2610.01205

作者:Jecia Z. Y. Mao,Sue M. Cho,Francis X. Creighton,Deepa Galaiya,Russell H. Taylor,Manish Sahu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing video-based approaches, video-based approaches predominantly, approaches predominantly rely, microsurgical technical skill, Objective assessment

备注:

点击查看摘要

Abstract:Objective assessment of microsurgical technical skill is essential for competency-based training and quality assurance, yet existing video-based approaches predominantly rely on RGB images and therefore overlook the 3D spatial relationships that characterize instrument-anatomy interactions. Although stereo operating microscopes provide complementary depth information, conventional stereo matching algorithms can produce sparse and unreliable depth estimates under high-magnification imaging conditions, limiting their use for automated skill assessment. This work presents a semantic RGB-Depth framework for surgical skill assessment from microscopic stereo videos. A regression-based depth fusion method combines sparse metric stereo depth with dense monocular depth estimates to generate a dense geometric representation of the surgical scene. This representation is integrated with semantically decomposed RGB streams corresponding to individual surgical instruments and surrounding anatomy. A hierarchical attention architecture jointly encodes these streams to capture discriminative patterns of instrument use and instrument-anatomy interaction across surgeons at different training levels. The framework was evaluated on 33 ex vivo transoral microlaryngeal procedures performed by six surgeons, comprising attending surgeons and surgical residents, using leave-one-surgeon-out cross-validation. The proposed semantic RGB-Depth model achieved an F1 score of 0.938 for skill-level classification, compared with 0.696 for semantic RGB and 0.929 for semantic depth. These results suggest that geometric information can improve automated surgical skill assessment from microscopic stereo videos. The learned spatial, temporal, and semantic attention patterns also support qualitative examination of the scene regions, video segments, and semantic streams emphasized by the model.

111. 【2610.01201】SEE: Object Permanence Through Self-Supervision

链接:https://arxiv.org/abs/2610.01201

作者:Pramish Paudel,Ajad Chhatkuli,Luc Van Gool,Danda Pani Paudel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:predict and plan, Object, keeping track, track, video representations

备注:

点击查看摘要

Abstract:Object permanence, keeping track of an object's identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object's hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot's attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object's position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86% of occlusions, against 32% for SlotContrast, and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM. The two streams also allow downstream planning, with the position stream as the action of a world model. Project page: this https URL

112. 【2610.01192】FlashBack: Knowing When to Remember in Streaming Vision-Language Models

链接:https://arxiv.org/abs/2610.01192

作者:Yi Chen,MingMing Yu,Rui-Qi Wang,Boran Wang,Xiaohang Cao,Chu Tang,Jingmin Chen,Jie Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:bounded compute budget, process continuously growing, continuously growing video, growing video streams, real-time perception

备注:

点击查看摘要

Abstract:Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.

113. 【2610.01191】Color Independent Word Segmentation From Transcribed Bangla Passages

链接:https://arxiv.org/abs/2610.01191

作者:Faias Satter,Noor Masrur,Sk. Md. Masudul Ahsan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:optical character recognition, making people jobs, people jobs easier, character recognition, making people

备注: 6 pages, 8 figures, 6 tables. Accepted version of the paper published in the 2023 6th International Conference on Electrical Information and Communication Technology (EICT)

点击查看摘要

Abstract:An optical character recognition(OCR) system can scan paper and extract text, making people's jobs easier. While numerous OCR systems are accessible in the software sector, finding a dependable equivalent solution for Bangla is tough. When it comes to handwritten texts, the case is even more rare. The first fundamental step to any OCR is to segment words from text images. If this stage fails, the total OCR's performance will be poor no matter how promising the later stages perform. This research aims to segment words in a handwritten Bangla text image. This research can be implemented on any smartphone-captured image, irrespective of the color and type of paper and ink. Furthermore, as smartphone-captured images can create shadow interferences, the custom dataset built for this research is created in such a way that every possible obstacle that can be faced is included. For 7374 words, a total of 7278 bounding boxes are generated, which have recall of 90.60 %, precision of 91.80 %, and F1-score of 91.20 %. The system can be further improved with nested operations on bounding boxes containing several words or by adjusting the adaptive thresholding and dilation filter sizes to a more precise level.

114. 【2610.01180】Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models

链接:https://arxiv.org/abs/2610.01180

作者:Yuliang Cai,Mohammad Rostami,Jesse Thomason

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:involve negated clauses, produce incorrect answers, models consistently struggle, questions involve negated, negated clauses

备注:

点击查看摘要

Abstract:Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.

115. 【2610.01166】CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

链接:https://arxiv.org/abs/2610.01166

作者:Kunyang Li,Hai Nguyen,Joshua Lowe,Chenguang Zhao,Peace C. Madueme,Mehdi Hedjazi Moghari,Mubarak Shah,Pegah Khosravi,Yuzhang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Cardiovascular magnetic resonance, Cardiovascular magnetic, including cine imaging, Cine CMR, magnetic resonance

备注: Code, benchmark resources, and model weights are available at [this https URL](https://github.com/AI-MIND-Lab/CineMR)

点击查看摘要

Abstract:Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at this https URL.

116. 【2610.01162】PhysicsLENS: Diagnosing Physical Property Blindness in Video Generation Models

链接:https://arxiv.org/abs/2610.01162

作者:Isaiah Milkey,Som Sagar,Aditya Taparia,Xinyuan Liu,Jiqing Wen,Ransalu Senanayake

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Reliable video world, provide scalable predictive, scalable predictive environments, Reliable video, robot learning

备注:

点击查看摘要

Abstract:Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, generated robot videos can violate physical principles and complete tasks through physically implausible behavior, limiting their reliability for robot learning and planning. Current video-generation benchmarks exclude physics that are inherently hidden by visuals (e.g., weight, viscosity, friction). Due to this, video models are evaluated on the fidelity of physics, not the underlying accuracy of physics. We introduce PhysicsLENS, a dataset and benchmark for evaluating plausibility of physical properties grounded in robotics. PhysicsLENS uses matched scenario pairs that hold the same conditioning frame and task, while varying underlying physics in the scene description. Scenarios are curated from public robot video sources and annotated across seven physical domains: collision, gravity, momentum, friction, deformation, fluid, and causality. We evaluate across four video generation models, producing over 400 human-annotated labels. Results show that plausible-looking videos often ignore the stated property (34 of 47), and that stating the property lowers plausibility only slightly and not significantly.

117. 【2610.01148】OptimusMesh: Compact Autoregressive Mesh Generation from Point Clouds via Sparse Latent Pivots

链接:https://arxiv.org/abs/2610.01148

作者:Mazhar Iqbal,Naoya Chiba,Xuanmeng Sha,Tomohiro Mashita,Yuki Uranishi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:point clouds remains, geometrically faithful, fundamental challenge, Generating compact, remains a fundamental

备注:

点击查看摘要

Abstract:Generating compact and geometrically faithful 3D meshes directly from point clouds remains a fundamental challenge. Point clouds are unordered and sparse, whereas meshes exhibit irregular structure and varying topology. As a result, many existing approaches rely on implicit representations followed by surface extraction or reconstruction. Although effective, these pipelines can produce dense or over-smoothed meshes, often requiring computationally expensive post-processing and simplification. We present OptimusMesh, a framework for direct compact triangle mesh generation from point clouds using sparse latent pivot conditioning. Our key idea is to compress $2{,}048$ oriented input points into only $16$ sparse latent pivots, reducing the geometric conditioning set by $128\times$. These pivots provide a compact structural representation shared across a two-stage autoregressive framework that first generates mesh vertices and then predicts triangular faces conditioned on the generated vertices and the same pivots. Compared with the evaluated recent point-cloud-conditioned autoregressive methods, which use $257$ decoder-conditioning tokens, OptimusMesh uses only $16$, yielding a $16.1\times$ shorter conditioning sequence. Experiments show that OptimusMesh produces the most compact outputs among the compared recent autoregressive methods, using $25.7\%$--$94.1\%$ fewer faces while maintaining competitive geometric fidelity and distributional quality.

118. 【2610.01135】he RSNA Intracranial Aneurysm (RSNA-ICA) Dataset

链接:https://arxiv.org/abs/2610.01135

作者:Maria Correia de Verdier,Rachit Saluja,Jason Sho,Maryam Vabarizad,Rennie Yung-Chieh Chen,Uyen N. T. Nguyen,Mona Alrehaili,Layal Aweidah,Deniz Bulja,Wesley C. Chan,Hernan Chaves,Madhavi Duvvuri,Huseyin Ekin Ergin,Undrakh-Erdene Erdenebold,Ekim Gumeler,Mohamed Sobhi Jabal,Chin-Chi Kuo,Fatima Mubarak,Sevde Nur Emir,Scott Riley K. Ong,Johanna Ortiz,Almudena Pérez-Lara,Andreas M. Rauschecker,Shayan Sirat Maheen Anwar,Charit Tippareddy,Tam Tran,Sorawis Visrutaratna,John Mongan,Adam E. Flanders,Robyn Ball,Greg Zaharchuk,Peter D. Chang,Felipe Kitamura,Errol Colak,Luciano Prevedello,Tyler Richards,Data Contributor Group,Dataset Annotator Group,Evan Calabrese,Jeffrey D. Rudie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:RSNA Intracranial Aneurysm, intracranial aneurysm detection, Intracranial aneurysm, Intracranial aneurysm rupture, detection remains challenging

备注: 48 pages (including supplementary material)

点击查看摘要

Abstract:Intracranial aneurysm rupture is associated with substantial morbidity and mortality, yet aneurysm detection remains challenging, particularly for small lesions and on routine non-angiographic imaging examinations. To support the development and evaluation of artificial intelligence (AI) algorithms for intracranial aneurysm detection and localization, the Radiological Society of North America (RSNA), in collaboration with the American Society of Neuroradiology (ASNR), the Society of Neurointerventional Surgery (SNIS), and the European Society of Neuroradiology (ESNR), curated the RSNA Intracranial Aneurysm (RSNA-ICA) Dataset. Developed for the 2025 RSNA Intracranial Aneurysm Detection Challenge, RSNA-ICA is a large, publicly available, expert-annotated dataset comprising 7202 CTA, MRA, and MRI series from 4278 adult patients collected across 21 institutions in 12 countries spanning five continents. The dataset includes 2566 CTA, 2166 MRA, and 2470 MRI series from patients with and without intracranial saccular aneurysms, providing substantial geographic and imaging diversity. Expert annotations indicate both aneurysm presence and location, and 178 series additionally include three-dimensional segmentations of challenge-defined vascular locations. RSNA-ICA was used to develop and evaluate algorithms in the 2025 RSNA Intracranial Aneurysm Detection Challenge. Of the 7202 image series, 5041 are publicly available through MIRA (this https URL), while the remainder were used for challenge public and private test sets. The dataset is freely available to the research community for noncommercial use and provides a comprehensive resource for advancing AI-based aneurysm detection across both angiographic and routine neuroimaging examinations.

119. 【2610.01134】Open Vocabulary Word Recognition From Transcribed Bangla Texts

链接:https://arxiv.org/abs/2610.01134

作者:Faias Satter,Sk. Md. Masudul Ahsan

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:optical character recognition, making people jobs, people jobs easier, handwritten Bangla word, character recognition

备注: 6 pages, 4 figures, 5 tables. Accepted version of the paper published in the 2023 26th International Conference on Computer and Information Technology (ICCIT). Code: [this https URL](https://github.com/FaiasPromit/Optical-Character-Recognition-From-Handwritten-Bangla-Texts)

点击查看摘要

Abstract:An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are available in the software industry, finding a reliable equivalent solution for Bangla takes much work. When it comes to handwritten texts, the situation is much more unusual. Recognizing words from word images is the most critical stage in any OCR process. It is the second stage after segmenting words from text pictures. If this stage fails, the overall performance of the OCR will be poor, regardless of how well the other phases perform. This study aims to recognize words using deep learning in a handwritten Bangla word image. Three object detection models, SSD with MobileNetV2, Faster R-CNN with InceptionResNetV2, and an ensemble model of these two, have been used to train and test handwritten word images. A modified Non-Maximum Suppression has been introduced to enhance the effectiveness of the models' results. A customized dataset of 9841 handwritten Bangla word images has been compiled, featuring diverse handwriting styles from various individuals. All three models' performances have been checked against the test dataset, and the ensemble model has been the most impressive, with an F1-score of 92.61%. Also, at the word level, the ensemble model correctly recognizes 96.12% of the words to some extent. The system can be further improved by introducing a post-processing phase to correct errors generated by the system.

120. 【2610.01114】Affine-Aligned Atlas for Canonical Gaussian Construction in Video Representation

链接:https://arxiv.org/abs/2610.01114

作者:Masaya Takabe,Hiroshi Watanabe,Sujun Hong,Tomohiro Ikai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:fast rendering capability, rendering capability, Gaussian splatting, canonical Gaussians, canonical Gaussian

备注:

点击查看摘要

Abstract:Gaussian splatting has recently emerged as an efficient representation for images and videos due to its explicit structure and fast rendering capability. Existing Gaussian-based video representations often decompose a video into canonical Gaussians and temporal deformation. However, when a video contains large global motion such as camera movement, the canonical representation may become misaligned with individual frames, increasing the burden on the temporal deformation model. In this paper, we propose an affine-atlas canonical Gaussian representation, which constructs canonical Gaussians in a larger affine-aligned atlas space. Frame-wise affine transforms absorb global motion before canonical Gaussian construction, reducing the gap between the canonical representation and target frames. Since the proposed method only modifies the canonical construction stage, it can be integrated into existing canonical-Gaussian-based methods with negligible additional parameter cost. Experiments show that our method improves reconstruction quality especially for sequences with large camera motion.

121. 【2610.01098】MVDG: Efficient Multi-view 3D Disambiguation on Unconstrained Real-World Images

链接:https://arxiv.org/abs/2610.01098

作者:Hanyuan Xiao,Gonglin Chen,Haolin Xiong,Wenbin Teng,Haiwei Chen,Yajie Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Illusory matches, visually similar, obstacle for large-scale, reconstruction and visual, visual localization

备注:

点击查看摘要

Abstract:Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inference complexity for downstream structure-from-motion (SfM). We present MVDG, a scalable multi-view disambiguation framework built on the 3D foundation model VGGT, which jointly reasons over an arbitrary number of multiview images. By incorporating 3D-aware multi-view features, our method reduces dependence on pairwise comparisons by encoding and decoding views in a single pass. We further observe that direct multi-view fine-tuning of VGGT can be unstable under noisy supervision; motivated by label ambiguity in Doppelgangers, we construct a pseudo-pairwise training set from AerialMegaDepth and show that fine-tuning on sampled subsets yields stable optimization and strong generalization to held-out scenes. Finally, because full SfM evaluation (even with faster pipelines such as GLOMAP) remains expensive, we process a pseudo-pairwise dataset for efficient validation; we derive a predictive relationship between regular SfM metrics and the classification accuracy on this pseudo-pairwise test. Experiments show that our method achieves comparable pairwise accuracy while improving both SfM accuracy and inference speed over baselines.

122. 【2610.01096】Dataset Identity, Not Novelty: The Source of an Inflated OOD Detection Gain

链接:https://arxiv.org/abs/2610.01096

作者:Donghoon Lee,Shinjin Kang

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:OOD dataset, OOD, trained classifier, OOD data, dataset

备注: 30 pages, 7 figures, 33 tables

点击查看摘要

Abstract:A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and returns a score. It fits that score on in-distribution data, and the benchmarks that evaluate it supply a second piece of OOD data for the fitting itself. Some detectors tune a constant on it. Others fit a direction in feature space or train a flexible combiner and report the number that it reaches as the gain that is still available. Every such fit is validated on held-out samples of the same OOD dataset. That check rules out memorizing individual images. It says nothing about a fit that has instead learned which dataset it is looking at, and a direction that recognizes one OOD dataset rather than novelty passes it perfectly. The detector that a practitioner installs meets OOD data from a source that nobody fitted it on, so the difference decides what the reported number is worth. We measure it by holding out the whole OOD dataset rather than a sample of it, and we call that gap the inflation. We read it across a range of combiners on ImageNet and CIFAR-100 backbones. Most of the gain that the usual protocol reports turns out to be dataset identity rather than novelty. The size of the fit does not move what survives, so the effect is not ordinary overfitting. The share depends instead on whether the input exposes class identity, and two controls that vary that property alone separate the inflation on every backbone of both benchmarks. A closed form accounts for the effect and computes it from the fitting rows, so a practitioner can tell which fits will inflate without running the hold-out protocol. One of these fits survives, namely the single constant that the field already picks on a designated validation dataset. Anything above it reports a gain that the hold-out protocol does not return, and on one benchmark what survives falls while what is reported climbs.

123. 【2610.01092】Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

链接:https://arxiv.org/abs/2610.01092

作者:Patrick Amadeus Irawan,Iskandar Muda Rizky Parlambang,Rava Maulana,Qinrong Cui,Erland Hilman Fuadi,Zayd M. K. Zuhri,Nanda Ryaas Absar,Ahmed Elshabrawy,Wilfried Ariel Mulyawan,Shoubin Yu,Yue Zhang,Mohit Bansal,Alham Fikri Aji

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:increasingly being explored, simulators for embodied, Video generation, Video generation models, embodied planning

备注: Preprint. 51 pages, 19 figures, 23 tables. Code, dataset and project website linked in the paper

点击查看摘要

Abstract:Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.

124. 【2610.01069】Overcoming Kernel Redundancy for Scaling Logic Gate Networks

链接:https://arxiv.org/abs/2610.01069

作者:Sejin Park,Hongjae Lee,Changwoo Han,Seung-Won Jung

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Differentiable logic gate, recently attracted attention, logic gate networks, Differentiable logic, logic gate

备注: NeurIPS 2026

点击查看摘要

Abstract:Differentiable logic gate networks, which operate using only logic gates, have recently attracted attention as an efficient alternative to conventional neural networks. However, despite their efficiency, the scaling behavior of logic gate networks remains underexplored. By contrast, scaling model capacity is a central design principle in deep neural networks and typically leads to improved performance. This discrepancy raises a key question: Can similar scaling benefits also be achieved in logic gate networks? In this work, we focus on width as a primary scaling axis and conduct a systematic analysis of its behavior in logic gate networks. We observe that naive width scaling often introduces redundancy among logic kernels, limiting the effective use of additional kernels and leading to performance saturation. To address this limitation, we propose a dynamic logic kernel framework that reorganizes kernel utilization by promoting specialization across kernel groups. This enables the network to better utilize increased width via input-dependent kernel routing, while ensuring that both routing and computation are implemented entirely with gate-level Boolean operations at inference time. We further find that kernel redundancy is most pronounced at the first gate level, motivating an early-stage dynamic logic kernel strategy that concentrates adaptation at this level. Experimental results demonstrate that our approach improves kernel utilization and increases kernel diversity, leading to higher accuracy with improved parameter efficiency.

125. 【2610.01056】HierGF: Hierarchical Gaussian Fields via Geometry-perception Message Passing for Sparse-view 3D Reconstruction

链接:https://arxiv.org/abs/2610.01056

作者:Bi'an Du,Zhimin Zhang,Daizong Liu,Baoquan Chen,Wei Hu

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:cultural heritage digitization, randomly captured views, virtual reality, multimedia applications, augmented reality

备注: Accepted to IEEE Transactions on Multimedia (TMM), 2026

点击查看摘要

Abstract:Sparse view 3D reconstruction is an important and common scenario in multimedia applications, such as augmented reality/virtual reality (AR/VR) content creation, cultural heritage digitization, and certain robotic applications, where only a limited number of randomly captured views may be available. However, sparse views contain only limited 3D information, posing two major challenges:1) too few images are available for matching, making it difficult to build multi-view consistency; 2) insufficient view coverage leads to a lack of information in under-sampled regions, resulting in missing parts of object structure. Existing methods mostly still rely on limited reprojection errors and regularization terms, which are prone to overfitting to a single view and inconsistent appearances across views. In geometrically under-sampled regions, they often rely on heuristic density control, lacking reliable guidance and often resulting in blurring and structural this http URL address these issues, this paper proposes Hierarchical Gaussian Fields (HierGF), which revisits sparse-view reconstruction from a hierarchical geometry-perception perspective and converts limited observations into reliable self-generated supervision beyond fixed priors and heuristic density control. In particular, we transform coarse 3D geometric information and additional 2D generative priors into structured pseudo-supervision through a two-stage geometry-perception backbone network, thereby enhancing multi-view consistency with very few input views. In addition, we introduce a learnable confidence network to guide gradients toward cross-view consistent content, and a geometrically consistent densification module to improve the reconstruction of multi-view alignment and under-sampled regions.

126. 【2610.01052】owards Subject Consistency over Dynamic Subject Sets in Video Generation

链接:https://arxiv.org/abs/2610.01052

作者:Tongcheng Zhang,Jun Zhu,Jianfei Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dynamic subject sets, longer durations, dynamic subject, subject sets, inconsistency events

备注: Project website: [this https URL](https://dynsc-paper.pages.dev/)

点击查看摘要

Abstract:We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82\% for Wan-2.1-1.3B and 5.66\% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.

127. 【2610.01039】Bootstrapping Video Interaction Generation with Synthetic State Transitions

链接:https://arxiv.org/abs/2610.01039

作者:Jiho Jang,Jinyoung Kim,Nojun Kwak,Kyungjune Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:resulting state transitions, recent video generative, synthesize high-fidelity videos, portray plausible physical, synthesize high-fidelity

备注: IJCAI 2026

点击查看摘要

Abstract:While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit `start' and `end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.

128. 【2610.01022】owards Automatic Video Annotation with ASH: Zero-Shot Open-Vocabulary Multi-Object Tracking and Segmentation

链接:https://arxiv.org/abs/2610.01022

作者:Arash Rocky,Q. M. Jonathan Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:design makes multi-category, makes multi-category open-vocabulary, Video Instance Segmentation, tracking computationally prohibitive, multi-category open-vocabulary tracking

备注:

点击查看摘要

Abstract:Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory requirements confine them to short video clips and their single-prompt inference design makes multi-category open-vocabulary tracking computationally prohibitive. This work introduces two contributions toward fully automated tracking annotation of arbitrary video. The Generalized Presence Token (GPT) reformulates SAM3's inference pipeline to process N text prompts simultaneously via virtual prompt batching, reducing image encoding cost from O(N) to O(1) with no modifications to any learned component. The Annotation and Segmentation Handler (ASH) extends any memory-attention VIS tracker to sequences of arbitrary length through overlapping temporal chunks with IoU-based inter-chunk identity matching, requiring no dataset-specific training. Instantiated on SAM3, the resulting pipeline -- SAM3-ASH -- achieves state-of-the-art HOTA on MOTS20 under fully zero-shot conditions and remains competitive with trained specialists across seven additional benchmarks, while peak GPU memory consumption stays below 25 GB, establishing a practical baseline for scalable, training-free automated video annotation.

129. 【2610.01019】FutureWorlds: Learning Robotic World Models from Alternative Futures

链接:https://arxiv.org/abs/2610.01019

作者:Hao Wu,Shengju Qian,Weiyan Wang,Fan Xu,Fan Zhang,Yuanpeng He,Qingsong Wen,Yuxuan Liang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:understanding action outcomes, Robotic world models, predict action-conditioned future, models predict action-conditioned, Robotic world

备注: 32 pages, including references and appendix. Code: [this https URL](https://github.com/Alexander-wu/FutureWorlds)

点击查看摘要

Abstract:Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: this https URL.

130. 【2610.01013】VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction

链接:https://arxiv.org/abs/2610.01013

作者:Junyi Wu,Fanqing Kong,Leyang Chen,Shaoqiu Zhang,Yulun Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieved remarkable progress, unifying camera estimation, vision models, remarkable progress, unifying camera

备注: 21 pages, including references and appendices

点击查看摘要

Abstract:Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. To address these limitations, we introduce VASC, a training-free sparse attention method combining value-aware block selection and execution-aware cross-layer memory. Our value-aware block selection integrates pooled query--key relevance with neighboring value contrast, reducing redundancy while preserving query-relevant and distinctive content. Cross-layer memory tracks unserved demand across layers and updates this state according to actual execution, enabling previously underserved blocks to compete under a fixed computation budget. Experiments on 7Scenes and NeuralRGB-D with VGGT and $\pi^3$ demonstrate improved pose estimation and reconstruction quality compared with FasterVGGT, together with up to $2.29\times$ faster inference than dense VGGT. Code is available at this https URL.

131. 【2610.01012】Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning

链接:https://arxiv.org/abs/2610.01012

作者:Gunwoo Lee,Yoori Oh,Yoseob Han

类目:Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)

关键词:ensuring phonetic accuracy, silent talking-face videos, silent talking-face, ensuring phonetic, synthesis aims

备注: Accepted to BMVC 2026

点击查看摘要

Abstract:Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: this https URL

132. 【2610.00994】VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

链接:https://arxiv.org/abs/2610.00994

作者:Xianda Du,Max Ku,Weiming Ren,Zhi Rui Tam,Chunlin Ren,Ping Nie,Min-Hung Chen,Wenhu Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Existing synthetic image, Existing synthetic, synthetic image evaluators, image evaluators typically, scalar quality score

备注: Preprint. Project page: [this https URL](https://tiger-ai-lab.github.io/VIEScore2/)

点击查看摘要

Abstract:Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.

133. 【2610.00981】NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields

链接:https://arxiv.org/abs/2610.00981

作者:Shota Kobayashi,Koki Seno,Daichi Yashima,Komei Sugiura

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:leveraging data collected, serve as embodiment-agnostic, motion-centric representations, representations for leveraging, robot flows

备注: Accepted at ACCV 2026

点击查看摘要

Abstract:We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at this https URL

134. 【2610.00973】Concept Driven Domain Adaptation: Finding an Abstract Needle in a Haystack

链接:https://arxiv.org/abs/2610.00973

作者:Haiming Zhao,Tai Wang,Kun Zhang,Xicheng Peng,Zhiyang Li

类目:Computer Vision and Pattern Recognition (cs.CV); Physics Education (physics.ed-ph)

关键词:Science teachers frequently, Science teachers, teachers frequently search, intend to teach, teachers frequently

备注: 19 pages, 10 figures

点击查看摘要

Abstract:Science teachers frequently search for documentary excerpts not by describing what appears on screen, but by querying the abstract concepts they intend to teach. This use case exposes a limitation of existing language-based video moment retrieval methods, which typically assume that queries describe observable events, whereas instructional search requires retrieving concrete visual phenomena that instantiate an underlying scientific principle. We study this setting as concept-to-example video retrieval, an abstract-needle-in-a-haystack problem where compact curriculum concepts must be grounded in temporally sparse documentary evidence. To bridge this abstraction gap, we propose Concept-Driven Domain Adaptation (CDDA), a three-stage framework for adapting two-tower vision-language models to concept-level retrieval. CDDA treats concepts as intermediate semantic anchors: it first structures the textual embedding space with textbook and teacher-handbook example-concept pairs, then transfers this concept-aware geometry to documentary visuals under a frozen visual encoder, and finally jointly adapts both encoders with sparse visual concept supervision. From a geometric perspective, this staged alignment reduces text-concept and vision-concept angular gaps, thereby encouraging concept-level adaptation while preserving the pretrained model's concrete image description alignment. On a curated middle-school physics retrieval benchmark, CDDA achieves stronger pedagogically oriented concept retrieval than several competitive multimodal baselines, including Qwen3-VL-Embedding-2B, while maintaining concrete image-text matching after adaptation.

135. 【2610.00970】RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation

链接:https://arxiv.org/abs/2610.00970

作者:Minsu Kim,Jaesung Choe,Jiwoo Lee,Yu-Chiang Frank Wang,Seon Joo Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:existing methods remain, methods remain confined, semantic scene understanding, neglecting spatial relations, Recent advances

备注: 10 pages, NeurIPS 2026 accepted (poster)

点击查看摘要

Abstract:Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.

136. 【2610.00960】Video-Index: A Curated Meta-Benchmark for Video Understanding

链接:https://arxiv.org/abs/2610.00960

作者:Enxin Song,Yinuo Xu,Shusheng Yang,Wenhao Chai,Jiatao Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:exploit answer options, partial visual evidence, claims to measure, answer options, visual evidence

备注: Blog: [this https URL](https://www.enxinsong.com/blog/video-index/) GitHub: [this https URL](https://github.com/Espere-1119-Song/Video-Index) Hugging Face: [this https URL](https://huggingface.co/datasets/Video-Index/Video-Index)

点击查看摘要

Abstract:A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: this https URL GitHub: this https URL Hugging Face: this https URL

137. 【2610.00953】wo Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold

链接:https://arxiv.org/abs/2610.00953

作者:Keuntae Kim,Yong Suk Choi

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:masked diffusion MLLM, diffusion MLLM, MLLM can stabilize, masked diffusion, MLLM

备注: NeurIPS 2026 Workshop on BeNTo (Beyond Next-Token Prediction - Diffusion Flow Models for Next-Generation Decoding)

点击查看摘要

Abstract:An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On V*Bench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron's overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/V*Bench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.

138. 【2610.00952】A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

链接:https://arxiv.org/abs/2610.00952

作者:Giyeong Oh,Junghun Park,Yuhan Bae,Youngjae Yu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:replacing sparse alt-text, captioners replacing sparse, language model, Recaptioned image-text corpora, dense descriptions

备注: initial commit

点击查看摘要

Abstract:Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($\pi$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $D_{\pi,V_c,C}$: at a fixed text budget of $B = 64$ it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by $+3.39$ to $+6.36$ under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ($\approx$ 490M) together with the audit-artifact bundle.

139. 【2610.00930】Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance

链接:https://arxiv.org/abs/2610.00930

作者:Mingrun Jiang,Yuejia Liu,Zishan Shao,Ting Jiang,Qinsi Wang,Hancheng Ye,Yixiao Wang,Rui-Feng Wang,Kangning Cui,Yixuan Chen,Fan Yang,Xiang Cheng,Hai Li,Yiran Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:increasingly exploits timestep, models increasingly exploits, Post-training quantization, exploits timestep, increasingly exploits

备注:

点击查看摘要

Abstract:Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating CFG structure into diffusion quantization, activation quantization still operates independently across conditional and unconditional coordinates, leaving cross-activation structure unexploited. We show that matched CFG activations form a strongly correlated two-dimensional source and that, under a fixed bit budget, the choice of branch coding basis materially affects quantization fidelity. Motivated by this observation, we introduce branch-space transform coding, which rotates matched CFG branches via an offline derived 2x2 orthogonal matrix, requiring minimal modifications to model parameters or the quantization pipeline. We further derive the Guidance-Correlation Branch Transform (GCBT), which jointly incorporates the CFG guidance direction and cross-branch second moments. Under an equal-rate quantization-noise surrogate, GCBT admits a closed-form per-layer solution without gradient optimization or angle search. Applied on top of existing diffusion PTQ methods, GCBT yields statistically significant fidelity gains in most evaluated comparisons with no statistically significant degradation, while leaving the underlying host quantization pipeline unchanged.

140. 【2610.00929】Platonic Task Arithmetic

链接:https://arxiv.org/abs/2610.00929

作者:Junghwan Park,Woojin Cho

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:common coordinate system, arithmetic stays confined, weight-space task arithmetic, task arithmetic stays, similar behavior

备注: NeurIPS2026

点击查看摘要

Abstract:Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so weight-space task arithmetic stays confined to a single model and cannot cross architectures without a structural correspondence. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows of one shared, model-agnostic object, which we call the platonic task vector. To make it operational for models that pair an image or audio encoder with a text encoder, we introduce Universal Task Descriptors: matrices whose shape is independent of architecture and embedding dimension, which record a task's functional effect and support addition and negation as matrix operations. Transferring a descriptor into a target means editing the target until it reproduces the descriptor on the task's unlabeled probe images and class-name prompts, requiring no per-image labels. We realize this edit in two ways. First, the descriptor factorizes into a shift field on image embeddings, so a single least-squares solve yields a linear operator that folds into the target's last layer as a weight edit; by linearity, a bank of such operators admits any composition at any strength as a signed sum. Second, a low-rank adapter trained on the same objective reaches every layer and fits compositions jointly, at the cost of one optimization per edit. Heterogeneous models share this object only partially, with a model-specific residual comparable in norm to the shared component, yet cross-model transfer still retains 74-80 percent of the gain of the target's own descriptors. Experiments across six model families, eight classification tasks, and an audio-text setting show that task knowledge transfers and composes across heterogeneous models under both realizations.

141. 【2610.00926】A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform

链接:https://arxiv.org/abs/2610.00926

作者:Chengkai Xu,Yiming Cui,Jiaqi Liu,Yicheng Guo,Cheng Qin,Geyuan Zhang,Xinwei Dong,Shiyu Fang,Peng Hang,Jian Sun

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:directly maps multimodal, maps multimodal sensory, multimodal sensory inputs, unified differentiable models, Autonomous driving

备注: 21 pages, 6 figures, accepted by IEEE transactions on intelligent transportation systems

点击查看摘要

Abstract:Autonomous driving is a cornerstone technology for the future of intelligent transportation, where end-to-end learning has emerged as a transformative paradigm that directly maps multimodal sensory inputs to driving actions through unified differentiable models. While offering advantages, the effectiveness of end-to-end autonomous driving (E2E-AD) is ultimately determined by the quality of its training ecosystem. This paper provides a comprehensive review of training methods and ecosystem for E2E-AD. We introduce a Data-Strategy-Platform taxonomy that conceptualizes training as an interdependent system. The data layer defines what can be learned, the strategy layer governs how learning aligns with driving objectives, and the platform layer supports scalability and continuous evolution. Within this framework, we survey recent advances across data-centric pipelines, learning paradigms, and training infrastructures, and analyze their interplay in shaping model performance, robustness, and deployability. Finally, we reflect on current limitations and articulate a forward-looking vision that emphasizes a shift from data quantity to data value, from isolated optimization to foundation-driven generalization, and from static training to integrated training-testing loops, aiming toward robust, scalable, and trustworthy autonomous driving systems. We maintain a continuously updated repository tracking cutting-edge literature and works at \href{this https URL}{Our Project Page}.

142. 【2610.00922】EyeTAG: Eye Trajectory-Aware Gaze Estimation

链接:https://arxiv.org/abs/2610.00922

作者:Jungmin Lee,Niamat Ullah,Yoseob Han

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:natural head-eye motion, head-eye motion underpins, motion underpins applications, human-computer interaction, natural head-eye

备注: Accepted to BMVC 2026

点击查看摘要

Abstract:Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0$^\circ$ on Gaze360 and performs on par with the strongest baseline on EVE (2.56$^\circ$ vs. 2.58$^\circ$). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at this https URL.

143. 【2610.00895】owards Fast and Disentangled Counterfactuals for Visual Foundation Models

链接:https://arxiv.org/abs/2610.00895

作者:Sidney Bender,Benedikt Kunz,Ahmed Zeid,Shinichi Nakajima,Klaus-Robert Müller,Marco Morik

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Clever Hans strategies, Clever Hans, Hans strategies, Foundation models, models remain vulnerable

备注:

点击查看摘要

Abstract:Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.

144. 【2610.00881】Machine Translation for Sign Languages

链接:https://arxiv.org/abs/2610.00881

作者:Ozge Mercanoglu Sincan,Anton Pelykh,Edward Fish,Harry Walsh,JianHe Low,Karahan Sahin,Oline Ranum,Sobhan Asasi,Steven Emery,Richard Bowden

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:past decade, evolving from isolated, Sign language machine, progressed substantially, isolated sign recognition

备注: Accepted for publication in the Annual Review of Linguistics, Volume 13

点击查看摘要

Abstract:Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation systems. Advances in pose estimation, transformer architectures, and large-scale dataset collection have driven progress, yet challenges remain. Datasets are limited compared to spoken-language resources; evaluation metrics inadequately capture the linguistic quality of output; and models must capture the simultaneous, multi-layered, and three-dimensional structure of sign languages. This manuscript provides a comprehensive review that seeks to balance technical challenges with stakeholder considerations. We examine the linguistic properties that make sign languages computationally unique, trace the evolution of recognition, translation, and production systems, and analyze ongoing technical challenges. Crucially, we address ethical considerations around data governance, community involvement, and appropriate use. Drawing on interdisciplinary perspectives spanning computer vision, sign language linguistics, and deaf studies, our analysis emphasizes that continued progress requires sustained collaboration across these fields and with deaf communities.

145. 【2610.00878】UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

链接:https://arxiv.org/abs/2610.00878

作者:Pengfei Qi,Haoran Lin,Sizhuang Chen,Kai Luo,Sirui Zhang,Xinqi Liu,Fei Cheng,Wenrui Chen,Liming Yin,Kailun Yang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:initial target azimuths, arbitrary initial target, General-purpose embodied robots, General-purpose embodied, dynamic person tracking

备注: The project page is at [this https URL](https://tw5775.github.io/UniTrackPLA)

点击查看摘要

Abstract:General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking. Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency. We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments. The project page is at this https URL.

146. 【2610.00864】Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

链接:https://arxiv.org/abs/2610.00864

作者:Jiawei Fan,Sifeng Wang,Yuqing Hou,Anbang Yao

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Robotic Foundation Models, Foundation Models, Robotic Foundation, aiming to overcome, generation in Robotic

备注: Project page: [this https URL](https://github.com/IntelChina-AI/K-MF)

点击查看摘要

Abstract:In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at this https URL.

147. 【2610.00861】Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints

链接:https://arxiv.org/abs/2610.00861

作者:Melika Shirian,Kianoosh Vadaei

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:budget requires deciding, requires deciding, perturbation, budget requires, limited budget

备注:

点击查看摘要

Abstract:Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, which restrict the extent to which perturbation can be concentrated on a small number of locations. We introduce an importance-guided allocation mechanism that uses a fixed clean-gradient prior to steer perturbation toward model-sensitive regions while leaving the feasible perturbation set unchanged. A centered allocation objective encourages perturbation at above-average importance locations and discourages unnecessary expenditure elsewhere, thereby redistributing rather than enlarging the available budget. Across ten robust model--dataset configurations under a common capacity-limited threat setting, the proposed method improves attack success over matched APGD- and PMA-based baselines by $2.52$ to $17.70$ percentage points. Allocation analysis shows that these gains are accompanied by substantially greater perturbation mass in high-importance regions without increased global $\ell_1$ consumption. Mechanism ablations further show that centered non-uniform redistribution provides part of the benefit, while model-derived importance yields an additional improvement. These results identify perturbation allocation as a distinct and practically relevant dimension of adversarial optimization under shared-budget, locally constrained threat models.

148. 【2610.00859】CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

链接:https://arxiv.org/abs/2610.00859

作者:Chensheng Peng,Wenhao Ding,Ran Tian,Zewei Zhou,Jef Packer,Maximilian Igl,Peter Karkus,Yan Wang,Masayoshi Tomizuka,Boris Ivanovic,Marco Pavone,Yuxiao Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:jointly predict actions, jointly predict, action, actions, video

备注: Project page: [this https URL](https://ctrl-wam.github.io/)

点击查看摘要

Abstract:World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: this https URL

149. 【2610.00855】Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

链接:https://arxiv.org/abs/2610.00855

作者:Cigdem Kokenoz,Amir Salarpour,Alkim Domeke,Christopher Salas,Pedram MohajerAnsari,Long Cheng,Mert D. Pesé,Bing Li

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:safe autonomous navigation, autonomous navigation, perception is critical, critical for safe, safe autonomous

备注: 9 pages, 3 figures, 4 tables. Submitted to IEEE ICRA 2027

点击查看摘要

Abstract:Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.

150. 【2610.00851】SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing

链接:https://arxiv.org/abs/2610.00851

作者:Thiru Thillai Nadarasar Bahavan,Yu Xia,Sachith Seneviratne,Saman Halgamuge

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Open Set Recognition, Open Set, Set Recognition, spherical representation learning, aims to enable

备注:

点击查看摘要

Abstract:Open Set Recognition (OSR) aims to enable models to accurately classify known classes while rejecting samples from unseen classes. A key challenge in OSR lies in the inability to model the unbounded distribution of unknown classes during training, often leading to the misclassification of samples from these classes. Rather than modeling unknowns, recent work shapes the feature space so that known classes are compact and well separated, and spherical representation learning methods have achieved strong results this way. Label smoothing has been identified as one of the key drivers of this success, yet it applies the same coefficient to every training sample, regardless of how well each sample is already embedded. We show that the spherical representation learning objectives used in OSR share a single alignment--uniformity structure in which labels enter only through the alignment term. Label smoothing therefore acts as an alignment dial, and a fixed coefficient sets this dial to the same value for every sample. We propose a plug-in, SmoothOperator (SmoothOP), which sets the smoothing coefficient of each sample from its \textbf{prominence}, an embedding-space signal measuring how clearly the sample's own class stands out against its strongest competing class. Our method integrates into four existing spherical representation learning methods at minimal training overhead. SmoothOP assigns strong smoothing to samples with high prominence, which reduces their alignment and relaxes their pull. On the Semantic Shift Benchmark, SmoothOP-augmented variants generally outperform their base objectives across datasets, degrees of semantic shift, and OSR post-processors, with gains of up to 4.7\% in AUROC, OSCR, and closed-set accuracy.

151. 【2610.00848】Geometric Similarity in VLM Low-Level Vision Representations

链接:https://arxiv.org/abs/2610.00848

作者:Shao-Jun Xia,Huixin Zhang,Zhen Lei,Anlan Sun,Yuner Zhang,Xiaoyang Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:architectures including autoregressive, representative architectures including, universal vision backbones, including autoregressive, diffusion transformers

备注: First version: 10 pages

点击查看摘要

Abstract:Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.

152. 【2610.00825】Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

链接:https://arxiv.org/abs/2610.00825

作者:Rui Liu,Bhavin Jawade,Haoqi Li,Shivam Mehta,Karan Saxena,Yinghong Lan,Cameron R. Wolfe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:quality control requires, speaker visible articulation, text line matches, Dubbing quality control, candidate text line

备注:

点击查看摘要

Abstract:Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce $\textit{Align Then Reason}$ (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.

153. 【2610.00812】Video Generation Models: A Survey of Post-Training and Alignment

链接:https://arxiv.org/abs/2610.00812

作者:Chaoyu Li,Xiaoyi Gu,Yogesh Kulkarni,Eun Woo Im,Mohammadmahdi Honarmand,Zeyu Wang,Juntong Song,Fei Du,Xilin Jiang,Kexin Zheng,Tianzhi Li,Fei Tao,Pooyan Fazli

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:complex spatiotemporal dynamics, video generation models, Video generation, progressed from short, low-quality clips

备注: Published in Transactions on Machine Learning Research (TMLR), 2026. Project page: [this https URL](https://github.com/people-robots/Awesome-Video-Generation-Post-Training)

点击查看摘要

Abstract:Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.

154. 【2610.00809】Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics

链接:https://arxiv.org/abs/2610.00809

作者:Zhixi Zhu,Kristina Gligoric

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:Vision-Language Models, costs accumulate quickly, processing a typical, accumulate quickly, requires millions

备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single $2\times8$ image grid built via simple shot-transition detection approaches full-video understanding ($\kappa$ within~.05), at $\sim 15\%$ of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.

155. 【2610.00785】VTV-FM: Flow Matching through Variational Terminal-Velocity Closure

链接:https://arxiv.org/abs/2610.00785

作者:Haoyang Jiang,Yuheng Li,Di Yang,Yanhai Xiong,Haipeng Chen,Yi He

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:simple source distribution, fitting continuous-time motion, learns generative transport, learns generative, Flow matching

备注: Accepted at NeurIPS 2026. Code: [this https URL](https://github.com/HaoyangJiang-WM/VTV-FM)

点击查看摘要

Abstract:Flow matching (FM) learns generative transport by fitting continuous-time motion from a simple source distribution to the data distribution. Most existing methods use first-order bridges: once a source and a target sample are paired, the path is a straight motion with constant velocity. FM with optimal transport (OT) improves the pairing, but the bridge itself remains linear, limiting its ability to model curved motion, acceleration, and changing directions. A natural remedy is to use second-order phase-space dynamics; however, learning the bridge requires target-side terminal-velocity information that static datasets do not provide. We propose Variational Terminal-Velocity Flow Matching (VTV-FM), a second-order FM framework that derives the missing velocity by minimizing acceleration energy, yielding a closed-form closure for static data. The same minimum-acceleration variational construction also defines the OT pairing cost and the acceleration targets used for training. Experiments on low-dimensional datasets, PDE-governed physical fields, and CIFAR-10 show that VTV-FM improves transport geometry and generation quality over first-order and high-order FM baselines.

156. 【2610.00757】Video Evidence Indexing: Learning Where to Look from Video Previews for Token-Budgeted Long-Video Question Answering

链接:https://arxiv.org/abs/2610.00757

作者:Haowen Guan,Shengzhi Li,Shichao Pei

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video Evidence Indexing, fixed context width, Long-video question answering, Video Evidence, Long-video question

备注:

点击查看摘要

Abstract:Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may require broad temporal coverage, but the answer is often supported by only a compact set of moments. To locate these moments efficiently, we propose token-budgeted Video Evidence Indexing (VEI): given a dense low-resolution Video Preview, the model constructs a compact high-resolution Evidence Set for final reasoning. We treat VEI as a policy that must jointly solve \textit{evidence localization}, which finds question-relevant moments, and \textit{budget planning}, which decides where to spend the limited high-resolution frame budget. We implement this idea with an inference pipeline: the Video Preview provides cheap global coverage, Video Evidence Indexing constructs the Evidence Set, and Answer Generation combines both inputs for final VQA. To address missing frame-level supervision, we adopt privileged self-distillation, where an answer-aware teacher guides the normal test-time policy on student-generated indexing traces. We explore previews at 1, 6, 12, and 24 visual tokens per frame, training a single policy that supports all four resolutions. Experiments show that Video Evidence Indexing improves accuracy under limited visual budgets, and self-distillation further improves both QA accuracy and temporal evidence localization.

157. 【2610.00753】Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning

链接:https://arxiv.org/abs/2610.00753

作者:Syon Mansur,Joel Zylberberg

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Neurons and Cognition (q-bio.NC)

关键词:dominant mode, coordination of parameter, parameter updates, updates across layers, training

备注: 10 pages, 5 figures

点击查看摘要

Abstract:End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.

158. 【2610.00751】Signal-Noise Factorization Isolates Nuisance Variation into Removable Subspaces

链接:https://arxiv.org/abs/2610.00751

作者:Sakin Kirti,Joel Zylberberg

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:Recent theoretical work, theoretical work identified, work identified fundamental, Recent theoretical, identified fundamental properties

备注:

点击查看摘要

Abstract:Recent theoretical work identified fundamental properties of representation geometry that shape inference ability of deep neural networks. These include signal-noise factorization (SNF), the ability to segregate signal from noise, and signal-signal factorization (SSF), the ability to segregate task-specific and task-irrelevant signals. Here, we built regularizers that reinforce these two properties during training. We compared networks trained with these regularizers to $L_2$-regularized baseline networks on the CIFAR-100 classification task to understand how our regularizers shape representation geometry and impact performance on a well-known computer vision baseline. Enhancing SNF via regularization improved model performance but enhancing SSF did not. Motivated by biomedical applications, we investigated how our regularizers affected performance on the BloodMNIST dataset treated with MedMNIST-C corruptions at five severity levels, and found even larger performance gains using the SNF regularizer. To understand the mechanism by which SNF-regularization produces improved performance, we analyzed the nuisance subspaces across regularization regimes, finding that the SNF-regularized models represent noise in distinct subspaces, separate from class-relevant signal. Because this geometry is explicit, the dominant corruption-induced directions can be estimated on held-out data and projected out of the representations. This manipulation led to a substantial gain in accuracy. These results show that regularizers that enforce signal-noise factorization can produce substantial improvements on computer vision tasks that contain out-of-distribution image distortions at inference time. They also highlight how shaping representations affects model performance: isolating nuisance variables from categorical ones is more important than maintaining factorized representations of categorical variables.

159. 【2610.00749】What Builds the Scene? Luminance Dominates Geometry Formation in 3D Gaussian Splatting

链接:https://arxiv.org/abs/2610.00749

作者:Rezvan Joshaghani,Steven Cutchin

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Gaussian Splatting, jointly from RGB, RGB supervision, making it difficult, difficult to isolate

备注: 28 pages, 6 figures, 7 tables

点击查看摘要

Abstract:Standard 3D Gaussian Splatting (3DGS) learns geometry and appearance jointly from RGB supervision, making it difficult to isolate how luminance and chroma contribute to the learned representation. We study this by training models under different channel supervision, freezing their non-appearance parameters (position, scale, rotation, and opacity), and re-estimating appearance with the same solver before comparing held-out reconstruction. Across eleven benchmark scenes with four independent runs each, geometry learned from luminance alone supports held-out reconstruction 0.085 dB below RGB-trained geometry on average. If chroma is deleted from a trained model, a sufficiently expressive solver can re-fit it on the frozen geometry to the original quality or slightly better. Higher-order spherical harmonics contribute much more reconstruction quality to luminance than to chroma, improving PSNR by 1.44 dB versus 0.19 dB on average, although on mirror-like surfaces hue does still change with viewpoint. The luminance advantage is even larger when geometry is being formed. Chroma-only supervision produces geometry 3.9-5.5 dB worse than luminance-only supervision after the same appearance solve; densification explains part of this gap. Overall, geometry formation in standard 3DGS is strongly luminance-dominated but not luminance-exclusive, and much of the chromatic appearance can be recovered after spatial support has formed.

160. 【2610.00737】Personalized Image Generation with Reasoning and Reflection

链接:https://arxiv.org/abs/2610.00737

作者:Bo Ni,Ngoc N. Tran,Qinwen Ge,Franck Dernoncourt,Seunghyun Yoon,Samyadeep Basu,Sungchul Kim,Puneet Mathur,Nedim Lipka,Tong Yu,Yu Wang,Ryan A. Rossi,Tyler Derr

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:remained narrowly focused, Personalized image generation, curated visual exemplars, remained narrowly, narrowly focused

备注:

点击查看摘要

Abstract:Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user's lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for personalized image generation from user histories. The benchmark comprises two complementary tasks and a multi-axis evaluation protocol that assesses target fidelity, visual quality, user distinguishability, semantic alignment with the user's history, and task-specific utility. Grounded in real-world e-commerce and social media settings, the benchmark includes: (1) Personalized Scene Generation, which places a given object in a scene that reflects a user's preferences and lifestyle, motivated by personalized product presentation; and (2) Personalized Creative Generation, which generates a novel image on a specified topic that is faithful to a user's aesthetic and visual identity, motivated by social media content creation. We further propose PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward. Across both tasks, PEARL outperforms strong baselines, achieving an average improvement of 15% across personalization metrics.

161. 【2610.00693】FedMAD: Modulation-Aware Directional Aggregation for Federated Learning in Remote Sensing Image Classification

链接:https://arxiv.org/abs/2610.00693

作者:Barış Büyüktaş,Begüm Demir

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:recently attracted increasing, attracted increasing attention, enables collaborative model, collaborative model training, requiring direct access

备注:

点击查看摘要

Abstract:Federated learning (FL) has recently attracted increasing attention in remote sensing (RS) since it enables collaborative model training across decentralized RS image archives without requiring direct access to local data. However, FL performance significantly degrades when the data distributions between clients are heterogeneous, which often occurs due to geographical differences, seasonal changes, and varying image acquisition and atmospheric conditions. To address this challenge, in this letter, we propose a novel personalized FL framework (denoted as FedMAD) for RS image classification problems. The proposed framework separates globally shared representation parameters from client-specific adaptation parameters to preserve client-specific features while maintaining globally transferable representations. This is achieved by integrating lightweight modulation modules and local batch normalization layers into the backbone network. Although globally shared parameters are collaboratively optimized between clients, client-specific parameters remain local to preserve domain-specific feature characteristics. In addition, FedMAD introduces a modulation-aware directional aggregation strategy that dynamically adjusts the importance of aggregation for each client according to the alignment of local modulation updates. This allows the global optimization process to suppress conflicting client updates originating from heterogeneous data distributions while enhancing the contribution of clients with consistent adaptation behaviors. The experimental results obtained on the BigEarthNet-S2 and EuroSAT datasets demonstrate the effectiveness of FedMAD compared to state-of-the-art FL algorithms under heterogeneous RS data distributions. The code of the proposed framework will be publicly available at this https URL.

162. 【2610.00691】Soundwich: Video Generation with Layered and Controllable Audio

链接:https://arxiv.org/abs/2610.00691

作者:Zhuo Ning,AmirHossein Naghi Razlighi,Sagi Polaczek,Daniel Cohen-Or,Ali Mahdavi-Amiri

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)

关键词:Recent joint audio-video, single mixed track, Recent joint, synthesize realistic videos, joint audio-video generative

备注: 35 pages. Code: [this https URL](https://github.com/CodyNing/Soundwich)

点击查看摘要

Abstract:Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at this https URL.

163. 【2610.00686】SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

链接:https://arxiv.org/abs/2610.00686

作者:Mikhail Dereviannykh,Vikram Voleti,Simon Donne,Mallikarjun Byrasandra Ramalinga Reddy,Shimon Vainer,Mark Boss

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Recent video-based world, Recent video-based, video-based world models, world models pair, scalability of autoregressive

备注: 29 pages, 22 figures, including references and appendix; 9 pages of main text

点击查看摘要

Abstract:Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.

164. 【2610.00680】Curvature Under Attack in hZACH-ViT: Gauge Symmetry, Boundary Saturation, and Adversarial Failure

链接:https://arxiv.org/abs/2610.00680

作者:Athanasios Angelakis,Marta Gomez-Barrero

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:learned logit temperature, compact Vision Transformer, learned logit, logit temperature, numerical safeguards

备注: 12 pages, 3 figures, 4 tables. Accepted at NeurReps 2026: Symmetry and Geometry in Neural Representations, NeurIPS 2026

点击查看摘要

Abstract:Curvature is often treated as an intrinsic property of a representation, although its empirical effect also depends on coordinate scale, learned logit temperature, and numerical safeguards. We study this interaction in hZACH-ViT, a compact Vision Transformer with Euclidean, Poincare, and spherical prototype heads. The backbone architecture, seed-specific initialization, 50-per-class training subset, and optimization protocol are matched across three MedMNIST datasets and five seeds. At the fixed comparison curvature $c=1$, Poincare has the lowest class-macro PGD attack-success rate in all 12 dataset-budget cells and under a stronger CE+DLR multi-restart attack on all three datasets, but it also has the lowest clean MacroF1. An end-to-end curvature intervention changes the interpretation. Reducing Poincare curvature to $c=0.1$ improves clean MacroF1 in every one of the 15 paired seed-dataset comparisons and removes hard boundary clipping, yet on OrganAMNIST it increases strong attack success from $89.7\%$ to $99.3\%$ (paired difference $+9.57$ points; 95\% hierarchical bootstrap CI $[+5.52,+14.03]$). At $c=1$, $40$-$47\%$ of clean Poincare features are hard-clipped, the radial Jacobian of the inherited map is nearly zero, and dimensionless attack trajectories are unusually long and inefficient. The spherical head provides a control: its curvature change is an exact scale gauge to floating-point precision and produces much smaller attack differences. These results do not establish intrinsic hyperbolic robustness. They identify an implementation-sensitive regime in which curvature, scale, and proximity to the Poincare boundary jointly organize clean recognition and adversarial representation motion.

165. 【2610.00677】Harnessing Vision-Language Models for Perceptual Quality Assessment and Autonomous Content Adjustment in Augmented Reality

链接:https://arxiv.org/abs/2610.00677

作者:Elias Rotondo(1),Lin Duan(1),Yanming Xiu(1),Sangjun Eom(1),Conrad Li(1),Maria Gorlatova(1) ((1) Duke University)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:foster innovative solutions, Advancements in augmented, healthcare delivery, augmented reality, continue to foster

备注: To be published in VRST 2026. Main Manuscript: 12 pages, 5 figures; Supplemental Materials: 7 pages, 10 figures. The accompanying public repository can be accessed by visiting [this https URL](https://github.com/Duke-I3T-Lab/RateAR)

点击查看摘要

Abstract:Advancements in augmented reality (AR) continue to foster innovative solutions, facilitating novel methodologies within educational systems, healthcare delivery, and risk-mitigation protocols. However, optimizing for end-user immersion and comfort remains challenging, as AR head-mounted displays contend with constrained scene geometry, spatial jitter, and temporal instability. User studies are the standard AR evaluation method for visual quality, but their cost, diminishing scalability, and inflexibility pose bottlenecks during iterative application design. To address this problem, we present an automated framework for AR content evaluation and refinement, built on vision-language models (VLMs), to evaluate and predict the visual fidelity of AR scenes as perceived by users. First, we introduce RateAR, a benchmark of AR images and videos collected across diverse scenes and environmental conditions, with good-to-excellent reliability (ICC(2,5) = .90) across perceptual factors, including object placement, scale, and shadow consistency. Subsequently, we evaluate eleven commercial VLMs on the crafted benchmark. Results support that VLM-based quality predictions strongly correlate with human subjective judgments, achieving Spearman's rank-order correlations of up to 0.8695. An ablation study further suggests that, compared to other prompting strategies, our contextual prompting yields better alignment with human ratings while balancing introduced complexity cues. Building on these findings, we construct an automated AR content adjustment system and conduct a 21-participant user study. More than 90% of participants found that the system improved placement and size coherence of virtual content.

166. 【2610.00666】VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

链接:https://arxiv.org/abs/2610.00666

作者:Vu Dinh Xuan,Duc-Hai Nguyen,Minh-Dung Dao,Vu Quynh Giao,Quang Hong Nguyen,Binh-Son Hua,Barry O'Sullivan,David Murphy,Hoang D. Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:computer vision papers, vision-language models, computer vision, comparison figures, Qualitative comparison figures

备注: 29 pages, 18 figures, 6 tables. Code: [this https URL](https://github.com/ReML-AI/visionq)

点击查看摘要

Abstract:Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introduce VisionQ, the first benchmark built from peer-reviewed CV comparison figures that grounds every judgment in a named visual criterion: each question states the criterion, and a judge is credited only when it selects the output the authors identify as best on that criterion. We call this task criterion-conditioned visual discrimination. VisionQ comprises (1) a corpus of 1,409 CVPR and ICCV papers with 1,800+ validated comparison figures and 3,911 hand-annotated data points linking method crops to author-stated visual claims; (2) a six-axis, 51-leaf taxonomy of the visual criteria behind qualitative judgment; (3) a criterion-conditioned evaluation protocol that hides method names, captions, and paper identity and reports accuracy per criterion; and (4) VisionQ-Judge, a DPO-tuned Gemma-4-E4B judge trained on symmetric evidence pairs, which reduces last-option predictions by 7.0pp and improves accuracy by 2.5pp on a held-out test set. Evaluating 20 open- and closed-source VLM judges, we find that the strongest reach only 63.1% accuracy (chance 32.2%) and that reliability varies sharply across criteria. Code: this https URL. Data: this https URL.

167. 【2610.00623】HAWK: Rethinking Multimodal Drafting for Speculative Decoding

链接:https://arxiv.org/abs/2610.00623

作者:Wenhan Yang,Anirudh Rao,Ashwin Chandra

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieved substantial lossless, lightweight drafters struggle, large vision-language models, substantial lossless speedups, Speculative decoding

备注:

点击查看摘要

Abstract:Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where lightweight drafters struggle to use rich multimodal information. A second limitation is that standard distillation supervises the drafter only along the original training trajectory, without modeling how target predictions shift after the drafter's own proposals. As drafting moves away from this trajectory, the drafter can increasingly disagree with the target, reducing acceptance in later steps. We propose HAWK to address both limitations. HAWK uses representation similarity to select informative target layers and learns how to combine their hidden states. For visual information, it directly provides the drafter with compressed visual hidden states from the target model instead of raw visual tokens, making the visual information easier for a shallow drafter to use. HAWK also trains the drafter to capture how target predictions change after its own proposals, improving its agreement with the target during multi-step drafting. On SmolVLM-256M across ten multimodal benchmarks, HAWK raises average acceptance length from 3.32 to 4.08 and speedup from 2.19x to 2.60x over EAGLE-3 under greedy decoding, and from 2.89 to 3.41 and 1.92x to 2.19x under sampling.

168. 【2610.00600】Just Align $\bm{x}$: Aligning Predictions, Not Representations

链接:https://arxiv.org/abs/2610.00600

作者:Yuyao Zhang,Yuwei Hu,Ziyang Mai,Yu-Wing Tai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:transfer reliably, aligns clean-image predictions, clean-image prediction, prediction, Representation alignment

备注:

点击查看摘要

Abstract:Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image prediction, creating a mismatch between the auxiliary objective and the denoising task. This suggests a different principle: auxiliary supervision should improve the prediction target itself rather than impose a separate representation target. We introduce JAx (Just Align x), a prediction-supervision method that aligns clean-image predictions across noise levels. JAx couples a noisier student observation with a cleaner observation through a Markov degradation that preserves the original JiT input distribution. Under this coupling, the oracle prediction from the cleaner state has the same conditional mean as the optimal JiT target, while its conditional target covariance is no greater. Thus, oracle prediction alignment preserves the population JiT objective up to a constant while providing a lower-variance training target. To make this construction practical with an imperfect EMA teacher, JAx combines ground-truth supervision with a reliability-gated coupling band that selects nearby teacher states based on prediction risk. On ImageNet 256x256, JAx consistently improves FID and accelerates convergence across JiT-B/16, L/16, and H/16, without an external encoder or changes to the architecture or sampling procedure. Gradient diagnostics further show reduced minibatch gradient variance, while ablations demonstrate that the gains cannot be explained by time reweighting alone. These results show that prediction-space supervision provides a simple and principled alternative to representation alignment for pixel-space generative models.

169. 【2610.00586】Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

链接:https://arxiv.org/abs/2610.00586

作者:Opegbemi Matthias Busoye,Tolulope Matthew Busoye,Eghonghon-aye Eigbe

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:limits CNN inference, limits CNN, CNN inference, constrained hardware, Gural and Murmann

备注: Extended version of a paper accepted at the NeurIPS 2026 Workshop on Global South in AI

点击查看摘要

Abstract:Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in $O(1)$ yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across $10{,}000$ random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1's 256 KB SRAM from six to nine.

170. 【2610.00582】Discrete Annotation, Continuous Preference: Rethinking Supervision for Accurate and Generalizable Aesthetic Image Cropping

链接:https://arxiv.org/abs/2610.00582

作者:Ziqing Zhang,Xiao Liu,Kai Liu,Jianze Li,Weihang Zhang,Linghe Kong,Yulun Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Aesthetic image cropping, Aesthetic image, image cropping aims, aesthetics and composition, aims to identify

备注: Code, model, and data are available at [this https URL](https://github.com/zzqingz/CPIC)

点击查看摘要

Abstract:Aesthetic image cropping aims to identify the optimal crop of an image in terms of aesthetics and composition. While supervision based on annotated data is fundamental, the field has been hindered by a long-standing problem: existing datasets suffer from (1) human subjectivity and (2) rigid discreteness confined to fixed sampling grids. These flawed annotations not only limit the accuracy and generalization of trained models but also severely distort fair evaluation. To overcome this, we propose to model human cropping preference as a multi-peaked, continuous, and sharp field over the crop space. We introduce the Continuous Preference Field (CPF), which recovers a dense preference landscape from discrete annotations through (1) peak clustering, (2) off-lattice refinement, (3) negative shaping, and (4) field assembly. Based on this, we train CPIC, a VLM-based cropping model optimized via GRPO with the CPF reward, which overcomes template collapse, achieving state-of-the-art performance and exceptional out-of-domain generalization. Finally, to resolve the long-standing benchmark evaluation crisis, we introduce CPICD, a comprehensive recalibration of existing ground-truth boxes. By leveraging the CPF to correct grid-bound artifacts across mainstream benchmarks, CPICD establishes a rigorous and reliable foundation for future cropping research. Extensive experiments and user studies demonstrate the superiority of our CPF, CPIC, and CPICD. Code, model, and data are available at this https URL.

171. 【2610.00576】Gestalt: Large Multimodal Interplay Model

链接:https://arxiv.org/abs/2610.00576

作者:Zequn Yang,Yu Miao,Haotian Ni,Ziheng Chen,Chengxiang Huang,Dongzhan Zhou,Kai Chen,Qi Zhang,Ji-Rong Wen,Yake Wei,Di Hu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multimodal model built, large multimodal model, large multimodal models, multimodal, large multimodal

备注: 17 pages, 7 figures

点击查看摘要

Abstract:In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence.

172. 【2610.00573】FORTE: Adaptive Scoring and Exact Keyframe Selection for Long-Video Question Answering

链接:https://arxiv.org/abs/2610.00573

作者:Haifeng Huang,Biyin Xu,Chunsheng Xin,Yang Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Query-aware keyframe selection, large language models, selection enables multimodal, enables multimodal large, multimodal large language

备注:

点击查看摘要

Abstract:Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typically search within a fixed, uniformly sampled candidate pool, preventing evidence outside this pool from ever being selected. Given a limited relevance-scoring budget, the key challenge is to allocate evaluations adaptively to promising frames while continuing to explore underrepresented temporal regions. We introduce FORTE, a training-free framework that addresses this challenge through two stages: adaptive relevance scoring and global keyframe optimization. Starting from sparse, uniformly distributed observations, our efficient Gaussian-process relevance predictor estimates relevance for unscored frames, exploiting temporal locality and the approximately banded kernel structure to reduce the core computation from cubic to linear time in the number of frames for fixed bandwidth. The scoring stage then selects which frames to score next by balancing predicted relevance with temporal coverage, prioritizing promising regions while also exploring less-represented parts of the video. The optimization stage selects the final keyframes by maximizing an objective that jointly captures measured relevance and temporal coverage. We derive an exact algorithm that leverages the logarithmic coverage structure to identify the optimal subset of the scored candidate pool in time linear in the pool size, for a fixed final-frame budget. Experiments on four long-video question-answering benchmarks show that FORTE achieves the highest observed mean accuracy among the compared selectors under every tested scoring budget. Further evaluations demonstrate its consistent effectiveness across different relevance scorers and downstream MLLMs.

173. 【2610.00559】PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

链接:https://arxiv.org/abs/2610.00559

作者:Xinge Peng,Yiting Lu,Tianwu Zhi,Wen Wen,Jianzhao Liu,Xin Li,Zhibo Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dynamics remains unclear, physical consistency underlying, consistency underlying real-world, shown strong multimodal, remains unclear

备注: Accepted at NeurIPS 2026 (Main Track)

点击查看摘要

Abstract:Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce PhysVista, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes event-level reasoning and scale-level reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments across a diverse set of VLMs reveal substantial limitations in physical reasoning and plausibility assessment, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.

174. 【2610.00544】Memorizon: Training World Models Beyond Their Context Window

链接:https://arxiv.org/abs/2610.00544

作者:Tingting Liao,Xuezhi Liang,Hao Li,Guangyi Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Streaming world models, Streaming world, render a place, place consistently, consistently across repeated

备注: Project page: [this https URL](https://tingtingliao.github.io/memorizon)

点击查看摘要

Abstract:Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last $k$ chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-$K$ latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by $kK$, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: this https URL

175. 【2610.00483】PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion

链接:https://arxiv.org/abs/2610.00483

作者:Lehan Yang,Daiqing Qi,Wenhao Zhang,Avery Li,Yiqing Yang,Yifan Li,Yu Kong,Haitian Zheng,Zhifei Zhang,Zhe Lin,Varun Jampani,Sheng Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:accelerates diffusion transformer, exclusively semantic encoders, Representation alignment, REPA targets, diffusion transformer training

备注: NeurIPS 2026

点击查看摘要

Abstract:Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models trained to predict that structure remain overlooked as REPA targets. In pixel-space diffusion, SAM2, Depth Anything v2, and Metric3D v2 each outperform the DINOv2-only GenEval baseline, with the two geometric teachers leading the segmentation teacher. A flat sum of all four teachers, however, lands below the best single geometric teacher, as semantic and geometric gradients compete for one denoiser projection. We introduce PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream, routes Depth Anything v2 and Metric3D v2 through a geometric projection stream, and adds a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces. All four teachers are frozen during training and dropped at inference. Applied to PixelGen and DeCo with a single recipe, PixelDense improves GenEval, DPG-Bench, and HPS v2.1, raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, and beats every single-teacher and unfactored multi-teacher variant. In partial-noise reconstruction, independent panoptic, depth, and surface-normal probes show up to 53.1% PQ gain and 36.0% depth AbsRel reduction at $\tau=0.5$ across COCO and Flickr30K. From random initialization, PixelDense also reaches the baseline's peak GenEval 1.23x faster. In SDEdit editing on PIE-Bench, PixelDense keeps more of the source background and layout at every edit strength, raising background PSNR by up to 2.2 dB.

176. 【2610.00451】PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

链接:https://arxiv.org/abs/2610.00451

作者:Rikhat Akizhanov(1),Yangsong Zhang(1),Nikolai Kaliazin(1),Peter Wolf(2),Yoshihiko Nakamura(1),Pascal Fua(3),Fabio Pizzati(1),Ivan Laptev(1) ((1) MBZUAI, (2) ETH Zürich, (3) EPFL)

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:typically separate visual, common physical laws, existing approaches typically, approaches typically separate, governed by common

备注: 31 pages, 12 figures. Project page: [this https URL](https://rihat99.github.io/PACT/)

点击查看摘要

Abstract:Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.

177. 【2610.00447】Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

链接:https://arxiv.org/abs/2610.00447

作者:Anson Y. Lam,Shuqing Li,Michael R. Lyu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Multimedia (cs.MM)

关键词:scene stays fixed, generated scene stays, stays fixed, generated scene, scene stays

备注: 26 pages, 6 figures

点击查看摘要

Abstract:Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment evaluators plus one perceptual-quality control, then test four targeted scene degradations. Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators, with prompt-bootstrap lower bounds above 1 for 11/19. Rankings are more stable than scores, yet 18/19 evaluators change their point-estimate winner under some configuration. Pairwise protocol margin envelopes show which comparisons keep their direction across the tested settings. Selected pairs have opposite pointwise intervals, but no reversal survives simultaneous inference over the full search. Thus the observed winner changes are descriptive, not confirmed changes in generator superiority. Sensitivity remains separate: no evaluator, even the prompt-free control, exceeds 67% tie-adjusted directional discrimination on layout scrambling, which is diagnostic rather than human-validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) with protocol-dependent comparisons and selection-aware uncertainty.

178. 【2610.00421】Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification

链接:https://arxiv.org/abs/2610.00421

作者:Bhanu Prakash Vangala,Sowmya Guda,Latha Peddi,Navya Vangala

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:heavily published application, Automated classification, benchmarks routinely exceeding, medical imaging, routinely exceeding

备注:

点击查看摘要

Abstract:Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.

179. 【2610.00414】From Image Latent Space to Fuzzy Rules: Interpretable Analysis of Gastrointestinal Foundation Model

链接:https://arxiv.org/abs/2610.00414

作者:Michael D. Vasilakakis(1),Dimitris K. Iakovidis(1) ((1) Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large-scale datasets demonstrate, datasets demonstrate strong, demonstrate strong transferability, medical imaging tasks, imaging tasks

备注: Accepted at the excv, ECCV 2026 Workshops. 17 pages, 4 figures, 7 tables

点击查看摘要

Abstract:Foundation models pretrained on large-scale datasets demonstrate strong transferability to medical imaging tasks. However, understanding how their latent representations encode clinically relevant information remains an open challenge in safety-critical domains. This study proposes a prototype-based fuzzy-rule framework that interprets the patch-level features produced by the inner layers of pretrained foundation models, without any fine-tuning. Class-specific prototypes are learned by clustering in the feature space, yielding compact visual patterns. Patch features are then expressed as prototype similarities and classified by fuzzy rules with linguistic IF-THEN conditions that are human readable. The framework is applied across the final two blocks of ViT-S/16 backbones pretrained on ImageNet-1K and GastroNet-5M, and benchmarked against k-nearest neighbours, kernel SVM, and linear probing under identical frozen features, on wireless capsule endoscopy classification, gastrointestinal endoscopy classification, and colonic polyp segmentation. The experimental analysis shows that the proposed method, without backbone fine-tuning, reaches accuracy comparable to these black-box classifiers, and that domain-specific pretraining yields features that are both discriminative and symbolically compressible. Because the resulting rules are extracted from real data and expressed in interpretable terms, they are further used as an instrument to investigate synthetic medical images, providing a human-readable account of which real prototypes and rules a generator reproduces or fails to reproduce, localising where a synthetic image departs from real tissue rather than summarising it with a single score. The framework thus offers a transparent, depth-resolved view of how foundation models organise clinically relevant structure, together with a practical downstream use of the extracted rules.

180. 【2610.00365】Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment

链接:https://arxiv.org/abs/2610.00365

作者:Jinho Chang,Jong Chul Ye

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, high-quality data, advances in distillation, distillation and flow-map, enabled deterministic

备注: 25 pages, 13 figures

点击查看摘要

Abstract:Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely on first-order gradient information, which is either inapplicable or suffers from instability and inefficiency in black-box reward scenarios. Here, we introduce ZeNOVA, a stable and efficient initial noise alignment method in a gradient-free manner. Specifically, we address existing algorithms' major challenge in black-box scenarios through annealed soft-value guidance, manifold-constrained hyperspherical Langevin dynamics, and Metropolis-Hastings jumping. Extensive experiments on image and video generative models show that ZeNOVA outperforms all evaluated zeroth-order baselines by optimizing the initial noise toward higher rewards substantially more stably while exploiting the geometry of the Gaussian prior, demonstrating its practical applicability to various black-box reward alignment.

181. 【2610.00360】DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

链接:https://arxiv.org/abs/2610.00360

作者:Haoyu Wang,Siyuan Qian,Yanjun Li,Zeyu Zhang,Yandong Guo,Boxin Shi,Hao Tang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:discover finger-object contacts, Reinforcement learning, Target success, PPO, dexterous manipulation

备注:

点击查看摘要

Abstract:Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our baseline PPO runs end near their initial action noise after 5M steps, motivating explicit control of exploration scale. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed. We study three policy-optimization settings: PPO, critic-free GRPO continuation, and a flow-parameterized PPO variant (FPO). Across five YCB objects and three training seeds, mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO). On a RealMan RM75 arm with an Inspire/RH56 hand, 360 trials over three objects raise mean Target success from 25.0% to 85.0% (FPO), 10.0% to 63.3% (GRPO), and 8.3% to 43.3% (PPO), with one trained model per object-method condition. PPO component screening favors noise control over the tested optimizer contraction; the selected PPO schedule yields higher mean Target success than linear decay with the same endpoints on three tested objects. Training return, deterministic Target success, and tolerance to execution noise dissociate; schedules should therefore be judged by terminal task success under the intended execution conditions, per task and policy-optimization setting. Code: this https URL. Website: this https URL.

182. 【2610.00359】Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength

链接:https://arxiv.org/abs/2610.00359

作者:Candi Zheng,Yuan Lan

类目:Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:reference image-guided editing, rapid progress, prompt and reference, reference image-guided, remain too coarse

备注:

点击查看摘要

Abstract:Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotations, while existing zero-shot methods often yield unsatisfactory results. We introduce SoftPaint, a new zero-shot sampling method that leverages soft masks to enable a continuous spectrum of edits, from fully preserving the original content to completely re-synthesizing the masked region. Going beyond zero-shot inpainting methods, we design a Langevin-iteration-based sampler that respects per-pixel soft mask strengths, which applies universally to image and video diffusion models, enabling tasks such as video editing. The method is gradient-free, memory-efficient, and achieves smooth, pixel-level edits across multiple image and video backbones.

183. 【2610.00350】Vmem-$φ$: Low-Compute Out-of-Distribution Detection in Spiking Neural Networks from Membrane-Potential Statistics

链接:https://arxiv.org/abs/2610.00350

作者:Arul Rana,Agrim Tripathi,Shoaib Ahmed Dipu,Md. Shaown Miah,Syed Ishtiaque Ahmed,Sayeed Shafayet Chowdhury

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Spiking Neural Networks, Spiking Neural, Neural Networks, processing event-camera data, detection remains challenging

备注:

点击查看摘要

Abstract:Spiking Neural Networks (SNNs) offer an energy-efficient approach to processing event-camera data, yet out-of-distribution (OOD) detection remains challenging in this setting. Existing OOD detection methods often depend on model outputs or computational components that are unavailable in object detection SNNs or are poorly suited to low-compute deployment. To that effect, we show that the subthreshold membrane potential \(V_{\mathrm{mem}}(t)\) provides a useful internal signal for detecting distribution shifts. Simple per-channel statistics derived from these membrane dynamics enable OOD detection. To evaluate this approach, we introduce Gen1-C, an event-camera corruption benchmark developed upon the Prophesee Gen1 automotive detection dataset, containing six sensor-motivated histogram-level stress tests at five severity levels. We further propose the Multi-Descriptor Deviation (MDD), a corruption-blind method that operates on membrane-potential statistics. At the highest corruption severity, MDD achieves an AUROC of more than 0.88 on five of the six corruptions using only a bounded 64-frame observation window. Notably, the remaining corruption is also the one that has the smallest effect on the underlying detector. These results show that the temporal membrane-potential dynamics can provide an effective and low-cost signal for OOD detection in SNN-based event perception.

184. 【2610.00341】UnifiedAttack: Evaluating the Safety of Large Multimodal Models in Synergistic Harmful Image-Text Generation

链接:https://arxiv.org/abs/2610.00341

作者:Bingjun Luo,Jialin Guo,Tony Wang,Siqi Li

类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Multimodal Models, transition toward natively, critical challenge, Large Multimodal, synergistic harmful image-text

备注:

点击查看摘要

Abstract:As Large Multimodal Models (LMMs) transition toward natively unified architectures, evaluating their safety in synergistic harmful image-text generation tasks becomes a critical challenge. Unlike unimodal threats, synergistic risks emerge when text and image modalities are coordinated to produce harm that significantly exceeds their individual components. We introduce UnifiedAttack, a novel benchmark designed to evaluate LMM safety in collaborative scenarios by focusing on the harmfulness gain achieved through cross-modal synergy. The benchmark incorporates samples filtered for their multimodal potential alongside a novel subset of synthesized disinformation queries. To verify identified vulnerabilities, we propose a synergistic hijacking framework featuring In-Context Reskinning (ICR) and Cognitive Planning Injection (CPI). ICR utilizes few-shot learning to wrap adversarial intent in benign virtual shells to desensitize safety filters, while CPI hijacks the reasoning path by enforcing a plan-then-execute paradigm. By compelling the system to commit to a neutral logical plan, we exploit its internal drive for consistency to induce the synchronized generation of harmful multimodal content. Extensive evaluations on state-of-the-art architectures demonstrate that UnifiedAttack consistently bypasses modern alignment. Our findings reveal that the structural helpfulness and logical coherence of unified models can be systematically weaponized, highlighting the urgent need for logic-aware defenses in synergistic generation tasks. Code is available at this https URL .

185. 【2610.00333】LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation

链接:https://arxiv.org/abs/2610.00333

作者:Jaeyun Shin,Hangeol Chang,Jong Chul Ye

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Multimodal on-policy distillation, strong reasoning capabilities, on-policy distillation, grounding, visual

备注:

点击查看摘要

Abstract:Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribution entangles its visual grounding signal with its own language prior, preventing the grounding information from being transferred independently. Conversely, increasing the strength of visual supervision can improve perception but may overemphasize visual evidence and degrade language reasoning. To address this trade-off, we introduce LEGO-OPD, which selectively composes factors from a Language Expert and a Grounding expert into One teacher distribution for multimodal OPD. Under a generalized Bayesian formulation, the language expert provides a prior over candidate tokens, while the grounding expert contributes a visual likelihood that updates this prior, rather than transferring its complete predictive distribution. This factorized composition allows language reasoning and visual grounding to be controlled independently. We further introduce adaptive calibration to determine how strongly the visual likelihood should update the language prior at each decoding prefix. Specifically, LEGO-OPD uses the grounding expert's image-induced prediction shift as a prefix-dependent reference, preventing both insufficient and excessive visual supervision. Experiments with Qwen3 models show that LEGO-OPD consistently outperforms the evaluated single- and multi-teacher OPD baselines on both multimodal and text-only reasoning tasks. Moreover, it improves the initial student's visual perception while preserving text-only reasoning.

186. 【2610.00330】Retrospective Open-Vocabulary Memory for Long-Term Object Search

链接:https://arxiv.org/abs/2610.00330

作者:Jiaming Wang,Zhiwei Xue,Chen Jizhuo,Peng Shiqi,Harold Soh

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:search requires learning, changing environment, requires learning, object search requires, uneven observations

备注: 25 pages, 5 figures

点击查看摘要

Abstract:Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate retrospective open-vocabulary memory as probabilistic inference from censored observations, where the key idea is to reason with evidence per opportunity: a detection or non-detection should influence belief only in proportion to the robot's opportunity to observe the corresponding location. We introduce ECROM, which uses this principle to estimate long-term prevalence for concepts specified only at query time and converts the resulting belief directly into an active-search prior. To evaluate this problem, we introduce a controlled long-term benchmark in ten HM3D homes that independently varies object placement and observation opportunity across repeated traversals. ECROM improves support-level AP on held-out queries by 4.5 points and search SPL by 4.2 points over the strongest competing memory in each metric. The benchmark, dataset, and code will be open-sourced.

187. 【2610.00319】EgoRefine: Ego-Referenced Predictive Alignment and Trajectory-Conditioned Reliability-Aware Fusion for Asynchronous Collaborative Perception

链接:https://arxiv.org/abs/2610.00319

作者:Lingzhao Kong,Yongsheng Zang,Yu Kang,Kailun Yang,Jie Fu,Yukun Zuo,Zhiyong Li

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO); Image and Video Processing (eess.IV)

关键词:extending sensing range, enables connected agents, share complementary observations, perception enables connected, Collaborative perception enables

备注: The source code will be made publicly available at [this https URL](https://github.com/godk0509/EgoRefine)

点击查看摘要

Abstract:Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating occlusion. Under asynchronous communication, however, cooperative features arrive with temporal delay. Existing prediction-based methods compensate for these features mainly from the transmitting agent's own history, leaving residual misalignment with the ego agent's current observation; subsequent fusion also often overlooks spatial variations in alignment quality. We propose EgoRefine, an ego-referenced predictive alignment and reliability-aware fusion framework for asynchronous collaborative perception. Its Ego-referenced Predictive Alignment module uses the current ego feature to guide cooperative trajectory-field prediction and refines the sampling offsets along an ego-referenced trajectory direction. Its Trajectory-conditioned Reliability-aware Fusion module treats the trajectory discrepancy between the ego and cooperative streams and the directional refinement magnitude as alignment cues, using them to condition the relation between aligned features and adaptively reweight the two streams before convolutional fusion. Experiments on V2V4Real and DAIR-V2X-Seq show that EgoRefine outperforms TraF-Align by 1.6 and 2.9 points on average in AP@0.5 and AP@0.7, respectively. The source code will be made publicly available at this https URL.

188. 【2610.00317】DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

链接:https://arxiv.org/abs/2610.00317

作者:Youngjun Jun,Kyumin Choi,Youngmin Kim,Seonghyun Jin,Sunwoo Park,Jangho Park,Jong Chul Ye

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:models increasingly rely, generate short action, short action chunks, models increasingly, receding-horizon control

备注: Preprint

点击查看摘要

Abstract:Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.

189. 【2610.00315】Beyond Pixel Reconstruction: Retrieval-Guided Glyph-Aware Restoration for Low-Resource Manchu Historical Documents

链接:https://arxiv.org/abs/2610.00315

作者:Ting Huang,Dongdong Wang,Mingqiu Liang,Siyang Lu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:paired training data, preserve invaluable linguistic, cultural heritage, training data, invaluable linguistic

备注: 8 pages, 7 figures

点击查看摘要

Abstract:Historical Manchu documents preserve invaluable linguistic and cultural heritage, yet their digitization is hindered by severe degradations and the scarcity of paired training data. Existing document restoration methods primarily optimize pixel-level reconstruction, which can produce visually plausible results while failing to preserve the structural identity of Manchu glyphs. To address this limitation, we propose a retrieval-guided glyph-aware restoration framework that goes beyond pixel reconstruction by explicitly incorporating glyph-level structural knowledge. Our method retrieves relevant glyph exemplars to provide structural guidance during restoration and integrates this information into the reconstruction process, improving the recovery of degraded character structures under low-resource conditions. Extensive experiments on Manchu historical documents demonstrate that the proposed approach improves both image restoration quality and glyph-level fidelity compared with existing restoration methods. These results highlight the importance of incorporating character-aware structural priors for reliable restoration of low-resource historical documents.

190. 【2610.00302】Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

链接:https://arxiv.org/abs/2610.00302

作者:Wenping Yin,Fabian Desuer,Ziqi Liu,Naixia Mou,Weijia Li,Pedram Ghamisi,Xiao Xiang Zhu,Hao Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:complementing conventional remote, conventional remote sensing, remote sensing imagery, crowdsourced disaster imagery, street-level observations

备注:

点击查看摘要

Abstract:Crowdsourced imagery provides timely, fine-grained, street-level observations for disaster mapping, complementing conventional remote sensing imagery (RSI) during emergency response. However, such imagery is often unstructured, spatially ambiguous, and lacks reliable geographic metadata, making manual geolocalization and interpretation labor-intensive and difficult to scale. This work proposes a multi-task Geospatial Reasoning Disaster mapping framework, namely GRDisaster, to examine the potential of vision-language models (VLMs) in understanding, geolocalizing, and reasoning over crowdsourced disaster imagery. GRDisaster is built on a newly curated benchmark dataset derived from PhotoMappers, comprising 26,340 images organized into human-validated volunteered geographic information (VGI), street-view imagery (SVI), RSI cross-view triplets covering multiple disaster events from 2018 to 2024. The framework combines deterministic and probabilistic cross-view geolocalization with multi-view fusion to associate VGI images with georeferenced SVI and RSI. It introduces two sets of spatial reasoning indicators for cross-view geolocalization validation and disaster damage assessment. These indicators use structural, environmental, and global-scene cues to validate cross-view correspondences and visually observable damage evidence with expert-verified annotations to assess disaster severity, improving the interpretability of VLM outputs. To our knowledge, this study provides the first systematic investigation and unified evaluation framework for examining how VLM-based spatial reasoning can transform crowdsourced disaster imagery into actionable geospatial artificial intelligence (GeoAI) through cross-view geolocalization validation, interpretable spatial reasoning, and damage-aware severity assessment.

191. 【2610.00294】LENS-GRF: Permutation-Invariant Lesion Evidence Network with Gated Residual Fusion for Acne Severity Grading and Multi-Rater Clinical Oracle Analysis

链接:https://arxiv.org/abs/2610.00294

作者:Muhammad Muhtasim Shahriar,M. F. Mridha

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Gated Residual Fusion, acne severity grading, fine-grained lesion evidence, Lesion Evidence Network, severity grading requires

备注: Submitted to Computer Methods and Programs in Biomedicine (Elsevier)

点击查看摘要

Abstract:Automated acne severity grading requires both whole-face context and fine-grained lesion evidence. We propose LENS-GRF (Lesion Evidence Network with Set-Transformer and Gated Residual Fusion), an interpretable multi-stage framework for four-class acne severity grading. The method combines Adaptive Facial Skin Segmentation and a global Vision Transformer prior with a permutation-invariant Lesion Set Transformer that encodes localized lesion patches and spatial geometry. Gated Residual Fusion adaptively controls the local residual contribution and reduces to the global prediction when the gate is zero. On ACNE04, fully automated LENS-GRF with YOLOv11s achieved 80.82% accuracy; with ground-truth lesion annotations, it achieved 95.89% +/- 0.59% accuracy and a Quadratic Weighted Kappa of 0.9753. A data-integrity audit identified 15 cross-split duplicate image pairs, including five with conflicting severity labels. In locked zero-shot evaluation on the full PLSBRACNE01 cohort (200 subjects, 600 views), automated LENS-GRF achieved 35.00% accuracy versus 42.50% for the global baseline. On the 148-subject common cohort used for three-dermatologist oracle analysis, ground-truth lesion inputs increased the best oracle accuracy to 47.97%, while the highest oracle QWK was 0.5799. Pairwise oracle agreement ranged from 49.32% to 66.22%, highlighting detector domain shift, annotation variability, and cross-criterion mismatch.

192. 【2610.00279】Multi-Resolution Feature Fusion U-Net for Magnetic Resonance Imaging Segmentation

链接:https://arxiv.org/abs/2610.00279

作者:Eirini Cholopoulou,Dimitrios E. Diamantis,Dimitris K. Iakovidis

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:monitoring disease progression, medical image segmentation, disease progression, Deep Learning, essential for clinical

备注:

点击查看摘要

Abstract:The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease progression. While Deep Learning (DL) architectures, such as U-Net and its extensions are very effective in medical image segmentation tasks, they often struggle with preserving fine-grained details and global contextual information. This is especially challenging for MRI data segmentation, where anatomical structures are characterized by irregular boundaries and variations in shape, contrast, and scale. To address this challenge, we propose a novel DL architecture for MRI segmentation across different anatomical structures. Specifically, the architecture introduces a module, named Multi-Resolution Feature Fusion (MRFF), that can be easily integrated into any U-Net-like architecture. The MRFF is integrated in all levels of an encode-decoder structure, along with attention mechanisms and skip connections to extract features at multiple resolutions, enabling the model to capture both fine-grained details and global contextual information. We evaluate the MRFFU-Net on two publicly available benchmark MRI datasets of different anatomical targets; one for Cerebrospinal Fluid (CSF) segmentation in spinal MR scans, and one for left atrium cardiac segmentation, from the Medical Segmentation Decathlon (MSD) challenge. Experimental results indicate that MRFFU-Net outperforms state-of-the-art models across multiple evaluation metrics, demonstrating its effectiveness in MRI segmentation.

193. 【2610.00204】Query Independent Variable Rate Visual Token Coding

链接:https://arxiv.org/abs/2610.00204

作者:Hongbo Zhang,Zihao Yang,Liuyang Song,Daqian Yang,Haoyang Yao,Yan Wen,Zhengtao Yao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual-token compression, language model pays, compression for vision, selection problem, discard the rest

备注:

点击查看摘要

Abstract:Visual-token compression for vision--language models is posed almost entirely as a selection problem: decide which tokens to keep and discard the rest. The criteria that work best rank tokens by the attention the language model pays them, which makes the ranking a function of the question being asked. That is invisible in a single-turn benchmark and decisive whenever a compressed representation is written once and read many times, as when it is cached across the turns of a conversation or transmitted between a device and a server. We take the other half of the classical transform-coding toolkit instead: keep every token and vary its rate. A transform code exposes each token's measured distortion--rate curve, and a fixed bit budget is distributed across tokens by exact integer rate--distortion optimisation on those curves. No text enters the pipeline, so one compressed representation serves any query. At equal bit budgets, on two datasets and two capacities, it preserves the model's output distribution and its answers better than uniform-rate coding, the closed-form water-fill and distortion-ranked pruning. It matches attention-ranked pruning on the question pruning was tuned for, and overtakes it once the compressed image must answer a different question about the same image.

194. 【2610.00196】GPEC: Efficient Pre-LLM Gaussian Process Embedding Correction for Cardiac Video Caption Generation

链接:https://arxiv.org/abs/2610.00196

作者:Arefeh Rezaei

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:shown strong potential, large language models, Gaussian Process Embedding, Multimodal large language, Gaussian Process

备注:

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have shown strong potential for video understanding and caption generation, but their performance may decline in specialized medical imaging domains such as echocardiography. This work introduces Gaussian Process Embedding Correction (GPEC), a modular and computationally efficient pre-LLM error-correction method that improves the visual representations used by VideoChat2 for cardiac ultrasound caption generation. GPEC is inserted between the visual projection layer and the language model and learns a residual correction that moves the projected visual representation toward an annotation-guided target. The target is constructed by converting structured video annotations into qualitative attributes, generating a fixed-format reference caption, and mapping it into the language-model embedding space. The correction is modeled using a sparse variational Gaussian Process with inducing points, natural-parameter variational updates, and a block-wise linear kernel, while the original VideoChat2 components remain this http URL method is evaluated using representation-level, caption-level, content-oriented, and execution-time metrics by comparing the original VideoChat2 with VideoChat2 + GPEC under identical input and reference conditions. Results show improved caption similarity and content alignment after applying the proposed correction. Furthermore, GPEC adds less than 0.05 s of inference-time overhead per video in the evaluated setting. These findings indicate that GPEC can improve caption generation in specialized medical video domains with minimal computational cost, without requiring end-to-end fine-tuning of the pretrained multimodal backbone.

195. 【2610.00195】GS-PQM: A Parameter-Domain Quality Metric for Compressed Gaussian Splatting

链接:https://arxiv.org/abs/2610.00195

作者:Pedro Martin,António Rodrigues,João Ascenso,Maria Paula Queluz

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Gaussian Splatting, enabled substantial reductions, Recent advances, advances in Gaussian, enabled substantial

备注:

点击查看摘要

Abstract:Recent advances in Gaussian Splatting (GS) compression have enabled substantial reductions in GS model size. Reliable objective quality assessment is therefore essential for comparing compression methods and guiding the development of more efficient GS codecs. Existing GS quality assessment typically relies on image and video quality metrics, requiring rendering of predefined viewpoints and making the quality estimate dependent on the selected views. This paper introduces GS-PQM, a novel full-reference quality metric for post-training GS compression that operates directly in the GS parameter domain. GS-PQM estimates perceptual quality from a set of parameter-domain distortion errors using a Support Vector Regression model. Experimental results show that GS-PQM outperforms 25 existing image, video, and point-cloud quality metrics in assessing compressed GS content, providing an accurate and computationally efficient alternative to rendering-based quality assessment.

196. 【2610.00188】Uncertainty-Aware RL-Controlled Adaptive 3D Mapping

链接:https://arxiv.org/abs/2610.00188

作者:Alpay Ozkan,Tunc Ozan Aydin,Marc Pollefeys,Jelena Trisovic,Daniel Barath

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Image and Video Processing (eess.IV)

关键词:Voxel-based volumetric mapping, remain inherently inefficient, grids remain inherently, Voxel-based volumetric, fixed-resolution grids remain

备注: To appear at BMVC 2026. Code available at [this https URL](https://github.com/alpayozkan/UnRL)

点击查看摘要

Abstract:Voxel-based volumetric mapping is fundamental to 3D reconstruction, yet fixed-resolution grids remain inherently inefficient - wasting memory in uniform regions and losing detail in complex ones. Existing adaptive methods, such as MAP-ADAPT, partially address this by varying resolution based on geometry and user-defined semantic class lists, but these heuristics require expert tuning, lack generalization to unseen objects, and provide no explicit mechanism to control memory usage. We propose an adaptive framework that refines voxels based on semantic entropy, which captures label uncertainty, together with geometric curvature and texture richness as scene complexity cues, yielding principled resolution allocation without reliance on semantic taxonomies. To make the accuracy-memory trade-off explicit and user-controlled, we further introduce a reinforcement learning agent that learns voxel subdivision policies under a user-specified target memory budget, replacing hand-tuned thresholds with a single intuitive control parameter. The resulting multi-resolution TSDF achieves higher geometric accuracy, better semantic consistency, and improved memory-accuracy trade-offs compared to MAP-ADAPT and fixed-resolution baselines on both synthetic and real-world datasets. Our code and models are available at this https URL.

197. 【2610.00141】Evaluating the Robustness of Anti-UAV Detection under Controlled Fog Degradation: Fog-Aware Training and Clear-Sky Tradeoff

链接:https://arxiv.org/abs/2610.00141

作者:Gur Levy Birkental,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision-based anti-UAV systems, previous robustness studies, robustness studies treat, Vision-based anti-UAV, treat adverse weather

备注:

点击查看摘要

Abstract:Vision-based anti-UAV systems must function in poor visibility, yet most benchmarks use only clear-sky footage, and previous robustness studies treat adverse weather as a simple present/absent condition. As a result, the impact of fog severity on ground-to-air UAV detection remains poorly understood. This work presents the first severity-controlled fog benchmark for this task: synthetic fog at ten severity levels is applied to the RGB modality of the Anti-UAV300 dataset, comparing a clear-trained YOLOv5m baseline to a fog-aware model trained on both clear and foggy images. Detection performance drops sharply and non-linearly: degradation is front-loaded across light-to-moderate fog (beta approximately 0.05-0.10), with a 96% reduction in mAP@0.5:0.95 from clear to thickest fog, mainly due to lost recall and confidence. On the comparable metric (mAP@0.5), this collapse exceeds the most extreme rain degradation reported in the closest prior benchmark. Fog-aware training boosts detection across all severities (up to +0.320 mAP@0.5:0.95) with only a 9.1% drop in clear-sky accuracy, raising the threshold for reliable detection while not preventing collapse under extreme fog. Since reliability is lost within a narrow visibility range, simple clear vs adverse tests underestimate operational risk.

198. 【2610.00125】A Comprehensive Review of One-Pixel Attack: Research Status, Taxonomy, Applications, Regulation Policy and Future Directions

链接:https://arxiv.org/abs/2610.00125

作者:Mirza Niaz Morshed,Md. Masudul Islam,Galib Muhammad Shahriar Himel,Md. Aslam Uddin,Hui Liu,Md. Shafiqul Islam

类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)

关键词:reliably induce high-confidence, induce high-confidence misclassification, autonomous driving, deep learning, medical diagnosis

备注:

点击查看摘要

Abstract:One-Pixel Attacks (OPAs) represent one of the most extreme demonstrations of adversarial fragility in deep learning, where modifying a single pixel can reliably induce high-confidence misclassification across domains such as medical diagnosis, autonomous driving, biometrics, and quantum communication. Despite their conceptual simplicity, OPAs remain underexamined in existing adversarial-attack surveys, which provide only fragmented or cursory coverage. This PRISMA-guided review synthesizes high-quality studies from 2017 to 2026 and delivers a unified, multi-axis taxonomy of OPA research spanning algorithmic foundations, black-box evolutionary optimization, emerging hybrid and program-synthesis attacks, defence mechanisms, interpretability tools, and domain-specific vulnerabilities. Our analysis reveals the dominance of Differential Evolution-based strategies, the rise of efficiency-optimized and saliency-guided methods, and persistent gaps in dataset diversity, transferability, and standardized evaluation. We summarized and assess defence paradigms including pixel restoration, anomaly detection, input-space transformations, and robust training highlighting their trade-offs in robustness, imperceptibility, and computational overhead. Building on these insights, we outline future research priorities involving selective pixel recovery, transformer-specific vulnerability analysis, saliency-driven optimization, and real-world domain-adaptive defences. We further propose a regulatory framework emphasizing robustness testing, incident disclosure, and AI security governance. This review establishes a comprehensive foundation for understanding, evaluating, and mitigating ultra-sparse adversarial threats in contemporary AI systems.

199. 【2610.00111】A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight

链接:https://arxiv.org/abs/2610.00111

作者:Rasul Khanbayov,Hasan Kurban

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multimodal reasoning models, supervise multimodal systems, filtering training data, shapes multimodal reasoning, supervise multimodal

备注:

点击查看摘要

Abstract:Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models. Trusting one means first checking that it uses its evidence, and that check is itself worth scrutinizing, so we ask whether a counterfactual probe of visual grounding measures what it claims to. The probe edits the image so the ground truth flips, holds the reasoning trace fixed, and asks whether the verdict follows. We formalize it as the Verdict Grounding Score and show it cannot be read the way such scores are read. A verdict responds only to an edit that reaches the judge's decision-relevant reading, so the score is capped by how perceptible the edit is, and unless editing makes the attribute easier to read, the error is one-sided: the score can only make a judge look less grounded than it is. The practical failure is therefore a false alarm, an auditor discarding a usable overseer. Under assumptions we state, we show this missing quantity is not merely bounded but identified from three quantities the same audit protocol already collects, which makes the false-alarm rate directly measurable rather than merely a concern. Auditing nine judges, we find the predicted ordering holds strictly across our entire primary pool, and the typical judge there acts on only about half of the edits whose attribute it can otherwise resolve. Applying a conservative rejection threshold certifies several cells as false alarms outright, the clearest being a judge that detects the injected error essentially every time while still scoring as if it had not used the image at all. The rule that follows is that an image-side counterfactual score should never be reported alone: a detection probe on the unedited image upper-bounds it, certifies its false alarms, and costs nothing extra to run.

200. 【2610.00097】DramaAgent: Agentic Storytelling Video Generation

链接:https://arxiv.org/abs/2610.00097

作者:Ting Huang,Biao Wu,Ronghao Chen,Zeyu Zhang,Tengfei Cheng,Qizhen Lan,Huacan Wang,Hao Tang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:audio remains challenging, aligned audio remains, Recent diffusion, producing coherent long-form, substantially improved

备注:

点击查看摘要

Abstract:Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross-scene continuity, and audio-visual mismatch over extended sequences. We propose DramaAgent, a hierarchical, agentic, and model-agnostic framework for long-form text-to-video-and-audio generation. Rather than improving the underlying video backbone itself, DramaAgent introduces an upper-level control layer that decomposes generation into story planning, persistent character conditioning, scene-wise synthesis, and reflection-guided targeted repair. The framework maintains reusable story and character states across scenes, diagnoses failures such as identity drift, missing scene semantics, temporal discontinuity, and cross-modal mismatch, and repairs problematic clips in a stage-specific manner. Experiments across multiple video generation backbones show that DramaAgent improves long-horizon coherence, character consistency, narrative fidelity, and scene-level audio-visual consistency over direct generation and strong baselines. These results suggest that hierarchical agentic control is a practical direction for controllable long-form audiovisual generation. Code: this https URL. Website: this https URL.

201. 【2610.00069】A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction

链接:https://arxiv.org/abs/2610.00069

作者:Vivek Chavan,Jörg Krüger

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:rich procedural evidence, Work Environment Model, exo data, process or retain, Procedural State Memory

备注: Accepted for oral and poster presentation at the ACVR Workshop, ECCV 2026. Non-archival abstract; not published in the workshop proceedings. 8 pages, 1 figure

点击查看摘要

Abstract:Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspired by event segmentation theory, we detect boundaries using changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, rather than fixed windows or visual novelty alone. Each segment is abstracted into an evidence-linked event card containing actor, interval, location, action, objects/tools, pre/post state, confidence, and provenance. These event cards incrementally update the WEM, enabling compact, auditable documentation and retrieval under on-premise privacy constraints. We instantiate the design with frozen DINOv2 and VJEPA-2 encoders and a local language model, and outline evaluation criteria for segmentation quality, memory compression, retrieval fidelity, and long-horizon QA.

202. 【2610.00067】Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation

链接:https://arxiv.org/abs/2610.00067

作者:Zhaoyang Wang,Haiyong Chen,Dongying Li,Yining Wang,Huapeng Wu,Xinwei Lv,Atik Shahariar

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reliable visual inspection, aero-engine blade manufacturing, imaging conditions, Reliable visual, production lines

备注: This manuscript is Accepted at conference PRCV 2026

点击查看摘要

Abstract:Reliable visual inspection is essential for quality assurance in aero-engine blade manufacturing, where defect appearance may vary across production lines, imaging conditions, blade poses, and surface backgrounds. Such domain shifts cause a mismatch between training and deployment data and degrade the reliability of deep defect detectors in online inspection. This problem is particularly challenging because aero-engine blade images usually contain sparse defects, making pseudolabel-based adaptation vulnerable to noisy or missing predictions. To address this issue, we propose Aero-engine Blade Defect Detector (ABDD), an online adaptive detection framework based on test-time adaptation. ABDD introduces a Dual-Alignment Strategy to jointly adapt global visual style and local defect morphology by combining feature-statistics alignment with pseudo-box alignment. To reduce error accumulation from unreliable pseudo labels, an Uncertainty-aware Box Filtering mechanism evaluates pseudo boxes using classification confidence, classification entropy, and localization entropy. In addition, a lightweight Sparse Dilated Mona module enables parameter-efficient delta tuning while limiting source-domain forgetting. ABDD is evaluated on CD-AeBD and HD-AeBD under multiple domain-shift scenarios, with TTA strategies compared under a unified RT-DETR + Swin-T architecture. Experiments show that ABDD consistently improves detection robustness under domain shifts, and its practicality is further validated on an industrial inspection platform.

203. 【2610.00064】Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition

链接:https://arxiv.org/abs/2610.00064

作者:Changyi Li,Yu Xiao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:part or tool, combine a manipulation, decomposition, compositional, action classifier assigns

备注: Accepted by BMVC 2026

点击查看摘要

Abstract:Assembly actions are compositional: they combine a manipulation with a part or tool. In deployment, systems routinely encounter novel combinations of familiar components, yet an atomic action classifier assigns every unseen combination exactly zero probability by construction. The prevailing solution is verb--noun decomposition, which predicts components separately and recombines them to reach unseen actions. While widely adopted, how decomposition generalizes under compositional shift remains poorly understood. We present a systematic analysis of verb--noun decomposition across three assembly datasets (MECCANO, HAViD, and IMPACT). Although decomposition escapes the atomic ceiling, its generalization extends only partially beyond it. Unseen-composition performance remains strongly tied to the co-occurrence structure of the training data, indicating that much of the observed gain arises from interpolation within densely supported regions of the compositional space rather than from unconstrained recombination. Across datasets, failures consistently concentrate on the larger-vocabulary component, and IMPACT's verb-heavy vocabulary reverses the bottleneck from nouns to verbs. We further show that shared-encoder training introduces component entanglement, encouraging reliance on co-occurrence patterns that transfer poorly to unseen compositions and trailing independent recombination by up to $6.0\times$ in harmonic mean. Taken together, these findings explain why decomposition achieves only partial compositional generalization in practice. By identifying primitive support, vocabulary asymmetry, and component entanglement as connected sources of error, we provide a portable diagnostic framework for studying compositional recognition beyond aggregate accuracy. Code: this https URL.

204. 【2610.00040】DSSR-3D: Decoupled Reasoning for View-Dependent Referring in 3D Gaussians

链接:https://arxiv.org/abs/2610.00040

作者:Thanh-Khoi Nguyen,Thien-Phuc Tran,Minh-Triet Tran

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, Gaussian Splatting, Splatting have enabled, distilling semantic knowledge, foundation models

备注:

点击查看摘要

Abstract:Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation models into 3D representations. However, existing referring fields embed language features in a globally view-invariant space, making them fundamentally unable to resolve observer-centric spatial relations (e.g., "to the left of") that depend on camera pose. We propose DSSR-3D, an inference-time framework for view-dependent referring segmentation on continuous 3D Gaussian fields, formalized as two interfaces - pose-invariant semantic localization and pose-conditioned spatial reasoning - such that any pair of functions satisfying these constraints yields a valid instantiation, requiring no retraining of the underlying semantic field and no reliance on discrete geometric proxies such as bounding boxes. We instantiate the two interfaces with a temperature-sharpened softmax localization mechanism and a projection-based directional scoring function, fused via a lightweight, training-free step, and show they transfer zero-shot to structurally distinct semantic fields without adaptation. We further propose ViewRef-GS, a benchmark isolating view-dependent segmentation on 3D Gaussian fields, evaluated jointly with an augmented Ref-LERF to provide a comprehensive testbed for viewpoint-dependent spatial grounding. Experiments show consistent gains over existing 3DGS-based referring methods, with no additional training beyond the base semantic field

205. 【2610.00031】Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing

链接:https://arxiv.org/abs/2610.00031

作者:Kaizhen Tan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:infer urban attributes, existing urban data, imagery is increasingly, predictive accuracy, photograph contributes

备注:

点击查看摘要

Abstract:Street-view imagery is increasingly used to infer urban attributes, but predictive accuracy alone does not reveal how much a photograph contributes beyond data already available for the same place. We compare image-based predictions with existing urban data across seven attributes from five public resources and three VLMs. The same urban units are evaluated using images, task context, nearby observations, and public records, while image replacements and conflicting records test source reliance. Existing urban data matched or exceeded image-only models for road damage, curb ramps, and house price, while neighbouring official statistics nearly matched the best image result for population. Images were more informative for building type, building function, and low-rise floor count. For floor count, image advantage increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. Models frequently followed conflicting records. OpenFACADES floor annotations were generated with OpenStreetMap floor values and showed the opposite height-dependent error pattern from image-only reruns. Street-view image value therefore depends on visual legibility and local data coverage. Comparing images with existing urban data can guide image collection and clarify the provenance of derived urban maps.

206. 【2610.00030】Domain generalization and synthetic data in object detection: the enabler, the probe, and the gap

链接:https://arxiv.org/abs/2610.00030

作者:Elfi I.S. Hofmeijer,Ella P. Fokkinga,Friso G. Heslinga,Klamer Schutte,Jörgen M. Karlholm

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:experience performance degradation, operational environment, weather type, experience performance, performance degradation

备注: Submitted to SPIE Sensors + Imaging 2026

点击查看摘要

Abstract:Object detection models often experience performance degradation when deployed under distribution shifts, caused by for example changes in weather type, operational environment, or object appearance. Domain Generalization (DG) aims to develop models that remain robust under such shifts and generalize well to unseen domains. DG research specifically focused on object detection models is scarce, although these models face additional challenges around localization and multi-scale representations. Synthetic data is a promising tool to support in DG, by enabling large-scale generation of diverse new samples. In this paper, we present an object detection-centric review of DG and examine the role of synthetic data from three complementary perspectives. First, synthetic data acts as an enabler of DG through diversification and alignment strategies that aim to improve robustness to distribution shifts. Second, it serves as a probe that enables controlled experimentation to identify and understand failure modes. Third, we discuss the synthetic-to-real gap, a particularly challenging form of domain shift that arises when models trained on synthetic imagery are deployed on real-world data. Through reviewing these perspectives, we identify limitations of current DG approaches for object detection and argue that future research requires representation-aware methods that explicitly address both localization and classification under domain shift.

207. 【2610.00024】Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

链接:https://arxiv.org/abs/2610.00024

作者:Genpei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:universal negative finding, vision-language model architectures, mid-layer interpretability, vision-language model, universal negative

备注: 13 pages, 4 figures

点击查看摘要

Abstract:Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p 1e-4). Despite the null, the errors separate operationally into three failure modes -- Perception Failure, Encoded-but-Disconnected, Prior-Override -- learnable above 60% on all three architectures, and the architecture's prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.

208. 【2610.00017】Spatial Lifting for Dense Prediction

链接:https://arxiv.org/abs/2610.00017

作者:Mingzhi Xu,Tao Zhou,Yong Li,Yizhe Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:present Spatial Lifting, Spatial Lifting, dense prediction tasks, present Spatial, lifting standard inputs

备注: 28 pages 5 figures

点击查看摘要

Abstract:We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this dimensionality lifting allows us to achieve good performance on benchmark tasks compared to conventional approaches, while reducing inference costs and \textbf{drastically lowering the number of model parameters}. The SL framework produces intrinsically structured outputs along the lifted dimension. This emergent structure facilitates dense supervision during training and enables single-forward-pass self-consistency-based quality and uncertainty estimation at test time. Spatial Lifting introduces a simple and general modeling strategy that offers a promising path toward more efficient, accurate, and reliable deep networks for dense prediction tasks in vision.

209. 【2610.00006】Emergent Object Binding Has a Finite Spatial Horizon

链接:https://arxiv.org/abs/2610.00006

作者:Mayank Singal

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Pretrained Vision Transformers, Vision Transformers encode, Pretrained Vision, Vision Transformers, Transformers encode

备注: 14 pages, 4 figures

点击查看摘要

Abstract:Pretrained Vision Transformers encode whether two image patches belong to the same object. This IsSameObject signal is decodable from frozen patch embeddings at high accuracy, which suggests that object binding emerges from self-supervised pretraining alone. We show that this single accuracy number hides the structure of the signal. Binding is local: the probability that two patches of the same object are decoded as bound falls off monotonically with the distance between them and levels off at a nonzero floor, a falloff well described by an exponential with a finite length scale. This decay holds across object sizes, across three families of probe, on both ADE20K and COCO, and across DINO and CLIP backbones, which indicates that it is a property of the representation rather than of the decoder. Reading binding as local spatial coherence with a finite range accounts for a set of behaviors that the aggregate score leaves unexplained: binding weakens on large objects, separates distinct objects of the same class less reliably than objects of different classes, and groups object parts with their wholes. It is, by contrast, unaffected by occlusion once object size is controlled. We map each behavior with confounds controlled. As a preliminary observation, the horizon and its floor are organized at different depths in DINOv2 and DINOv3, which we report as suggestive given the small number of layers probed and the confound between the two models.

210. 【2610.00003】STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

链接:https://arxiv.org/abs/2610.00003

作者:Animesh Varma

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)

关键词:Vision models pretrained, Vision models, infer hidden physical, hidden physical properties, properties from motion

备注: 17 pages, 7 figures, 3 tables. Preprint

点击查看摘要

Abstract:Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth. In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%. These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.

211. 【2609.39564】A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

链接:https://arxiv.org/abs/2609.39564

作者:Seonho Lee,Wonryeol Jeong,Alberto Cereser,Inha Kang,Hyeonjong Kim,Seungmin Kwak,Dongmin Park

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Delegating complete application, simply producing plausible, producing plausible outputs, Delegating complete, Game Design Documents

备注:

点击查看摘要

Abstract:Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at this https URL.

212. 【2610.00860】MorphoBranch: A Fine-Structure-Preserving Workbench for Morphometric Analysis of Branched Cellular Structures

链接:https://arxiv.org/abs/2610.00860

作者:Song Zhiying,Ling Hanyi,Wu Junyi,Jiang Yangbo

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Quantitative Methods (q-bio.QM)

关键词:Fluorescence-labeled cellular arbors, Background and Objectives, bias skeleton-based measurements, weakly labeled processes, Fluorescence-labeled cellular

备注: 12 pages, 9 figures

点击查看摘要

Abstract:Background and Objectives: Fluorescence-labeled cellular arbors provide readouts of neuronal and microglial morphology, but fine and weakly labeled processes are prone to fragmentation and false connections that bias skeleton-based measurements. We present MorphoBranch, a fine-structure-preserving, human-reviewable workbench for morphometry of branched cellular structures. Methods: MorphoBranch combines a deterministic Morphometry Engine with an LLM-assisted Refinement Engine. The Mor- phometry Engine implements an image-to-graph workflow integrating multiscale structural evidence extraction, hysteresis segmen- tation, evidence-constrained skeleton refinement, and graph-based morphometry. The Refinement Engine maps natural-language requests to registered actions for parameter adjustment, preview execution, metric reporting, and unsupported-request handling, while image processing and quantitative computation remain deterministic and reviewable. Results: MorphoBranch was evaluated on two public neuronal axon datasets, AxonMIP and AxonStack, and the in-house Cell- Morph dataset of microglial fluorescence images. It achieved the highest Skeleton F1 and clDice and the lowest length-estimation error among the evaluated methods on all three datasets, while also achieving the highest Dice and IoU on AxonMIP and Axon- Stack. Across 150 natural-language tasks, the Refinement Engine achieved a 94.0% end-to-end success rate. Conclusions: These results demonstrate that MorphoBranch provides a reproducible, human-reviewable workflow for mor- phometric analysis of branched cellular structures. It supports fine-structure-preserving quantification across neuronal axon and microglial fluorescence images while maintaining inspectable and reproducible analysis workflows.

213. 【2610.00805】Spatially Gated Diffusion for Localized Counterfactual Chest Radiograph Editing

链接:https://arxiv.org/abs/2610.00805

作者:Kamran Ullah Afaq,Basit Raza

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:preserving unrelated content, chest radiograph requires, radiograph requires completing, unrelated content, chest radiograph

备注: 16 pages, 1 figure, 10 tables

点击查看摘要

Abstract:Editing a chest radiograph requires completing the requested change while preserving unrelated content. We study a latent diffusion editor with an instruction-independent source trajectory and an instruction-conditioned editing trajectory. A learned gate mixes their post-sampler candidates at each executed step, and a separate image-space mask composites the decoded proposal with the source. On 2,400 MIMIC-derived requests, 2,244 outputs met the joint target, preservation, quality, and coverage rubric (93.5%), and target completion was 97.4%. Joint validity exceeded a matched composition-only control by 1.7 percentage points (paired patient-cluster 95% interval, 0.6--2.8). At a fixed learned mask, the learned-gate proposal improved joint validity by 1.6 points (0.8--2.4); at a fixed learned proposal, the learned mask improved it by 2.6 points (1.7--3.5). Across three training seeds, mean joint validity was 93.5% with a 0.3-point sample standard deviation. A blinded 240-request assessment yielded adjudicated joint validity of 93.3% for the full editor and 91.7% for composition only. Protected-region mean absolute error decreased from 0.0190 in the raw proposal to 0.0075 after composition. These findings distinguish recurrent-gating effects on the proposal from preservation through final composition in the assessed cohort.

214. 【2610.00384】RIQE: a NIQE-style reference model for Computed Tomography

链接:https://arxiv.org/abs/2610.00384

作者:Fabio Mattiussi

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Image Quality Evaluator, Natural Image Quality, Quality Evaluator, Radiology Image Quality, Image Quality

备注: 17 pages, 6 figures, 6 tables. Code and model: [this https URL](https://github.com/Metallogik/RIQE) , archived at doi: [https://doi.org/10.5281/zenodo.23055559](https://doi.org/10.5281/zenodo.23055559)

点击查看摘要

Abstract:The Natural Image Quality Evaluator (NIQE) scores an image by its statistical distance from a model fitted on pristine images, and its distributed model is fitted on photographs. We release the Radiology Image Quality Evaluator (RIQE), a NIQE-style model fitted on 3,792 full-dose slices from 158 patients of the public LDCT-and-Projection-data collection, with a declared intensity mapping, a manifest of every slice and a script that reproduces the fit. On 40 held-out patients, RIQE ranks reduced-dose reconstructions, simulated by projection-domain noise insertion, worse than the full-dose reconstruction of the same slice in 240 of 240 chest and 230 of 240 abdominal pairs, and ranks images with 20% more noise worse than their source in 97.5-100% of cases. Its preferences among filtered images, however, do not follow lesion signal. With a 4 mm, +10 HU lesion inserted in noisy abdominal slices, RIQE prefers bilateral filtering to the unfiltered image in every image up to a 32 HU residual, at which 29% of the lesion's matched-filter signal remains and its detectability index falls from 0.51 to 0.33; it never prefers Gaussian smoothing, which at the same 32 HU residual leaves 70% of the signal and a detectability index of 0.48. Fitted on photographs with the parameters published for NIQE, the same code ranks every simulated reduced-dose abdominal image better than its full-dose counterpart. RIQE is suited to ranking a degraded image against its source; under the conditions tested it should not be the sole criterion for selecting, comparing or tuning denoisers.

215. 【2610.00318】LensBridge: Frequency-Guided Compound Degradation Adaptation for Lens Aberration Correction and Veiling Glare Removal

链接:https://arxiv.org/abs/2610.00318

作者:Xiaolong Qian,Zhonghua Yi,Qi Jiang,Kailun Yang,Shuhang Xie,Shaohua Gao,Kaiwei Wang

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Optics (physics.optics)

关键词:exhibit residual lens, Point Spread Functions, spatially varying blur, Veiling Glare, residual lens aberrations

备注: All code will be available at [this https URL](https://github.com/XiaolongQian/LensBridge)

点击查看摘要

Abstract:Simplified optical systems often exhibit residual lens aberrations and Veiling Glare (VG), resulting in spatially varying blur and contrast reduction. Large-scale Lens Libraries (LensLib) enable reusable aberration correction models by covering diverse Point Spread Functions (PSFs), but their aberration-only training distribution does not include target-specific veiling glare. Extending such foundations to compound degradation is challenging because realistic target-system compound pairs are difficult to obtain. To address this challenge, we propose LensBridge, a two-stage framework that first establishes a reusable aberration correction foundation and then adapts it to compound optical degradation using only a few unpaired target observations. In Stage I, we build a PSF-aware one-step diffusion foundation by constructing discrete degradation priors from LensLib PSFs and learning to retrieve them directly from aberrated images, enabling PSF-aware correction without requiring explicit PSF at inference. In Stage II, we adapt this foundation to compound degradation through frequency-domain guidance. At the data level, Frequency-guided Degradation Completion (FDC) transfers target low-frequency characteristics to LensLib aberrated images while preserving aberration structures to synthesize compound training pairs; at the model level, Frequency-guided Pseudo Decomposition (FPD) forms aberration- and VG-dominant pseudo observations to condition separate adaptation branches. Extensive experiments across multiple optical systems demonstrate that LensBridge effectively extends reusable aberration correction foundations to joint aberration correction and veiling glare removal without target-system paired supervision. All code will be available at this https URL.