本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新655篇论文,其中:

  • 自然语言处理119
  • 信息检索25
  • 计算机视觉112

自然语言处理

1. 【2609.04199】Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

链接https://arxiv.org/abs/2609.04199

作者:Yuntian Deng,Pengyu Nie,Stuart Shieber

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:input introduces repeated, large remote model, recurring text functions, introduces repeated cost, implement with rules

备注: EMNLP 2026 System Demonstrations. Demo: [this https URL](https://programasweights.com)

点击查看摘要

Abstract:Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which the Program-as-Weights fast compiler produced no exact matches, compile by training reaches 83.6% semantic accuracy. This higher accuracy comes with a higher compile-time cost: roughly a minute rather than seconds for the fast compiler. We deploy the compiler in a public interactive service and demonstrate compiled functions in a multi-site website helper, a language-controlled 3D avatar, and a bidirectional English-Claudish translator.

2. 【2609.04197】ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

链接https://arxiv.org/abs/2609.04197

作者:Lihao Liu,Peng Tang,Kunwar Yashraj Singh,Shabnam Ghadar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Evolutionary prompt optimizers, iteration appends rules, Evolutionary prompt, rules and caveats, iteration appends

备注: EMNLP 2026

点击查看摘要

Abstract:Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; Propose generates candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection. On seven public NLP benchmarks - Tweet, MMLU, GSM8K, HotpotQA, ScoNe, HoVer, and PUPA - ESPO improves average accuracy by $+$3.76 pp over the state-of-the-art (74.67% vs 70.91% for GEPA), matching or exceeding GEPA on every dataset while producing prompts 47% shorter (1,004 vs 1,878 chars) and faster at inference. Cross-model experiments across four additional student models (Gemma 3 12B, Mistral 14B, Qwen3 32B, Claude Haiku 4.5) show ESPO yields the best average accuracy on every model tested, with the largest gap on Qwen3 GSM8K (15.00% $\to$ 91.40%). A generalization bound (Appendix) grounds each phase in a corresponding term of the test-time gap, and the ablation confirms a key prediction: adding diversity without bootstrap selection actually hurts performance ($-$1.20%).

3. 【2609.04194】Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

链接https://arxiv.org/abs/2609.04194

作者:Kevin Du,Alexander Hoyle,Laura Ruis,Acyr Locatelli

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:offer a legible, legible window, Reasoning, model arrives, reasoning step

备注: Published at COLM 2026

点击查看摘要

Abstract:Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.

4. 【2609.04180】Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

链接https://arxiv.org/abs/2609.04180

作者:Joseph Lee,Yidi Huang,Dokyoon Kim,Shu Yang,Li Shen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:auxiliary views, Gaps remain, acquire knowledge, large language models, knowledge

备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.

5. 【2609.04173】Last Translation Benchmark

链接https://arxiv.org/abs/2609.04173

作者:Vilém Zouhar,Niyati Bafna,Mukund Choudhary,Maike Züfle,Sara Rajaee,Pinzhen Chen,Jannis Vamvas,Sara Papi,Ona de Gibert,Bhavitvya Malik,Eliya Habba,Orfeas Menis Mastromichalakis,Patrícia Schmidtová,Michelle Wastl,Sheriff Issaka,Leshem Choshen,Stella Biderman,Antonis Anastasopoulos,Jan Niehues,Rico Sennrich,Mrinmaya Sachan,Ondřej Bojar,Kenton Murray,Jörg Tiedemann,Alham Fikri Aji,Philipp Koehn,Christof Monz,Alexandra Birch,Sowmya Vajjala,Chalamalasetti Kranti,Cristina España-Bonet,Nobin Sarwar,David Kaczér,Shunta Asano,Malik Marmonier,Daban Q. Jaff,Vaisakhi Mishra,Hend Al- Khalifa,Gabriele Sarti,Sourajit Saha,Nils Rehlinger,Juan Daniel Cuervo Villa,Jonathan Tonglet,Saugata Purkayastha,Dominik Macháček,Jagannathan Ramanujam,Heejin Do,Zuzana Nadova,Fred Philippy,Fabian Retkowski,Maria Lymperaiou,Silvia Casola,Hanna Yukhymenko,Shubhashis Roy Dipta,Sangwon Ryu,Andrés Jerez,Ron Keinan,Shuaib Shuaib Yusuf,Avantica Vempati,Maria Carmen Staiano,Sukannya Purkayastha,Adrian Cosma,Vitalii Babenko,Erivan Inan,Aviral Nigam,Wafa Aissa,Fatima Haouari,Venkata Prasanth Kumar Gummadi,Mehdi Jafarzadeh,Valentin Scourneau,Lukas Edman,Kaiser Sun,Shaomu Tan,Mohammad Sadegh Gholizadeh,Johannes-Rudolf David,Dipankar Srirag,Javier García Gilabert,Ruta Binkyte,Manar Ali,Ana-Maria Bucur,Sabry E. Farrag,Youssef Saber,Yihong Liu,Jean Maillard,Cojocaru Nicoleta,Xiaochuang Yuan,Sina Ahmadi,Philipp Mondorf,Kaustubh Dhole,Roman Wixinger,Shenbin Qian,Manuel Tuor,Sergey Troshin,Jonathan Yahav,Fida Mohammad Thoker,Amir Arsalan Rezapour,Lance Calvin Lim Gamboa,Manon Reusens,Kätriin Kukk,Koel Dutta Chowdhury

类目:Computation and Language (cs.CL)

关键词:test the limits, methods that inform, translation, Translation Benchmark, scientific progress

备注: typeset in Typst

点击查看摘要

Abstract:For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

6. 【2609.04172】Rethinking On-Policy Distillation of Large Language Models II: One Training Example

链接https://arxiv.org/abs/2609.04172

作者:Zixuan Fu,Bingxiang He,Yuxin Zuo,Haohuan Huang,Jinqian Zhang,Ruhang Xiao,Cheng Qian,Qinyu Luo,Huan-ang Gao,Yudong Wang,Zhiyuan Liu,Ning Ding,Chaojun Xiao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:On-policy distillation, combines student-generated rollouts, dense token-level supervision, OPD, combines student-generated

备注: 29 pages, 20 figures

点击查看摘要

Abstract:On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.

7. 【2609.04148】rminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

链接https://arxiv.org/abs/2609.04148

作者:Jie Wu,Zhenru Zhang,Beichen Zhang,Xuwu Wang,Yuhui Su,Mouxiang Chen,Peng Wang,Zhihai Wang,Que Shen,Hao Zhou,An Yang,Fei Huang,Yujiu Yang,Dayiheng Liu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:environments remain scarce, executable environments remain, terminal-based code agents, remain scarce, terminal-based code

备注

点击查看摘要

Abstract:As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.

8. 【2609.04108】Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

链接https://arxiv.org/abs/2609.04108

作者:Boyan Li,Bingsen Chen,Chenghao Yang,Ping Nie,Chen Zhao,Xi Ye

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:post-training reasoning LLMs, Reinforcement learning, on-policy distillation, verifiable rewards, OPD

备注

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@$k$ behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to this http URL provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.

9. 【2609.04083】CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

链接https://arxiv.org/abs/2609.04083

作者:Tingyu Song,Mingxin Li,Yanzhao Zhang,Dingkun Long,Chu Liu,Pengjun Xie,Yilun Zhao,Shu Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:models remain limited, MLLM-based embedding models, embedding models remain, attribute-object bindings, remain limited

备注

点击查看摘要

Abstract:MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.

10. 【2609.04061】When Models Edit Too Much: On the Fidelity of Minimal Code Edits

链接https://arxiv.org/abs/2609.04061

作者:Tongyao Zhu,Wei Hern Lim,Min-Yen Kan

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:original implementation, Large language models, edit existing code, existing code, Large language

备注: EMNLP 2026 (Main)

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.

11. 【2609.04048】ranslation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

链接https://arxiv.org/abs/2609.04048

作者:Hasan Alkhder,Mohammad Abboush,Igor Tchappi,Ahmet Zengin,Amro Najjar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:systems typically produce, Neural machine translation, Neural machine, alternative decision trajectories, decision trajectories implicitly

备注

点击查看摘要

Abstract:Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may differ in lexical authenticity, register, and structural stability. We propose reframing translation as a structured decision space explored by autonomous translation agents. Instead of analyzing a single output, we model distinct translation pathways as agents operating over a shared multilingual backbone. Inter-agent divergence is treated not as error but as an interpretable behavioral signal. We conduct an empirical study on Turkish--Syrian Arabic translation using three agents: (1) zero-shot direct translation, (2) dialect-stabilized translation via lightweight fine-tuning, and (3) pivot translation through English. Evaluation is performed on 5,000 dialogue sentences, while stabilization is trained on 5,000 additional Turkish--Syrian sentence pairs drawn from television dialogue and MADAR-Turk resources. Rather than optimizing for conventional performance metrics, we quantify structured behavioral displacement using dialect marker frequency, lexical proximity to standardized Arabic, and structural variance. Lightweight stabilization nearly doubles dialect marker usage, increasing it from 0.2266 to 0.4988, while significantly reducing structural instability. Pivot mediation introduces normalization pressure and measurable compression effects, whereas zero-shot translation exhibits the highest decision variance. We argue that translation divergence across agents reveals latent decision flexibility within multilingual models and we provide a principled interpretability framework for low-resource dialect generation.

12. 【2609.04047】he Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations

链接https://arxiv.org/abs/2609.04047

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:LLM brand recommendations, audit stochastic variation, standardized protocol exists, Researchers increasingly, Dice Roll Method

备注: 30 pages, 2 figures, 19 tables. Substantially revised; supersedes the Research Square preprint [https://doi.org/10.21203/rs](https://doi.org/10.21203/rs) . [this http URL](http://3.rs) -8883056/v1. Includes a pre-registered external validation on three independent corpora (Motoki et al., Rozado, llm-stability)

点击查看摘要

Abstract:Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.

13. 【2609.04034】Editable Visual Design

链接https://arxiv.org/abs/2609.04034

作者:Junyan Ye,Wei Liu,Dongzhi Jiang,Zichen Wen,HaoDong Li,Zhutao Lv,Jiaxin Lin,Jinhua Yu,Jun He,Zilong Huang,Rui Chen,Weijia Li

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:precluding layer-wise post-editing, Nano-Banana exhibit remarkable, inherently yields flattened, yields flattened bitmaps, remarkable visual expressiveness

备注

点击查看摘要

Abstract:While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

Cite as:
arXiv:2609.04034 [cs.CV]

(or
arXiv:2609.04034v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.04034

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
14. 【2609.04024】Instruction Duplication as an Inference-Time Control Primitive

链接https://arxiv.org/abs/2609.04024

作者:Victor Lavrenko(PeaceTech VC, Israel)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:controllable language-model systems, Procedural instruction, basic requirement, requirement for controllable, controllable language-model

备注: 7 pages, 2 tables. Code and frozen reproduction artifacts: [this https URL](https://github.com/victorlavrenko/answer-engineering/releases/tag/instruction-duplication-arxiv-v1)

点击查看摘要

Abstract:Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, and 16,800 scheduled generations, moving from one to two copies raises the deterministic All-8 diagnostic--responses passing all eight observable tests--from 90.22% to 93.17% (+2.95 percentage points), eliminating 30.2% of the failures remaining after one copy. Pre-provisional TF-IDF recall rises from 73.44% to 74.81% (+1.38 points; Holm-adjusted p .001), while final-answer accuracy remains exactly 60.21%. Premature commitment increases from 1.52% to 2.30% (p_Holm = .00536). A blinded challenge audit yields 10/30 directional confirmations, 20/30 perceptual ties, and no reversals; its prespecified 28/30 confirmation criterion is not met. Yet this distinction can matter operationally when a downstream system acts on the generated trajectory. In Answer Engineering (AE), where explicit trajectory state determines local repair, the published reason-first no-editing SSNHL endpoint was 25.1%; system-only AE was later reproduced at 84.2%, and the same trailing duplicate raised it to 97.1%. For conductive diagnostic branch preservation, the corresponding values are 58.9% published without editing, 78.6% with reproduced AE, and 73.8% with AE plus duplication--a within-AE decrease, but still 14.9 points above the no-editing baseline. Instruction duplication is therefore a low-complexity, placement-sensitive control whose practical value can emerge through the downstream system that consumes the exposed trajectory.

15. 【2609.04022】Representational alignment yields generalizable safety in language models

链接https://arxiv.org/abs/2609.04022

作者:Lingyu Li,Yan Teng,Yingchun Wang,Xia Hu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Aligning large language, Aligning large, large language models, safe deployment, large language

备注

点击查看摘要

Abstract:Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.

16. 【2609.03992】Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

链接https://arxiv.org/abs/2609.03992

作者:Sanyuan Chen,Min-Jae Hwang,Sho Inoue,Anna Sun,Bokai Yu,David Kant,Dongmin Hyun,Dorian Desblancs,Gregory Antonovsky,Oleg Repin,Peng-Jen Chen,Xutai Ma,Zehai Tu,Juan Pino,Wei-Ning Hsu

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:full-duplex dialogue synthesis, present Alignment-Free Text-Audiobox, dialogue synthesis, Diffusion Transformer trained, full-duplex dialogue

备注

点击查看摘要

Abstract:We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.

17. 【2609.03985】IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition

链接https://arxiv.org/abs/2609.03985

作者:Nazim-E-Alam,Tarek Rahman,Md Kishor Morol

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Zero-shot vision-language models, training-free species recognizers, visual species knowledge, vision-language models, multilingual Jina CLIP

备注

点击查看摘要

Abstract:Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.

18. 【2609.03967】Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes

链接https://arxiv.org/abs/2609.03967

作者:Revathy Venkataramanan,Aditya Luthra,Venkatesan Nadimuthu,Amit Sheth

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Large Language, yielding positive outcomes, ability of Large, yielding positive

备注

点击查看摘要

Abstract:Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positive outcomes. These models can process natural language inputs and leverage learned knowledge from their pretraining to generate meal plans. In this work, we investigate the ability of LLMs to analyze the suitability of given recipes for diabetes. The primary challenge for LLMs is to retrieve relevant dietary guidelines for diabetes, decompose recipes into ingredients and cooking methods, and apply these guidelines to determine the recipe's suitability. To study these challenges, we employ three kinds of prompts namely, (i) Direct Query Prompt (ii) Context-Guided Prompt, and (iii) Exemplary Context Prompt that incorporate different levels of diabetes dietary guidelines from medical sources. We introduce a benchmark dataset curated for this investigation consisting of 7607 recipes that include 3807 recipes suitable for diabetes and 3800 recipes not suitable for diabetes. Our results demonstrate that most LLMs are cautious in predicting recipes as suitable to prevent detrimental outcomes. Further, the models that can reason using the dietary guidelines performed better in predicting the suitability of recipes for diabetes. Overall, Mistral-7B and Llama 70B showed superior performance to their counterparts.

19. 【2609.03960】FiMI Banking: A Sovereign Model for Indian Retail Banking

链接https://arxiv.org/abs/2609.03960

作者:NPCI AI Research Team:Aman Kumar,Asit Desai,Chandra Bhushan,Harsh Sharma,Harshit Bhushan,Hrithik Kadam,Keyur Doshi,Kolisetty Sai Kapardheeswar,Krishanu Adhikary,Nadeem Shaik,Navya Prakash,Nitin Kukreja,Prashant Devadiga,Shamanth MH,Shantanu Pandey,Suvradip Paul,Yatharth Dedhia

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:answer product questions, Banks need conversational, product questions, account-related requests, regulatory constraints

备注

点击查看摘要

Abstract:Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fall short when a task requires grounded information, correct tool use, or cautious handling of bank-specific sensitive situations. We introduce FiMI Banking, a controlled Indian retail-banking setting. We build it from vetted banking documents, structured ground truth, synthetic customer backgrounds, and banking tools. We evaluate two post-training approaches: preference optimization for response-level behavior, and reinforcement learning with verifiable rewards for multi-turn tool-use tasks. Preference optimization improves safe behavior substantially: out-of-scope refusal rises from 52% to 80%. Reinforcement learning improves edge-case performance from 0.509 to 0.718 and order-sensitive task performance from 0.590 to 0.679, while using 29% fewer generated tokens. These results show that preference optimization and verifiable-reward reinforcement learning address complementary requirements for reliable banking agents.

20. 【2609.03955】wo-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

链接https://arxiv.org/abs/2609.03955

作者:Jiacheng Xu,Wentao Zhang,Zhiyi Lyu,Fuxiang Zhang,Chaojie Wang,Yang Liu,Bo An

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:substantially advanced code, Reinforcement learning, large language models, test cases, advanced code generation

备注: 21 pages, 7 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

点击查看摘要

Abstract:Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as counterexamples, depending on the solver's current failure modes. We propose Test Cases Scaling (TCS), a two-stage RL framework for effective test generation. Both stages train a test generator from a rolling policy-aligned buffer: Stage 1 generates tests consistent with the reference solution, and Stage 2 restricts the buffer to current failure modes and learns counterexample tests. Across TACO and LiveCodeBench, TCS improves both pass@1 and inference-time answer selection according to generated tests. We find the learned test generator also enables effective selection among other LLM outputs.

21. 【2609.03953】Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

链接https://arxiv.org/abs/2609.03953

作者:Joe Cecil,Marjorie Freedman

类目:Computation and Language (cs.CL)

关键词:Understanding the frequency, determining chatbot safety, evaluating systems, systems that detect, critical for determining

备注: 34 pages, 6 figures, to be published in Findings of the ACL: EMNLP 2026

点击查看摘要

Abstract:Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on its own: It misses factual errors that annotators catch. We also find disagreement among adjudicators, suggesting that adjudication over multiple candidate sources can improve benchmark completeness, but does not eliminate the need to apply judgment and expertise. Applied to an existing benchmark, this technique reveals a similar pattern of missing annotations. Together, these results suggest that in the settings examined here, single-pass hallucination benchmarks may achieve scale at the cost of undercounting factual errors. Multi-pass adjudication can improve coverage, but inferences drawn from the benchmarks are still sensitive to the judgment, expertise, and evidence used to determine error presence.

Comments:
34 pages, 6 figures, to be published in Findings of the ACL: EMNLP 2026

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.03953 [cs.CL]

(or
arXiv:2609.03953v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.03953

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
22. 【2609.03949】VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch

链接https://arxiv.org/abs/2609.03949

作者:WenJie Fan

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Kimi Linear, cache, NoPE MLA model, NoPE, Kimi

备注: 12 pages, 4 figures, 10 tables

点击查看摘要

Abstract:The problem. A long-lived KV cache must be compressed before the queries that will read it exist; selection by observed attention (H2O, SnapKV) collapses there (0.00-0.33 needle retrieval on a NoPE MLA model), because a token's importance has not yet been observed. The method. On Kimi Linear, VestigeKV evicts by a query-independent signal the cache already carries: the 64-dimensional decoupled branch, a vestige of RoPE that NoPE training repurposes into a salience channel. Reading 11% of each row, it partitions the cache: the top-m rows stay in the attended tier; every other row moves -- exactly, never deleted -- to a GPU-resident archive reachable per step by a certified trigger. No training, no quantization, no weight or kernel change. Cost. Nothing measurable: retrieval holds at 1.00 under 8x and 0.92 under 32x from 8k to 65k context, zero gap to full-row selection. The attended tier is 0.25 KB of Kimi Linear's 8.1 KB per-token cache at 32x; the archive stays bit-exact and GPU-resident, with host offload as the VRAM-reclaiming variant. The recall tier -- the standard configuration -- holds 128x at 1.00. Kimi K3 is reported to use a NoPE Gated-MLA variant; if its cache layout matches, the method plausibly extends there -- we make no claim beyond the measured model. NoPE exclusivity. The identical operator on a RoPE MLA collapses to 0.08 (plain eviction: 0.42); query-independent salience itself exists only without rotation (top-1 targets span 2.3-6.7% of tokens vs. 10.2-46.8%), and query-universal exact merging is provably impossible under RoPE. All thresholds were frozen before data; 20 archived verdicts and 8 closed routes accompany the paper.

23. 【2609.03943】More Criticism Does Not Make a Better Review: EquiReview-R

链接https://arxiv.org/abs/2609.03943

作者:Zexing Zhang,Jichao Li,Tianyang Lei,Yude Fu,Yang Kewei

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:specific criticisms, produce many specific, review, Abstract, criticisms

备注

点击查看摘要

Abstract:AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with omission and overcritique treated as separate risks. Building on this formulation, we introduce EquiReview-R, which resolves existing concerns against localized evidence, searches for missing issues from independent and review-conditioned perspectives, and returns stop, continue, or defer. To expose the failure mode that motivates this design, we construct an evidence-linked trajectory corpus. Its retrospective analysis shows why revision must precede further search: nearly all concerns in a high-recall review lack a definitive evidential disposition, while an earlier refinement mechanism cannot revise them. On a frozen cohort of previously unseen papers, EquiReview-R satisfies the prespecified non-inferiority criterion for major omission, reduces major overcritique from 15.5% to 8.1%, and attains a one-sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation-matched controls, controlled pairs, and ablations show that the gain comes from revision rather than extra inference or shorter output. We release the corpus as ReviewTrace, an evidence-linked resource for studying review revision, disagreement, and provenance.

24. 【2609.03941】Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO

链接https://arxiv.org/abs/2609.03941

作者:Hyun Bin Park,Du-Seong Chang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:repeated fresh rollout, RL-based post-training, models is increasingly, increasingly bottlenecked, bottlenecked by repeated

备注: 51 pages, 25 figures, 17 tables. Accepted at COLM 2026

点击查看摘要

Abstract:RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates them by compatibility with the current policy. The fresh on-policy stream remains unchanged, and the method adds no auxiliary generation or training machinery. Across mathematical reasoning, multimodal reasoning, and Agentic Search benchmarks, this single intervention outperforms naive replay and matches or exceeds broader replay methods on Avg Mean@32. In Agentic Search, where environment interaction dominates cost, it delivers comparable quality at materially lower wall-clock time.

25. 【2609.03930】Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords

链接https://arxiv.org/abs/2609.03930

作者:Yelingyun Zhang,Atis Kapenieks,Marina Platonova

类目:Computation and Language (cs.CL)

关键词:Suffix Dependency Ratio, Existing research, Fixed Suffix Dependency, tendency for English, fixed morphological rules

备注

点击查看摘要

Abstract:Existing research has repeatedly observed the tendency for English loanwords to cluster in the masculine gender across different recipient languages, yet the origin of this pattern remains difficult to determine, as fixed morphological rules and default assignments are frequently analysed together. This study proposes the Fixed Suffix Dependency Ratio (FSDR) to quantify the degree of reliance on fixed derivational suffixes across different genders, and to distinguish between morphological anchoring and free-choice in distribution. By examining 1,832 Latvian noun lemma types, the results reveal a significant FSDR asymmetry within the loanword system: feminine loanwords rely significantly more on fixed derivational suffixes, while masculine loanwords are more concentrated in the free-choice zone. This pattern exhibits loanword specificity and has become more pronounced in contemporary usage. FSDR therefore provides a quantitative framework for testing default gender and shows how masculine default can be activated and reinforced under language contact.

26. 【2609.03923】Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

链接https://arxiv.org/abs/2609.03923

作者:Muneeb Khan,Frederic Kirstein,Terry Ruas,Bela Gipp

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:LLM agents fail, online meeting delegation, Collaborative Agent Predictive, fail to recognize, Agent Predictive Architecture

备注: Accepted at EMNLP 2026 Main

点击查看摘要

Abstract:In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking opportunities on the AMI corpus. We present CAPA (Collaborative Agent Predictive Architecture), an architecture for online meeting delegation. A Perceiver updates the meeting state from each observed turn. A Predictor forecasts how the conversation will continue. A Controller decides whether to speak and which proposition to surface. A Generator phrases the chosen contribution in the participant's style. Two judges score the forecast and the action against the next observed turn. A Recalibrator updates the meeting state from those verdicts for future decisions. To evaluate online delegation, we introduce an episode-level protocol that scores whether, when, and what a delegate contributes around the participant's actual idea units. The protocol's schema-constrained LLM judges align with human annotations at Cohen's kappa = 0.71. On 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery (26.1 -- 52.2), and keeps hallucination at 0.6%. The failure mode shifts from omission to selection, with each residual near-miss attributable to a specific module of the architecture. Mechanism ablations identify the meeting state as the lever that closes the recognition gap, where raw-context scaling alone does not.

27. 【2609.03915】RuleMem: Active Rule Memory for Long-Term Conversational Agents

链接https://arxiv.org/abs/2609.03915

作者:Xingyuan Zeng,Zuohan Wu,Quanming Yao,Yue Wang,Wei Liu,Libin Zheng,Jiuke Wang,Jian Yin

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Question answering agents, dispersed dialogue histories, temporally dispersed dialogue, Question answering, reason over massive

备注

点击查看摘要

Abstract:Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue histories. However, existing memory mechanisms primarily treat past information as \textit{passively} stored facts, leading to semantic gaps and unreliable reasoning. To address this limitation, we propose RuleMem, a rule-based memory framework that induces reusable logical rules from historical interactions to \textit{actively} guide both evidence retrieval and reasoning. Specifically, RuleMem constructs natural-language Horn clauses from conversations and validates them via a Rule Perplexity Consistency (RPC) mechanism. These induced rules enable the retrieval of semantically distant evidence while providing an explicit logical structure for answer generation. We conducted a comprehensive evaluation of RuleMem on two long-term conversational benchmarks, LoCoMo and LongMemEval_s*. In a rigorous comparison against 14 baselines on LoCoMo, RuleMem achieved the highest accuracy, exceeding the baseline average by 27.47 points (a 54.3% relative improvement).

28. 【2609.03894】CROCODIL: Cross-Model Code Editing with LLMs

链接https://arxiv.org/abs/2609.03894

作者:Linghan Zhong,Aditya Thimmaiah,Jayanth Srinivasa,Milos Gligoric,Junyi Jessy Li

类目:Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:ubiquitous tools, code, Large language models, foreign code, LLMs

备注: EMNLP 2026 Findings

点击查看摘要

Abstract:Large language models (LLMs) have become ubiquitous tools for code generation and editing. However, development teams often use multiple LLM assistants. Different developers may prefer different models, and individual developers may switch between models across different coding sessions. Because of this, the edits any one model makes are frequently applied to foreign code originally generated by another model. These LLMs are often trained on different datasets, and as a result have different stylistic preferences. Do LLMs behave differently when they edit foreign code originally written by a different LLM with a different coding style? We find that models tend to make more, and often excessive, edits on foreign code. We introduce CROCODIL (Cross-model Code Editing with LLMs), a post-training framework for reducing excessive edits while preserving functional correctness. CROCODIL's similarity reward penalizes large changes, while its execution reward scores build and test success. We use the product of these two rewards to encourage the policy to decrease the edit size without decreasing the edit task success rate. CROCODIL is available at this https URL.

29. 【2609.03887】Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

链接https://arxiv.org/abs/2609.03887

作者:Hoang Cuong Nguyen,Mark Dras,Usman Naseem

类目:Computation and Language (cs.CL)

关键词:refuse harmful requests, harmful requests shape, train language models, train language, refuse harmful

备注: 27 pages, accepted at EMNLP 2026 Main Conference

点击查看摘要

Abstract:How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in this https URL.

30. 【2609.03844】Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference

链接https://arxiv.org/abs/2609.03844

作者:Simone Ceppi,Ignacio Sanchez

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large Language Models, per-token Bernoulli trials, independent per-token Bernoulli, introduce Stateless Bernoulli, Large Language

备注: Accepted at EMNLP 2026 Main Conference

点击查看摘要

Abstract:We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token against a counter-based random number generator, reducing membership complexity to $O(1)$ and enabling single-kernel execution with zero intermediate allocations. We prove that this formulation preserves the same detection guarantees as fixed-size green lists: the z-score test remains $\mathcal{N}(0,1)$ under the null. The stateless architecture enables capabilities unavailable to existing methods: full-vocabulary self-salt watermarking (over 6000$\times$ faster than KGW's self-salt and 2$\times$ faster than SynthID despite biasing the entire vocabulary with candidate-dependent seeding) and architectural compatibility with distributed inference. In end-to-end generation benchmarks, SBW adds less than 1\% overhead at all batch sizes. We additionally identify hash function design as a previously unexplored axis for watermark quality, showing that a GPU-native Jenkins hash improves null calibration by 1.8$\times$ while producing more diverse text. Experiments across two seeding schemes and eight $(\gamma, \delta)$ configurations confirm statistical equivalence with ROC-AUC differences below 0.01.

31. 【2609.03820】Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

链接https://arxiv.org/abs/2609.03820

作者:Prakhar Khatri

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Long-video language models, small fixed slice, hour sampled, Orthogonal Matching Pursuit, Long-video language

备注: 16 pages, 6 figures. Code and data: [this https URL](https://github.com/codeprakhar25/omp-keyframe-sampling)

点击查看摘要

Abstract:Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.

32. 【2609.03814】Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

链接https://arxiv.org/abs/2609.03814

作者:Danting Zhang,Bei Peng,Robert Loftin

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, Large, content, moderation

备注: Accepted by EMNLP Findings 2026

点击查看摘要

Abstract:Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures at the criterion level. Models struggle most when correct decisions depend not on overall harmfulness, but on the specific aspect of the content that the criterion requires them to assess. Our results highlight a key limitation of current content moderation benchmarks: strong performance on aggregated labels does not provide sufficient evidence that LLMs can reliably evaluate content with respect to individual moderation criteria. These findings call for the development of evaluation methods that explicitly measure criterion-conditioned behaviour.

33. 【2609.03811】VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence

链接https://arxiv.org/abs/2609.03811

作者:JoyIndustrial VisCAD Team:Linxin Cai,Qiuhe Hong,Zhichao Huang,Guanlin Li,Hongsen Liu,Ziqi Liu,Yichen Long,Luya Wang,Yuchen Wang,Wenxiang Wu,Huimu Yu,Ning Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:AI-assisted computer-aided design, AI-assisted computer-aided, challenging phases, involves two challenging, CAD

备注: Technical report from JoyIndustrial's AI CAD project

点击查看摘要

Abstract:AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to executable programs in a CAD domain-specific language. Assembly-level generation must additionally handle interacting parts, plan mating relations, estimate poses, and place all parts correctly. Existing specialized CAD models are commonly trained on narrow input domains, such as renders or texts, and often generalize poorly, while general-purpose frontier models cover broader inputs but perform inconsistently across CAD domains. We present VisCAD, a foundation model suite designed to provide both broad generalization and strong CAD capability for realistic industrial products. At its core is VisCAD-M1, a 27B model trained through mid-training and post-training for part-level design generation. On PubCADBench and RealCADBench, VisCAD-M1 achieves the highest average part-level score among the evaluated models, reaching 0.5540 compared with 0.5496 for the strongest frontier model. Reusing VisCAD-M1 as a test-time verifier can further raise the score to 0.5797, an approximately 5 percent relative improvement over the previous state of the art. VisCAD also includes a domain-specific harness that leverages frontier models for complex assembly generation and demonstrates advantages over general-purpose harnesses in both quantitative and qualitative evaluations.

34. 【2609.03797】ransfiver: Human-AI Co-Inference through a Shared Editable State

链接https://arxiv.org/abs/2609.03797

作者:Minji Park,Seunghyun Yoon,Hyuk Lim

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)

关键词:Long-term human-AI interaction, Long-term human-AI, guides inference, inference is updated, updated implicitly

备注

点击查看摘要

Abstract:Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user. We introduce the TRANSparent Framework for Interactive, Verifiable, Editable Representation (Transfiver), an architecture for human-AI co-inference through a shared editable state. Its central idea is that interaction-specific information is maintained in a single persistent state $(S_t)$ that both the model and the human update. Transfiver distinguishes two modes of state evolution. In an implicit stream update, the model interprets ongoing interaction and decides whether new information revises an existing state item or creates a new one. In an explicit directed edit, a human inspects and modifies an addressed item. Both act on the same underlying state, so a human correction changes the state that subsequent computation reads, rather than adding another instruction or separate record. The architecture separates shared parameters $(\theta)$, learned before ordinary use, from the persistent state $(S_t)$, which evolves during deployment without parameter retraining. Extending Transfiver to rich natural-language, relational, and large-scale shared states remains open.

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)

Cite as:
arXiv:2609.03797 [cs.AI]

(or
arXiv:2609.03797v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.03797

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
35. 【2609.03788】A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval

链接https://arxiv.org/abs/2609.03788

作者:Santiago Poveda-Gutiérrez,Hideki Nakayama,Mayumi Bono

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Isolated Sign Language, Sign Language Recognition, Language Recognition, Japanese Sign Language, Sign Language

备注: 4 pages, 2 figures, 1 table. Extended version of an abstract presented at the BU-SHI workshop (Broadening the Users: A Cross-Disciplinary Roadmap for Social Humanoid Interaction), IEEE RO-MAN 2026, Kitakyushu, Japan, 28 August 2026. The workshop is non-archival; no proceedings

点击查看摘要

Abstract:Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning a sign-level clip into a free-form procedural description of the articulation with an open-weight vision-language model, and (2) retrieving the closest entry from a vocabulary of target descriptions with a multilingual sentence encoder: a reverse sign language dictionary that needs no gloss supervision and admits an open vocabulary. On 1,300 sign-level segments from a Japanese Sign Language (JSL) dialogue corpus annotated with procedural descriptions (against a 2% top-10 chance floor over the 503-entry target vocabulary), fine-tuning the captioner substantially improves seen-class retrieval: language and vision tower fine-tuning raises top-10 retrieval on seen classes from 4.5% (untrained) to 49%, becoming statistically indistinguishable from a standard supervised closed-set classifier (I3D) on two of the three test sets where a closed-set classifier can be evaluated at all. More importantly, unseen-class retrieval also improves significantly over the untrained pipeline (11.5% - 21.0% top-10, p=0.0094), a regime in which the closed-set classifier cannot participate. A matcher-side empirical upper-bound analysis shows the sentence encoder already recovers close to 100% of paraphrased gold descriptions, locating a gap in captioning quality that we aim to address in future work. To our knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.

36. 【2609.03781】IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

链接https://arxiv.org/abs/2609.03781

作者:Saikat Mondal,Mamta,Deeksha Varshney,Oana Cocarascu,Asif Ekbal

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:primarily in English, Large language models, Large language, evaluated primarily, Large

备注: 38 pages, 7 figures, 33 tables. Accepted to Findings of EMNLP 2026. Contains examples of harmful model outputs

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and risk categories. Our analysis shows that the model does not behave equally safely across all languages and prompt styles. Instead, safety performance depends strongly on both the languages used and the way a request is phrased using persuasive cues. We further observe that different risk categories exhibit different levels of vulnerability, with some types of harmful content being significantly more susceptible to persuasion-based jailbreaks than others. These findings reveal important limitations of current safety evaluations, which are largely English-centric, and underscore the need for multilingual and persuasion-aware benchmarking frameworks to more accurately assess real-world LLM safety. Our implementation is available at this https URL. Warning: this paper contains example data that may be offensive or harmful.

37. 【2609.03775】ypological Feature Prediction with Large Language Models: An In-Context Learning Approach

链接https://arxiv.org/abs/2609.03775

作者:Qianwen Wang,York Hay Ng,Aditya Khan,En-Shiun Annie Lee

类目:Computation and Language (cs.CL)

关键词:holds downstream utility, multilingual NLP, features holds downstream, downstream utility, holds downstream

备注: Accepted to EMNLP 2026

点击查看摘要

Abstract:Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels and feature types remains underexplored. Given LLMs' abilities in meta-linguistic reasoning and in providing rationales, we investigate LLMs' performance in typological feature prediction via an in-context learning approach with linguistic data from URIEL+ and Glottolog. We find that zero-shot prompting is insufficient, but when given phylogenetic and geographic neighbour evidence, LLMs substantially outperform all baselines without disadvantaging low-resource languages. We further find that most LLM rationales are consistent with the provided evidence, offering a step toward explainable typological feature prediction.

38. 【2609.03773】RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

链接https://arxiv.org/abs/2609.03773

作者:JoyIndustrial VisCAD Team:Linxin Cai,Qiuhe Hong,Zhichao Huang,Guanlin Li,Zongzhen Li,Hongsen Liu,Yichen Long,Wei Wang,Yuchen Wang,Dongyue Yang,Huimu Yu,Xianwen Zhong

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Parametric computer-aided design, Parametric computer-aided, CAD modeling, CAD, Existing CAD benchmarks

备注: Benchmark from JoyIndustrial's AI CAD project

点击查看摘要

Abstract:Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.

39. 【2609.03770】OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education

链接https://arxiv.org/abs/2609.03770

作者:Elakkiya Rajasekar

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:curriculum analytics report, computation informs decisions, practising outcome-based education, outcome attainment routinely, education compute learning

备注: 14 pages, 6 figures, 7 tables. Submitted to IEEE Transactions on Learning Technologies

点击查看摘要

Abstract:Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews of curriculum analytics report an absence of evidence on how that computation informs decisions. This paper presents OBER+, an extension of a deployed institutional attainment platform that computes the step from a measured shortfall to an evaluated corrective action. Five connected stages accumulate attainment across deliveries of a course, signal a shortfall and a persistent shortfall, grade it on cutoffs the regulator already uses, record the decision against a catalogue of practices annotated with their evidence, log the change, and quantify the subsequent movement in the shortfall. A further rule compares successive statements of an outcome, so attainment is never read as a series across a point at which the outcome changed. Applying the rules to the live record of two real courses produced three results. Every outcome of a core course was substantively redefined between consecutive deliveries, with subject matter moving between outcome numbers, so a naive reading would have reported a twenty-five point collapse between quantities that do not refer to the same learning. Recomputing the platform's figures from its documented rule showed six of ten differing by more than rounding explains, in a pattern that identified a defect since reported to the institution. Across fifteen statement pairs from three transitions, five were identical character for character, and among the ten that were not, the outcome carrying a given number was nearest to a differently numbered earlier outcome in six, a result resting on an ordering of similarities and requiring no threshold and no labelling. The contribution is a computational design for outcome-based reporting, stated as rules any attainment platform can implement, with evidence of what they make visible in a live institutional record.

40. 【2609.03749】Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG

链接https://arxiv.org/abs/2609.03749

作者:Alexandr Goultiaev Tolstokorov,Kyriakos Mouratidis,Javad Dogani,Nikolaos Laoutaris

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:Third-party retrieval-augmented generation, retrieval-augmented generation, marketplaces create, reused without compensation, RAG

备注: Accepted to EMNLP 2026 Main Conference

点击查看摘要

Abstract:Third-party retrieval-augmented generation (RAG) marketplaces create a new auditing problem: data providers may license corpora to a RAG operator, yet later have no visibility into whether their documents are being reused without compensation. Auditing this misuse is difficult because the operator is non-cooperative, answers are paraphrased by the generator, and one response may combine evidence from many providers. We propose DirBucket, a provider-side semantic watermarking and black-box auditing framework for document-level reuse in multi-provider RAG. DirBucket watermarks documents by meaning-preserving paraphrases whose embeddings are biased toward provider-bucket secret directions, enabling detection from black-box answers while preserving retrieval utility. On a challenging benchmark that reflects mixed-provider reuse under black-box access, DirBucket is the only method that consistently achieves strong target detection with no non-target activation, detecting non-compliance in every audit within 23 audited answers on our primary benchmark. The watermark survives adversarial post-answer laundering, and none of the evaluated evasion strategies simultaneously defeats detection while preserving user-perceived answer quality. Detection transfers unchanged to a second benchmark built from real clinical, cyber-threat-intelligence, and legal provider corpora. These results suggest that embedding-space watermarking can make document reuse in third-party RAG statistically auditable.

41. 【2609.03742】KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

链接https://arxiv.org/abs/2609.03742

作者:Yi Xu,Yifan Hou,Xiaoyu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:overwhelm novice learners, valuable educational resources, novice learners, dense and lengthy, lengthy formats

备注: This work is published on EMNLP 2026 (Findings). Our code and dataset are available at [this https URL](https://github.com/yixu-cityu/KnowVis) and [this https URL](https://huggingface.co/datasets/yixu-cityu/KnowVis)

点击查看摘要

Abstract:Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe cognitive overload for novice learners lacking prior domain knowledge. Existing video summarization methods fail to resolve this mismatch, as they primarily produce text-heavy, linear condensations that still demand high cognitive effort. To bridge this gap, we propose KnowVis, a framework that transforms linear video lectures into pedagogically grounded visual narratives. KnowVis first extracts a detailed concept map from multimodal video content to identify important and challenging threshold concepts, then constructs structured knowledge units, and finally synthesizes engaging visual summaries. Alongside the framework, we introduce a curated dataset of 125 educational videos across 10 academic disciplines, paired with 1,079 generated visual summaries. Extensive automated evaluations and a human study demonstrate that, compared to state-of-the-art baselines, KnowVis generates more accurate and clear visuals that successfully reduce cognitive load and significantly improve student learning effectiveness and knowledge retention.

42. 【2609.03734】Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

链接https://arxiv.org/abs/2609.03734

作者:Oline Ranum,Edward Fish,Simon Hadfield,Richard Bowden

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:adequately reflect sign, sign language translation, sign language proficiency, evaluating sign language, reflect sign language

备注

点击查看摘要

Abstract:BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning stronger sign representations. In this paper, we evaluate the relationship between spatio-temporal understanding and BLEU-4 across six SLT models on Phoenix-2014T and CSL-Daily, showing that gains in BLEU-4 are not on their own evidence of better sign language understanding. This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation. It aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4. Applied to SLT, this protocol targets content transfer, is more robust to train-test overlap, and gives a different picture of the field: the five gloss-free systems are largely within noise of one another on Phoenix-2014T, while the gloss-supervised system stands 9.3 points higher, a gap invisible to BLEU-4.

43. 【2609.03719】Opening mind by opening architecture: analysis strategies

链接https://arxiv.org/abs/2609.03719

作者:Francesco Vitucci,Giuseppe Silvi,Daniele Giuseppe Annese,Francesco Scagliola,Anthony Di Furia

类目:Computation and Language (cs.CL)

关键词:closed-architecture audio processor, audio processor model, numerical signal processing, research environments caused, electroacoustic composition

备注: Presented at the 7th International Csound Conference (ICSC 2024), Vienna, September 2024

点击查看摘要

Abstract:In numerical signal processing for electroacoustic composition, the progressive loss of specific development and research environments caused by the increasing use of digital market tools has favoured the dominance of the closed-architecture audio processor model. This model, while powerful, envisions the possibility of describing output data about its perceived characteristics, but at the cost of ignoring its internal process and interacting systems, which become complex, powerful environments but closed in an inscrutable black box, a loss we must consider. Any digital signal processing technique tells a story. Just as the words of a language incorporate social, historical and technical polysemic layers, a signal processor has its own story of implementation, a gradual technological achievement with its inevitable aesthetic consequences. Through the looking-glass of literature, one can access those environments with renewed awareness by reestablishing a scientific method and an attitude to research. In this specific case, starting from the case study of Manfred Schroeder's historical reverbs, we illustrate the process of building analytical evaluation tools, as well as practical implementation, at the basis of a conscious study path.

44. 【2609.03718】What Do CAE Simulation Agents Really Need Beyond a Generic Harness?

链接https://arxiv.org/abs/2609.03718

作者:Jiasheng Shi,Tianhan Zhang

类目:Computational Engineering, Finance, and Science (cs.CE); Computation and Language (cs.CL); Computational Physics (physics.comp-ph)

关键词:COMSOL takes real, Computer-aided engineering, real expertise, demanding areas, CAE simulation agent

备注

点击查看摘要

Abstract:Computer-aided engineering (CAE) simulation is among the largest and most demanding areas of engineering, where setting up a solver such as OpenFOAM, FEniCS, or COMSOL takes real expertise. Large language model (LLM) agents promise to turn a natural-language request into a working simulation, and recent CAE agents add simulation-specific machinery: multi-agent decomposition, domain retrieval, and scripted reflection. That machinery suited weak base models; modern harnesses already supply multi-turn reasoning, tool use, and execution feedback. We ask what a CAE simulation agent still needs beyond a generic harness. With information access and repair budget held fixed, a single-agent harness matches or beats multi-agent specialized systems (FoamBench 96.4\% vs.\ 88.2\%). Ablations trace this to capabilities the harness already provides: execution-feedback repair lifts FoamBench from 71.8\% with no repair round to 96.4\%, while scripted reflection adds nothing. The one input that still helps is domain knowledge supplied as solver tutorials, our largest measured gain (80.9\% to 96.4\%).

Subjects:

Computational Engineering, Finance, and Science (cs.CE); Computation and Language (cs.CL); Computational Physics (physics.comp-ph)

Cite as:
arXiv:2609.03718 [cs.CE]

(or
arXiv:2609.03718v1 [cs.CE] for this version)

https://doi.org/10.48550/arXiv.2609.03718

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
45. 【2609.03687】A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities

链接https://arxiv.org/abs/2609.03687

作者:Anh Danh,Rick Nouwen,Massimo Poesio

类目:Computation and Language (cs.CL)

关键词:contextual reasoning, important task, task in contextual, Coreference resolution, plural

备注

点击查看摘要

Abstract:Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechanism for representing and retrieving singular and plural entities for plural reference. We use a combination of mechanistic interpretability and attention pattern analysis to study the process in which LLMs predict a pronoun to refer back to previously mentioned entities. Using a range of causal intervention techniques, we find a set of attention heads that are responsible for (1) representing coreference information in the input, (2) identifying entities that form a plural reference, (3) transferring the information to the component that is responsible for selecting the antecedents and predicting the pronoun. We also find that LLMs align with humans in preference for plural pronoun. Specifically, entities in a plural construction are more likely to be referred to as a plural entity if they are ontologically similar and are linked by the conjunction "and".

46. 【2609.03677】Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

链接https://arxiv.org/abs/2609.03677

作者:Julian Truetsch,Felix Hauser,Christoph Stiller,Frank Bieder

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Robotics (cs.RO)

关键词:Understanding the composition, essential for safety, composition of large-scale, reliable operation, large-scale autonomous driving

备注: 9 pages, 5 figures, submitted to the IEEE Open Journal of Intelligent Transportation Systems (OJ-ITS), our implementation and benchmark dataset are available at [this https URL](https://github.com/KIT-MRT/AD-Diff)

点击查看摘要

Abstract:Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned with the training data, resulting in potentially dangerous performance degradation. Yet, existing data analysis pipelines largely rely on metadata, predefined labels, or manual inspection, which provide limited semantic insight or do not scale. This paper studies set difference captioning: given two subsets of images, the goal is to produce a natural-language hypothesis describing differences between the target and reference set. Building on a two-stage formulation, we adapt the method to autonomous driving by focusing on object-centric patches derived from object detection, which simplifies aggregation and enables attribution of differences to specific object instances or categories. To evaluate this setting in-domain, we introduce a new benchmark, AD-Diff Bench. Low-concentration experiments assess the suitability of set-difference-captioning approaches to sparse, real-world differences. We restrict our experiments to open-weight models to support reproducibility and ease of deployment. The proposed benchmark and analysis provide a step towards practical, human-interpretable dataset introspection for autonomous driving datasets. Our implementation and benchmark dataset are available at this https URL

47. 【2609.03654】Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

链接https://arxiv.org/abs/2609.03654

作者:Arianna Miola,Bruno Spaccavento,Lorenzo Silotto,Marco Bianchetti,Luca Cagliero

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Information Retrieval (cs.IR)

关键词:statements poses significant, poses significant challenges, answering systems due, banks' financial statements, financial statements poses

备注

点击查看摘要

Abstract:The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdictions and institutions. We introduce FinRAG-QA, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023. Unlike prior financial QA benchmarks, which centre on U.S. filings and single-institution analysis, FinRAG-QA targets cross-institutional retrieval over documents averaging 198k words, longer than any existing financial QA resource. On this benchmark we evaluate a multi-stage RAG pipeline and isolate the contribution of each component. Contextual chunk enrichment combined with a retrieval-optimised embedding model raises NDCG@10 from 0.322 to 0.710; conditional on the ground truth being retrieved, a reasoning-optimised generator raises answer accuracy from 44.6% to 79.0% (+34.4 percentage points), at roughly 20x the generation latency. We further show that cross-encoder reranking degrades retrieval when the first-stage ranking is already strong, and that a single top-ranked chunk outperforms larger contexts at generation time. Experiments were run in late 2024-early 2025 with the models available at that time.

48. 【2609.03652】he Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

链接https://arxiv.org/abs/2609.03652

作者:Sara Sorahi,Kevin Tang,Reza Kazemian

类目:Computation and Language (cs.CL)

关键词:addressing class imbalance, imbalance in NLP, British National Corpus, real training data, common strategy

备注

点击查看摘要

Abstract:Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in the context of discourse pragmatic function classification, a task where data sparsity is a structural feature rather than a collection artefact. Using 410 manually annotated instances of the English word look drawn from the British National Corpus, spanning four functions: Attention Signal, Directive, Discourse Marker, and Interjection. We generate synthetic training examples with Llama 3.1 and partition them by their cosine distance from real training data in RoBERTa embedding space. We compare six training conditions that differ in the placement of synthetic examples relative to the empirical decision boundary, while holding augmentation quantity constant across conditions. All augmented conditions improve macro F and accuracy over the real only baseline, but core proximal examples (NEAR) yield the largest gains in macro F (0.113), while a distance balanced mix achieves the highest accuracy (0.748). No condition improves AUC, indicating that augmentation shifts the decision boundary rather than improving the model's underlying probability estimates. These findings suggest that where synthetic examples land in representation space matters as much as how many are generated, with implications for low resource pragmatic classification more broadly.

49. 【2609.03633】/think Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

链接https://arxiv.org/abs/2609.03633

作者:Seunghee Koh,Sungjae Choi,Minchan Kwon,Sunghyun Baek,Junmo Kim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:reasoning improves large, improves large reasoning, spurious CoT termination, produces long, large reasoning models

备注: Accepted to EMNLP 2026 Main Conference

点击查看摘要

Abstract:Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that injects an end-of-think token (EoT, /think) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a clean answering phase. Answering-phase generation can continue before the model regenerates another EoT, with the span preceding this regenerated EoT scaling with the reasoning tokens saved by early exit and exhibiting continued reasoning behavior. We call this spurious CoT termination, where reasoning-like generation continues into the answering phase. We hypothesize that insufficient attention to the injected EoT contributes to spurious CoT termination and probe this hypothesis with Exit-token Attention Biasing (EAB). Across four LRMs, five benchmarks, and two early-exit methods, increasing attention to the injected EoT reduces spurious CoT termination and answering-phase length. These results reveal a limitation of controlling LRMs by externally matching their explicit think-block format. Inserting the EoT token conforms to this format but does not by itself guarantee the intended reasoning-to-answering transition. Our code is available at this https URL.

50. 【2609.03619】Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

链接https://arxiv.org/abs/2609.03619

作者:Xuanfa Jin,Zhijian Ma,Yongcheng Zeng,Xinyu Cui,Haifeng Zhang,Jun Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:large language models, multiple agents iteratively, agents iteratively refine, improves the reasoning, responses through discussion

备注: EMNLP 2026 Findings, 24 pages, 4 figures

点击查看摘要

Abstract:Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R$^2$-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R$^2$-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R$^2$-MAD achieves consistent improvements over existing single-agent and MAD baselines.

51. 【2609.03597】KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records

链接https://arxiv.org/abs/2609.03597

作者:Tasmiad Hasan,Arafat Zaman Ratul,Sarker Sadman Saalim,S.M. Shah Nawaz Hossain,Khan Raiyan Ibne Reza,Sumaiya Tabassum Nimi

类目:Computation and Language (cs.CL)

关键词:dedicated Unicode glyphs, Unicode glyphs, dedicated Unicode, OCR pipeline, mainstream font

备注: 12 pages, 5 figures, 11 tables, NLLP Workshop @ EMNLP 2026. Dataset: [this https URL](https://huggingface.co/datasets/RaiyanKhaan/KhatianDoc)

点击查看摘要

Abstract:Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten records that carry these fractions, RS Khatians, are the authoritative title record for millions of parcels and a frequent subject of civil litigation, yet no benchmark has asked whether a machine can read one. We introduce KhatianDoc, a four-task benchmark built from 107 real RS Khatian records from the Vumi (land) Office of Munshiganj, Bangladesh: symbol recognition, base-16-to-decimal conversion, structured field extraction, and legal document question answering over 1,634 QA pairs. Ground truth was transcribed by hand, verified by a land-law practitioner to full agreement, and anonymized through positional tokens that keep the referential distinctions multi-hop questions depend on. We evaluate six multimodal LLMs (8B to 72B+, open and closed) under a fixed zero-shot protocol. Five QA categories, 39.3% of our stratified set, return zero correct answers from every model; on the arithmetic task, every model that emits a number does worse than a constant-mean baseline, with exact- and near-match scores coinciding: decorrelation, not approximation. Auditing our own metrics surfaced two artifacts in opposite directions: we correct a refusal-scoring bug and report the fixed scores beside the originals, and flag an inflated metadata metric as an upper bound. KhatianDoc documents not a performance gap but the absence of a capability, with verified ground truth for future systems. Code and data, with a redacted image release, are publicly available.

52. 【2609.03595】How Far Can Synthetic Data Take Thai OCR?

链接https://arxiv.org/abs/2609.03595

作者:Kunat Pipatanakul

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Thai OCR model, OCR model adapted, makes synthetic OCR, real Thai documents, real Thai document

备注: 20 pages, technical report

点击查看摘要

Abstract:We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on training granularity, with in-domain reconstruction approaching real printed supervision under page-level training (1.82% versus 1.31% median character error rate) but underperforming out-of-domain reconstruction under crop-level training (15.59% versus 5.52%). Guided by these findings, we adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero using 45,723 synthetic pages: relative to its base checkpoint, it reduces median character error rate from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting and outperforms Typhoon OCR v1 7B on all five evaluation sets, showing that synthetic-only training can be competitive.

53. 【2609.03580】HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

链接https://arxiv.org/abs/2609.03580

作者:Tzu-Ling Lin,Dong-Ting Yao,Teng-Fang Hsiao,Wei-Chih Chen,Hong-Han Shuai

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large Language Models, Language Models, Large Language, undermine review reliability, academic peer review

备注: Accepted to EMNLP Findings 2026

点击查看摘要

Abstract:The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review, where verification requires grounding claims in long, technical papers. We introduce HalluPeer, a benchmark for detecting hallucinations in scientific peer reviews, providing aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Our pipeline induces a peer-review-specific hallucination taxonomy, identifies review contexts, and injects hallucinations with automated filtering. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, while evaluation on authentic reviews demonstrates that HalluPeer-defined hallucination patterns occur in real peer reviews, highlighting the critical need for source-aware verification. Our project page can be found in this https URL

54. 【2609.03577】Language, Language Models, and What We're Talking About

链接https://arxiv.org/abs/2609.03577

作者:Malvina Nissim

类目:Computation and Language (cs.CL)

关键词:linguistic worlds conveyed, Language models, models, Language, commonly discussed

备注: In Proceedings of the Twelfth Italian Conference on Computational Linguistics (CLiC-it 2026)

点击查看摘要

Abstract:Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want language models to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between language models designed as technical products and language models designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.

55. 【2609.03511】Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

链接https://arxiv.org/abs/2609.03511

作者:Karthika Nhayakkat,Rajat Verma,Maharaj Brahma,Vetcha Gnana Mahesh,Maunendra Sankar Desarkar,Ganesh Ramakrishnan,Rohit Saluja

类目:Computation and Language (cs.CL)

关键词:Large Language Models, free word-order languages, Large Language, Language Models, demonstrate strong multilingual

备注: Accepted at EMNLP 2026 (Findings - Long paper)

点击查看摘要

Abstract:Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the structural sensitivity of multilingual LLMs using two linguistically grounded perturbation settings in Hindi and Malayalam: constrained constituent reordering and active-passive voice transformation. We introduce a benchmark dataset IndicReStruct, with two variants, GSM8K-Reordered and GSM8K-Voice, constructed from GSM8K while preserving semantic meaning. Across six state-of-the-art LLMs and multiple prompting strategies, we observe consistent and significant degradation in mathematical reasoning performance under structurally perturbed inputs. To further understand these failures, we perform qualitative error analysis and mechanistic interpretability experiments using residual-stream activation patching. Our analyses show that reasoning failures frequently arise from disruptions in entity-quantity alignment and that intermediate transformer layers contribute most strongly toward reasoning restoration. Overall, our findings suggest that current multilingual LLMs remain highly sensitive to surface syntactic realization and lack robust compositional invariance under structurally different but semantically equivalent inputs.

56. 【2609.03502】Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

链接https://arxiv.org/abs/2609.03502

作者:Kunat Pipatanakul,Potsawee Manakul,Warit Sirichotedumrong,Sittipong Sripaisarnmongkol,Pakorn Nathong,Phatrasek Jirabovonvisut

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:typically requires choosing, TTS typically requires, deploying TTS typically, large voice-cloning model, Challenge-Set Keyword Accuracy

备注: 18 pages, technical report

点击查看摘要

Abstract:In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.

57. 【2609.03487】Pattern Over-Generalization of Knowledge Graph Embedding

链接https://arxiv.org/abs/2609.03487

作者:Junsik Kim,Kangil Kim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Knowledge graph embedding, low-dimensional vector space, Knowledge graph, knowledge graphs, KGE models

备注: Accepted to EMNLP 2026, 22 pages, 9 figures

点击查看摘要

Abstract:Knowledge graph embedding (KGE) demonstrates its effectiveness for predicting missing links in knowledge graphs (KGs) by projecting entities and relations into a low-dimensional vector space. It is crucial for KGE models to effectively capture inference patterns (patterns) inherent in KGs, such as symmetry/antisymmetry, inversion and composition. Although recent KGE models exhibit strong capabilities in modeling such diverse patterns, they suffer from inherent limitations stemming from pattern over-generalization, where embeddings learned from only a single pattern instance inevitably generalize that pattern to all related instances, i.e., generalize the pattern universally. To address this issue, we propose PogRE (Pattern Over-Generalization Robust Embedding), a simple but effective method that utilizes dense linear transformations and compound operations for relation representation. Our theoretical analysis demonstrates that a dense linear transformation allows a pattern to become progressively universal as more triples are observed in the pattern. Furthermore, after observing d+1 linearly independent entities (d+1 denotes the dimension of entity), the linear transformation guarantees universal generalization of the pattern across all related instances. Experimental results on three standard benchmark datasets show that PogRE outperforms existing state-of-the-art KGE models in link prediction. Moreover, our empirical results indicate that PogRE effectively addresses the negative impact of over-generalization.

58. 【2609.03467】When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

链接https://arxiv.org/abs/2609.03467

作者:Wen-Yu Chang,Yun-Nung Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, motivating growing interest, Large language, long-horizon conversational agents, language models

备注

点击查看摘要

Abstract:Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory benchmark derived from Lo- CoMo with four query styles: dialog, implicit, counterfactual, and composed. Across five rep- resentative memory systems, we evaluate both retrieval recall and end-to-end response qual- ity. Our experiments show that conversational framing exposes substantial retrieval gaps over- looked by QA benchmarks, especially on im- plicit and composed queries, which multi-facet query rewriting narrows for raw-turn mem- ory but not abstractive memory. We further find that strong retrieval does not fully trans- late into response quality, and that implicit queries exhibit silent grounding, where mem- ory improves contextual grounding without ex- plicitly surfacing the gold fact. These results point to reasoning-based memory elaboration as a promising direction, and we release aux- iliary supportive_memory annotations captur- ing conversationally useful context beyond the original gold evidence.

59. 【2609.03454】When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA

链接https://arxiv.org/abs/2609.03454

作者:Hyunseo Oh,Chong-Kwon Kim,Yoonhyuk Choi

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:combine emotional distress, large language model, mental-health question answering, Retrieval-augmented generation, single-turn mental-health question

备注: 8 pages, 3 figures. Presented at the KDD 2026 Undergraduate Consortium

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model responses, but its effect is not uniformly beneficial in single-turn mental-health question answering, where user queries often combine emotional distress, treatment concerns, and safety-sensitive needs. We study when retrieval helps or hurts mental-health QA, and whether a lightweight selective retrieval policy can better control this trade-off. We operationalize retrieval need using three draft-conditioned utility dimensions: psychoeducational need, coping need, and response specificity, together with a rule-based safety trigger. Following psychotherapy-grounded RAG systems such as coTherapist, we construct a compact and controllable guideline corpus comprising coping-strategy, psychoeducational, and safety resources. We fine-tune an instruction-tuned generator on MentalChat16K using QLoRA and compare Closed-book, Always Retrieval, and Selective Retrieval settings on CounselBench-Eval and CounselBench-Adv. Experiments show that retrieval is not uniformly beneficial in this domain. Always Retrieval improves specificity but lowers overall quality and introduces additional safety-sensitive failures. Selective Retrieval preserves closed-book behavior for low-need cases while avoiding the additional degradation caused by unconditional retrieval, supporting the view that retrieval activation is a safety-sensitive control decision.

60. 【2609.03450】Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

链接https://arxiv.org/abs/2609.03450

作者:Kazuki Nakayashiki

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Study, archived source record, agent that inherits, inherits six one-line, one-line memories

备注: 46 pages, 7 figures, 35 tables. Twelve registered studies (14,760 attempted episodes) on one instrument lineage; every package was frozen, timestamped and externally deposited before its first confirmatory call. Manuscript, LaTeX source, all episode files, frozen packages, analyzers and the generator of every number are archived at Zenodo: doi: [https://doi.org/10.5281/zenodo.22267221](https://doi.org/10.5281/zenodo.22267221)

点击查看摘要

Abstract:An agent that inherits six one-line memories may pull at most one archived source record before acting; a directive written into the store can steer that choice: a pointer to the record, a criterion that identifies it, or both. Across twelve registered studies on one instrument lineage (14,760 attempts) we measured where the request goes under each form. On six direct-provider models a length-matched criterion exceeded a bare id by +35.0 points [+31.2, +38.8] (Study D); the contrast failed its registered superiority rule on a nine-model OpenRouter-served panel (Study E). Appending the id cancelled the criterion on three Claude models (Opus 5: 40/40 to 0/40; Study F-x); six byte-matched edits gave each exact string its own effect (Study G), and a re-run at eighty runs per cell left fifteen of thirty replication contrasts within the margin, fifteen unresolved and none beyond (Study G'). A ratification line (+96.0 points on Opus 5) and a budget of two credits restored the target on all three (Study J); across five criterion strings the suffix's cancellation held for four of the five wordings on Opus 5 and all five wordings on Fable 5.1 (Study H2); in a second store every model followed the criterion (Study H1). Continued into a decision, the criterion moved the choice toward the current record (+100.0 points, Opus 5) and away from it on Fable 5.1 (Study I). A one-character plan pointer's effect (+78.0 points; Study B, after a correction of its first repository report) returned the same verdict under a prospectively registered re-run (+81.7 points; Study B'). All results are descriptive effects of exact edits on fixed panels with registered intervals and no mechanism claim.

61. 【2609.03436】It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

链接https://arxiv.org/abs/2609.03436

作者:Yigit Utku Bulut

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, Reasoning traces, moments and early-legible, early-legible fates, traces of large

备注: 25 pages, 11 figures, 4 tables. Also available at doi: [https://doi.org/10.5281/zenodo.22261107](https://doi.org/10.5281/zenodo.22261107) . Code and pre-registered protocols: [this https URL](https://github.com/bulutyigit/problem-not-path)

点击查看摘要

Abstract:Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.

62. 【2609.03432】Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

链接https://arxiv.org/abs/2609.03432

作者:Xiangyu Wang,Jin Wu,Xiaoyu Li,Chanjin Zheng,Yifeng Zhou

类目:Computation and Language (cs.CL)

关键词:tasks remains challenging, LLM is susceptible, creativity tasks remains, verbosity bias, leniency bias

备注: Accepted to EMNLP 2026

点击查看摘要

Abstract:Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, and wide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at this https URL.

63. 【2609.03430】Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

链接https://arxiv.org/abs/2609.03430

作者:Heng Wang,Jielin Qiu,Wenting Zhao,Cheng Qian,Liangwei Yang,Jiawei Han,Heng Ji,Silvio Savarese,Shelby Heinecke,Huan Wang

类目:Computation and Language (cs.CL)

关键词:Large language models, severe memory bottleneck, achieve superior performance, language models achieve, models achieve superior

备注

点击查看摘要

Abstract:Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks it matches the strongest prior evictor while serving 32-43% higher throughput than it in vLLM deployment. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them. Our code is publicly available at this https URL.

64. 【2609.03426】Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations

链接https://arxiv.org/abs/2609.03426

作者:Yunao Zheng,Bin Wen,Xiaojie Wang

类目:Computation and Language (cs.CL)

关键词:requiring repeated dense, local static patterns, repeated dense computation, reuse local static, Transformers lack

备注

点击查看摘要

Abstract:Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to high parameter and activation costs. We propose Lngram v2, which decouples the number of routes, memory dimension, and backbone width, and introduces a context-aware grouped-query attention readout to scale memory capacity independently. A zero-value Sink and counterfactual surrogate gradients further improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments across vision--language models (VLMs) of different scales show consistent improvements, including successful scaling to a 30B-parameter model. Compared with Lngram v1, Lngram v2 substantially reduces both total and activated memory parameters while maintaining or improving language modeling performance. Further analysis shows that its discrete IDs preserve substantial semantic structure of continuous hidden states, enabling semantic recovery from IDs alone and stable ID--semantic associations across datasets. These results establish Lngram v2 as an efficient and scalable latent conditional memory mechanism whose discrete addresses also provide a structured interface for analyzing internal model representations.

65. 【2609.03410】o What Extent Do Large Language Models Understand Bangla Idioms?

链接https://arxiv.org/abs/2609.03410

作者:Mousumi Akter,Md. Faiyaz Abdullah Sayeedi,Nurul Labib Sayeedi,Swakkhar Shatabda

类目:Computation and Language (cs.CL)

关键词:reflecting cultural nuances, posing unique challenges, reflecting cultural, meaning identification, integral part

备注

点击查看摘要

Abstract:Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms, complemented by a synthetic multiple-choice question (MCQ) dataset for idiom meaning identification. We conduct a comprehensive evaluation of recent large language models (LLMs) across three idiom-related tasks: paraphrasing, idiom span detection, and meaning identification, leveraging zero-shot and few-shot prompting strategies. Our results reveal substantial variability in model performance, with no single LLM consistently outperforming others across all tasks. Notably, Phi-4-mini-instruct excels in paraphrasing, Kimi-K2-32b-instruct in span detection, and Gemini-2.5-flash in meaning identification. We believe that our datasets and analyses will provide valuable resources to guide future research in improving LLM comprehension of idiomatic expressions, particularly in Bangla and other low-resource languages.

66. 【2609.03395】abScope: Question-Adaptive Scope Selection for Table Question Answering

链接https://arxiv.org/abs/2609.03395

作者:Yuxiang Wang,Junhao Gan,Jianzhong Qi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Large Language, Language Models, table size increases, shown strong performance

备注: conference paper preprint

点击查看摘要

Abstract:Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affected by irrelevant table content, while questions requiring broader evidence may still benefit from full-table reasoning. Based on this observation, we propose a question-adaptive framework that dynamically selects between localized and full-table reasoning. The framework constructs question-specific sub-tables through operation-aware table decomposition and uses the predicted question type to determine the appropriate reasoning mode. We further introduce silver reference sub-tables for evaluating evidence selection and construct SLQA, a benchmark based on real-world long tables. Experiments on WikiTQ and SLQA show that localization is particularly effective for lookup and local reasoning questions, while adaptive selection between localized and full-table reasoning achieves the best overall performance. These results highlight that long-table QA requires deciding not only how to localize, but also when to localize. Our code and datasets will be made available upon publication of the paper.

67. 【2609.03394】Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

链接https://arxiv.org/abs/2609.03394

作者:Divyesh Bommana,Mohammad Saim,Tianyu Jiang

类目:Computation and Language (cs.CL)

关键词:single shared event, missing many real-world, shared event, Emotion, real-world scenarios

备注: Accepted to EMNLP 2026 (Main Conference) Dataset and code: [this https URL](https://github.com/cincynlp/Chiaro)

点击查看摘要

Abstract:Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.

68. 【2609.03370】FrameBench:A Language Understanding Benchmark Based on Frame Semantics

链接https://arxiv.org/abs/2609.03370

作者:Chihiro Yano,Ryohei Sasano

类目:Computation and Language (cs.CL)

关键词:relating lexical meaning, background knowledge called, knowledge called semantic, unstated information, assumed to proceed

备注: Accepted in EMNLP Findings 2026

点击查看摘要

Abstract:In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieved strong performance across a wide range of downstream tasks. However, it remains unclear whether they can reproduce the kinds of implicit enrichment that humans naturally make during comprehension. To address this question, we introduce FrameBench, a benchmark grounded in frame semantics. FrameBench consists of multiple-choice questions that test whether models distinguish the frames evoked by the same verb across contexts. We construct the benchmark for English and Japanese using FrameNet-style resources and a generation-and-verification pipeline with native-speaker judgments. Our experiments on a diverse set of models reveal challenges for small models, while several large models surpass the human reference scores. We release the constructed FrameBench dataset and the code for dataset construction and evaluation at this https URL.

69. 【2609.03366】Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT

链接https://arxiv.org/abs/2609.03366

作者:Zikai Zhou,Yufei Jin,Yilin Xu,Yu-Chiang Wang,Chieh-Ju Chao,Monica S. Lam

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Logic in Computer Science (cs.LO)

关键词:decision, pivotal conditions, Satisfiability Modulo Theories, Accountability, pivotal

备注: Accepted to EMNLP 2026 (Main Conference). 46 pages, 6 figures, 28 tables. Code and prompts: [this https URL](https://github.com/stanford-oval/clinical-trial-matching)

点击查看摘要

Abstract:Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified rationales (how was the decision reached), assumptions (what was assumed rather than known), policy consistency (the same treatment for the same facts), and pivotal conditions (what would change the outcome). We introduce self-faithfulness as an automatic test of accountability: changing the pivotal conditions should change the decision. We examine accountable AI through clinical trial matching, a high-stakes task central to evidence-based medicine. Although LLM-based matchers match patients to trials reasonably accurately, they apply decision policies inconsistently and produce rationales that are unfaithful to their own decisions. We introduce VERDICT, an LLM-based agent that translates a decision task, its constraints, and its policy into Satisfiability Modulo Theories (SMT), then derives the decision with SMT and MaxSMT solvers -- so policies are applied consistently and decisions are accountable by construction. Across a SIGIR 2016-derived dataset and TREC 2021, VERDICT achieves the strongest decision accuracy among LLM-only and neurosymbolic baselines, applies policies with perfect consistency, and produces clinician-preferred rationales grounded in explicit assumptions and pivotal conditions, with improved counterfactual self-faithfulness.

Comments:
Accepted to EMNLP 2026 (Main Conference). 46 pages, 6 figures, 28 tables. Code and prompts: this https URL

Subjects:

Computation and Language (cs.CL); Computers and Society (cs.CY); Logic in Computer Science (cs.LO)

ACMclasses:
I.2.7; I.2.4; J.3

Cite as:
arXiv:2609.03366 [cs.CL]

(or
arXiv:2609.03366v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.03366

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Zikai Zhou [view email] [v1]
Thu, 3 Sep 2026 04:49:02 UTC (3,208 KB)

70. 【2609.03350】From Zero to Hero: An Open LLM Ecosystem for Armenian

链接https://arxiv.org/abs/2609.03350

作者:Erik Arakelyan,Khatun Avetisyan,Meri Davtyan,Heghine Grigoryan,Nane Khachatryan,Hayk Shahsuvaryan,Henrik Sergoyan,Vahan Martirosyan

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:open Armenian LLM, Armenian, Armenian LLM, open Armenian, low-resource language

备注: 18 pages, 4 figures, 13 tables. Data and model: [this https URL](https://huggingface.co/collections/COPA-AI/armenian-llm-ecosystem) . Code: [this https URL](https://github.com/COPATeam/armenian_llm_ecosystem)

点击查看摘要

Abstract:Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.

71. 【2609.03331】FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models

链接https://arxiv.org/abs/2609.03331

作者:Jiayuan Ma,Yuqi Lu,Weiyang Guo,Chenrui Wang,Junyi Shu,Xuebo Liu,Min Zhang,Jing Li

类目:Computation and Language (cs.CL)

关键词:describe visual content, Vision-language models, incorrect assumptions, increasingly deployed, deployed in multi-turn

备注: Accepted at EMNLP2026 Main Conference. For code and data, see [this https URL](https://github.com/lab-klc/FPCO-Dialog)

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dialogue turns. We introduce FPCO-Dialog, a benchmark for evaluating correction and cooperation behavior in VLMs under repeated false premises. FPCO-Dialog contains 1,080 images and 10,800 question turns, stratified by visual complexity, object category, and false-premise class, and uses a 10-turn protocol in which a correct dialogue prefix is followed by repeated false-premise referring expressions. We evaluate 20 commercial and open-source VLMs with a model-agnostic protocol and CorrTP@K, a correction-rate metric over false-premise turns, scored by two independent detectors. FPCO-Dialog reveals substantial and persistent cross-model differences in aggregate correction tendency, model-specific turn-wise dynamics, and systematic variation across false-premise types under the benchmark's substitution distribution. The dataset, evaluation protocol, model outputs, detector labels, and code are available.

72. 【2609.03330】Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour

链接https://arxiv.org/abs/2609.03330

作者:Huixiang Fu,Marian-Andrei Rizoiu

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Social and Information Networks (cs.SI)

关键词:large language models, poor cross-domain generalization, weak rationale grounding, foundation detection systems, Moral language plays

备注

点击查看摘要

Abstract:Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MA\textbf{C}- and \textbf{H}ate-speech-\textbf{A}ware \textbf{R}ationale-aligned \textbf{M}oral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\% subsample of the MFTC, MFRC, and News training pools together with the richer supervision in MFTCXplain, CHARM improves AUC by up to 15.3\% in-domain, surpasses the supervised baselines on every out-of-domain dataset in both AUC and F1, and offers a scalable, low-cost alternative to prompting-based LLM detectors. We further apply CHARM to large-scale COVID-19 discourse on Twitter and show that moral value alignment is strongly associated with online endorsement behavior. By making moral framing measurable at scale, CHARM offers a practical tool for studying the spread of morally charged misinformation.

73. 【2609.03322】How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

链接https://arxiv.org/abs/2609.03322

作者:Dun Li Chan,Emily Liu,Niyathi Allu,Christian Hoang

类目:Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:models encounter typos, Language models encounter, disrupted token order, decoder-only language models, Language models

备注: 12 pages, 4 figures

点击查看摘要

Abstract:Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints by analyzing layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores are especially associated with activation-patching recovery under token substitution and shuffling. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.

74. 【2609.03321】Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue

链接https://arxiv.org/abs/2609.03321

作者:Yihang Li,Chenhui Chu

类目:Computation and Language (cs.CL)

关键词:Finite State Machine, Neural Finite State, Neural Finite, next-token prediction objective, low fine-tuning cost

备注: EMNLP 2026 Main Conference

点击查看摘要

Abstract:The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializing turn-taking control and response generation onto a single causal tape under the standard next-token prediction objective, thereby preserving semantic prowess at a low fine-tuning cost. However, its reliance on synthetic text data fundamentally limits turn-taking naturalness, as Large Language Models (LLMs) cannot faithfully simulate the fine-grained acoustic temporal dynamics of real human dialogues. In this work, we propose a decoupled data approach that learns turn-taking from real Human-Human (HH) spoken dialogues while shaping semantic behavior through configurable Human-Agent (HA) text dialogues. To operationalize this approach, we introduce a rule-based event-guided data transformation method that serializes HH spoken dialogues into FSM tapes by classifying turn-taking events and applying deterministic mapping rules, enabling scalable supervision without LLM-generated annotations. We further propose a Source-Aware Calibrated (SAC) Loss that jointly calibrates the long-tailed distribution of state transition tokens and channels each data source toward the capability it best supervises. Experiments show that our approach substantially improves turn-taking proficiency while recovering the foundation LLM's semantic capability. Our code and model are available at this https URL.

75. 【2609.03293】PACE: Towards Surfacing Hidden Conflicts in User Requests

链接https://arxiv.org/abs/2609.03293

作者:Yoojin Kim,Jihyoung Jang,Hyounghun Kim

类目:Computation and Language (cs.CL)

关键词:user current circumstances, Personalized assistants, current circumstances, introduce Personalized Assistants, user requests

备注: EMNLP 2026 (59 pages); Code: [this https URL](https://github.com/p2chp2t/pacemaker)

点击查看摘要

Abstract:Personalized assistants should not only comply with user requests but also assess whether those requests are appropriate given the user's current circumstances. However, prior work has primarily focused on accurately executing requests, overlooking the need for assistants to account for context and engage in conflict-based refusal. Furthermore, while existing work on conflict or safety detection relies on explicitly provided factors, real-world scenarios often involve implicit factors that must be retrieved from a knowledge base (KB). To this end, we introduce Personalized Assistants for Conflict Evaluation (PACE), a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate. PACE pairs user requests grounded in well-defined personas with egocentric KB facts, requiring models to integrate contextual evidence to determine whether a request is conflicting. This implicit retrieval setting hinders the direct association between user requests and conflict-inducing knowledge, making it difficult for existing models to identify relevant user-specific facts. To address this challenge, we further propose PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence. Experiments on PACE evaluate both evidence retrieval quality and conflict decision accuracy, showing that PaceMaker consistently outperforms existing approaches.

76. 【2609.03273】Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers

链接https://arxiv.org/abs/2609.03273

作者:Karthikeyan A,Jaya Nirmala S,Sangeetha Sivanesan,Indhu R,Pranav Kumar,Bharat Jude Johnson,Vishnu Ram

类目:Computation and Language (cs.CL)

关键词:rich verbal morphology, agglutinative low-resource language, Minimum Edit Distance, phonetic transformation, distinct letters

备注

点击查看摘要

Abstract:Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct letters. Prior work targets word-level surface errors with rule-based methods, statistical n-gram models, Minimum Edit Distance, or hybrid pipelines with a transformer re-ranker; such methods cannot reliably handle contextual errors - subject-verb agreement, tense consistency, or cross-word sandhi - which require sentence-level understanding. We propose an end-to-end sequence-to-sequence formulation and fine-tune mT5-small and mBART-50 on a synthetic corpus of up to 657,720 noisy-clean Tamil sentence pairs spanning ten error categories. Both backbones follow the same four-stage progressive schedule, each stage targeting one weakness: surface noise (v2), contextual grammar (v3), single-site sandhi (v4), and multi-site cross-word sandhi (v5). On a 1,000-sentence balanced diagnostic set verified disjoint from all training data, our best model, mBART-50 v5, reaches 69.3% top-1 exact-match accuracy, with 87.5% on sandhi and 43.5% on subject-verb agreement. The schedule is what produces these gains: subject-verb accuracy rises from 1.0% to 52.5% once contextual pairs are introduced, and sandhi from 0% to 87.5% once multi-site sandhi pairs are. We additionally quantify a precision-recall trade-off this literature has not reported: sandhi recall is paid for monotonically in identity accuracy. Finally, Tamil-LLaMA-7B-Instruct reaches 19.0% zero-shot and 24.7% with three demonstrations against a 20.0% copy baseline, showing that a Tamil-adapted instruction model does not transfer to specialised sentence-level correction without task-specific supervision.

77. 【2609.03261】MedQA-MM: Shortcuts Behind Medical Visual Reasoning

链接https://arxiv.org/abs/2609.03261

作者:Benlu Wang,Yifan Zhang,Jiaqing Yu,Chin Siang Ong,Juncheng Huang,Zhuohao Li,Zhenyu Zhang,Arman Cohan,Hong Yu,Zonghai Yao

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:benchmark score credits, score credits final, credits final answers, benchmark score, score credits

备注

点击查看摘要

Abstract:A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

78. 【2609.03254】What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation

链接https://arxiv.org/abs/2609.03254

作者:Daisuke Kikuta

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large Language Models, Large Language, Language Models, users generate artifacts, iterative cycles

备注: Accepted at EMNLP 2026 Industry Track. The code is available at [this https URL](https://github.com/ntt-dkiku/llm-revision-propagation)

点击查看摘要

Abstract:Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies and propagate the revision to all affected parts of the artifact. This paper studies this ability of LLMs on conversationally generated artifacts, where the artifact context and its dependencies may be embedded in the conversation history. Toward practical use, we also explore cost-effective test-time compute for this new setting. Specifically, we introduce a new benchmark for this setting, and evaluate nine revision methods, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark. The results show that baselines achieve accuracies of 68.3--93%, and the most cost-effective method is selecting from three parallel samples using either LLM-based or medoid selection, which improves accuracy by 2.2--9.7%. Our code and dataset are available at this https URL.

79. 【2609.03235】SGD-KV: Summarization Guided KV Cache Compression

链接https://arxiv.org/abs/2609.03235

作者:Zeyu Liu,Woomin Song,Xuandi Fu,Sai Muralidhar Jayanthi,Vivek Govindan,Aram Galstyan,Sravan Babu Bodapati,Srikanth Ronanki

类目:Computation and Language (cs.CL)

关键词:Large language models, linearly growing size, Large language, face severe memory, severe memory bottlenecks

备注: Accepted in NeurIPS2026 Efficient Reasoning Workshop

点击查看摘要

Abstract:Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different attention heads. We present SGD-KV (Summarization-Guided KV Cache Compression), a head-aware framework that leverages a novel chunk-summarization diagnostic task to systematically identify and prioritize attention heads specialized in hierarchical information aggregation. Experiments on Qwen2.5-7B-1M and Qwen3-32B across diverse long-context benchmarks demonstrate that SGD-KV achieves state-of-the-art performance with contexts up to 1M tokens, while reducing KV cache memory usage by up to 75%. Our findings show that strategically allocating the KV cache budget based on the summarization score distribution of attention heads yields a superior efficiency-accuracy trade-off for long-context inference.

80. 【2609.03221】Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

链接https://arxiv.org/abs/2609.03221

作者:Rohith Reddy Bellibaltu,Manpreet Singh,Deepak Parashar,Rahul Joshi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:treats demographically distinct, agent treats demographically, identical patients differently, clinically identical patients, standard tool

备注: 13 pages, 1 figure, 2 tables. Code and data: [this https URL](https://github.com/rohithreddybc/FairMedAgent)

点击查看摘要

Abstract:Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-running an identical condition ten times over sixteen vignettes (same narrative, same descriptor string, nothing varied) moved a clinical agent's action in 8.7% of outcome-vignette cells, and instability was heterogeneous across actions by a factor of eight, from 0.022 for ICU escalation to 0.179 for controlled-substance caution. No demographic contrast in our data was distinguishable from that floor. A second model gives a pooled floor of 6.7% and ranks the six actions almost identically (Spearman 0.94, exact p=0.017), so the floor is not one system's artefact. Majority-vote aggregation over five draws removes 39% of it and then flattens, and a null simulation attributes the residue to heterogeneous per-cell rates, so replication mitigates without eliminating. Any counterfactual fairness estimate reported without a per-action floor beside it therefore cannot be read as evidence of disparity. The measurements were taken with FairMedAgent, an evaluation harness for disparity in the actions of clinical LLM agents whose estimand, the within-range counterfactual flip rate, counts only flips between actions a published decision rule admits and a clinician has adjudicated. That estimand requires band adjudication, which is under way; no disparity result is claimed here. Each synthetic vignette runs a six-stage trajectory (five model-facing decisions around a deterministic environment step) under fixed-form conditions spanning race, sex, age, insurance, English proficiency, and their intersections. The harness, the floor protocol, and every analysis script are released.

81. 【2609.03218】he Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

链接https://arxiv.org/abs/2609.03218

作者:Ahmed Asaad,Amr Mohamed,Yang Zhang,Omneya Abdelsalam

类目:Computation and Language (cs.CL); Computational Engineering, Finance, and Science (cs.CE); Portfolio Management (q-fin.PM)

关键词:Large Language Models, Large Language, personalize their responses, Language Models, prompts to personalize

备注

点击查看摘要

Abstract:Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.

82. 【2609.03215】SWIM: Student Writing Simulation via Proficiency-Conditioned Generation

链接https://arxiv.org/abs/2609.03215

作者:Heejin Do,Jakub Kontak,Mrinmaya Sachan

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:students develop content, organize ideas, choose words, develop content, Writing

备注: EMNLP 2026 Findings

点击查看摘要

Abstract:Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.

83. 【2609.03213】LLMs Learn Better In-Context from Rules than from Examples

链接https://arxiv.org/abs/2609.03213

作者:Xiang Fu,Seungmin Cho,Yukyung Lee,Najoung Kim

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, exhibit in-context learning, in-context learning capabilities, learning

备注

点击查看摘要

Abstract:Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from prompt contexts without weight updates. We compare the learning efficacies of two prominent modes of in-context learning: (1) learning from descriptions of rules (instruction following); and (2) learning from examples of input-output demonstrations (few-shot prompting). Through five learning tasks that cover diverse domains (games, arithmetic, linguistic inferences), we compare two modes of learning (rules vs. examples) specifying the same underlying task. We furthermore explore model and task properties that modulate the learning efficacies. We find that models generally learn more reliably from rules than from examples alone, and additional examples on top of rules or simply scaling up the number of examples do not lead to consistent and significant gains. Instruction tuning amplifies the benefit of rule-based learning while keeping example-based learning capacities intact. Surprisingly, we find no privileged effect of example-based learning in base models, and rules still lead to gains in algebraic task domains. Overall, the comparative efficacy of rules over examples is larger when the task recruits algebraic abstractions and computations, and smaller when the task requires distributional sensitivity and/or recruits parametric knowledge.

84. 【2609.03206】Learning to Zoom Efficiently with a Contrastive Curriculum

链接https://arxiv.org/abs/2609.03206

作者:Falko Helm,Iryna Gurevych

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:modern visual agents, important foundational part, involving high-resolution images, efficiently handle tasks, handle tasks involving

备注: EMNLP 2026

点击查看摘要

Abstract:Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images. Most previous methods need an extensive warm-start supervised fine-tuning phase for teaching models zoom-in. We show that this is not necessary by proposing a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT. Our InfoNCE-style reward uses a curriculum of increasingly hard negative tool calls as a contrastive training signal. Empirical experiments on $V^*$, HRBench and MME-RealWorld show that our approach is competitive while being more efficient. When used as a drop-in replacement for SFT, we even outperform all baselines. To directly measure the zoom-in ability of models, we further introduce the scalable synthetic MuffinChihuahua (MC) dataset. Each image consists of a grid with every cell either showing a muffin or chihuahua. Leveraging the MC dataset's unique region of interest labels, we find that recall is the metric that most strongly correlates the zoom-in region with final task performance. Our model and code for reproduction is publicly available under this https URL

85. 【2609.03203】VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis

链接https://arxiv.org/abs/2609.03203

作者:Mengzhe Geng

类目:ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)

关键词:Expressive speech systems, Expressive speech, utterance is delivered, Expressive, speech systems make

备注

点击查看摘要

Abstract:Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet downstream audio scores rarely reveal whether those choices were licensed by the source record, a source-use failure that occurs before any waveform exists. VoxReason makes that pre-synthesis decision measurable as a listener-free task for source-grounded speech planning. Before synthesis, VoxReason measures whether delivery choices are grounded in cited source records. Systems output a source-cited speaking-plan with evidence citations, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema validity, and one-cue counterfactual locality. On 1,440 checked source-label cases, shortcut controls show why slot accuracy alone is unsafe: a key-lookup oracle reaches 1.000 plan-slot accuracy on seen keys, while an emotion prior still reaches 0.958 slot accuracy on source-key-disjoint cases without citing intensity or identity. In a separate 100-case learned source-key-disjoint comparison, a 7B locality SFT+CF repair improves plan-slot accuracy/locality from 0.684/0.141 to 0.919/1.000, and removing source records lowers citation-required grounded score by 0.488. Rendered waveform quality remains outside the present evaluation.

86. 【2609.03201】MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval

链接https://arxiv.org/abs/2609.03201

作者:Meriem Yacoubi,Pia Schmidt,Nenad Petrovic,Ahmed Frikha,Martin Kirchhoff,Alois Knoll

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:distinguishing repeated evidence, Long-term LLM agents, agents must preserve, preserve information, information across interactions

备注: 8 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Long-term LLM agents must preserve information across interactions while distinguishing repeated evidence, historical states, updates, and unresolved contradictions. Existing textual memory systems retrieve semantically relevant memories efficiently but often leave these relationships implicit, whereas richer structured approaches model them through global graphs, hierarchical abstractions, or reflection at greater complexity. We introduce MemoryLACE (MemLACE), a lightweight memory framework that explicitly models the lifecycle of textual evidence through sparse merge, supersession, and contradiction relations while preserving atomic natural-language memories and their provenance. Rather than retrieving memories independently, MemLACE reconstructs relation-aware evidence units that expose current, historical, supporting, and conflicting evidence for downstream reasoning. Across BEAM and StructMemEval, using open-weight and proprietary LLM backbones, MemLACE achieves the highest overall performance in same-backbone comparisons while reducing end-to-end runtime on BEAM by 66.6% relative to Hindsight, the strongest reported reflective-memory baseline. Ablation studies identify lifecycle expansion and temporal awareness as the principal contributors to these gains. Together, the results demonstrate that explicitly modeling the local lifecycle of textual evidence is sufficient to substantially improve long-term memory reasoning without requiring comprehensive knowledge graphs or global reflection.

87. 【2609.03181】Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

链接https://arxiv.org/abs/2609.03181

作者:Alejandro Barón García,Feng Wang,Emilia Garcia Casademont,Han Xiao

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:parsing model built, built to serve, document parsing model, document parsing, parsing model

备注: 15 pages, 5 figures, 8 tables. Model at [this https URL](https://huggingface.co/jinaai/jina-ocr-v1)

点击查看摘要

Abstract:We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at this https URL.

88. 【2609.03160】No country for old linguists: LLM-brain alignment underdetermines neural computation

链接https://arxiv.org/abs/2609.03160

作者:Elliot Murphy

类目:Computation and Language (cs.CL)

关键词:alignment, Nastase, Abstract, language, shared

备注

点击查看摘要

Abstract:Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both rely on distributed, context-sensitive representations shaped by statistical learning. Their rejection of simple cortical "boxology" is persuasive, and they articulate a strong case for the value of LLM-brain alignment research. The key question is what kind of inference LLM-brain alignment licenses. My claim here will be narrow: representational alignment can in principle constrain mechanistic hypotheses, but it does not by itself identify a mechanism. Nastase et al. acknowledge that an encoding model can capture features represented in neural activity without establishing a shared architecture or algorithm. Yet the authors sometime move from alignment to "shared computational principles" and ultimately to LLMs as mechanistic models of natural language. Indeed, their methodological caveat that alignment does not establish a shared architecture or algorithm sits uneasily with their conclusion that LLMs might instantiate the same computational principles as biological brains and provide a "fully mechanistic model" of language. I discuss what I consider to be problems of logical, causal, and computational underdetermination in Nastase et al.'s (2026) proposal.

89. 【2609.03158】Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

链接https://arxiv.org/abs/2609.03158

作者:Qingchan Zhu,Weihang You,Hanqi Jiang,Changdi Yang,Tianming Liu,Geng Yuan

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Visual token pruning, Visual token, vision-language models, reduces the inference, inference cost

备注: Accepted to EMNLP 2026 main

点击查看摘要

Abstract:Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.

90. 【2609.03150】Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning

链接https://arxiv.org/abs/2609.03150

作者:Mehreen Hossain Chowdhury,Nowshin Mahjabin,Ahmed Shafin Ruhan,Md Azam Hossain,Abu Raihan Mostofa Kamal,Md Tahmid Rahman Laskar

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Multi-domain fine-tuning, separates domain-specific updates, token-level routing separates, routing separates domain-specific, assuming that token-level

备注: 13 pages, 1 figure, 16 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

点击查看摘要

Abstract:Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates domain-specific updates. We test this assumption in MoE+LoRA using Python code paired with biomedical text and mathematical reasoning. Although these domains show near-disjoint expert routing, adding biomedical data substantially increases code perplexity, indicating that routing separation alone may not prevent negative transfer. To localize the failure, we introduce Jaccard routing overlap and adapter-gradient cosine similarity, which measure expert sharing and update compatibility, respectively. These diagnostics indicate that interference arises mostly from nearly orthogonal domain gradients competing within the same low-rank adapter subspace. We address this issue with SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts when adapter-level contention is detected, while keeping the router fixed. We evaluate SpawnLoRA on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and find that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA. These results demonstrate that structural separation inside experts provides benefits beyond routing or rank expansion alone.

91. 【2609.03148】Large Language Models in Resolving Contextual Knowledge Conflicts

链接https://arxiv.org/abs/2609.03148

作者:Xinye Yang,Zhenyang Liu,Ruisi Li,Yuanyuan Lei

类目:Computation and Language (cs.CL)

关键词:externally provided context, prior works focused, internal parametric knowledge, LLM internal parametric, provided context

备注: Accepted to EMNLP 2026

点击查看摘要

Abstract:Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual conflicts (factual, inferential, temporal, granularity, perspective, and ambiguity) and contribute a comprehensive dataset ContextConflict for this setting. The dataset contains 5,781 samples, covers both reasoning and summarization tasks, and includes both explicit contradictions and implicit conflicts that require multi-step reasoning. Experiments on nine LLMs show that current models still fall short in resolving contextual knowledge conflicts. We further provide mechanistic interpretability insights into how LLMs process such conflicts, revealing their latent awareness of conflicts and the representational geometry underlying conflict processing. In addition, our analysis uncovers a consistent model bias towards earlier evidence, and this positional preference serves as a key obstacle to effective conflict resolution. Motivated by these findings, we further propose a simple training-free, label-free steering method that steers activations to encourage a more comprehensive incorporation of evidences for better conflict resolution. On our dataset, the method consistently improves accuracy on reasoning tasks and generates higher-quality, more balanced summaries for summarization tasks.

92. 【2609.03047】SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

链接https://arxiv.org/abs/2609.03047

作者:Michael J. Bommarito II

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:archives manage large, manage large collections, Libraries and archives, Evaluating LLM Fitness, archives manage

备注: 17 pages, 15 tables. Code available at [this https URL](https://github.com/mjbommar/shelf-benchmark) ; data available at [this https URL](https://huggingface.co/datasets/mjbommar/SHELF)

点击查看摘要

Abstract:Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Synthetic Harness for Evaluating LLM Fitness, addresses this gap. It is a Python system that turns labelled taxonomies, writing specifications, and a generation budget into controlled benchmark data and evaluation tasks. This first release contains 62,899 model-written documents based on Library of Congress vocabularies, with tasks for classification, clustering, retrieval, pair classification, and instruction retrieval. We compare TF, TF-IDF, BM25, popular encoders, and, on subject classification only, zero-shot decoders; each method appears only on tasks that support it. Subject classification reaches 0.8887, while genre-form classification reaches only 0.2605, and several pair and clustering tasks remain near chance. Sparse methods remain competitive on classification, while TF-IDF is the fastest measured arm in the subject timing experiment. SHELF also varies bibliographic facets independently and can generate new, verifiably unseen documents after a model's training cutoff. Comparisons with LCSHBench and Project Gutenberg show that model rankings transfer more reliably than absolute scores, but SHELF scores do not estimate accuracy on production catalogue data. We release all source code and data under permissive licenses on GitHub and Hugging Face.

93. 【2609.03005】Unifying Conformal Language Tasks with In-Context Ensembles

链接https://arxiv.org/abs/2609.03005

作者:Xiao Shi Huang,Chen-Yuan Lin,Bruce Kuwahara,Kin Kwan Leung,Jesse C. Cresswell

类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:extractive question answering, retrieving relevant content, question answering, reduce to retrieving, retaining enough pertinent

备注: Findings of EMNLP 2026. Code is available at [this https URL](https://github.com/layer6ai-labs/conformal-relevance)

点击查看摘要

Abstract:Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and conciseness, removing as much irrelevant information as possible. Conformal prediction methods have been used to guarantee coverage, and must be optimized for conciseness through design of a score function. State-of-the-art scoring functions use hand-engineered LLM prompts asking the model to rate the importance of content, but manual prompt engineering is labor-intensive and task-specific. We introduce the Conformal Relevance framework which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input. We demonstrate this framework's application on seven NLP tasks, and also theoretically study the impact of diversity for ensembled conformal scores, giving a complementarity condition that characterizes when ensembling improves worst-case sentence scores, and a saturation bound on ensemble improvement.

94. 【2609.02998】Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

链接https://arxiv.org/abs/2609.02998

作者:Zhiwei Zhang,Zechen Sun,Fei Zhao,Kang Peng,Bin Liang,Huayu Deng,Yao Hu,Kam-Fai Wong,Mu Chuan

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:providing dense token-level, accelerates post-training, On-policy distillation, post-training by providing, TGOPD

备注: 17 pages, 6 figures, 7 tables

点击查看摘要

Abstract:On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.

95. 【2609.02959】he Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors

链接https://arxiv.org/abs/2609.02959

作者:Toni J.B. Liu,Jiajun Bao,Yizhou Liu,Gurbir Arora,Nicolas Boullé,Raphaël Sarfati,Christopher J. Earls

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:language model predict, final prediction state, prediction state, language model, model predict

备注

点击查看摘要

Abstract:What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the model falls back on when uncertain. This structure --- which we term the \emph{direction of ignorance} --- appears in all four model families examined (\texttt{Llama}, \texttt{Qwen}, \texttt{Gemma}, and \texttt{Pythia}), ranging from 0.4B to 405B parameters. Projecting the final prediction state onto this direction yields a per-token \emph{prior loading factor} $\lambda$, which, empirically, declines steadily as the context becomes more informative. Formally, the same projection decomposes the prediction state into two orthogonal vectors that correspond exactly to the two factors of a tempered Bayesian update: a unigram prior raised to the exponent $\lambda$ and a context-driven likelihood. This geometric-probabilistic interpretation calibrates $\lambda$, making it meaningfully comparable across model sizes and families, with larger models generally exhibiting lower prior reliance in the high-context limit. Finally, we show that the direction of ignorance is causally active: raising or lowering $\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence.

96. 【2609.02954】LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

链接https://arxiv.org/abs/2609.02954

作者:Huiyuan Xie,Yuqin Huang,Zhicheng Hao,Yida Cai,Shaochun Wang,Zhenghao Liu,Yuxiao Ye

类目:Computation and Language (cs.CL)

关键词:legal issue, legal, legal issues, issue, legal issue identification

备注

点击查看摘要

Abstract:Identifying the issues disputed between litigating parties is a crucial component of real-world litigation. However, legal issues remain comparatively underexplored in legal AI research. In this work, we study the computational modelling of legal issue identification in litigation. We introduce a legally grounded hierarchical schema that represents legal issues through both free-form issue descriptions and structured legal categories, and formulate legal issue identification as two complementary tasks: legal issue generation and legal issue classification. Based on this formulation, we construct LexIssue, a benchmark containing 430 real-world Chinese civil litigation cases and 1,303 expert-annotated disputed legal issues. We further develop an issue-centric legal knowledge base spanning 27 causes of action and 441 candidate legal issue entries to support retrieval-augmented reasoning. Experimental results across a diverse set of models show that retrieval-augmented generation using the constructed legal issue knowledge base consistently improves performance in identifying disputed legal issues and their corresponding legal attributes.

97. 【2609.02947】Privacy-Preserving Heterogeneous Multi-LLM Federated Inference for Cognitive Diagnosis

链接https://arxiv.org/abs/2609.02947

作者:Yagna Manasa Boyapati,Chong Yu,Tianyu Jiang,Justin Zhan

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Significant challenges remain, accurate cognitive diagnosis, Significant challenges, AI-driven educational systems, cognitive diagnosis

备注: 16 pages, 3 figures, Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Significant challenges remain in AI-driven educational systems in balancing privacy preservation with accurate cognitive diagnosis. To overcome this, we propose a federated inference framework in which several commercial LLM APIs collaborate without requiring access to raw student data or proprietary model internals. Using multiple federated entities, such as LLaMA-3.3-70B, GPT-4o-mini, and Claude-3-Haiku, our framework builds upon a heterogeneous multi-LLM architecture. The predictions generated by these entities are combined with epsilon-local differential privacy by adding Laplace noise locally to each entity's prediction output before aggregation, while residual-based aggregation mitigates model heterogeneity. Our approach is predicated on an honest-but-curious trust paradigm in which API providers are presumed not to abuse submitted queries, and our differential privacy mechanism shields the published diagnostic results from external inference. We conduct rigorous privacy-utility analysis showing strong privacy guarantees with minimal accuracy loss, and extensive real-world evaluations across three educational benchmarks confirm the framework's practical usability and cross-domain generalizability.

98. 【2609.02942】Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

链接https://arxiv.org/abs/2609.02942

作者:Anshul Bagaria,Sowmya S Sundaram,Gokul S Krishnan,Balaraman Ravindran

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:evaluate AI-generated text, pipelines are increasingly, evaluate AI-generated, judgments arise, arise from reasoning

备注: Accepted for publication at EMNLP 2026. 5 pages, 6 figures

点击查看摘要

Abstract:LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggests that rubric formulations encode recoverable evaluative signals, allowing scores to be partially anticipated independently of model outputs. Finally, counterfactual perturbations reveal that judges often fail to reliably update their decisions when either the candidate response or the rubric criterion is reversed. Our findings raise concerns about the reliability of rubric-based LLM evaluation and highlight the need for further methodological study of automated evaluation via LLMs.

99. 【2609.02941】SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

链接https://arxiv.org/abs/2609.02941

作者:Eunseo Choi,Hyunku Kang,Chanwoo Kim

类目:ound (cs.SD); Computation and Language (cs.CL)

关键词:Speech emotion recognition, emotion recognition systems, emotion recognition, Speech emotion, Speaker-Invariant Speech Emotion

备注: Accepted to INTERSPEECH 2026

点击查看摘要

Abstract:Speech emotion recognition (SER) faces two fundamental challenges: scarcity of labeled data and inter-speaker variability, both of which hinder generalization of emotion recognition systems. While prior adversarial approaches address speaker variability, they fall short in leveraging powerful pre-trained representations. We propose SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme. wav2vec 2.0 provides rich self-supervised representations that alleviate dependency on large labeled datasets, while ECAPA-TDNN enables suppression of speaker identity via a stronger adversarial signal than shallow classifiers. Evaluated on IEMOCAP, SISER achieves a UA of 60.63%, outperforming the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%), with ablation emphasizing that the choice of speaker classifier architecture is a key factor.

100. 【2609.02940】Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

链接https://arxiv.org/abs/2609.02940

作者:Chan-Jan Hsu,Jaeyeon Kim,Chao-Han Huck Yang,Shinji Watanabe,Hung-yi Lee,Carlos Busso

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Recent automatic speech, automatic speech recognition, systems increasingly integrate, increasingly integrate large, integrate large language

备注: 24 pages, 8 figures, 2 tables. Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two strategies remains underexplored. In this work, we refine warm-initialized LLM-based ASR models by leveraging their own pre-adaptation base LLMs, focusing on LoRA-adapted settings where the base LLM is preserved. To achieve this, we propose Hybrid Search, a targeted correction strategy motivated by two observations. First, interaction features that characterize the relationship between LLM-based ASR hidden states and base-LLM hidden states provide informative signals about a token's degree of semantic dependence. Second, selectively refining targeted tokens with high semantic dependence improves ASR performance far beyond naive global LLM-correction methods including rescoring and late fusion. Our analysis suggests that, even after semantic knowledge transfer through warm initialization, LLM-based ASR models can still leverage their base LLM to further improve inference-time performance.

101. 【2609.02902】RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

链接https://arxiv.org/abs/2609.02902

作者:Ram Narayanan,Harshit Rajgarhia,Abhishek Mukherji

类目:Computation and Language (cs.CL)

关键词:Deploying task-oriented dialogue, user behaviour evolves, behaviour evolves faster, robust training requires, task-oriented dialogue agents

备注

点击查看摘要

Abstract:Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to annotate, while user behaviour evolves faster than labelling pipelines can keep pace. We present RL-ADA (Reinforcement Learning with Adversarial Dialogue Agents), a co-evolutionary training framework that eliminates this bottleneck by replacing human labels with \emph{world feedback}: consequence-based reward signals derived directly from measurable interaction outcomes. A Customer Support Agent (DA, 3B parameters) and an Adversarial Customer Agent (CA, 7B parameters) co-evolve in an adversarial arena guided by a fixed automated judge: the DA is rewarded for correctly handling multi-turn customer conversations to successful resolution, while the CA is rewarded for producing realistic, intent-concealing utterances that cause misroutes, creating asymmetric adversarial pressure through opposing but independently structured rewards. An isolation gym iteratively retrains the weaker agent on prior-failure transcripts, requiring no human annotation at any stage. In a banking customer support proof of concept, tool-routing errors are eliminated and the strict end-to-end PASS rate doubles over five co-evolutionary cycles, driven solely by automated arena reward with no labelled data. We additionally observe the emergence of \textbf{Contextual Camouflage}, an adversarial strategy in which the CA learns to embed intent within dense realistic customer detail purely from reward pressure, with direct implications for enterprise red-teaming and robustness evaluation.

102. 【2609.02901】Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

链接https://arxiv.org/abs/2609.02901

作者:Fengrun Zhang,Li Fu,Wangjin Zhou,Lu Fan,Youzheng Wu,Xiaodong He

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:Modern automatic speech, Modern automatic, automatic speech recognition, inverse text normalization, scenarios require

备注: Submitted to IEEE SLT 2026

点击查看摘要

Abstract:Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically produced by cascaded modules, where a spoken-form ASR output is rewritten by a separate ITN component, making written-form ASR-ITN vulnerable to recognition errors and decoupling normalization from acoustic-contextual modeling, especially for semantically dependent numeric expressions. In this paper, we propose Dual-Form ASR (DF-ASR), a framework that extends spoken-form ASR capability to semantics-aware written-form ITN through paired spoken-form and written-form supervision while retaining prompt-level selection between transcript forms. The dual-form supervision is constructed via a large language model (LLM)-driven generate-and-judge workflow, and training is further enhanced by ITN-MWER, a sequence-level objective that assigns higher cost to errors on normalization-sensitive spans. We also introduce a decision-aware REQUIRE-ITN/\FORBID-ITN protocol to separately measure required normalization and forbidden-span preservation. On manually annotated Chinese subsets from SpeechIO, DF-ASR consistently outperforms open-source ASR-ITN systems, remains competitive with strong closed-source references, and preserves reliable prompt-level control between spoken-form and written-form outputs.

103. 【2609.02899】Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

链接https://arxiv.org/abs/2609.02899

作者:Xingyao Xiao(Stanford University),Yihong Cheng(City University of Macau)

类目:Computation and Language (cs.CL); Applications (stat.AP); Methodology (stat.ME)

关键词:large language model, reliability of large, large language, contamination, LLM

备注: 21 pages, 4 figures, 3 tables. Code and data: [this https URL](https://github.com/DoriaXiao/anchor-dif)

点击查看摘要

Abstract:Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC, GSM8K, HellaSwag, MMLU), we first calibrate the measure against ground truth: it recovers injected contamination dose-responsively (a corrected effect of +0.187 accuracy points for test-set leakage) and never flags a negative-control model trained only on the legitimate training split (-0.012). We then quantify leaderboard impact: the rank correlation between a standard leaderboard and a paraphrase-controlled leaderboard is 0.997, and a sensitivity analysis shows that the observed differential contamination is far below the level needed to move rankings, with only 3 of 188 model-by-benchmark cases showing differential contamination corroborated across two references. Contamination among these public models is therefore largely uniform: it inflates absolute scores without reordering the leaderboard, and ranking distortion requires the rare case of differential contamination. We provide a calibrated invariance audit, released as a reference implementation, and recommend that leaderboards report paraphrase-controlled rankings alongside confidence intervals.

104. 【2609.02898】Distilled Rapid Embedding Transfer (DRET): Parameter-Efficient Biomedical Domain Adaptation via Priority-Based Embedding Transfer

链接https://arxiv.org/abs/2609.02898

作者:Girish Sundaram,Daniel Berleant

类目:Computation and Language (cs.CL)

关键词:ClinicalBERT achieve strong, computational demands make, Large domain-specific language, biomedical NLP tasks, domain-specific language models

备注: 11 pages, 5 figures, 6 tables

点击查看摘要

Abstract:Large domain-specific language models such as BioBERT and ClinicalBERT achieve strong performance on biomedical NLP tasks, but their computational demands make them impractical for many real-world deployments. General-purpose, parameter-efficient models such as DistilBERT are lightweight yet lack the domain knowledge required for specialized tasks such as PICO (Population, Intervention, Comparison, Outcome) classification. We introduce Distilled Rapid Embedding Transfer (DRET), a knowledge-transfer paradigm that injects biomedical domain knowledge from large specialized models into a smaller general-purpose model without retraining on the original specialized corpora. DRET is developed as an iterative family of strategies: a unified tokenizer-merge strategy (DRET 1.x), hybrid embedding averaging (DRET 2.0), and a priority-based embedding-transfer mechanism (DRET 3.x) that hierarchically selects embeddings from the most authoritative source models, further combined with embedding-layer freezing, differential learning rates, label propagation, and imbalance-aware loss functions (DRET 4.x). We evaluate DRET on token-level PICO classification using the EBM-NLP corpus under severe class imbalance, across a twelve-metric battery. DRET-enhanced DistilBERT (66M parameters) attains balanced accuracy, recall, and ROC-AUC competitive with, and on several class-wise metrics exceeding, models an order of magnitude larger, while retaining DistilBERT's efficiency. We further show that transfer occurs at the embedding level through cosine-similarity, semantic-shift, and t-SNE analyses. DRET offers a scalable, resource-efficient route to near-domain-expert performance for biomedical text mining, with direct application to automated systematic literature reviews and clinical decision support.

105. 【2609.02897】Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

链接https://arxiv.org/abs/2609.02897

作者:Oszkár Urbán,Young D. Kwon,Stylianos I. Venieris,Cecilia Mascolo

类目:Computation and Language (cs.CL)

关键词:accelerates LLM inference, decoding accelerates LLM, accelerates LLM, LLM inference, drafting candidate tokens

备注

点击查看摘要

Abstract:Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted token to its top-1 probability exceeds a threshold with no dependence on draft length or underlying drafter architecture. A per-step tree policy adjusts the draft tree's depth, width, and node count directly from a fused signal of draft top-1 confidence and a rolling acceptance history capturing recent draft-target agreement, allowing the total draft count to vary rather than only be redistributed. The two adaptations operate on orthogonal axes and compound in effect. Implemented on the SGLang production-grade serving engine, AdaptiveSpec improves throughput over the state-of-the-art autoregressive speculative decoding method EAGLE-3 by up to 56%, recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three target models (DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, Qwen3-8B).

106. 【2609.02896】PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction

链接https://arxiv.org/abs/2609.02896

作者:Jiaxin Duan,Fengyu Lu,Junfei Liu

类目:Computation and Language (cs.CL)

关键词:attracted considerable attention, Medical relation extraction, recent years, attracted considerable, considerable attention

备注

点击查看摘要

Abstract:Medical relation extraction (MRE) is commonly known for extracting entities and their relations jointly from a medical text, which has attracted considerable attention in recent years. Previous studies treat MRE as a sequence tagging task, which results in either a challenging design of the tagging schema or a failed extraction of multiple relations, due to intricate relationships among medical entities. In this work, we review the task from the linguistic perspective and propose a novel pipeline framework, PiPMRE, developed on language models to enhance MRE performance. Specifically, PiPMRE consists of a relation generator and a relation filter. Given a text, the generator first yields multiple relational triplets, and then the filter scores each triplet and retains only those that pass the borderline as the final results. Implementing PiPMRE requires no tagging schema; instead, we use a simple template to reformulate the input text, ensuring that entities and relations are generated in a contextual order. Extensive experimental results on two public datasets demonstrate the advancement of PiPMRE. It surpasses the previous state-of-the-art by an average of 5.6 recall points and 4.4 accuracy points. PiPMRE's superiority is also demonstrated in few-shot settings.

107. 【2609.02895】BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events

链接https://arxiv.org/abs/2609.02895

作者:Parth Bramhecha,Smit Deshmukh,Sairaj Bodhale,Adwait Borate,Raviraj Joshi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:posing substantial risks, Large-scale public events, political rallies, religious festivals, posing substantial

备注

点击查看摘要

Abstract:Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social cohesion. While automated fake news detection has seen significant methodological progress, existing benchmarks frequently fail to capture the socio-cultural nuances and event-specific dynamics characteristic of the Indian context. This paper introduces BharatGather, a curated, multi-source dataset specifically engineered for binary misinformation classification within the ecosystem of Indian mass gatherings. The corpus comprises 14,646 records constructed through a hybrid pipeline involving systematic web scraping of prominent fact-checking platforms, multimedia transcript extraction, and Large Language Model (LLM)-mediated synthetic augmentation to ensure narrative diversity. By providing a resource tailored to the unique complexities of event-aware misinformation in India, this work facilitates the development of culturally informed detection systems and establishes a rigorous benchmark for evaluating their performance in high-stakes public environments.

108. 【2609.02894】R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG

链接https://arxiv.org/abs/2609.02894

作者:Yucan Guo,Miao Su,Saiping Guan,Long Bai,Zhongni Hou,Zixuan Li,Xiaolong Jin,Jiafeng Guo,Xueqi Cheng

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Large Language Models, enhancing Large Language, Language Models, Large Language, Retrieval-Augmented Generation

备注

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently handles simple queries but struggles with relational or multi-hop reasoning. Graph-based RAG alleviates this issue but incurs higher inference complexity and latency. In practice, user queries can differ significantly in their complexity, rendering a fixed RAG strategy suboptimal. However, existing hybrid text-graph RAG methods typically rely on heuristic and LLM-based routing, resulting in unnecessary overhead and strong dependence on the underlying LLM. To address these challenges, we propose R$^{2}$Adapter, a lightweight plug-in Routing and Rewriting Adapter designed to allocate queries between vanilla and graph-based RAG dynamically. By routing only the queries that genuinely benefit from graph-based reasoning, R$^{2}$Adapter reduces unnecessary graph retrieval overhead. Additionally, uncertain graph-routed queries are rewritten to better expose their multi-hop reasoning requirements, improving retrieval quality without additional supervision. Extensive experiments on three multi-hop QA benchmarks demonstrate that R$^{2}$Adapter reduces graph-based RAG usage by up to 59% while maintaining comparable answer accuracy. This adapter is model-agnostic and can be seamlessly integrated into diverse vanilla and graph-based RAG pipelines, providing an efficient and adaptive solution for hybrid RAG systems.

109. 【2609.02893】Probe Generalization as Subspace Selection for OOD Deception Detection

链接https://arxiv.org/abs/2609.02893

作者:Daniel Yoo,Adrians Skapars

类目:Computation and Language (cs.CL)

关键词:concepts inside language, inside language model, language model activations, Linear probes, detect behaviors

备注

点击查看摘要

Abstract:Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datasets, we find that projecting inputs onto a small subset of principal components (PCs) from the training distribution of activations enables cross-domain transfer that nearly matches the performance of probes trained directly on the test distribution. Furthermore, we find that PC interpretations can be used to find a subset of those transferable PCs. By using an LLM judge to score each PC on whether its most/ least activating examples imply a transferable deception direction, then probing on the highest-scoring PCs, we close the baseline-to-oracle gap by 78% on Insider Trading Report and by 25% on Sandbagging. The directions a source probe weights heavily appear to encode source-specific surface features, while the directions that actually transfer appear to encode the same contrast more abstractly, in a way natural language descriptions can capture. Broadly, our results suggest that the OOD robustness of probes is largely determined by subspace selection.

110. 【2609.02892】Counterexamples as Feedback for Agent Self-Correction

链接https://arxiv.org/abs/2609.02892

作者:Sidhesh Badrinarayan,Adithya Parthasarathy

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Single-turn code-generation metrics, code-generation metrics understate, Single-turn code-generation, receiving concrete feedback, code-generation metrics

备注

点击查看摘要

Abstract:Single-turn code-generation metrics understate a central property of deployed agents: whether they can repair a wrong artifact after receiving concrete feedback. This paper presents A-CEGIS, a lightweight framework that uses counterexamples as feedback for evaluating multi-turn refinement in natural-language-to-regex synthesis. An agent proposes a regex, a deterministic oracle checks it under full-match semantics, and compact false-positive or false-negative witnesses guide the next turn. On 30 NL-RX-Turk tasks, diagnostic counterexample feedback solves 90\% of tasks within a four-turn ablation budget, compared with 17% for zero-shot generation, 27% for generic self-correction, and 23% for error-only feedback. In a full diagnostic run with hardening, all tasks are solved on the hidden set by the final turn, with mean time-to-success of 2.7 turns and robust success of 77% after targeted probing. These results show that A-CEGIS measures how efficiently an agent improves across turns while adding a practical robustness check beyond the original held-out cases.

111. 【2609.02890】Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent

链接https://arxiv.org/abs/2609.02890

作者:JaeHa Yoon,Minjun Park,Seoyeon Kim,Jiwoo Lee,Hyunwoo Choi,Dohyun Kang

类目:Computation and Language (cs.CL)

关键词:personalized language agent, user interaction history, inference time, personalized language, request at inference

备注

点击查看摘要

Abstract:A personalized language agent must convert a user's interaction history into behavior on each new request at inference time. Two strategies dominate. Retrieval pulls a few of the user's most relevant past items into the prompt, which is accurate but pays a per-query selection and context cost that grows with the history. Distillation instead compresses the history once into a compact natural-language persona, which is bounded, query-independent, and interpretable, but is widely assumed to sacrifice accuracy. Whether, and on which tasks, a distilled persona can match retrieval has not been characterized cleanly. We introduce PersonaLink, a training-free method that distills a user's history into a bounded three-field persona and recursively refines it: each pass self-evaluates the frozen agent on a held-out slice of the user's own labeled history, rewrites the persona from its errors, and keeps the result only when it does not regress on that slice. Because every comparison shares one frozen 7B backbone and differs only in what is placed in context, the design isolates the effect of representation from that of the model. The result is a clear task-type asymmetry. On 200 users of LaMP-2 (15-way news categorization), PersonaLink reaches 0.745-0.755 accuracy, statistically indistinguishable from BM25 retrieval (0.760-0.765).

112. 【2609.02889】Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

链接https://arxiv.org/abs/2609.02889

作者:Michael Nguyen,Wei Chen Tan,Nurul Aisyah Hassan,Arvind Raman,Li Hua Lim,Ahmad Faiz Razak

类目:Computation and Language (cs.CL)

关键词:large language models, including persona, format rules, growing body, body of work

备注: 17 pages

点击查看摘要

Abstract:A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflective prompt-evolution methods usually optimize this harness as one flat string. We instead ask where the optimization value actually resides. We introduce HARNESSEVO, which decomposes the harness into four separately evolvable slots: role, task-strategy, tool/format-rules, and reflection/control. Using the same reflective optimizer under an iso-budget setting, we pair this decomposition with leave-one-in and leave-one-out attribution to measure the contribution of each slot. On ALFWorld with a frozen 7B backbone, HARNESSEVO does not significantly improve the overall binary success rate over either the stock harness or flat-string evolution: 0.657 versus 0.642 and 0.642, respectively. However, the slot-level analysis reveals that nearly all useful optimization value is localized in the reflection/control slot, which achieves a leave-one-in gain of +0.119. The other slots are individually null. We further show that uniform budget splitting is harmful: allocating 64 rollouts across four slots leaves only 16 per slot, below the optimizer's effective search floor, causing every slot to freeze at its empty seed. Concentrating the budget on the high-credit control slot recovers the lost gain, reaching 0.761 with half the split budget. The effect is task-contingent. On WebShop, all slots freeze empty and all methods tie, indicating a genuine absence of recurrent, verbalizable control failures rather than budget starvation. Overall, our results suggest that harness value is localized, uniform budget splitting can be actively harmful, and credit assignment should precede structured agent-evolution.

Comments:
17 pages

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.02889 [cs.CL]

(or
arXiv:2609.02889v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.02889

Focus to learn more

              arXiv-issued DOI via DataCite</p>
113. 【2609.01865】ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval

链接https://arxiv.org/abs/2609.01865

作者:Aaryan Kapoor,Md Abdullah Al Hafiz Khan

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Embedding-based code retrieval, retrieving lexically similar, retrieval-augmented code generation, lexically similar code, retrieving correct code

备注: Accepted to EMNLP 2026 (Main Conference). Camera-ready version. 17 pages, 6 figures, 8 tables

点击查看摘要

Abstract:Embedding-based code retrieval is a core component of coding agents and retrieval-augmented code generation, where retrieving correct code matters more than retrieving lexically similar code. Existing code-retrieval benchmarks do not plant controlled, execution-verified single-edit variants of each query's canonical implementation in the search pool, leaving the question of whether embeddings can functionally discriminate correct from near-clone-but-incorrect code unanswered in a retrieval setting. Resolving this requires a benchmark whose search pool itself contains the relevant counterfactuals -- execution-verified buggy variants near-identical to each canonical -- so that a retriever's rank ordering can be directly tested for functional discrimination rather than topical or identity overlap. We introduce ExecRetrieval, 939 Python tasks each paired with one execution-verified canonical implementation and up to four execution-verified buggy distractors, each generated by a mechanical mutation making a single targeted edit, and evaluate 23 dense embedding configurations plus BM25 under provider-native invocation with paired McNemar tests and query-level bootstrap intervals. With near-clone counterfactuals in the pool, the top hosted system reaches exec@10 = 1.00 but only exec@1 = 0.331; rank-1 misses are paired buggy variants 91.5-99.4% of the time across the four leading systems, and the canonical scores below at least one of its four paired distractors in 67-78% of queries on the leading systems. The full dataset, execution oracle, embedding matrices, environment snapshot, and pairwise statistical tests are released at the URL in Appendix D.

114. 【2608.29517】LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

链接https://arxiv.org/abs/2608.29517

作者:Veerendra Kumar Sunkavalli

类目:Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:Large language models, Large language, learning analytics, evaluated almost exclusively, graders in learning

备注: 14 pages, 2 figures, 8 tables. Under review at the Journal of Learning Analytics (LAK27 research track). Artifact: [this https URL](https://anonymous.4open.science/r/lak-27-A618)

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-registered rater-effects battery (many-facet Rasch severity, residual halo, generalizability/decision studies, cross-version shifts, differential functioning) on public corpora in two languages (ENEM/Essay-BR; ASAP): 2,377 essays, 12 judges, 4 providers, 5 version contrasts, replicated cells, released as a score tensor. Judge severity spans 219 points on ENEM's 0-1000 scale; on ASAP the panel spread is 15-33% of the score range against a between-trained-human gap near 1%. Judge-human correlations sit in an undiscriminating .47-.56 band. All five version contrasts shift severity beyond a family-wise permutation null (up to 133 points), and one judge was deprecated mid-study, caught by identity canaries. Two pre-registered tests returned honest nulls: severity-adjusted leaderboard reversals did not survive a permutation null, and "silent drift" was refuted: agreement moved with severity in four of five contrasts. Replication yields self-consistency (phi=.80 at k=2) but not human-level accuracy, and a same-instrument check overturned our own halo comparison: matched on instrument and calibration, we find no credible evidence that judge halo exceeds the trained-human range.

115. 【2608.24683】One Timeline, Many Renderings: A Wolfram Language Paclet for heterogeneous musical output

链接https://arxiv.org/abs/2608.24683

作者:Francesco Vitucci,Michele Lorusso,Francesco Scagliola

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:engraved notation, algorithmic composition, composition may require, real-time control, Wolfram Language paclet

备注: Accepted at the International Csound Conference (ICSC) 2026

点击查看摘要

Abstract:One algorithmic composition may require a Csound score, engraved notation, real-time control, and a rehearsal click. Authored separately, their timelines drift. Temporal System is a Wolfram Language paclet that instead compiles one immutable store of typed entities on a rational beat timeline through backend-specific contracts. It emits Csound synthesis, beta MusicXML 4.0, OSC control, and click artifacts that remain synchronized because they share that store. Conversion to seconds, samples, or hertz occurs only at render time. Csound notes use stable named instruments in external .orc files; curves become k-rate signals declared against score p-fields. The click backend derives rehearsal audio from the same meter and tempo and reuses the Csound serializer. We describe the temporal, semantic, and rendering-contract layers, their practical trade-offs, and the limits of this proprietary authoring environment within an otherwise open-source ecosystem. The archived supplement exposes the reported outputs pending paclet release.

116. 【2608.24660】From local kernels to global form: modeling the emergence of musical content

链接https://arxiv.org/abs/2608.24660

作者:Francesco Vitucci,Michele Lorusso,Francesco Scagliola

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:including non-homogeneous formulations, Markov models, including non-homogeneous, non-homogeneous formulations, models are established

备注: Accepted at XXV CIM - Colloquio di Informatica Musicale, L'Aquila, 2026

点击查看摘要

Abstract:Markov models are established tools for symbolic music, including non-homogeneous formulations. The narrower contribution examined here is an observation-driven estimation mechanism: overlapping sliding windows derive a trajectory of local transition kernels from one symbolic sequence rather than from an exogenous formal partition. We test this mechanism on 273 logical note events from Debussy's Syrinx (1913), using the often-proposed A-B-A' reading as a reference rather than ground truth. We apply the same validation to absolute-pitch and notated-duration kernels. At $L=6$, both reference boundaries attain the Jensen--Shannon maximum in both dimensions; the duration plateau is substantially narrower (64 of 267 comparisons) than the pitch plateau (210 of 267). Because the theoretical maximum for consecutive sliding-window comparisons is set by window geometry and equals $1/\sqrt{L-1}$ for maximal turnover of the entering/leaving transition, the pitch value at $L=6$ and its broad plateau are not, by themselves, strong evidence. Their cross-dimensional alignment is consistent with boundary sensitivity, while the broad plateaus preclude treating either curve alone as a unique automatic segmenter. Five-hundred-draw re-synthesis experiments quantify departure from the source in both dimensions and expose an exact-copy degeneracy at $L=2$.

117. 【2609.03355】ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

链接https://arxiv.org/abs/2609.03355

作者:Quang Hoang Trung,Quang Huu Hieu,Nguyen Van Hoang Phuc,Vo Nguyen Le Duy

类目:Machine Learning (stat.ML); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Logit-based knowledge distillation, autoregressive language models, Logit-based knowledge, autoregressive language, language models

备注

点击查看摘要

Abstract:Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives. Existing local approaches often select candidate tokens from either the teacher or the student alone. Teacher-only selection can miss tokens that the student considers likely, while student-only selection can rely on an inaccurate ranking early in training. We propose Adaptive Local Relational Alignment (ALRA), a position-specific framework combining student proposals with teacher guidance. At each valid prediction position, the student proposes likely tokens, while the teacher's most probable token is included as an anchor. ALRA adjusts the number of selected tokens according to how broadly the teacher distributes probability within this candidate set relative to the current batch. Adaptive Local Divergence retains the mass-matching term and separately matches the relative token distributions within the selected and remaining vocabulary regions. Unlike the exact full-vocabulary decomposition, it replaces the teacher-mass coefficients of the two conditional terms with unit coefficients, preventing either term from being downweighted solely because its region has low teacher probability. Student-Weighted Pairwise Relational Alignment emphasizes high-probability token pairs with small student probability gaps and gives less weight to unlikely or clearly separated pairs. Experiments on The Pile with randomly initialized 200M- and 500M-parameter students across nine zero-shot benchmarks yield average accuracies of 36.62% and 37.40%. ALRA exceeds the strongest competing distillation baseline by 0.94 and 0.83 percentage points and improves over pre-training without distillation by 2.31 and 2.91 points, respectively.

118. 【2609.02911】A Spectral Phase Admissibility Certificate for Complex Linear Maps

链接https://arxiv.org/abs/2609.02911

作者:Snigdha Chandan Khilar

类目:General Physics (physics.gen-ph); Computation and Language (cs.CL)

关键词:Kontsevich Segal Witten, Segal Witten criterion, linear maps Standard, maps Standard techniques, Standard techniques analyze

备注

点击查看摘要

Abstract:The paper imports the Kontsevich Segal Witten criterion from quantum gravity into machine learning to evaluate complex linear maps Standard techniques analyze magnitude or positive definiteness whereas this method exclusively limits the collective phase of a spectrum The researchers create three distinct differentiable certificates comprising a determinant sector a subset product envelope and the full criterion The subset envelope prevents all exterior power eigenvalues from touching the negative real axis This constraint precisely matches the accept or reject choices of an exponential minor enumeration while reducing processing expenses drastically The team provides a differentiable enforcement application via a Schur parameterization The document also identifies crucial boundaries regarding where this system works The constraint cannot balance deep linear propagation since restricting the phase budget damages eigenvector conditioning Furthermore the technique remains completely blind to magnitude based targets like normalizing flow likelihoods Thus researchers must restrict this tool specifically to models that process the argument of a spectral product

119. 【2609.02900】DisclosureBeta: A Measurement-Channel Theory for Regime-Conditioned Betas from LLM-Read Risk Disclosures

链接https://arxiv.org/abs/2609.02900

作者:Ping Kuen Wong

类目:Risk Management (q-fin.RM); Computation and Language (cs.CL); General Finance (q-fin.GN)

关键词:error budget, recent listing, strong empirical IPO, empirical IPO accuracy, text-based competitor Breitung

备注: 8 pages, 0 figures. Theory preprint; empirical evaluation forthcoming in a companion paper

点击查看摘要

Abstract:The problem is the beta a desk needs when a firm's price history is too short to trust: an S-1 filer, a recent listing, or a name just past a regime break. The state of the art collapses to a comparable-firm peer beta with no error budget, and the recent text-based competitor Breitung (2025) reports strong empirical IPO accuracy but no identification theory, no error budget, and no lower bound. We fill that gap. We model a large language model as a noisy measurement channel on a firm's latent risk characteristics and write its channel noise into the asset-pricing error budget. In a piecewise-stationary Fama-French five-factor model the loadings are a function of latent risk characteristics and an inferred regime. We prove identification and consistency of the regime-conditional loading function under explicit assumptions on the channel, the detector, and within-regime sampling, and give a matching lower bound showing that the disclosure-noise and detector-misclassification terms are unavoidable for any estimator that observes only returns, factors, LLM features, and a regime estimate. A disclosure-incentive corollary makes estimation precision monotone in a firm-level disclosure-incentive measure (DIM). An adaptive convex combination of the text-based and rolling-window estimators is never worse than either component and shifts its weight toward text exactly when price history is short, stale, or straddles a detected regime break. The empirical evaluation on a frozen, pre-registered panel of price-history-thin firms is forthcoming; this preprint records the theory and the pre-registered design so priority is established independently of the empirical outcome.

信息检索

1. 【2609.04083】CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

链接https://arxiv.org/abs/2609.04083

作者:Tingyu Song,Mingxin Li,Yanzhao Zhang,Dingkun Long,Chu Liu,Pengjun Xie,Yilun Zhao,Shu Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:models remain limited, MLLM-based embedding models, embedding models remain, attribute-object bindings, remain limited

备注

点击查看摘要

Abstract:MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.

2. 【2609.04047】he Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations

链接https://arxiv.org/abs/2609.04047

作者:Dmitrij Żatuchin

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:LLM brand recommendations, audit stochastic variation, standardized protocol exists, Researchers increasingly, Dice Roll Method

备注: 30 pages, 2 figures, 19 tables. Substantially revised; supersedes the Research Square preprint [https://doi.org/10.21203/rs](https://doi.org/10.21203/rs) . [this http URL](http://3.rs) -8883056/v1. Includes a pre-registered external validation on three independent corpora (Motoki et al., Rozado, llm-stability)

点击查看摘要

Abstract:Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.

3. 【2609.03971】he WebKurator.de Platform: Combined Regional and Topical Web Curation

链接https://arxiv.org/abs/2609.03971

作者:Michael Dinzinger,Natanael Arndt,Ben Böck,Jelena Mitrović,Michael Granitzer

类目:Information Retrieval (cs.IR)

关键词:regionally relevant content, relevant content, remains a central, central challenge, challenge for national

备注

点击查看摘要

Abstract:The systematic curation of the Web remains a central challenge for national libraries and memory institutions that aim to preserve culturally and regionally relevant content. Existing directory-based approaches such as Curlie implement a predominantly topic-centric, one-dimensional hierarchy, where geographic aspects are intertwined with topical and linguistic categories. To address this limitation, we present this http URL, a collaborative platform for combined regional and topical web curation, initially focused on the German web. WebKurator introduces a two-dimensional curation model that explicitly separates topical categorization and geographic annotation. The system integrates LLM-based topic classification and imprint-based address extraction with geocoding, and supports user suggestions together with moderated review. The platform is bootstrapped from the German Imprints Dataset, a large-scale collection of 5.54 million websites. Among them, 3.14 million contain imprint pages, for which we successfully extracted and geocoded postal addresses. Of these, 2.58 million (85.17%) are located in Germany and also have an assigned topic label. These websites form the initial foundation of this http URL and can be continuously extended through user suggestions.

4. 【2609.03915】RuleMem: Active Rule Memory for Long-Term Conversational Agents

链接https://arxiv.org/abs/2609.03915

作者:Xingyuan Zeng,Zuohan Wu,Quanming Yao,Yue Wang,Wei Liu,Libin Zheng,Jiuke Wang,Jian Yin

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Question answering agents, dispersed dialogue histories, temporally dispersed dialogue, Question answering, reason over massive

备注

点击查看摘要

Abstract:Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue histories. However, existing memory mechanisms primarily treat past information as \textit{passively} stored facts, leading to semantic gaps and unreliable reasoning. To address this limitation, we propose RuleMem, a rule-based memory framework that induces reusable logical rules from historical interactions to \textit{actively} guide both evidence retrieval and reasoning. Specifically, RuleMem constructs natural-language Horn clauses from conversations and validates them via a Rule Perplexity Consistency (RPC) mechanism. These induced rules enable the retrieval of semantically distant evidence while providing an explicit logical structure for answer generation. We conducted a comprehensive evaluation of RuleMem on two long-term conversational benchmarks, LoCoMo and LongMemEval_s*. In a rigorous comparison against 14 baselines on LoCoMo, RuleMem achieved the highest accuracy, exceeding the baseline average by 27.47 points (a 54.3% relative improvement).

5. 【2609.03901】Comparing Retrieval Methods for Academic Advisor Discovery: A Six-Method Study of 768 CS Faculty Profiles Across 9 US Universities

链接https://arxiv.org/abs/2609.03901

作者:Biraj Subedi

类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:academic advisor discovery, research interest statement, applicant research interest, information retrieval methods, advisor discovery

备注: 16 pages, 2 figures, 9 tables. Code and demo at [this https URL](https://github.com/subedibiraj/academic-discovery)

点击查看摘要

Abstract:We present a comparative evaluation of six information retrieval methods for the task of academic advisor discovery: ranking CS faculty members by relevance to a graduate applicant's research interest statement. The methods span sparse lexical matching (Jaccard overlap, TF-IDF, BM25), dense semantic retrieval (all-MiniLM-L6-v2 sentence embeddings), hybrid score fusion, and learning-to-rank. Evaluation uses a new domain-specific collection: 768 faculty profiles scraped from 9 US CS departments, with 162 graded relevance judgments (grade 0/1/2) across 5 queries representing distinct graduate student research profiles. Across all five queries, Reranked achieves the highest mean NDCG@10 (0.477, std 0.138), followed by Semantic (0.450), Hybrid (0.421), BM25 (0.406), Jaccard (0.303), and TF-IDF (0.246). After Bonferroni correction across all 15 pairwise comparisons, TF-IDF is significantly worse than BM25, Semantic, Hybrid, and Reranked; no other pairwise difference survives correction at 5 queries. A field ablation reveals that biography alone (NDCG 0.634) outperforms the full model combining biography with research area tags (0.593). A controlled experiment shows that concatenating arXiv paper abstracts reduces NDCG@10 by 0.176, motivating a late-fusion architecture. All code, scrapers, and relevance labels are released openly.

6. 【2609.03857】GRASP: Graph-Retrieval Automated Scoring Pipeline for Label-Free Multi-Topic Essay Grading

链接https://arxiv.org/abs/2609.03857

作者:Aafreen Husain,Samar Shailendra,Saad Sajid Hashmi

类目:Information Retrieval (cs.IR)

关键词:Automated Scoring Pipeline, exams consisting solely, Graph-Retrieval Automated Scoring, research has historically, historically focused

备注: Accepted at ICONIP 2026

点击查看摘要

Abstract:Automated short-answer grading research has historically focused on exams consisting solely of questions pertaining to a single topic. Automatic grading of exams containing questions about more than one topic remains less explored. In this work, a Graph-Retrieval Automated Scoring Pipeline (GRASP) is introduced for grading label-free multi-topic science exams. Label-free exams are short-answer exams in which a student's responses to several distinct topics are merged into a single paragraph, with no markup labels or segmentation indicating which span answers which question. Reference answers for each question are encoded into a FAISS vector index via Sentence-BERT, and a semantic similarity graph is constructed over this set of reference answers. At grading time, sentence count heuristics, with a large language model used to resolve ambiguous cases, are first applied to predict how many distinct topics were answered in the student essay. This process is performed without training data or domain-specific example essays. Candidate reference nodes, each storing one (question, reference answer, concatenation of both) from the reference index, are then retrieved through cosine similarity based Retrieval-Augmented Generation (RAG) and Graph Retrieval-Augmented Generation (GRAG). GRAG operates by taking the top cosine matches as seed nodes and then performing a graph traversal over strong edges to find additional reference nodes that may have been missed by RAG. The Hungarian algorithm is then used to optimally assign one reference node per question segment such that no reference is duplicated. Each segment is then graded against its assigned reference independently using GPT-4.1-mini. This experiment is performed to show the effect of retrieval quality on grading accuracy and the benefit of graph-augmented retrieval versus strict cosine similarity methods at various levels of essay complexity.

7. 【2609.03810】Unified Pitch Graphs for Diagnosing Pitching Strategy

链接https://arxiv.org/abs/2609.03810

作者:Kichang Lee,JeongGil Ko

类目:Information Retrieval (cs.IR)

关键词:common representations collapse, Unified Pitch Graphs, aggregate statistics, Pitching strategy, baseball is expressed

备注: 10 pages, 9 figures, 3 tables

点击查看摘要

Abstract:Pitching strategy in baseball is expressed through both physical execution and the ordered context in which pitches are used, yet common representations collapse pitches into discrete types or aggregate statistics. We present Unified Pitch Graphs (UPG), a hierarchical graph representation for retrospective analysis of sequential spatiotemporal events. UPG preserves each pitch as an exact event with reconstructed three-dimensional trajectory and context, connects consecutive pitches through directed sequence edges, and organizes the same events across semantic and temporal resolutions. A support-adaptive mechanism backs off from fine, long sequences when repeated evidence is insufficient, while retaining exact event lineage. We evaluate UPG on 3.94 million MLB Statcast pitches from 2021 to 2026. Nominally identical pitch sequences exhibit distinct physical executions, and ordered structure becomes increasingly evident in longer context-conditioned paths. Support-adaptive backoff increases held-out path coverage from 18.9\% to 94.9\% while improving execution reconstruction from $R^2=0.495$ to $0.685$. UPG also reliably localizes controlled execution changes that discrete pitch-mix and sequence representations cannot detect. These results demonstrate that UPG provides a traceable, multi-scale representation for identifying recurring strategy patterns without conflating retrospective associations with causal or future-performance claims.

8. 【2609.03674】LLM4AIGQ: LLM-based AI Guidance Query Generation Framework for Multi Interest Mining

链接https://arxiv.org/abs/2609.03674

作者:Xiangchen Pan,Jiayi Xu,Jing Wang,Xing Fang,Lingyun Zhu

类目:Information Retrieval (cs.IR)

关键词:provide search queries, Guidance queries stimulate, queries stimulate user, playing a crucial, e-commerce field

备注

点击查看摘要

Abstract:Guidance queries stimulate user consumption by extracting preferences to provide search queries with guidance value, playing a crucial role in the e-commerce field. Traditional AI-generated queries (AIGQ) generation primarily relies on a two-stage "Query-to-AI-Generated-Query" (Q2AIGQ) association paradigm, first recalling user primary search queries from user profiles, historical behavior sequences, item-side information, and the current query through multi-path retrieval, then generalizing AIGQ via rule-based methods. This approach suffers from semantic drift due to information cascade loss; additionally, primary search query derivation heavily depends on "user-item" co-occurrence relationships, lacking exploration of user multi-interests, resulting in guidance queries with low value and mismatched purchase intent. To address the expressive limitations of traditional co-occurrence-based retrieval, we propose LLM4AIGQ, an LLM-based solution for generating AI guidance queries tailored to users' multi-interests. This approach segments user interests by integrating user profiles and historical interaction sequences, infers specific consumption intents for each sub-interest, and subsequently generates corresponding AIGQ. In terms of model training, we employ a post-training pipeline comprising Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and Direct Preference Optimization (DPO) to enhance the model's capability in generating AIGQ. We also introduce a multi-level reward design to satisfy the requirements of multi-objective optimization and long-chain reasoning in practical applications. Regarding deployment, we adopt a nearline-generation and online-read architecture to meet latency constraints. Extensive experimental analyses demonstrate that our model achieves robust performance in both offline evaluations and online A/B tests.

9. 【2609.03654】Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

链接https://arxiv.org/abs/2609.03654

作者:Arianna Miola,Bruno Spaccavento,Lorenzo Silotto,Marco Bianchetti,Luca Cagliero

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Information Retrieval (cs.IR)

关键词:statements poses significant, poses significant challenges, answering systems due, banks' financial statements, financial statements poses

备注

点击查看摘要

Abstract:The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdictions and institutions. We introduce FinRAG-QA, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023. Unlike prior financial QA benchmarks, which centre on U.S. filings and single-institution analysis, FinRAG-QA targets cross-institutional retrieval over documents averaging 198k words, longer than any existing financial QA resource. On this benchmark we evaluate a multi-stage RAG pipeline and isolate the contribution of each component. Contextual chunk enrichment combined with a retrieval-optimised embedding model raises NDCG@10 from 0.322 to 0.710; conditional on the ground truth being retrieved, a reasoning-optimised generator raises answer accuracy from 44.6% to 79.0% (+34.4 percentage points), at roughly 20x the generation latency. We further show that cross-encoder reranking degrades retrieval when the first-stage ranking is already strong, and that a single top-ranked chunk outperforms larger contexts at generation time. Experiments were run in late 2024-early 2025 with the models available at that time.

10. 【2609.03522】EPIC: Explicit Posterior Item Conditioning for Semantic ID Diffusion Recommendation

链接https://arxiv.org/abs/2609.03522

作者:Tuan-Binh Tran,Thanh Tam Nguyen,Quoc Viet Hung Nguyen,Dung D. Le,Tung Kieu,Thanh Trung Huynh

类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:generative recommendation predicts, generating a short, short tuple, tuple of discrete, generative recommendation

备注: 11 pages, 7 figures, 3 tables

点击查看摘要

Abstract:Semantic ID (SID) generative recommendation predicts the next item by generating a short tuple of discrete tokens. Recent masked-diffusion methods improve this process through bidirectional context and flexible decoding, yet recommendation ultimately requires selecting among complete catalog items. At each denoising step, a partial SID can correspond to multiple feasible items, while existing methods primarily reason through position-wise token predictions. We propose Explicit Posterior Item Conditioning (EPIC), which introduces explicit item-level competition into SID denoising. EPIC constructs a personalized posterior over feasible candidate items using the current generation context and the user's recent interactions, then projects this distribution back to unresolved SID positions to guide subsequent token decisions. The pretrained backbone remains frozen and requires no additional decoder forward pass. Experiments on four Amazon benchmarks show consistent improvements over strong baselines, while diagnostic analyses indicate that the gains primarily arise from personalized transition evidence that preserves promising item hypotheses during denoising.

11. 【2609.03482】From Topical Relevance to Answerability: Entailment Distillation for Conversational Retrieval

链接https://arxiv.org/abs/2609.03482

作者:Shuai Qin,Guojia An,Weikang Guo,Pei Ke,Jiwei Wei,Yang Yang,Jie Zou

类目:Information Retrieval (cs.IR)

关键词:Existing conversational retrievers, retrievers commonly treat, conversational retrievers commonly, commonly treat topical, Existing conversational

备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Existing conversational retrievers commonly treat topical relevance as a proxy for answerability. However, a passage that closely matches the dialogue context is not necessarily the one that supports the correct answer. We identify this mismatch as a systematic answerability gap. To address this issue, we propose CLEAR, a framework that shifts conversational retrieval from topical relevance to answerability. The core of CLEAR is entailment distillation, which transfers answer-passage entailment supervision into a cross-encoder reranker so that the reranker discriminates answer-supporting passages from topical distractors at inference time, without requiring answers. CLEAR is complemented by a passage-centric abductive recall module that brings low-similarity yet answerable passages into the candidate pool by inferring answerable queries from passages with an LLM. Across TopiOCQA, QReCC, and out-of-domain TREC CAsT datasets, CLEAR consistently improves top-ranked precision over strong query-rewriting and dense-retrieval baselines, with the largest gains observed in conversations involving heavier topical noise. Moreover, applying our reranker on top of an LLM-driven query rewriter yields further gains.

12. 【2609.03470】ExplainRoute: A Pre-Deployment Audit Framework for Non-Answer-Giving Programming Tutors

链接https://arxiv.org/abs/2609.03470

作者:Yiming Gai,Yingying Zhang,Xuefei Huang

类目:Information Retrieval (cs.IR)

关键词:immediately providing model, providing model answers, Programming tutors, support learners', immediately providing

备注

点击查看摘要

Abstract:Programming tutors should support learners' own explanations rather than immediately providing model answers. We present ExplainRoute, a pre-deployment audit framework for non-answer-giving programming tutors. Given a code line and a learner explanation, it estimates the explanation state and selects one of two bounded responses: a Feynman-style self-explanation prompt or a Socratic scaffold. The framework exposes its state, strategy, cited code fragment, and leakage risk through a machine-checkable contract. Unlike benchmarks that rank tutors by fluency alone, ExplainRoute audits information boundaries, response polarity, failure closure, and the value of learner-explanation visibility before classroom deployment. We evaluate it offline on the 1,770-pair SelfCode corpus using a code-group split, with 443 pairs reserved in 11 untouched holdout groups. The evaluation compares direct answers, fixed open self-explanation, fixed Socratic scaffolding, adaptive routing, and an adaptive no-state ablation. Contract validity reaches 100% for all pedagogical conditions. Adaptive routing matches the frozen reference rule on 60.5% of records, with state macro-F1 of 0.238 (Open: 0.229; Socratic: 0.246), showing no reliable adaptive advantage. An independent language-model judge scores adaptive responses 4.516/5, outperforming the no-state ablation (2.819/5) but slightly below fixed open self-explanation (4.598/5) and Socratic scaffolding (4.658/5). A blinded rubric evaluation on a stratified 40-row subset confirms that visible learner explanations improve information value while adaptive routing does not outperform fixed strategies. The contribution is a validated audit protocol and a boundary finding, rather than evidence of improved learning, retention, or causal instructional effectiveness.

13. 【2609.03454】When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA

链接https://arxiv.org/abs/2609.03454

作者:Hyunseo Oh,Chong-Kwon Kim,Yoonhyuk Choi

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:combine emotional distress, large language model, mental-health question answering, Retrieval-augmented generation, single-turn mental-health question

备注: 8 pages, 3 figures. Presented at the KDD 2026 Undergraduate Consortium

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model responses, but its effect is not uniformly beneficial in single-turn mental-health question answering, where user queries often combine emotional distress, treatment concerns, and safety-sensitive needs. We study when retrieval helps or hurts mental-health QA, and whether a lightweight selective retrieval policy can better control this trade-off. We operationalize retrieval need using three draft-conditioned utility dimensions: psychoeducational need, coping need, and response specificity, together with a rule-based safety trigger. Following psychotherapy-grounded RAG systems such as coTherapist, we construct a compact and controllable guideline corpus comprising coping-strategy, psychoeducational, and safety resources. We fine-tune an instruction-tuned generator on MentalChat16K using QLoRA and compare Closed-book, Always Retrieval, and Selective Retrieval settings on CounselBench-Eval and CounselBench-Adv. Experiments show that retrieval is not uniformly beneficial in this domain. Always Retrieval improves specificity but lowers overall quality and introduces additional safety-sensitive failures. Selective Retrieval preserves closed-book behavior for low-need cases while avoiding the additional degradation caused by unconditional retrieval, supporting the view that retrieval activation is a safety-sensitive control decision.

14. 【2609.03450】Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

链接https://arxiv.org/abs/2609.03450

作者:Kazuki Nakayashiki

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Study, archived source record, agent that inherits, inherits six one-line, one-line memories

备注: 46 pages, 7 figures, 35 tables. Twelve registered studies (14,760 attempted episodes) on one instrument lineage; every package was frozen, timestamped and externally deposited before its first confirmatory call. Manuscript, LaTeX source, all episode files, frozen packages, analyzers and the generator of every number are archived at Zenodo: doi: [https://doi.org/10.5281/zenodo.22267221](https://doi.org/10.5281/zenodo.22267221)

点击查看摘要

Abstract:An agent that inherits six one-line memories may pull at most one archived source record before acting; a directive written into the store can steer that choice: a pointer to the record, a criterion that identifies it, or both. Across twelve registered studies on one instrument lineage (14,760 attempts) we measured where the request goes under each form. On six direct-provider models a length-matched criterion exceeded a bare id by +35.0 points [+31.2, +38.8] (Study D); the contrast failed its registered superiority rule on a nine-model OpenRouter-served panel (Study E). Appending the id cancelled the criterion on three Claude models (Opus 5: 40/40 to 0/40; Study F-x); six byte-matched edits gave each exact string its own effect (Study G), and a re-run at eighty runs per cell left fifteen of thirty replication contrasts within the margin, fifteen unresolved and none beyond (Study G'). A ratification line (+96.0 points on Opus 5) and a budget of two credits restored the target on all three (Study J); across five criterion strings the suffix's cancellation held for four of the five wordings on Opus 5 and all five wordings on Fable 5.1 (Study H2); in a second store every model followed the criterion (Study H1). Continued into a decision, the criterion moved the choice toward the current record (+100.0 points, Opus 5) and away from it on Fable 5.1 (Study I). A one-character plan pointer's effect (+78.0 points; Study B, after a correction of its first repository report) returned the same verdict under a prospectively registered re-run (+81.7 points; Study B'). All results are descriptive effects of exact edits on fixed panels with registered intervals and no mechanism claim.

15. 【2609.03376】Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings

链接https://arxiv.org/abs/2609.03376

作者:Peichun Hua,Yunming Xiao

类目:Cryptography and Security (cs.CR); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:standard building block, Retrieval-Augmented Generation, made dense retrieval, large document collections, building block

备注: 22 pages, 10 tables, 6 figures

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) has made dense retrieval over large document collections a standard building block. Organizations increasingly outsource vector indexes to untrusted clouds, exposing proprietary corpora and user queries. Cryptographic protection is challenging because each query searches corpus-scale state, causing computation, correlated randomness, and communication to grow with the corpus. At million-document scale, a naive secure implementation takes minutes and about 90 GB of communication per query. Even recent optimized systems require 10--22 seconds. We propose Spruce (Scalable Private Outsourced Retrieval Using Compact Embeddings), which co-designs representations with the cryptographic protocol. Spruce learns compact binary codes that preserve candidates for full-precision reranking, replacing corpus-wide embedding scoring with efficient Hamming-distance computation under two-server multi-party computation (MPC). A corpus-calibrated fixed-radius protocol avoids multi-round candidate selection while preserving retrieval quality. Spruce also provides private cluster pruning, which trades minor quality loss for substantially less computation, and a one-core owner-operated dealer that removes cloud OT preprocessing bottlenecks. Across four corpora containing 383K--5.42M documents, Spruce preserves the original search quality with median candidate sets of only 382--1,952. At 10 Gbps inter-server bandwidth, full scans take 0.21--2.97 seconds, $4.8$--$6.7\times$ faster than the closest measured prior work. Private pruning takes 0.06--1.09 seconds, achieves $13.1$--$22.9\times$ speedups, and retains $93.9\%$--$97.3\%$ of full-float NDCG. On the largest corpus, pruning and the dealer jointly improve sustained throughput by $31.5\times$ at 1 Gbps per link.

Comments:
22 pages, 10 tables, 6 figures

Subjects:

Cryptography and Security (cs.CR); Information Retrieval (cs.IR); Machine Learning (cs.LG)

Cite as:
arXiv:2609.03376 [cs.CR]

(or
arXiv:2609.03376v1 [cs.CR] for this version)

https://doi.org/10.48550/arXiv.2609.03376

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
16. 【2609.03369】HypRQ-VAE: Hyperbolic Item Indexing for Long-Tail-Aware Generative Recommender Systems

链接https://arxiv.org/abs/2609.03369

作者:Longfeng Wu,Tong Zeng,Giovanni Seni,Zhimin Peng,Bhanu Pratap Singh Rawat,Si Zhang,Yao Zhou,Lecheng Zheng,Bo Ji,Yujun Yan,Dawei Zhou

类目:Information Retrieval (cs.IR)

关键词:Sequential recommender systems, language modeling task, Sequential recommender, large language models, model user behavior

备注: Accepted for publication in the 2026 IEEE International Conference on Data Mining (ICDM 2026)

点击查看摘要

Abstract:Sequential recommender systems model user behavior as item ID sequences, while recent generative methods cast recommendation as a language modeling task using large language models (LLMs). While this paradigm incorporates rich textual semantics, it introduces a fundamental mismatch: LLMs operate on text tokens, whereas recommender systems depend on discrete item indices. This misalignment often leads to hallucinations in generative recommendations. Existing methods attempt to bridge this gap by learning item vocabularies in Euclidean space, but they struggle to model the inherent long-tail distribution of real-world catalogs, where a small number of head items dominate, and a vast number of tail items reflect users' niche preferences. To address this issue, we introduce Hyperbolic Residual-Quantized Variational AutoEncoder (HypRQ-VAE), the first framework to learn item indexing in hyperbolic space. HypRQ-VAE leverages the unique properties of hyperbolic geometry, whose exponential volume expansion naturally accommodates the power law structure of user-item interactions. This allows the model to encode rich textual semantics while preserving the representational fidelity of sparse, long-tail items. Experiments on three benchmark datasets show that HypRQ-VAE significantly improves the performance of recommendation, particularly in recommending tail items. Our analysis attributes these gains to the superior capacity of hyperbolic space to model item hierarchies and sparsity in generative recommendation. Our code and data are available at: this https URL.

17. 【2609.03338】SciLENS: RL-Driven Autonomous Agents for Scientific Localized Evidence Navigation and Synthesis

链接https://arxiv.org/abs/2609.03338

作者:Leqi Zheng,Jinbo Su,Yuying Li,Chaokun Wang,Weiping Wang,Haitao Li,Jiajun Zhang,Shannan Yan,Zhaolu Kang,Rong Fu,Jie Wu,Fang Niu,Hang Zhang

类目:Information Retrieval (cs.IR)

关键词:proprietary online services, agents increasingly rely, Scientific literature synthesis, Localized Evidence Navigation, Scientific Localized Evidence

备注

点击查看摘要

Abstract:Scientific literature synthesis agents increasingly rely on proprietary online services, limiting reproducibility, privacy, and offline deployment. To address this challenge, we introduce SciLENS Scientific Localized Evidence Navigation and Synthesis), a fully local autonomous agent framework operating on a dual-tier infrastructure indexing approximately 12 million academic records. SciLENS pioneers the integration of structural visualization as an actionable tool within the reasoning loop, enabling the agent to compress complex citation topologies into validated data-driven charts and thereby mitigate context exhaustion during macro-level synthesis. To train the agent without human annotation, we develop an automated data synthesis pipeline that extracts multi-hop subgraphs from a citation knowledge graph, verified by cross-model consensus among 20 frontier models. The agent is subsequently aligned through a reverse-decomposition rubric strategy that provides fine-grained process rewards for early planning and strict evidence grounding. Evaluations across six scientific benchmarks encompassing standard QA, citation accuracy, factual reasoning, and structural synthesis demonstrate that SciLENS significantly outperforms open-source baselines and achieves performance comparable to GPT-5.2 and Gemini-3.0-pro. Our source code and data are released at this https URL.

18. 【2609.03313】SelfDR: Self-Distillation from Reasoning for LLM-Based Recommendation

链接https://arxiv.org/abs/2609.03313

作者:Chumeng Jiang,Jiayin Wang,Xinjie Lin,Zhiqiang Guo,Hengliang Luo,Min Zhang

类目:Information Retrieval (cs.IR)

关键词:Large Language Models, Large Language, Language Models, recently emerged, emerged as powerful

备注: 12 pages, 5 figures, CIKM'26

点击查看摘要

Abstract:Large Language Models (LLMs) have recently emerged as powerful backbones for recommendation. To better elicit their capabilities, reasoning has been widely incorporated to help LLMs interpret rich textual signals and improve recommendation accuracy. However, explicitly generating intermediate reasoning traces often incurs substantial computational costs, which limits practical deployment in real-world recommender systems. To address this challenge, we propose SelfDR, a Self-Distillation from Reasoning framework for LLM-based Recommendation. SelfDR distills an LLM's own reasoning-enhanced predictions to produce recommendations directly, improving recommendation effectiveness while maintaining inference efficiency. All components in the framework are built on the same base LLM, without relying on any external models. Specifically, the teacher recommender is constructed by training a reasoner with downstream performance as the reward, enabling it to generate targeted rationales that are later incorporated into the teacher's input. A student recommender for direct recommendation, with the same underlying model, then learns from the teacher through self-distillation with a dynamic weighting strategy. Extensive experiments on three public datasets validate the effectiveness, rationality, and efficiency of SelfDR. Codes are available at this https URL.

19. 【2609.03311】DoPR: Reusable Compressed Document Prefixes for Efficient LLM Reranking

链接https://arxiv.org/abs/2609.03311

作者:Beiya Dai,Yifan Wei,Guang Yang,Xing Shi,Xinbing Wang,Zhouhan Lin

类目:Information Retrieval (cs.IR)

关键词:causing substantial redundant, Large language models, Large language, pointwise reranking repeatedly, reranking repeatedly processes

备注: 15 pages, 6 figures

点击查看摘要

Abstract:Large language models (LLMs) are effective rerankers, but pointwise reranking repeatedly processes the same document across different queries, causing substantial redundant document-side computation. We propose \textbf{DoPR}, a compressed document prefix framework that decouples offline document processing from online reranking. DoPR first selects query-independent document representations and converts them into compressed document prefix states, which are precomputed offline and reused whenever the document is retrieved. During online reranking, the model scores each query-document pair by processing only the query and scoring token, with document information supplied by the stored prefix states. This design reduces online cost through both document-side compression and cross-query prefix-state reuse. Experiments on TREC DL, BEIR, and BRIGHT with Qwen3 models from $0.6$B to $8$B show that DoPR achieves up to 8.0$\times$ online document-side memory reduction and up to 8.04$\times$ latency speedup, while retaining \textbf{97.1\%-99.5\%} of the average NDCG@10 of matched full-document rerankers.

20. 【2609.03290】UniCon: A Unified Context-Centric Modeling Paradigm for CTR Prediction

链接https://arxiv.org/abs/2609.03290

作者:Jiajun Cui,Zhengqi Xu,Fan Zhang,Zhangteng,Gu Tang,Honghong Zhu,Mengxi Wu,Yulin Liang,Xingxing Wang

类目:Information Retrieval (cs.IR)

关键词:industrial click-through rate, click-through rate, major direction, direction for industrial, industrial click-through

备注: 10 pages, 5 figures, 2 tables

点击查看摘要

Abstract:Unified modeling has become a major direction for industrial click-through rate (CTR) prediction. Existing approaches typically unify sequential and non-sequential signals at the token level, model their interactions in a shared backbone, and increase model capacity to improve scaling behavior. However, this division originates from legacy feature-engineering practice and is misaligned with the underlying decision process. User behavior is inherently a sequence of homogeneous context units; at the level of input organization, historical behavior and the current request differ only in whether their outcomes are observed or remain to be predicted. Treating them as heterogeneous signals obscures structural dependencies within the user's decision context, limiting both scaling efficiency and prediction quality. This limitation is particularly pronounced in context-rich scenarios such as e-commerce shelves and waterfall feeds. To address this, we propose UniCon, a unified context-centric modeling architecture that treats the request context as the basic modeling unit and organizes history and prediction targets as homogeneous context units. Intra-context attention captures local coupling among items within a context (Locality), while inter-context attention models the dynamic evolution of decision states across contexts (Dynamics). This organization bridges the structural gap between history and target and supports more effective scaling of unified CTR models. Context-unit-level sequence compression further reduces deployment overhead. On Meituan search advertising, UniCon improves offline AUC by 0.0139 over a strong production baseline and achieves statistically significant online lifts of 3.09% in RPM, 2.07% in CTR, and 2.95% in revenue.

21. 【2609.03047】SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

链接https://arxiv.org/abs/2609.03047

作者:Michael J. Bommarito II

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:archives manage large, manage large collections, Libraries and archives, Evaluating LLM Fitness, archives manage

备注: 17 pages, 15 tables. Code available at [this https URL](https://github.com/mjbommar/shelf-benchmark) ; data available at [this https URL](https://huggingface.co/datasets/mjbommar/SHELF)

点击查看摘要

Abstract:Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Synthetic Harness for Evaluating LLM Fitness, addresses this gap. It is a Python system that turns labelled taxonomies, writing specifications, and a generation budget into controlled benchmark data and evaluation tasks. This first release contains 62,899 model-written documents based on Library of Congress vocabularies, with tasks for classification, clustering, retrieval, pair classification, and instruction retrieval. We compare TF, TF-IDF, BM25, popular encoders, and, on subject classification only, zero-shot decoders; each method appears only on tasks that support it. Subject classification reaches 0.8887, while genre-form classification reaches only 0.2605, and several pair and clustering tasks remain near chance. Sparse methods remain competitive on classification, while TF-IDF is the fastest measured arm in the subject timing experiment. SHELF also varies bibliographic facets independently and can generate new, verifiably unseen documents after a model's training cutoff. Comparisons with LCSHBench and Project Gutenberg show that model rankings transfer more reliably than absolute scores, but SHELF scores do not estimate accuracy on production catalogue data. We release all source code and data under permissive licenses on GitHub and Hugging Face.

22. 【2609.02944】Reflect-SQL: A Self-Reflection Based Framework for Text-to-SQL

链接https://arxiv.org/abs/2609.02944

作者:Anupreksha Jain,Manish Shrivastava

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Democratizing data access, real-world complexities, Democratizing data, crucial goal, goal for modern

备注: Accepted in PAKDD 2026

点击查看摘要

Abstract:Democratizing data access through natural language is a crucial goal for modern enterprises, but the practical adoption of Text-to-SQL is critically hindered by real-world complexities: 1. Obscure and large database schemas, 2. Ineffective retrieval of relevant tables and columns due to structured setting of schemas and vague user query, 3. Generation of syntactically or logically flawed SQL due to a lack of robust validation and correction mechanism. To address these systemic challenges, we introduce Reflect-SQL, a novel framework for Text to SQL, grounded in multi-stage self-reflection approach to develop understanding of obscure schema using a knowledge base, setup a process for effective retrieval and system to generate syntactically/semantically SQL. Instead of a single-pass attempt, our system employs an LLM-as-a-judge driven scoring mechanism within interconnected feedback loops to iteratively refine the results at every stage. A feedback-driven retrieval loop refines the user's natural language query, while a synthesis loop validates and corrects the SQL and finally, an entailment loop optimizes the end-to-end process and continuously enriches the knowledge base. By integrating these layers of reflection, Reflect-SQL bridges the critical gap between user intent and complex data. On the challenging BIRD benchmark, our framework achieves an execution accuracy of 72.03%, significantly outperforming state-of-the-art baselines, demonstrating a major leap in reliability for enterprise applications.

23. 【2609.02913】CHSR-RRF: A curriculum-gated hybrid retrieval framework with reciprocal rank fusion and leakage-aware benchmarking for educational RAG

链接https://arxiv.org/abs/2609.02913

作者:Terence Ateya,Zavier Ndum Ndum,Jicheng Fu,Kelly Tendongkeng

类目:Information Retrieval (cs.IR)

关键词:standard retrievers optimize, retrievers optimize topical, Retrieval-augmented generation, educational question answering, optimize topical relevance

备注: 45 pages, 14 figures

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) is increasingly used in educational question answering, but standard retrievers optimize topical relevance without enforcing curriculum validity. In school settings, a passage can be relevant yet inappropriate if it comes from the wrong subject, level, or examination context; we call this failure mode curriculum leakage. We present CHSR-RRF, a curriculum-gated hybrid retrieval framework that applies metadata constraints before retrieval, then combines sparse and dense search with reciprocal rank fusion and deterministic reranking. We also introduce CERB, a 126-case benchmark for curriculum-constrained retrieval with hierarchy-aware relevance labels and explicit leakage annotations. On a 61-case pilot, pre-retrieval gating reduces leakage by 4.6x ($p0.001$) while preserving ranked recall, whereas applying the same constraints after retrieval collapses recall and exact-scope success to zero ($p=0.039$). A full-benchmark lower-bound analysis further shows that many remaining failures arise from corpus and metadata gaps rather than retrieval design alone. These results show that retrieval in structured educational domains should be treated as constrained selection, with validity enforced when the candidate pool is formed rather than after ranking.

24. 【2609.02894】R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG

链接https://arxiv.org/abs/2609.02894

作者:Yucan Guo,Miao Su,Saiping Guan,Long Bai,Zhongni Hou,Zixuan Li,Xiaolong Jin,Jiafeng Guo,Xueqi Cheng

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Large Language Models, enhancing Large Language, Language Models, Large Language, Retrieval-Augmented Generation

备注

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently handles simple queries but struggles with relational or multi-hop reasoning. Graph-based RAG alleviates this issue but incurs higher inference complexity and latency. In practice, user queries can differ significantly in their complexity, rendering a fixed RAG strategy suboptimal. However, existing hybrid text-graph RAG methods typically rely on heuristic and LLM-based routing, resulting in unnecessary overhead and strong dependence on the underlying LLM. To address these challenges, we propose R$^{2}$Adapter, a lightweight plug-in Routing and Rewriting Adapter designed to allocate queries between vanilla and graph-based RAG dynamically. By routing only the queries that genuinely benefit from graph-based reasoning, R$^{2}$Adapter reduces unnecessary graph retrieval overhead. Additionally, uncertain graph-routed queries are rewritten to better expose their multi-hop reasoning requirements, improving retrieval quality without additional supervision. Extensive experiments on three multi-hop QA benchmarks demonstrate that R$^{2}$Adapter reduces graph-based RAG usage by up to 59% while maintaining comparable answer accuracy. This adapter is model-agnostic and can be seamlessly integrated into diverse vanilla and graph-based RAG pipelines, providing an efficient and adaptive solution for hybrid RAG systems.

25. 【2609.01865】ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval

链接https://arxiv.org/abs/2609.01865

作者:Aaryan Kapoor,Md Abdullah Al Hafiz Khan

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Embedding-based code retrieval, retrieving lexically similar, retrieval-augmented code generation, lexically similar code, retrieving correct code

备注: Accepted to EMNLP 2026 (Main Conference). Camera-ready version. 17 pages, 6 figures, 8 tables

点击查看摘要

Abstract:Embedding-based code retrieval is a core component of coding agents and retrieval-augmented code generation, where retrieving correct code matters more than retrieving lexically similar code. Existing code-retrieval benchmarks do not plant controlled, execution-verified single-edit variants of each query's canonical implementation in the search pool, leaving the question of whether embeddings can functionally discriminate correct from near-clone-but-incorrect code unanswered in a retrieval setting. Resolving this requires a benchmark whose search pool itself contains the relevant counterfactuals -- execution-verified buggy variants near-identical to each canonical -- so that a retriever's rank ordering can be directly tested for functional discrimination rather than topical or identity overlap. We introduce ExecRetrieval, 939 Python tasks each paired with one execution-verified canonical implementation and up to four execution-verified buggy distractors, each generated by a mechanical mutation making a single targeted edit, and evaluate 23 dense embedding configurations plus BM25 under provider-native invocation with paired McNemar tests and query-level bootstrap intervals. With near-clone counterfactuals in the pool, the top hosted system reaches exec@10 = 1.00 but only exec@1 = 0.331; rank-1 misses are paired buggy variants 91.5-99.4% of the time across the four leading systems, and the canonical scores below at least one of its four paired distractors in 67-78% of queries on the leading systems. The full dataset, execution oracle, embedding matrices, environment snapshot, and pairwise statistical tests are released at the URL in Appendix D.

计算机视觉

1. 【2609.04203】mporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

链接https://arxiv.org/abs/2609.04203

作者:Shravan Venkatraman,Wenshuai Zhao,Mohammad Hassan Vali,Arno Solin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:fully self-contained framework, Self-Distillation over Time, continuous video state, Self-Supervised Self-Distillation, fully self-contained

备注

点击查看摘要

Abstract:We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S$^3$T improves VSTAT accuracy by $+1.74$ as a single model, $+2.38$ with souping, and $+2.70$ with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by $+7.95$ on VSTAT-YouTube state-tracking questions and $+4.50$ on MVBench Action Count.

2. 【2609.04202】okenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation

链接https://arxiv.org/abs/2609.04202

作者:Adeela Islam,Zorah Lähner,Vittorio Murino,Vladislav Golyanik

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:strong non-isometric deformations, non-isometric deformations remains, deformations remains challenging, substantial progress, recently seen substantial

备注: 25 pages, 13 figures and 12 tables; project page: [this https URL](https://4dqv.mpi-inf.mpg.de/TokenMatch/)

点击查看摘要

Abstract:While data-driven 3D shape correspondence estimation has recently seen substantial progress, robust matching under partial observations and strong non-isometric deformations remains challenging. Existing learning-based approaches often rely on hand-crafted descriptors or template-based representations, whereas recent generative models over functional maps suffer from high inference cost, limited interpretability, and poor generalisation to partial shapes. In response to these limitations, this paper introduces TokenMatch, a new transformer-based unified model for estimating 3D shape correspondences. Our feed-forward approach trained exclusively on BeCoS, a challenging non-isometric partial-to-partial shape-matching dataset, can generalise to matching full shapes without retraining or fine-tuning. TokenMatch uses self- and cross-attention mechanisms to efficiently learn patch-level and point-level relations as well as dense correspondences between shape pairs. Our core insight is that meshes can be adaptively tokenised into patches using shape curvature guidance, enabling effective learning of shape-specific geometric descriptors for correspondence estimation. We evaluate TokenMatch on standard benchmarks for partial and full shape matching, including CP2P, PSMAL, BeCoS, FAUST, SCAPE, and SHREC'19. Our method achieves consistently high performance, in most cases outperforming existing methods for partial and full shape matching in the mean geodesic error and intersection-over-union metrics, while also running faster at sub-second inference speeds.

3. 【2609.04201】Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

链接https://arxiv.org/abs/2609.04201

作者:Chin-Yang Lin,Yang-Che Sun,Cheng Sun,Fu-En Yang,Min-Hung Chen,Yen-Yu Lin,Wei-Chen Chiu,Yu-Lun Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models perform poorly, reconstruction models perform, long videos, models perform, perform poorly

备注: ECCV 2026. Project page: [this https URL](https://linjohnss.github.io/scal3r/)

点击查看摘要

Abstract:Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into significant geometric collapse. However, we observe that per-frame depth remains stable throughout this failure. The backbone's local geometry remains intact; only the global pose head breaks down. Motivated by this decoupling, we introduce Scal3R. This approach reformulates online reconstruction as multi-reference relative pose querying. We use lightweight learnable tokens, which make up about ~1% of the parameters, and inject them into a completely frozen backbone via asymmetric attention. This setup queries poses relative to multiple past keyframes. An online pose-graph optimization system with loop closure suppresses long-range drift. Scal3R reaches convergence in 8 hours on a single GPU. It reduces the average ATE by over 60% on KITTI compared to the online baseline. It also achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. Project page: this https URL

4. 【2609.04200】Principia: Relational Physics Tests for Video Models

链接https://arxiv.org/abs/2609.04200

作者:Varun Varma Thozhiyoor,Shivam Tripathi,Venkatesh Babu Radhakrishnan,Anand Bhattad

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Evaluating physical reasoning, absolute motion measurements, motion measurements depend, Evaluating physical, frame rate

备注: Project Page: [this https URL](https://principiabench.github.io/)

点击查看摘要

Abstract:Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.

5. 【2609.04196】Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

链接https://arxiv.org/abs/2609.04196

作者:Kang Liao,Yihang Luo,Xiao-Ming Wu,Linyi Jin,Size Wu,Chunyu Lin,Yao Zhao,Fei Wang,Wei Li,Chen Change Loy

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:external offline modules, unified multimodal architecture, integrates physical understanding, spatial simulation, offline modules

备注: Project Page: [this https URL](https://kangliao929.github.io/projects/puffin-world/)

点击查看摘要

Abstract:We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

6. 【2609.04190】One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

链接https://arxiv.org/abs/2609.04190

作者:Adheesh Sunil Juvekar,Onkar Kishor Susladkar,Kiet A. Nguyen,Muntasir Wahed,Nabeel Bashir,Xiaona Zhou,Tianjiao Yu,Vedant Shah,Ismini Lourentzou

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Video editing spans, framework remains challenging, achieving high-quality instruction-guided, diverse editing paradigms, single unified framework

备注: [this https URL](https://plan-lab.github.io/editvid)

点击查看摘要

Abstract:Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.

7. 【2609.04183】Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

链接https://arxiv.org/abs/2609.04183

作者:Ye-Chan Kim,Seunghee Choi,SeungJu Cha,Si-Woo Kim,Hwiseon Kim,Hyungee Kim,Dong-Jin Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Weakly-Supervised Dense Video, Dense Video Captioning, describe multiple events, Video Captioning aims, Dense Video

备注: Accepted to EMNLP 2026 (main, long)

点击查看摘要

Abstract:Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.

8. 【2609.04174】Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

链接https://arxiv.org/abs/2609.04174

作者:Denis M. Akola,David F. Fouhey

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:predicting rich unified, VGGT have recently, Foundation Models, rich unified representations, feed-foward transformers

备注: Accepted to the European Conference on Computer Vision (ECCV) 2026. Project page: [this https URL](https://akola-mbey-denis.github.io/Z3D-page/)

点击查看摘要

Abstract:3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predict realistic depth maps for new views across multiple datasets.

9. 【2609.04151】Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

链接https://arxiv.org/abs/2609.04151

作者:Mengwei Ren,Xuaner Zhang,Zhihao Xia

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:support precise edits, produce high-quality images, follow complex instructions, identity, follow complex

备注

点击查看摘要

Abstract:Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surrounding scene changes. Existing subject-driven methods make fundamentally different choices about where identity is represented: through the input context (GPT-Image-2, NB2), as trainable subject-specific model parameters (LoRA), or as a persistent identity layer (PHOTA IDENTITY) reusable across generations and edits. We systematically benchmark these paradigms across subject-driven generation, editing, restoration, and multi-subject settings, with tasks designed to increasingly stress identity preservation. Our results show that identity preservation remains a distinct limitation of current generative foundation models: strong image quality and instruction following do not necessarily imply strong identity fidelity, and identity degradation becomes more pronounced under iterative edits, small subject scales, severe image degradation, and multi-subject composition. Persistent identity substantially reduces this degradation across generation, editing, and restoration, consistently improving identity preservation when applied to different foundation models while maintaining comparable instruction adherence and perceptual image quality. These results suggest that identity does not simply emerge from increasingly capable generative models, but can instead be represented as persistent subject knowledge that is composed independently with the underlying generative model.

10. 【2609.04131】Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

链接https://arxiv.org/abs/2609.04131

作者:Hongyu Qu,Guangming Yao,Ling Xing,Xiaobin Hu,Rongxing Ding,Guibin Zhang,Fan Zhang,Yi Yuan,Xiangbo Shu,Shuicheng Yan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, understanding requires multimodal, requires multimodal large, multimodal large language, process continuous visual

备注

点击查看摘要

Abstract:Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as additional visual context. Though effective, this store-and-retrieve paradigm keeps historical evidence as external visual context, preventing it from being internalized into a compact, evolving latent memory that can continuously guide streaming reasoning. To bridge this gap, we introduce LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize. Specifically, LatentStream comprises three coordinated components. First, Query-agnostic Hierarchical Streaming Memory organizes visual history into short-, mid-, and long-term levels under a fixed memory budget through Jenks-guided adaptive consolidation. Once a query arrives, Hierarchical Latent Memory Evolution equips groups of latent memory tokens with progressively expanding memory receptive fields, enabling them to iteratively retrieve historical evidence from their corresponding scopes and internalize it into a compact, fixed-length latent memory. Finally, Progressive Confidence-guided Latent Memory Optimization constructs a hierarchical progression reward from group-wise predictive entropy and jointly refines the latent memory tokens and retrieved evidence, encouraging increasingly confident streaming reasoning. Extensive experiments demonstrate that LatentStream achieves new state-of-the-art results on existing online and offline video benchmarks.

11. 【2609.04120】BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

链接https://arxiv.org/abs/2609.04120

作者:Wei Zhang,Xin Li,Peishu Shi,Jialin Gao,Xuekang Peng,Zhichao Lian,Yeying Jin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generate realistic videos, Video virtual try-on, pseudo data, aims to generate, try-on

备注: 23 pages, 18 figures, 8 tables. Accepted to ACM Multimedia 2026 (MM '26)

点击查看摘要

Abstract:Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to localize try-on regions, making them vulnerable to large motions and severe occlusions. Although mask-free image-based try-on methods have shown promising results by leveraging large-scale pseudo data, extending this paradigm to videos remains difficult, as constructing video-level pseudo data is prohibitively expensive. Furthermore, coarse keyframe sampling and the scarcity of multi-view try-on data limit existing keyframe-driven methods in maintaining garment consistency and handling diverse try-on tasks. To address these challenges, we propose BooM-VVT, a mask-free VVT framework built upon the keyframe-driven paradigm. To achieve mask-free VVT, we introduce a multi-stage training strategy that leverages image-level pseudo data for mask-free localization learning, substantially reducing the need for costly video-level pseudo data. To improve garment consistency, we propose Garment-Sensitive Keyframe Sampling, which selects keyframes based on garment-relevant body regions to better capture garment appearance. We further introduce Frame-Shared 3D-RoPE to establish spatiotemporal correspondences between keyframes and target video frames for accurate garment-detail transfer. Finally, we construct OmniView, a large-scale multi-view try-on dataset to support reliable try-on video generation under complex camera viewpoints and diverse try-on tasks. Extensive experiments demonstrate that BooM-VVT achieves superior temporal consistency and garment fidelity over existing methods. Project page: this https URL.

Comments:
23 pages, 18 figures, 8 tables. Accepted to ACM Multimedia 2026 (MM '26)

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.04120 [cs.CV]

(or
arXiv:2609.04120v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.04120

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Journalreference:
Proceedings of the 34th ACM International Conference on Multimedia (MM '26), 2026

Related DOI:

https://doi.org/10.1145/3767308.3835712

Focus to learn more

            DOI(s) linking to related resources</p>
12. 【2609.04110】he Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

链接https://arxiv.org/abs/2609.04110

作者:Yumeng Shi,Quanyu Long,Yin Wu,Wenya Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:representing time, frames in order, temporal, Abstract, video

备注: Accepted to EMNLP 2026 (main)

点击查看摘要

Abstract:Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This mismatch allows models to learn temporal answers from shortcuts such as objects, scenes, and language priors, without requiring internal video representations to capture event progression. To address this, we propose VT-Contrast, a representation-level temporal counterfactual objective for VideoLMs. Its design asks where temporal supervision should act and what temporal differences it should expose. VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance. It requires no architectural changes, is compatible with diverse VideoLM training tasks, and improves overall performance across temporal understanding benchmarks. Our code is available at this https URL.

13. 【2609.04096】Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

链接https://arxiv.org/abs/2609.04096

作者:Sixu Yan,Shikang Wang,Binhua Huang,Xuanlai Tang,Guohua Fan,Fan Huang,Haoxuan Li,Yongkang Li,Yuhan Li,Bencheng Liao,Zeyu Zhang,Wenyu Liu,Hangxin Liu,Xinggang Wang

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:paper proposes AdaRoboVLG, supports generalizable grasp, grasp synthesis, Unlike existing VLG, grasp

备注

点击查看摘要

Abstract:This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at this https URL

14. 【2609.04088】Efficient Semantic Understanding from Digital Foveation

链接https://arxiv.org/abs/2609.04088

作者:Caterina Caccavella,Vittorio Fra,Andreas Ziegler,Giulia D'Angelo,Yulia Sandamirskaya

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:allocates computational resources, computational resources uniformly, entire image, task relevance, resources uniformly

备注: Accepted at the 3rd Human-inspired Computer Vision Workshop at ECCV 2026

点击查看摘要

Abstract:Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We introduce a lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation. Beyond conventional dense prediction metrics, we use object-level evaluation to measure semantic understanding under sparse observations. On ADE20K-Object, a single foveated observation achieves 95.9% of the baseline Top-1 accuracy and 96.9% of the baseline Top-3 accuracy while requiring only 4.7% of the computational cost. At the scene level, semantic accumulation recovers 90.6% of the baseline object recall while using 58.6% of the computation. These results suggest that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively, highlighting active vision as an efficient alternative to uniform dense processing and motivating evaluation protocols beyond conventional pixel-wise segmentation metrics.

15. 【2609.04083】CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

链接https://arxiv.org/abs/2609.04083

作者:Tingyu Song,Mingxin Li,Yanzhao Zhang,Dingkun Long,Chu Liu,Pengjun Xie,Yilun Zhao,Shu Wu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:models remain limited, MLLM-based embedding models, embedding models remain, attribute-object bindings, remain limited

备注

点击查看摘要

Abstract:MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.

16. 【2609.04071】AP-Path: Task-Adaptive Structural and Token Pruning for Efficient and Trustworthy Pathology Foundation Models

链接https://arxiv.org/abs/2609.04071

作者:Mehedi Hasan,Ashfak Yeafi,Md Khairul Islam

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:high inference cost, transferable representation learning, improve transferable representation, inference cost, transferable representation

备注

点击查看摘要

Abstract:Pathology foundation models improve transferable representation learning for histopathology, but recent gains often rely on encoders with hundreds of millions of parameters and high inference cost. We propose TAP-Path, a task-adaptive compression framework that directly restructures a pretrained Virchow2 encoder rather than distilling it into a separate student. TAP-Path combines validation-driven transformer-block selection, physical removal of redundant blocks, input-adaptive patch-token pruning, multi-depth feature recovery, and a lightweight gated task head. The final model retains 24 of 32 transformer blocks and 70% of patch tokens after pruning, reducing encoder parameters by 24.96% (631.24M to 473.70M) and analytical encoder compute by 35.20% (340.13G to 220.40G FLOPs). Across three task-head optimization seeds, TAP-Path achieved $87.98 \pm 0.067%$ test accuracy, $81.26 \pm 0.49%$ balanced accuracy, and $82.38 \pm 0.48%$ macro-F1 on a 32-class histopathology benchmark, compared with 86.89% for full Virchow2 and 87.67% for UNI2-h. TAP-Path achieved a Brier score of $0.1800 \pm 0.0005$ and failure-detection AUROC of $0.9047 \pm 0.0060$. A validation-only rare-aware objective improved rare-class balanced accuracy in a secondary operating analysis. Frozen external evaluation on 433 CPTAC samples yielded $91.22 \pm 0.83%$ accuracy and $91.10 \pm 0.81%$ balanced accuracy. These results show that task-adaptive structural and token sparsification can improve the accuracy-efficiency trade-off of large pathology foundation models while preserving reliability under internal and external evaluation.

17. 【2609.04070】Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving

链接https://arxiv.org/abs/2609.04070

作者:Ruoyu Yao,Yusen Xie,Qingzhao Liu,Pei Liu,Zewei Yang,Yipeng Zhu,Xiaolong Wang,Jun Ma

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Vision-Language Models, autonomous driving remains, physics-constrained nature, significant challenge, reasoning of Vision-Language

备注: 8 pages, 5 figures

点击查看摘要

Abstract:Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrained nature of autonomous driving remains a significant challenge. In this work, we introduce LaPla, a unified Vision-Language-Action (VLA) framework featuring latent-aligned planning to seamlessly ground semantic understanding in precise motion execution. We first design an action tokenizer based on a residual vector-quantized variational autoencoder (VQ-VAE), capturing vehicle kinematics and encoding trajectory features into a structured latent space. Rather than discrete codebook lookups that inevitably introduce quantization errors, LaPla repurposes this representation as a physical prior to bridge the modality gap between high-dimensional semantics and the raw action space. Specifically, given multimodal inputs integrating multi-view images, historical actions, and textual instructions, LaPla incorporates concurrent action queries to causally attend to the multimodal context in a single forward pass, projecting hidden states directly into the pretrained VQ-VAE latent space. The frozen decoder then translates these continuous latents into actions, effectively eliminating quantization errors and ensuring physically plausible trajectories while bypassing time-consuming autoregressive generation. Extensive experiments on the nuScenes benchmark demonstrate that LaPla achieves competitive open-loop performance, reducing long-horizon L2 error by 15.52% compared to state-of-the-art VLA methods. Closed-loop evaluations on the NVIDIA AlpaSim simulator further confirm its superior capability in ensuring smooth driving progress, improving the success rate by 33.34 percentage points with significantly reduced inference latency.

18. 【2609.04034】Editable Visual Design

链接https://arxiv.org/abs/2609.04034

作者:Junyan Ye,Wei Liu,Dongzhi Jiang,Zichen Wen,HaoDong Li,Zhutao Lv,Jiaxin Lin,Jinhua Yu,Jun He,Zilong Huang,Rui Chen,Weijia Li

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:precluding layer-wise post-editing, Nano-Banana exhibit remarkable, inherently yields flattened, yields flattened bitmaps, remarkable visual expressiveness

备注

点击查看摘要

Abstract:While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

Cite as:
arXiv:2609.04034 [cs.CV]

(or
arXiv:2609.04034v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.04034

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
19. 【2609.04031】DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation

链接https://arxiv.org/abs/2609.04031

作者:Shuaiting Li,Zelin Gao,Haibin Shen,Yujun Shen,Haotong Qin,Yinghao Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:hinder practical deployment, achieved impressive progress, computational costs hinder, costs hinder practical, practical deployment

备注: Project page: \url{ [this https URL](https://robbyant-research.github.io/DSAQuant/) }; Code: \url{ [this https URL](https://github.com/robbyant-research/DSAQuant) }

点击查看摘要

Abstract:Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressing and accelerating advanced generative models without runtime overhead at inference. However, existing QAT methods suffer from a distinctive challenge in VDMs: while they often preserve prompt semantics, global layout, and coarse motion, the quantized model severely degrades visual details, texture fidelity, and sharpness. In this paper, we trace this degradation to the timestep-agnostic design of conventional quantization pipelines, which overlooks the stage-wise functionality of video denoising. In VDMs, early denoising steps mainly establish global structure and motion, whereas middle and late steps refine local appearance and high-frequency details. Based on this insight, we propose DSAQuant, a Denoising-Stage-Aligned Quantization-aware training framework for VDMs. During training, Denoising-Stage Oriented Supervision preserves teacher distillation in early steps for stable structure planning, while shifting later steps toward target-driven optimization to enhance detail reconstruction. During inference, Denoising-Stage Gated Guidance disables CFG in the final denoising steps to prevent it from amplifying quantization-induced errors into high-frequency artifacts. Extensive experiments on the Wan and CogVideoX families under W4A4 and W3A3 settings show that DSAQuant consistently outperforms the SOTA QAT baseline, improving the VBench average score by up to 6.60 under aggressive W3A3 quantization while preserving strong text-video alignment. These results demonstrate that effective VDM quantization requires not only reducing quantization error, but also aligning quantization training and inference with the stage-wise nature of video diffusion.

20. 【2609.04026】Stable and Scalable Bundle Adjustment of Holistic 3D Structures

链接https://arxiv.org/abs/2609.04026

作者:Shaohui Liu,Rémi Pautrat,Daniel Barath,Richard Hartley,Viktor Larsson,Marc Pollefeys

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Bundle Adjustment, computer vision, benefited from decades, decades of advances, sparse optimization

备注: To appear at ECCV 2026. Code available as part of the LIMAP toolbox at [this https URL](https://github.com/cvg/limap/)

点击查看摘要

Abstract:Bundle Adjustment (BA) is a cornerstone of 3D computer vision and has benefited from decades of advances in sparse optimization and numerical methods. It was originally developed for jointly optimizing camera intrinsics, poses and sparse 3D points. While extensions incorporate lines and other primitives, integrating richer geometric structures such as parallelism, coplanarity, or wireframes often introduces significantly increased computational cost and reduced numerical stability. In this paper, we propose a unified framework that extends bundle adjustment to jointly optimize geometric features and higher-order relations. We first introduce a taxonomy that distinguishes scalable geometric features with direct 2D measurements (e.g., points and lines), from groups encoding higher-order relations (e.g., coplanarity, parallelism, etc.), where we show that groups can be modeled as camera-like entities within the bundle adjustment framework. Building on this formulation, we propose that both group constraints and cross-feature relations (i.e., point-line associations) can be expressed through 2D reprojection measurements. By formulating group-induced and cross-feature reprojection errors, we preserve the sparsity structure of classical point-based BA under Schur elimination, while avoiding direct 3D regularization that degrades the conditioning and stability. Experiments on both real-world and synthetic datasets demonstrate runtime performance comparable to classical point-only bundle adjustment, while producing significantly richer 3D structures and improved geometric accuracy.

21. 【2609.04009】he Blind Spot in 2D Infants' Pose Estimation:Robust Learning from Noisy Annotations

链接https://arxiv.org/abs/2609.04009

作者:Emanuele Cardinale,Marco Proietti,Alessandro Cacciatore,Maria Francesca Spadea,Lucia Migliorelli,Sara Moccia

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:supervised deep learning, neural networks rely, impair model performance, high-quality labeled data, severely impair model

备注

点击查看摘要

Abstract:Noisy annotations pose a significant challenge for supervised deep learning, as neural networks rely on large-scale, high-quality labeled data whose corruption can severely impair model performance. Although robustness to label noise has been extensively studied for classification tasks, it remains relatively underexplored in Pose Estimation (PE). This limitation becomes critical in clinical contexts, including neonatology, where PE of preterm infants is used to support the assessment of spontaneous motility, a key indicator of neurodevelopmental trajectories. In such settings, infants' images labeling is further hindered by visual challenges (e.g., keypoint self-occlusions, caregiver interference), making the annotation process inherently susceptible to errors. To tackle noisy annotations in PE, we introduce REliable keypoint selection via Memory of traINing Dynamics (REMIND), a clustering-based keypoint-selection strategy that exploits keypoint-wise training dynamics to identify noisy labels without assuming any prior knowledge of the noise distribution, thus enabling noise-free model training. When evaluated on the proprietary NeoPose dataset, comprising 46 videos of 46 preterm infants recorded in real clinical settings, REMIND correctly identifies noisy annotations across multiple corruption scenarios, achieving up to 93\% Area Under the Curve (AUC) with three different PE architectures used in the relevant literature. To our knowledge, this is the first study to explicitly address label noise in preterm infants' PE, paving the way for the design of trustworthy learning-based algorithms for infants'monitoring support when data quality cannot be guaranteed.

22. 【2609.03995】Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition

链接https://arxiv.org/abs/2609.03995

作者:Abilash Philip Madavath,Chandra Yuvesh Aubeeluck,Augustin Raju,Nicolas Pyschny,Felix Hackelöer,Florian Zwanzig

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:carbide rotary burrs, rotary burrs conform, error-prone quality assurance, Verifying that manufactured, quality assurance task

备注: Extended abstract not yet published to a conference or journal

点击查看摘要

Abstract:Verifying that manufactured batches of milling tools or carbide rotary burrs conform to production order sheets remains a largely manual and error-prone quality assurance task. Automating this process with computer vision faces a critical cold-start constraint since no labelled imagery is available, leaving manufacturer catalogue photography as the sole source of supervision. We investigate how far catalogue supervision can support an industrial recognition pipeline under domain shift, explicitly measuring the gap between catalogue separability and performance on held-out field photographs. Our findings reveal three key insights. First, off-the-shelf frozen feature extractors do not reliably separate the two task attributes, head shape and tooth profile, motivating targeted representation learning. Second, metric learning produces near-perfect unsupervised cluster discovery on catalogue images (adjusted Rand index 0.94--0.97), but less than half of this gain transfers to field photographs. Third, the largest transfer gains do not come from model scale or representation complexity, but from simple changes that reduce domain sensitivity: converting images to grayscale (+0.22) and constraining retrieval using the known order sheet via Hungarian assignment (+0.11). We therefore treat catalogue photography as a useful cold start rather than a deployment-ready training domain, and provide empirical baselines and an evaluation protocol for catalogue-to-field transfer in precision tool manufacturing.

23. 【2609.03985】IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition

链接https://arxiv.org/abs/2609.03985

作者:Nazim-E-Alam,Tarek Rahman,Md Kishor Morol

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Zero-shot vision-language models, training-free species recognizers, visual species knowledge, vision-language models, multilingual Jina CLIP

备注

点击查看摘要

Abstract:Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.

24. 【2609.03981】Sharpening the Ensemble: An SSIM-Aligned Residual Refiner for Brain-MRI Inpainting Post-Processing

链接https://arxiv.org/abs/2609.03981

作者:Kubilay Kağan Kömürcü,İlkay Öksüz

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:plausible healthy tissue, anatomically plausible healthy, Brain-MRI inpainting replaces, analysis tools built, healthy tissue

备注: Accepted at the MICCAI BraTS Local Synthesis of Brain Tissue Inpainting Challenge (Task 4), MICCAI 2026. 12 pages, 2 figures

点击查看摘要

Abstract:Brain-MRI inpainting replaces a masked region of a scan with synthesized, anatomically plausible healthy tissue, so that analysis tools built for healthy brains can be applied to images they would otherwise reject. On the BraTS local-synthesis benchmark, which ranks submissions on the structural similarity index (SSIM), the peak signal-to-noise ratio, and the mean squared error (MSE) jointly, the strongest recent models are accurate, but several report blurry synthesized regions and attribute this to the mean-seeking behavior of the $\ell_1$ and MSE terms in their training losses. We address this in post-processing, forming a deep ensemble of the two co-first-place 2025 models and training a lightweight residual refiner on the ensemble's own outputs under an $\ell_1$ loss augmented with a structural-similarity term whose weight $\lambda$ we vary. At a moderate $\lambda$ the refiner improves SSIM over the ensemble, from $0.8767$ to $0.8780$ on a held-out reproduction of the official scorer and from $0.8555$ to $0.8572$ on the official validation leaderboard, with essentially no change in MSE. The gain is small but consistent, improving $62.6\%$ of the held-out cases with a signed-rank $p=2.2\times10^{-7}$, whereas over-weighting the structural term reverses it. Two ablations bound the effect. Adding any third model to the two-model ensemble degrades it, and classical unsharp masking fails to improve SSIM at any strength (best $0.8765$ against $0.8767$), so the gain reflects learned rather than indiscriminate sharpening. The result is a cheap, reproducible post-processing stage that improves an already strong ensemble without any large-scale retraining.

25. 【2609.03956】RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting

链接https://arxiv.org/abs/2609.03956

作者:Tomas Guija-Valiente,Blanca Rodriguez-Gonzalez,Norberto Malpica,Angel Torrado-Carvajal

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Medical image inpainting, automated brain MRI, brain MRI analysis, improve automated brain, brain MRI inpainting

备注: 11 pages, 2 figures. Preprint version corresponding to the initial submission prior to peer review, submitted as part of our participation in the BraTS 2026 Challenge. The final accepted version will be openly available in the official MICCAI proceedings on the conference website

点击查看摘要

Abstract:Medical image inpainting has the potential to improve automated brain MRI analysis by reconstructing healthy tissue within pathological regions. We introduce RARF, a task-agnostic region-aware rectified flow framework for masked data generation. We instantiate the framework for 3D brain MRI inpainting as our submission to the BraTS Inpainting Challenge 2026. RARF restricts the stochastic interpolation process to the inpainting region, while the observed voxels remain fixed and provide patient-specific anatomical context. A three-dimensional neural network receives the partially voided image, with Gaussian noise filling the missing region, together with the inpainting mask and the corresponding timestep. The model is trained using masked flow-matching and reconstruction-consistency objectives, combined with mask-aware preprocessing and data augmentation. During inference, the learned velocity field transports the initial noise toward a plausible reconstruction of the missing tissue, which is then combined with the unchanged observed anatomy. Experiments under the BraTS evaluation protocol show that the proposed approach produces competitive reconstructions while maintaining anatomical consistency. Source code is available at: this https URL.

26. 【2609.03952】WorldReward: Reward Modeling for Camera-Conditioned World Models

链接https://arxiv.org/abs/2609.03952

作者:Yibin Wang,Zehan Wang,Junshu Tang,Zhimin Li,Yujie Zhou,Jiazi Bu,Pengyang Ling,Feng Han,Zhixiong Zhang,Long Xing,Shengyuan Ding,Ziang Li,Cheng Jin,Yuhang Zang,Jiaqi Wang,Tianyu Pang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dynamics remain coherent, temporal dynamics remain, models generate interactive, generate interactive videos, remain coherent

备注: Website: [this https URL](https://codegoat24.github.io/WorldReward)

点击查看摘要

Abstract:Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.

27. 【2609.03931】Sparse auto-regressive modeling for scene generation from multi-view images

链接https://arxiv.org/abs/2609.03931

作者:Thomas Lucas,Maxime Pietrantoni,Philippe Weinzaepfel,Wonjune Cho,Bardienus Pieter Duisterhof,Vincent Leroy,Jerome Revaud

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:remaining computationally tractable, Generating complete, vision which requires, computationally tractable, fundamental challenge

备注: Accepted at ECCVV 2026

点击查看摘要

Abstract:Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generative modeling is hindered by the high computational cost of dense volumetric representations and the scarcity of large-scale 3D supervision. We introduce SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision. Our key insight is to formulate 3D scene generation in a structured, compact, voxel-aligned 3D latent space where only occupied voxels are represented. We learn this sparse latent space directly from multi-view images using photometric supervision via differentiable 3D Gaussian Splatting. Given a partial set of observed voxels encoded from sparse input views, scene completion reduces to predicting the missing latent tokens and their spatial support within the voxel grid. To this end, we train a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent generation of unseen regions. We demonstrate the effectiveness of our method on synthetic indoor scenes, achieving higher novel-view quality than prior work. We further validate its generalization on RealEstate10k, highlighting its applicability to real-world data.

28. 【2609.03919】OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping

链接https://arxiv.org/abs/2609.03919

作者:Zelong Lv,Sicheng Xu,Jianfeng Xiang,Ruicheng Wang,Yue Dong,Yu Deng,Guangzhong Sun,Jiaolong Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:video diffusion framework, framework with persistent, generating explorable, diffusion framework, high-fidelity visual scenes

备注: Accepted to ECCV 2026 Project page: [this https URL](https://maxtirerror.github.io/octworldpage/)

点击查看摘要

Abstract:We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWorld performs stable autoregressive world generation along user-specified camera trajectories. We focus on long-range generation, characterized by extended camera paths and wide viewpoint coverage, where preserving spatial consistency is particularly challenging when previously generated regions are revisited. To address this problem, we introduce OctMap, an extensible and spatially adaptive 3D memory that progressively fuses generated visual observations and their corresponding depth maps into a global representation. OctMap employs TSDF fusion within a dynamic sparse octree whose spatial resolution adapts to image evidence. This design preserves geometric and appearance details across diverse scene scales while maintaining low memory overhead. Experiments demonstrate that OctWorld generates long-range, spatially consistent videos and outperforms prior methods on both existing benchmarks and challenging long-range generation settings. OctMap also provides clear advantages over point-based caches and fixed-resolution TSDF volumes. Project page: this https URL

29. 【2609.03895】Concept of a Sensor Test Environment for Dusty Agricultural Conditions

链接https://arxiv.org/abs/2609.03895

作者:Peter Buckel,Johannes Hermann,Jonas Wollmann,Thomas Dietmueller,Timo Oksanen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous agricultural machinery, agricultural machinery, significant challenge, challenge for autonomous, autonomous agricultural

备注

点击查看摘要

Abstract:Dust in agriculture presents a significant challenge for autonomous agricultural machinery. Dust can impair the performance of sensors and algorithms. This work, therefore, presents a concept for a proving ground consisting of an indoor and outdoor area. The indoor area comprises a laboratory test bench where dust circulates in a closed system and a test hall where life-size objects can be placed. The outdoor area features dedicated test setups that enable reproducible data to be recorded with and without dust during real-world agriculture work. The proving ground and the setups are visualized in 3D.

30. 【2609.03892】GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

链接https://arxiv.org/abs/2609.03892

作者:Junqing Du,Fernando Ropero,Erkin Turkoz,Yanfeng Zhang,Lu Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:current multimodal large, multimodal large language, large language models, reasoning underpins understanding, physical world

备注

点击查看摘要

Abstract:3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.

31. 【2609.03829】he impact of phase information for few-shot fine-grained image classification

链接https://arxiv.org/abs/2609.03829

作者:Ruiling Liu,Linyue Zhang,Wenyi Zeng,Jiamiao Lu,Weichuang Zhang,Changming Sun,Zejun Zhang,Xiao Zhao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Few-shot fine-grained image, Few-shot fine-grained, fine-grained image classification, classify similar images, aims to classify

备注

点击查看摘要

Abstract:Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critical yet underutilized role of phase information in capturing structural relationships within an image. This study introduces a novel plug-and-play amplitude-phase integration (API) module that effectively combines local and global frequency amplitude and phase information for obtaining more comprehensive feature descriptors. Additionally, a dedicated network, named PSF-Net, is proposed that adaptively fuses phase-based spatial and frequency information for FSFGIS. The designed PSF-Net can be easily integrated into standard episodic training architectures for end-to-end training from scratch. Extensive experiments on five public datasets demonstrate that the method outperforms existing state-of-the-art benchmarks.

32. 【2609.03824】VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

链接https://arxiv.org/abs/2609.03824

作者:Ernesto Lozano,Alberto Jaenal,Javier Civera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:predicting camera poses, strong zero-shot generalization, showcasing strong zero-shot, foundation models, excel at predicting

备注

点击查看摘要

Abstract:3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.

33. 【2609.03820】Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

链接https://arxiv.org/abs/2609.03820

作者:Prakhar Khatri

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Long-video language models, small fixed slice, hour sampled, Orthogonal Matching Pursuit, Long-video language

备注: 16 pages, 6 figures. Code and data: [this https URL](https://github.com/codeprakhar25/omp-keyframe-sampling)

点击查看摘要

Abstract:Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.

34. 【2609.03816】When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

链接https://arxiv.org/abs/2609.03816

作者:Xinjian Zhao,Wei Pang,Zhixuan Yu,Xiangru Jian,Xiaozhuang Song,Yaoyao Xu,Zhongkai Xue,Dingshuo Chen,Shu Wu,Philip Torr,Tianshu Yu

类目:ocial and Information Networks (cs.SI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Graph Neural Networks, fundamental data structure, data structure underlying, Graphs, Graph

备注: IJCAI Survey Track, 2026

点击查看摘要

Abstract:Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: chemists read molecular diagrams and social scientists inspect network visualizations. Despite decades of work on graph visualization, most graph learning pipelines still treat graphs purely as symbolic structures, rarely leveraging the visual form of graphs. We argue that this gap deserves renewed attention in the era of powerful vision and vision-language models. This survey provides a first systematic overview of the emerging area we term vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning. We organize existing work into three threads. Vision for Graph Reasoning studies how models can use visual depictions of graphs to understand structure and carry out multi-step reasoning. Vision for Graph Learning explores how visual features can complement or augment graph encoders beyond known limitations of message passing. Scientific Graphs examines domains where standardized depiction conventions support both reasoning and learning. Our goal is to clarify what current methods can and cannot do, and to outline a path toward foundation models that perceive and reason about graphs as scientists do.

35. 【2609.03813】SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution

链接https://arxiv.org/abs/2609.03813

作者:Federico Putamorsi,Leonardo Zini,Marcella Cornia,Lorenzo Baraldi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Real-world image super-resolution, Diffusion Transformer, Real-world image, relies on Diffusion, image super-resolution

备注

点击查看摘要

Abstract:Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose internal activations can be dominated by a small number of massive channels. Yet improving perceptual quality in these models still typically requires fine-tuning the network or attaching additional adapters, leaving this structured activation space largely unexplored for adaptation. We investigate whether dominant channels can instead serve as a compact adaptation interface for frozen DiT-based SR models. We first characterize their behavior in pretrained SR backbones and show through controlled interventions that they strongly affect reconstruction quality. Building on this observation, we introduce SPARK, a lightweight input-conditioned controller that predicts bounded per-channel affine transformations for only the selected channels, while keeping the SR backbone and VAE frozen. Dominant channels are identified through an online activation-ranking procedure, and only a small predictor conditioned on the low-resolution VAE latent is optimized. Experiments on three DiT-based SR backbones across DIV2K, RealSR, and DRealSR show consistent gains in both fidelity and perceptual quality while modulating only eight channels per stream and block. Controlled comparisons further show that these gains cannot be explained by parameter budget or access to the selected channels alone.

36. 【2609.03811】VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence

链接https://arxiv.org/abs/2609.03811

作者:JoyIndustrial VisCAD Team:Linxin Cai,Qiuhe Hong,Zhichao Huang,Guanlin Li,Hongsen Liu,Ziqi Liu,Yichen Long,Luya Wang,Yuchen Wang,Wenxiang Wu,Huimu Yu,Ning Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:AI-assisted computer-aided design, AI-assisted computer-aided, challenging phases, involves two challenging, CAD

备注: Technical report from JoyIndustrial's AI CAD project

点击查看摘要

Abstract:AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to executable programs in a CAD domain-specific language. Assembly-level generation must additionally handle interacting parts, plan mating relations, estimate poses, and place all parts correctly. Existing specialized CAD models are commonly trained on narrow input domains, such as renders or texts, and often generalize poorly, while general-purpose frontier models cover broader inputs but perform inconsistently across CAD domains. We present VisCAD, a foundation model suite designed to provide both broad generalization and strong CAD capability for realistic industrial products. At its core is VisCAD-M1, a 27B model trained through mid-training and post-training for part-level design generation. On PubCADBench and RealCADBench, VisCAD-M1 achieves the highest average part-level score among the evaluated models, reaching 0.5540 compared with 0.5496 for the strongest frontier model. Reusing VisCAD-M1 as a test-time verifier can further raise the score to 0.5797, an approximately 5 percent relative improvement over the previous state of the art. VisCAD also includes a domain-specific harness that leverages frontier models for complex assembly generation and demonstrates advantages over general-purpose harnesses in both quantitative and qualitative evaluations.

37. 【2609.03806】SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

链接https://arxiv.org/abs/2609.03806

作者:Marco Cipriano,Leonardo Zini,Alexandra Schild,Valentin Teutschbein,Afsana Mimi,Marcella Cornia,Lorenzo Baraldi,Gerard de Melo

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Scalable Vector Graphics, attracting increasing attention, generative models improve, Vector Graphics, Scalable Vector

备注

点击查看摘要

Abstract:Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment. We introduce \textbf{\ours}, a human-aligned evaluation framework for text-to-SVG generation. Through controlled caption and image perturbations, we first show that CLIP-based scores barely react to the errors SVG generators actually make, such as wrong colors, counts, and spatial relations, and that off-the-shelf Vision-Language Model (VLM) judges, while more sensitive, respond unevenly across error types and SVG styles. We then introduce a human-annotated dataset for \textit{Semantic Alignment}, measuring how faithfully a generated SVG reflects its caption. Building on it, we develop two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning, for more expressive and interpretable assessment. Using both, we benchmark major open-source, commercial, and optimization-based SVG generators on an independent caption set.

38. 【2609.03804】Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications

链接https://arxiv.org/abs/2609.03804

作者:Minwei Zhao,Weiming Zhang,Jiawang Du,Qiming Liu,Weiming Zhuang,Pei Nie,Cai Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:shape urban form, Communities are fundamental, fundamental spatial units, Greater Bay Area, social life

备注: ECCV 2026 camera-ready version

点击查看摘要

Abstract:Communities are fundamental spatial units that shape urban form and social life. Whether a residential compound is spatially open or enclosed affects mobility, access to public services, and equity, yet studies of Chinese fengbi xiaoqu remain largely qualitative or small-scale, limiting reproducible city-scale analysis. We address this gap by introducing GBA-GCs, a metropolitan-scale multimodal benchmark for locally grounded gated/open community recognition in China's Greater Bay Area, covering 37,444 residential compounds with aligned boundary polygons, high-resolution satellite imagery, Chinese metadata, and structured attributes, together with expert-verified labels, inter-annotator reliability, and official evaluation splits. Built on this benchmark, we present Multimodal Classifier for Gated Community (MCGC), a vision-centric multimodal framework based on DINOv3-SAT that fuses imagery, text, and structured cues via modality-aware cross-attention and adaptive gating to mitigate modality imbalance. MCGC consistently outperforms strong unimodal and multimodal baselines. Finally, we apply the validated model to metropolitan-scale mapping and report equity-oriented findings including spatial clustering of GCs, privatized green space, and reduced pedestrian connectivity. The benchmark, code, and release documentation are available at this https URL.

39. 【2609.03796】LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

链接https://arxiv.org/abs/2609.03796

作者:Chuyan Chen,Haoxing Chen,Kun Chen,Zhenglin Cheng,Long Cui,Ruishan Fang,Zhangxuan Gu,Zhicheng Huang,Zhenzhong Lan,Yuanting Lei,Haoquan Li,Jianguo Li,Rongchuan Li,Sidu Li,Tao Lin,Deyuan Liu,Jiacheng Liu,Lin Liu,Yuxuan Lou,Zhisheng Lu,Yuxin Ma,Shuheng Shen,Peng Sun,Chaoyang Wang,Hongjun Wang,Xiaomei Wang,Yongxin Wang,Chengzhang Wu,Hongru Wu,Jun Xie

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Diffusion Transformer, frozen vision-language understanding, vision-language understanding module, understanding module built, diffusion language model

备注

点击查看摘要

Abstract:We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.

40. 【2609.03788】A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval

链接https://arxiv.org/abs/2609.03788

作者:Santiago Poveda-Gutiérrez,Hideki Nakayama,Mayumi Bono

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Isolated Sign Language, Sign Language Recognition, Language Recognition, Japanese Sign Language, Sign Language

备注: 4 pages, 2 figures, 1 table. Extended version of an abstract presented at the BU-SHI workshop (Broadening the Users: A Cross-Disciplinary Roadmap for Social Humanoid Interaction), IEEE RO-MAN 2026, Kitakyushu, Japan, 28 August 2026. The workshop is non-archival; no proceedings

点击查看摘要

Abstract:Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning a sign-level clip into a free-form procedural description of the articulation with an open-weight vision-language model, and (2) retrieving the closest entry from a vocabulary of target descriptions with a multilingual sentence encoder: a reverse sign language dictionary that needs no gloss supervision and admits an open vocabulary. On 1,300 sign-level segments from a Japanese Sign Language (JSL) dialogue corpus annotated with procedural descriptions (against a 2% top-10 chance floor over the 503-entry target vocabulary), fine-tuning the captioner substantially improves seen-class retrieval: language and vision tower fine-tuning raises top-10 retrieval on seen classes from 4.5% (untrained) to 49%, becoming statistically indistinguishable from a standard supervised closed-set classifier (I3D) on two of the three test sets where a closed-set classifier can be evaluated at all. More importantly, unseen-class retrieval also improves significantly over the untrained pipeline (11.5% - 21.0% top-10, p=0.0094), a regime in which the closed-set classifier cannot participate. A matcher-side empirical upper-bound analysis shows the sentence encoder already recovers close to 100% of paraphrased gold descriptions, locating a gap in captioning quality that we aim to address in future work. To our knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.

41. 【2609.03773】RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

链接https://arxiv.org/abs/2609.03773

作者:JoyIndustrial VisCAD Team:Linxin Cai,Qiuhe Hong,Zhichao Huang,Guanlin Li,Zongzhen Li,Hongsen Liu,Yichen Long,Wei Wang,Yuchen Wang,Dongyue Yang,Huimu Yu,Xianwen Zhong

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Parametric computer-aided design, Parametric computer-aided, CAD modeling, CAD, Existing CAD benchmarks

备注: Benchmark from JoyIndustrial's AI CAD project

点击查看摘要

Abstract:Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.

42. 【2609.03756】ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

链接https://arxiv.org/abs/2609.03756

作者:Javier del Pino(SperidLabs),Salvador Rodríguez(SperidLabs),Alejandro Garabito(SperidLabs),Javier Álvarez(SperidLabs),Chema Garabito(SperidLabs)

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Text-promptable segmentation models, present ENEAS, semantic, high-quality semantic tracking, text-promptable method

备注: 19 pages, 5 figures, 6 tables. Code and models: [this https URL](https://github.com/speridlabs/eneas)

点击查看摘要

Abstract:We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at this https URL

Comments:
19 pages, 5 figures, 6 tables. Code and models: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

ACMclasses:
I.4.6; I.2.10; I.4.8

Cite as:
arXiv:2609.03756 [cs.CV]

(or
arXiv:2609.03756v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.03756

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
43. 【2609.03742】KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

链接https://arxiv.org/abs/2609.03742

作者:Yi Xu,Yifan Hou,Xiaoyu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:overwhelm novice learners, valuable educational resources, novice learners, dense and lengthy, lengthy formats

备注: This work is published on EMNLP 2026 (Findings). Our code and dataset are available at [this https URL](https://github.com/yixu-cityu/KnowVis) and [this https URL](https://huggingface.co/datasets/yixu-cityu/KnowVis)

点击查看摘要

Abstract:Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe cognitive overload for novice learners lacking prior domain knowledge. Existing video summarization methods fail to resolve this mismatch, as they primarily produce text-heavy, linear condensations that still demand high cognitive effort. To bridge this gap, we propose KnowVis, a framework that transforms linear video lectures into pedagogically grounded visual narratives. KnowVis first extracts a detailed concept map from multimodal video content to identify important and challenging threshold concepts, then constructs structured knowledge units, and finally synthesizes engaging visual summaries. Alongside the framework, we introduce a curated dataset of 125 educational videos across 10 academic disciplines, paired with 1,079 generated visual summaries. Extensive automated evaluations and a human study demonstrate that, compared to state-of-the-art baselines, KnowVis generates more accurate and clear visuals that successfully reduce cognitive load and significantly improve student learning effectiveness and knowledge retention.

44. 【2609.03740】Fill My Mirror: Geometry-Constrained Mirror Inpainting

链接https://arxiv.org/abs/2609.03740

作者:Ofek Basson,Shimon Vainer,Yacov Hel-Or,Ohad Fried

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models remains challenging, producing geometrically consistent, remains challenging, mirror, generative models remains

备注

点击查看摘要

Abstract:Mirrors are common in real-world images, yet producing geometrically consistent reflections with generative models remains challenging. Unlike most objects, mirror appearance depends on scene geometry and viewpoint, making it hard to synthesize using learned appearance priors alone. We address this in the mirror inpainting setting, where the scene is fixed and only the mirror region is generated. Our key insight is that much mirror content is geometrically constrained by the visible scene and need not be hallucinated. We estimate scene geometry and project visible content into the mirror to recover reflection regions determined by geometry. A generative model then completes the mirror region via a two-mask diffusion strategy balancing geometric constraints with the model's learned priors, reducing projection artifacts and improving reflection consistency. The method is training-free and applicable to complex real-world scenes. We evaluate on MirrorBench-V2 (synthetic) and real images. Using standard and geometry-aware metrics, we show that explicitly using scene geometry improves consistency.

45. 【2609.03729】Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

链接https://arxiv.org/abs/2609.03729

作者:Yijun Yang,Shenghe Zheng,Wenbo Li,Jianhui Liu,Haoze Sun,Yanbing Zhang,Jiaxiu Jiang,Lin Song,Haoyang Huang,Nan Duan,Lei Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:general multimodal tasks, Vision-Language Models, multimodal tasks, remain fundamentally, physical world

备注: Accepted by ECCV 2026

点击查看摘要

Abstract:Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence ($XY$), depth consistency ($Z$), and temporal reversibility ($T$). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

46. 【2609.03695】SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

链接https://arxiv.org/abs/2609.03695

作者:Sobhan Asasi,Ozge Mercanoglu Sincan,Richard Bowden

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:challenging problem due, sign language learners, Sign language dictionaries, query video, remains a challenging

备注

点击查看摘要

Abstract:Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a sign from a dictionary, given only a query video, remains a challenging problem due to the natural variability between signers. Existing sign representation learning methods are built for closed-set recognition, producing embeddings that do not generalise to the open-set, signer-independent setting that retrieval demands. \textbf{SignSeek} closes this gap by contrastively learning sign representations with saliency-guided articulator masking. A contrastive objective aligns same-gloss signs across signers, while our Articulator Saliency-Guided Masking (ASGM) pinpoints the single most critical articulator per sign. This drives two complementary objectives, a masked contrastive alignment (MAC) loss that sees the sign through a single articulator and a masked prediction (MAP) loss that reconstructs it in latent space from the surrounding spatio-temporal context. Pretrained on 266K samples ($\sim$5,700 glosses) across multiple sign languages, \textbf{SignSeek} sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning. Strikingly, it achieves zero-shot generalisation to an entirely unseen British Sign Language (BSL), surpassing methods explicitly trained on BSL, and transfers seamlessly to isolated sign recognition and subtitle alignment, outperforming prior skeleton-based methods.

47. 【2609.03694】Observation-Conditioned Latent Energy Priors for Sparse Implicit Neural Shape Completion

链接https://arxiv.org/abs/2609.03694

作者:Paul Büschl,Ezequiel de la Rosa,Julia Wolleb,Julian McGinnis,César Nombela-Arrieta,Bjoern Menze

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Implicit neural representations, Implicit neural, shared coordinate decoder, per-instance latent codes, neural representations

备注: Accepted for publication in the MICCAI 2026 Workshop Proceedings (LNCS) as part of Off-Grid, the 1st Workshop on Continuous Representations and Grid-Free Methods in Medical Imaging. 12 pages, 4 figures, 3 tables

点击查看摘要

Abstract:Implicit neural representations (INRs) can model continuous 3D shapes with a shared coordinate decoder and per-instance latent codes. At test time, autodecoder-style models commonly freeze the decoder and optimize a new latent code from sparse off-grid SDF samples. When these samples underconstrain inference, the latent can drift toward regions that fit the observations but decode implausible unobserved geometry. We propose a post-hoc observation-conditioned latent energy prior for frozen INR decoders. The energy scores standardized latents conditioned on a permutation-invariant encoding of the sparse observation set and is used as a residual expert alongside an L2 latent prior selected on validation data. We evaluate on a controlled cell-nucleus SDF dataset and a public MedShapeNet-derived SDF completion dataset. The proposed L2 objective augmented with conditional energy improves consistently over a validation-selected L2 baseline in the sparsest cell-nucleus regimes and, on MedShapeNet, outperforms both L2 and a six-component GMM latent-density prior across all reported readouts. A shuffled-context ablation is consistently weaker than matched context, supporting an observation-specific contribution. These results suggest that lightweight conditional energies can make pretrained INR decoders more observation-aware without retraining.

48. 【2609.03690】MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

链接https://arxiv.org/abs/2609.03690

作者:Chenguang Zheng,Le Xue,Yichi Zhang,Wenbo Zhang,Zehui Ling,Gang Feng,Xin Gao,Yuan Qi,Yuan Cheng,Zixin Hu,Mei Tian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:whole-body PET, structure is essential, essential for clinical, clinical diagnosis, comprehensive whole-body PET

备注

点击查看摘要

Abstract:The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehensive whole-body PET/CT analysis. In this work, we introduce MetaStructAtlas, a large-scale dataset for grounded whole-body PET/CT interpretation that synthesizes multimodal imaging with integrated anatomical, metabolic, and semantic annotations. MetaStructAtlas provides 490 co-registered 3D PET and CT volumes with 50,470 organ-level segmentation masks and grounded radiology reports. To facilitate interactive reasoning, we further developed MetaStructVQA, a standardized 3D grounded visual question-answering benchmark containing 100,565 QA pairs. This framework explicitly links diagnostic queries to visual evidence across modalities, encompassing anatomical, morphological, and metabolic characteristics. Finally, we evaluate state-of-the-art 3D medical VLMs on MetaStructVQA, establishing a robust foundation for multimodal representation learning and integrated whole-body reasoning in nuclear medicine.

49. 【2609.03689】Semantic-Aware Subgraph State Space Model for WSI Classification in Histopathology

链接https://arxiv.org/abs/2609.03689

作者:Feixing Chen,Hao Lu,Lin Luo,Yan Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Histopathological subtyping relies, characteristic histological patterns, Subgraph State Space, Histopathological subtyping, State Space

备注

点击查看摘要

Abstract:Histopathological subtyping relies on the recognition of characteristic histological patterns. These patterns may be expressed by individual tissue structures or by the spatial distribution and co-occurrence of multiple structures, and they often span irregularly shaped tissue regions, termed semantic units in this work. However, conventional patch-based representations may fragment such units and fail to explicitly preserve their internal spatial organization, while efficiently modeling relationships among numerous spatially separated units remains challenging. To address these limitations, we propose the Semantic-Aware Subgraph State Space Model (SASG-SSM), a flexible and efficient framework for whole slide image (WSI) classification. Semantic-Aware Subgraphs (SASGs) first approximate irregularly shaped semantic units by adaptively grouping spatially connected patches guided by class-agnostic visual-semantic priors. By representing patches as graph nodes with adjacency edges, SASGs preserve their internal spatial organization rather than treating them as an unordered set. A Subgraph State Space Module (SG-SSM) subsequently combines a graph neural network encoder for intra-subgraph topology encoding with a Mamba-based state space encoder for efficient contextualization across large numbers of subgraphs. This module integrates local structural information within semantic units with global contextual information arising from their distribution and co-occurrence across the WSI, while efficiently modeling a large number of spatially distributed regions. Extensive experiments across four WSI subtyping datasets demonstrate consistent advantages over representative state-of-the-art methods. Further evaluations under small-cohort and few-shot settings demonstrate robustness and data efficiency under limited training data. Code will be released at this https URL.

50. 【2609.03688】oPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Models

链接https://arxiv.org/abs/2609.03688

作者:Juntao Xu,Shihong Li,Hoi Fan Au,Ning Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:rank complete images, Pairwise preference labels, Pairwise preference, labels rank complete, Token-Oriented Preference Optimization

备注: 37 pages, 11 figures

点击查看摘要

Abstract:Pairwise preference labels rank complete images, yet Diffusion-DPO applies their effect over many spatial and denoising-time coordinates. For attention-based, noise-prediction latent diffusion, ToPO (Token-Oriented Preference Optimization) constructs a per-minibatch, detached, separable spatial-temporal route from branchwise squared-residual contrast in a frozen reference denoiser. Preferred-branch cross-attention uses content tokens to modulate the spatial factor, and an auxiliary pixel-midpoint ordering term is added without local labels or a learned reward model. In matched three-seed retrainings with a shared update schedule, ToPO has higher endpoint estimates than Diffusion-DPO on all five reported SD-1.5 metrics and on HPSv2, ImageReward, and CLIP for SDXL. It also receives larger raw win shares in an aggregate blind SDXL A/B study. These findings are scoped to the reported equal-update U-Net protocols rather than an equal-compute comparison.

51. 【2609.03680】DropClick: Semi-Automated One-Click Segmentation for Agricultural Robotic Data

链接https://arxiv.org/abs/2609.03680

作者:Patrick Zimmer,Michael Halstead,Chris McCool

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Labelling vision datasets, Labelling vision, laborious and costly, stymies novel developments, Labelling

备注: Accepted to ICRA 2026

点击查看摘要

Abstract:Labelling vision datasets, especially for segmentation tasks, is a laborious and costly process that stymies novel developments in agricultural robotics. In this paper, we present DropClick, a click-guided segmentation tool that simplifies the annotation process. Our system utilises single-click inputs on objects to generate pseudo-labels, which can replace manual annotations. DropClick stands out as it is a semi-automated approach and does not require a click for every object in the scene. It can therefore further reduce the required amount of user input drastically. We evaluate our method on two challenging agricultural robotic datasets, SB20 and BUP20 for plant and fruit segmentation, respectively. DropClick is first trained on a small subset of just 5 images from the original training data. This DropClick model can then be deployed as a one-click segmentation system and achieves comparable or higher performance than other one-click methods achieving an mIoU of 70.0 and 72.6 points, for SB20 and BUP20 respectively. DropClick then excels at maintaining high performance when clicks are not given (e.g. dropped); when 50% of the clicks are missing it still maintains an mIoU of 68.9 and 71.3 points, for SB20 and BUP20 respectively. We validate DropClick as a pseudo-labelling approach by taking its outputs to train a Mask2Former instance-based segmentation model in a semi-supervised manner. In this process, partially removing user input from DropClick yields similar high performance when compared to providing all clicks, at 70.1 vs 70.7 points AP50 for SB20 and no difference for BUP20 at 77.0 for both models; at the same time saving 46.3% of total input for SB20 and 31.9% for BUP20.

52. 【2609.03677】Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

链接https://arxiv.org/abs/2609.03677

作者:Julian Truetsch,Felix Hauser,Christoph Stiller,Frank Bieder

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Robotics (cs.RO)

关键词:Understanding the composition, essential for safety, composition of large-scale, reliable operation, large-scale autonomous driving

备注: 9 pages, 5 figures, submitted to the IEEE Open Journal of Intelligent Transportation Systems (OJ-ITS), our implementation and benchmark dataset are available at [this https URL](https://github.com/KIT-MRT/AD-Diff)

点击查看摘要

Abstract:Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned with the training data, resulting in potentially dangerous performance degradation. Yet, existing data analysis pipelines largely rely on metadata, predefined labels, or manual inspection, which provide limited semantic insight or do not scale. This paper studies set difference captioning: given two subsets of images, the goal is to produce a natural-language hypothesis describing differences between the target and reference set. Building on a two-stage formulation, we adapt the method to autonomous driving by focusing on object-centric patches derived from object detection, which simplifies aggregation and enables attribution of differences to specific object instances or categories. To evaluate this setting in-domain, we introduce a new benchmark, AD-Diff Bench. Low-concentration experiments assess the suitability of set-difference-captioning approaches to sparse, real-world differences. We restrict our experiments to open-weight models to support reproducibility and ease of deployment. The proposed benchmark and analysis provide a step towards practical, human-interpretable dataset introspection for autonomous driving datasets. Our implementation and benchmark dataset are available at this https URL

53. 【2609.03675】CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

链接https://arxiv.org/abs/2609.03675

作者:Jing Jiang,Yiran Ling,Ruonan Li,Dimitrios Stamoulis,Jie Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires Vision Language, Vision Language Models, tight latency constraints, Streaming video understanding, Vision Language

备注

点击查看摘要

Abstract:Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downstream token pruning alone cannot substantially reduce end-to-end latency because the expensive frame encoding cost has already been incurred. We propose CoFiE, a Coarse-to-Fine Evidence Selection framework that decouples evidence selection into a coarse, query-agnostic filtering stage before the vision encoder and a fine, query-specific refinement stage during LLM prefill. CoFiE introduces Novelty-Guided Frame Filtering to retain visually distinctive candidate frames and Query-Specific Evidence Refinement to select the frames most relevant to the user query. This design removes substantial redundancy before frame encoding while preserving query-specific refinement once semantic information becomes available. Experiments show that CoFiE establishes a new state-of-the-art accuracy-efficiency trade-off across multiple video understanding benchmarks, reaching 78.86% accuracy on StreamingBench and 68.72% on OvO-Bench, with improvements of up to 3.15% over prior methods. Even with up to 80% evidence-frame filtering, CoFiE outperforms strong open-source multimodal models while improving end-to-end inference latency by up to 2.54 times.

54. 【2609.03673】Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation

链接https://arxiv.org/abs/2609.03673

作者:Yingmao Miao,Pengfei Zhang,Chaoran Xu,Meng Yu,Jing Tang,Xiangxiang Chu,Chao Shen,Chenhao Lin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autoregressively extending chunks, composing shorter parts, generators build long, build long videos, extending chunks

备注

点击查看摘要

Abstract:Video generators build long videos by composing shorter parts, either by generating segments one after another or by autoregressively extending chunks. Each new part usually depends on memories of historical observations, such as recent frames, selected key frames, memory banks, or cached features. These memories preserve visible evidence from the past, but current generators do not reliably turn such evidence into a world-state interface: what holds in the video world after previous actions and how it should change under the next prompt. A past frame remains valid history, but it may not describe the state needed by the next segment; some states must instead be inferred from occluded or implicit changes rather than copied from a directly observed frame. This creates a simple but overlooked question for video continuation: given a previous video, its prompt, and a new prompt, can a model generate a continuation that reflects the state determined by both the historical video and the new prompt? To answer this question, we introduce Statebench, a benchmark that targets this gap by testing continuations over three state categories: past-visible states, occluded-process states, and complex-transition states. We further propose Stateagent, which explicitly maintains an entity-state representation, updates it under the new prompt, grounds the predicted post-action state as a future end frame, and renders the next video. Experiments show that our method improves controlled video continuation by raising the all-case state score (SCS-All) from 45.2 to 69.3, and also benefits story generation at the one-minute scale. Code is avaliable at this https URL.

55. 【2609.03668】ARCOS: Zero-shot Boundary Localization for Corneal Layer Segmentation Across Optical Coherence Tomography Devices

链接https://arxiv.org/abs/2609.03668

作者:Nuno Vivas Brás,Benjamin Memmi,Maëlle Bouhassane,Cristina Georgeon,Vincent Borderie,Karsten Plamann,Anatole Chessel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:optical coherence tomography, Accurate segmentation, coherence tomography, disease or surgery, optical coherence

备注: 18 pages, 9 figures, 7 tables

点击查看摘要

Abstract:Accurate segmentation of corneal layers in optical coherence tomography (OCT) is essential for quantitative assessment of corneal morphology, including layer thickness and structural changes associated with disease or surgery. However, automatic segmentation remains challenging because corneal interfaces are thin, affected by speckle noise, and variable across acquisition devices. In this work, we propose ARCOS, a patch-based zero-shot boundary localization framework for corneal layer segmentation in clinical anterior-segment OCT images. Rather than performing conventional region classification, the method predicts boundary heatmaps for the main corneal interfaces from overlapping native-resolution patches. Patch-level predictions are stitched across the full B-scan and converted into boundary locations to obtain continuous, anatomically ordered layer segmentations. The network combines multi-scale feature fusion with a self-conditioned refinement module that uses intermediate boundary information to improve local heatmap predictions while preserving spatial detail. The method was evaluated on clinical OCT images acquired from multiple devices and compared with representative segmentation baselines using boundary localization and derived thickness metrics. The proposed method achieved an off-by-one boundary localization accuracy of 95.1% and a mean absolute boundary error of 0.514 pixels on the matched-device test set. In zero-shot cross-device evaluation, it maintained an average off-by-one accuracy of 84.3% and a mean absolute boundary error of 0.855 pixels across unseen acquisition devices, outperforming the baseline models. Thickness estimates derived from the predicted boundaries showed low error across corneal regions, supporting the method's use for quantitative corneal OCT analysis.

56. 【2609.03663】Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography

链接https://arxiv.org/abs/2609.03663

作者:Louis Chen,Torbjörn E. M. Nordling

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)

关键词:Background, Salience-guided Faithfulness Coefficient, coverage, Intuition, SaCo

备注: 92 pages, 38 figures, incl. supplementary

点击查看摘要

Abstract:Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanations have rested on inspecting heatmaps rather than on quantitative evidence about where a model reads it. We quantified the explanations and asked whether such explanations transfer between datasets and track model performance. Method. We trained eight condition-specific RhythmFormer models on NCKU-rPPG, recorded under three illumination levels, speaking, rotation, and cycling, estimated one heart rate per 5.12-second clip, and set them beside a UBFC-rPPG reproduction. Raw attention, rollout, attention flow, and Beyond Intuition were assessed by skin coverage and the Salience-guided Faithfulness Coefficient (SaCo). Results. Beyond Intuition ranked highest on both datasets, at median coverage 0.789 and SaCo 0.837 on Static level 3 against 0.826 and 0.917 on UBFC-rPPG; lower ranks differed. Within one participant of one condition, neither measure was related to a clip's heart-rate error, waveform correlation, or signal-to-noise ratio on either dataset: 186 of the 252 coefficients fell below $|\rho|=0.10$ and 28 reached $p0.05$ against the 13 expected by chance. Across the eight scenarios only Beyond Intuition's coverage followed the three performance measures, at $\rho=-0.43$, $+0.57$, and $+0.43$, while the attention-only methods' SaCo ran opposite to each. It failed at 40 lux alone, its median coverage falling to 0.180 and its median SaCo to $-0.178$, whereas motion degraded the estimates far more without such a drop. Conclusions. Skin coverage and SaCo carry information complementary to the performance measures rather than a proxy for them: attributing to the skin does not guarantee an accurate estimate. What an attribution reveals about a condition is where the model looks rather than how faithfully its map is ordered.

57. 【2609.03657】Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

链接https://arxiv.org/abs/2609.03657

作者:Onat Şahin,Mohammad Altillawi,George Eskandar,Carlos Carbone,Ziyuan Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:suffer severe artifacts, Gaussian Splatting, suffer severe, scene representations, Splatting

备注

点击查看摘要

Abstract:3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view settings. Recent generative 3D artifact fixers attempt to address this, but rely on paired corrupted and clean renders requiring costly, per-scene reconstructions across varying view configurations. While 2D image augmentations act as instant regularizers, no explicit equivalents exist for 3D representations to preserve spatial consistency across views, an essential property for 3D-aware training. We propose 3D Morphological Perturbations as an optimization-free regularizer that preserves spatial consistency. Leveraging explicit 3DGS, we treat each Gaussian as a fundamental building block - analogous to a 2D pixel - and apply perturbations across its morphological parameter space via scale, rotation, and pruning. Our method eliminates per-scene 3DGS optimization loops from dataset curation while enabling models to learn stronger geometric priors than sparse-view baselines in diagnostic ablations conducted on a lightweight video diffusion sandbox. Scaled to a 14B-parameter video model via ControlNet, our approach maintains visual fidelity while reducing mean depth error by 12.5% over state-of-the-art image-to-image 3D artifact refiners, ultimately boosting downstream robotics policy success rates by up to 8.0% across 3 of 4 manipulation tasks.

58. 【2609.03655】PL-SCEA: Reconfiguring Pretrained Attention for Few-Shot Industrial Anomaly Detection

链接https://arxiv.org/abs/2609.03655

作者:Xiaoyu Yang,Qixing Wu,Huixian Zhao,Changlong Jin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision Foundation Models, Vision Foundation, Foundation Models, pretraining objectives centered, provide transferable patch

备注

点击查看摘要

Abstract:Vision Foundation Models (VFMs) provide transferable patch representations for few-shot industrial anomaly detection, but their attention computation is typically inherited from pretraining objectives centered on semantic aggregation. This creates a potential mismatch: token relations that support semantic recognition may not adequately expose the localized texture and structural deviations required for anomaly localization. We therefore investigate the hypothesis that the attention computation of a frozen VFM can be reconfigured as a task-relevant component of anomaly detection. We instantiate this idea with Power-Law Self-Correlation Enhanced Attention (PL-SCEA), which retains the semantic context of pretrained query-key attention while constructing token-adaptive self-correlations over contextualized value features. Positive-correlation filtering and power-law reweighting then emphasize relations that are salient relative to each token's relational background, without introducing additional trainable attention projections. The resulting features are modeled by a lightweight variational autoencoder that provides a fixed-size reconstruction-based representation of category-specific normality. The two stages serve complementary roles: attention reconfiguration shapes how local relational deviations are represented, while reconstruction-based modeling converts deviations from learned normality into anomaly scores. Across MVTec AD and VisA, the complete framework achieves competitive image-level detection and consistently strong pixel-level localization across the evaluated few-shot settings. Ablations further show that PL-SCEA improves localization with either the VAE or a memory bank under the tested setting. These results support the view that task-aligned attention reconfiguration can improve the anomaly-localization capability of frozen pretrained representations.

59. 【2609.03641】ree-Structured Vector Quantization For Efficient And Progressive Image Compression

链接https://arxiv.org/abs/2609.03641

作者:Xinkun Wang,Tianyi Xu,Qingyu Luo,Mingming Ma,Changzhe Jiao,Fu Li,Yi Niu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vector-quantization based image, Vector-quantization based, achieved strong rate, achieved strong, produce a separate

备注

点击查看摘要

Abstract:Vector-quantization based image compression has achieved strong rate--distortion performance, yet most of them still produce a separate compressed representation for each target bitrate. Such variable-rate behavior allows one model to operate at multiple rates, but it does not necessarily provide a progressive bitstream whose prefixes are themselves decodable and can be refined by appending additional bits. We propose \textbf{Tree-VQ}, a progressive tree-structured vector quantization framework for learned image compression. Tree-VQ organizes discrete codewords as a hierarchical binary tree and represents each latent token by a routed root-to-leaf path. Crucially, every prefix of this path corresponds to a valid quantized representation, so shallow internal nodes serve as coarse reconstruction codes and deeper nodes provide successive refinements. This allows a compressed image to be decoded from an early prefix and progressively improved as more branch symbols are received, rather than being re-encoded for different target rates. To make this structure practical for compression, we introduce a prefix-compatible tree entropy model that codes progressive continuation decisions and routed branch refinements using only causally available decoded contexts. We further use rate-aware refinement scheduling to decide which spatial blocks should receive additional tree bits under a given prefix budget, and hierarchical prefix supervision to ensure that internal nodes are directly decodable at low rates. Experiments show that Tree-VQ achieves a superior performance--efficiency trade-off, delivering the best perceptual compression results with much fewer parameters and lower latency than competing methods.

60. 【2609.03639】Stabilizing Camera-Controlled Novel View Synthesis at Inference Time

链接https://arxiv.org/abs/2609.03639

作者:Prajwal Singh,Arjun Badola,Seema Kumari,Hajime Nagahara,Shanmuganathan Raman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:long generation horizons, pre-trained video diffusion, video diffusion models, single image, image using pre-trained

备注

点击查看摘要

Abstract:Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion models often becomes unstable under large camera motion and long generation horizons. Existing approaches commonly combine several inference-time components, making it unclear which design choices are most important for stability. We show that the main source of stability is simple. Decomposing camera motion into small autoregressive steps limits per-step geometric distortion and reduces error accumulation. A controlled camera-step study shows that performance remains stable for small motions and degrades more strongly as the per-step motion approaches $18$-$20^\circ$. We further evaluate geometry-constrained spatial attention and low-frequency appearance anchoring as supporting refinements, together with an efficient registration-free warping pipeline. Across RealEstate10K and MegaScene, CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines. The method remains effective for 56-frame generation and under substantial controlled depth corruption. These results show that careful control of camera motion at inference time can substantially improve the stability of camera-controlled novel view synthesis without retraining or modifying the diffusion backbone.

61. 【2609.03629】EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders

链接https://arxiv.org/abs/2609.03629

作者:Xinghao Wang,Dong Li,Wei Yu,Yingwei Pan,Tao Gong,Qi Chu,Nenghai Yu,Ting Yao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:remarkable generative capabilities, demonstrated remarkable generative, loosely curated training, curated training data, training data raises

备注

点击查看摘要

Abstract:Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities, yet their reliance on loosely curated training data raises pressing safety and copyright concerns. Concept erasure offers a principled remedy by removing unwanted semantics from pretrained models while preserving remaining concepts. However, existing approaches typically operate at a coarse granularity misaligned with the fine-grained, distributed nature of concept representations, leading to incomplete removal or degraded generation quality. We argue that surgical erasure fundamentally requires intervention at the level of monosemantic features, where each unit encodes a single interpretable concept. To this end, we propose EraseSAE, a novel framework that leverages sparse autoencoders to achieve surgical concept erasure in DiT-based T2V diffusion models via a principled decompose-attribute-erase pipeline. We first introduce the Partitioned Convolutional Sparse Autoencoder, which decomposes dense spatiotemporal activations into disentangled, interpretable sparse features while preserving spatiotemporal coherence. A contrastive attribution mechanism then contrasts activations from paired prompts to isolate concept-specific feature kernels. At inference, timestep-resolved spatiotemporal masks derived from the identified kernels confine erasure to regions where the target concept is active, leaving unrelated content intact. Extensive experiments across diverse diffusion models and concept erasure tasks demonstrate that EraseSAE achieves precise and robust concept removal with minimal quality degradation, substantially outperforming state-of-the-art methods. The code is available at this https URL.

62. 【2609.03615】Auditing Patient Privacy in Medical Generative Models: Scalable Memorization Detection with DeepSSIM++

链接https://arxiv.org/abs/2609.03615

作者:Antonio Scardace,Francesco Guarnera,Sebastiano Battiato,Daniele Ravì

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reproduce training samples, training samples raises, medical image synthesis, patient confidentiality, deep generative models

备注

点击查看摘要

Abstract:While deep generative models offer new opportunities for medical image synthesis and data sharing, their ability to memorize and reproduce training samples raises serious concerns about patient confidentiality. Detecting such memorization at scale remains challenging: traditional pixel-based metrics are sensitive to generation artifacts, whereas generic embedding-based metrics often lack the anatomical sensitivity required for medical data. To address this challenge, we introduce DeepSSIM++, a self-supervised similarity metric for scalable memorization auditing in medical generative models. By leveraging multi-scale feature aggregation and anatomy-preserving augmentations, DeepSSIM++ learns an embedding space where cosine similarity approximates the Structural Similarity Index (SSIM), eliminating the need for exact pixel-level registration. Compared with state-of-the-art baselines, DeepSSIM++ achieves an average Macro F1 improvement of 33 percentage points under ideal alignment and 46 percentage points under realistic spatial and intensity perturbations. Furthermore, it accelerates large-scale similarity computation by several orders of magnitude compared with analytical SSIM. By combining anatomical sensitivity and computational efficiency, DeepSSIM++ provides an open-source tool for scalable memorization auditing in medical generative AI. Code and data are publicly available at: this https URL.

63. 【2609.03602】SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

链接https://arxiv.org/abs/2609.03602

作者:Jinyang Wang,Shiwei Li,Junjian Wang,Zhiqiang Deng,Jianbin Gao,Yihang Zhao,Liu Liu,Yongjia Zhao,Jinlong Chen,Huirui Xu,Yifeng Pan,Kangwei Liu,Fan Ren,Ji Tao,Minghao Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:demonstrated strong potential, future scene dynamics, World models, learning predictive representations, scene dynamics

备注: 23 pages, 16 figures

点击查看摘要

Abstract:World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.

64. 【2609.03595】How Far Can Synthetic Data Take Thai OCR?

链接https://arxiv.org/abs/2609.03595

作者:Kunat Pipatanakul

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Thai OCR model, OCR model adapted, makes synthetic OCR, real Thai documents, real Thai document

备注: 20 pages, technical report

点击查看摘要

Abstract:We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on training granularity, with in-domain reconstruction approaching real printed supervision under page-level training (1.82% versus 1.31% median character error rate) but underperforming out-of-domain reconstruction under crop-level training (15.59% versus 5.52%). Guided by these findings, we adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero using 45,723 synthetic pages: relative to its base checkpoint, it reduces median character error rate from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting and outperforms Typhoon OCR v1 7B on all five evaluation sets, showing that synthetic-only training can be competitive.

65. 【2609.03585】xt2Thermal: Physics-Aware Thermal Image Synthesis from Textual Priors

链接https://arxiv.org/abs/2609.03585

作者:Tayeba Qazi,Brejesh Lall,Prerana Mukherjee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:datasets remain scarce, motivating extensive work, translating abundant RGB, abundant RGB images, infrared imaging offers

备注: 15 pages, 6 figures

点击查看摘要

Abstract:Thermal infrared imaging offers reliable perception in darkness and adverse weather, but thermal datasets remain scarce, motivating extensive work on translating abundant RGB images into thermal. Such translation is fundamentally ill-posed as thermal appearance is governed by surface emissivity and object temperature, neither of which is observable in the visible spectrum, so a single RGB image is consistent with many valid thermal outputs. We argue that language offers a natural means of resolving this ambiguity, and propose Text2Thermal, a framework for physics-aware thermal image synthesis from textual priors. Rather than inferring the unobservable radiometric factors from RGB, we supply them explicitly through thermally grounded captions encoding material, weather, time-of-day, and heat-emission state, and adapt a pretrained Stable Diffusion backbone to the thermal domain. Because the radiometric content is determined entirely by the prompt, Text2Thermal synthesizes thermal imagery without requiring a registered RGB image at inference; where spatial guidance is desired, an optional control signal imparts scene geometry without disturbing the prompt-specified radiometry. Experiments on M3FD, FLIR, and FMB show that Text2Thermal achieves state-of-the-art FID among thermal image synthesis methods while offering text-level control that translation-based approaches cannot provide.

66. 【2609.03572】Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

链接https://arxiv.org/abs/2609.03572

作者:Zhaoxin Fan,Tianbao Zhang,Wenjun Wu,Xiaofeng Wang,Yeying Jin,Jian Zhao,Zheng Zhu,Shuicheng Yan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:support action generation, action generation, World models offer, offer a promising, promising paradigm

备注: 14 pages

点击查看摘要

Abstract:World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficult to simultaneously achieve long-horizon anticipation and responsive, observation-grounded decision making. We present Drive-HWM, a hierarchical slow--fast world modeling framework that organizes future representation prediction and action generation at complementary temporal scales. The slow world model predicts multi-step future representations to capture extended scene evolution. To explicitly model the abundant motion dynamics in driving environments, we introduce Dynamic-Aware Latents learned through optical-flow prediction. Guided by these future representations, the fast model uses a lightweight multimodal backbone and an autoregressive expert to jointly predict the next frame and the immediate action from the latest observation. Next-frame prediction encourages the fast model to capture imminent scene evolution, while one-step action generation allows decisions to be continuously updated as new observations arrive. Extensive experiments on NAVSIM v1 and v2 demonstrate the strong driving performance of Drive-HWM. Comprehensive ablation studies further validate the effectiveness of the hierarchical slow--fast design, dynamics-aware future representations, and joint next-frame and action prediction.

67. 【2609.03569】Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

链接https://arxiv.org/abs/2609.03569

作者:Birgit Nierula,Karam Tomotaki-Dawoud,Mert Akguel,Mustafa Tevfik Lafci,David Przewozny,Anna Hilsmann,Peter Eisert,Sebastian Bosse

类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:expression analysis incomplete, real-time affective assessment, Head-mounted displays, render conventional image-based, requiring real-time affective

备注: Joint Proceedings of the ACM Intelligent User Interfaces (IUI) Workshops 2026

点击查看摘要

Abstract:Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face video with facial electromyography (EMG) from the occluded upper face to classify seven emotional categories (six basic emotions plus neutral). We introduce a synchronized multimodal dataset from 20 participants, pairing lower-face video with seven-channel upper-face EMG elicited by validated emotion stimuli. Under subject-independent test, our proposed late-fusion architecture merging convolutional visual embeddings with RBF-kernel EMG representations achieves 51% macro-F1, outperforming both image-only (41%) and EMG-only (43%) baselines. These results demonstrate that upper-face EMG provides robust complementary information under HMD-induced visual occlusion and establish a foundation for multimodal emotion recognition in naturalistic VR environments. This approach facilitates affect-adaptive applications, including communication training and therapeutic interventions. The dataset will be shared upon request under an ethical-use agreement.

68. 【2609.03563】FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

链接https://arxiv.org/abs/2609.03563

作者:Byeongjun Park,Byung-Hoon Kim,Hyungjin Chung

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generative rendering framework, generative rendering, present FlashRender, few-step generative rendering, framework that retakes

备注: Project page: [this https URL](https://byeongjun-park.github.io/FlashRender/)

点击查看摘要

Abstract:We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.

69. 【2609.03557】Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

链接https://arxiv.org/abs/2609.03557

作者:Haoyu Wang,Songchun Zhang,Haoran Li,Haoyang Huang,Zeyue Xue,Nan Duan

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:models require large-scale, video models require, resulting scene transitions, Action-conditioned video models, require large-scale visual

备注: 16 pages, 5 figures

点击查看摘要

Abstract:Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is difficult to obtain from ordinary real-world video because the actions that caused each visual change are typically unknown. We present a large-scale synthetic data production pipeline built on Unreal Engine for generating action-conditioned, multi-view video. To accommodate the different execution requirements of real-time physics and high-quality offline rendering, the pipeline executes trajectory generation and final rendering in two stages: Stage I runs real physics in PIE and records per-frame character states, control inputs, and camera states into an intermediate trajectory representation; Stage II replays those trajectories in a new engine process and renders them offline with Movie Render Queue (MRQ). Around this core, we develop a distributed production system with cache-aware task partitioning, node-local slot scheduling, automated scene screening, aesthetic and luminance filtering, partial-output recovery, asynchronous upload, and continuous cluster health monitoring. The production cluster contains 25 servers with eight NVIDIA RTX 5090 GPUs per server. From 2,384 asset packs, 429 levels were retained for production together with a pool of 40 humanoid characters. The pipeline has produced 2,691 hours of 1080p video and 6,076 hours of 720p video. We describe the system architecture, the implementation decisions that emerged from production failures, and the limitations of using perceptual quality proxies for world-model data curation. The pipeline described in this report constitutes the Unreal Engine synthetic-data production component used in EchoWM.

70. 【2609.03554】WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval

链接https://arxiv.org/abs/2609.03554

作者:Teng Guo,Xin Wang,Jiayou Xu,Keying Zhou,Jifeng Shen,Haoxin Ruan

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:demonstrated significant success, unifying representation learning, demonstrated significant, significant success, success by unifying

备注: Accepted to the 34th ACM International Conference on Multimedia (ACM MM 2026). 10 pages, 5 figures

点击查看摘要

Abstract:Generative retrieval has demonstrated significant success by unifying representation learning and search into a single sequence-to-sequence generation task. However, extending this paradigm to cross-modal retrieval reveals a critical challenge arising from the inherent information asymmetry across different modalities, such as the gap between concise text queries and dense visual candidates. This structural mismatch causes the autoregressive decoder to suffer from forced hallucination when generating identifiers via standard trie-constrained beam search, where the model is severely penalized for failing to guess fine-grained details absent from the query, allowing irrelevant candidates to hijack top rankings. To address this issue, we propose Wildcard Inference with Dynamic Expansion (WIDE). WIDE employs Adaptive Entropy Thresholding (AET) to calibrate layer-specific uncertainty boundaries offline. During the decoding generation phase, Asymmetry-aware Wildcard Decoding (AWD) detects semantic blind spots and emits wildcards instead of forced deterministic identifiers, dynamically expanding the search space without incurring log-probability penalties. Finally, Blind-Spot Re-ranking (BSR) evaluates the expanded candidate pool using a hybrid scoring mechanism that combines discrete generation confidence with continuous semantic similarity. Extensive experiments on the M-BEIR benchmark demonstrate that WIDE outperforms state-of-the-art generative retrieval methods, effectively suppressing forced hallucination while maintaining compact index structures.

71. 【2609.03544】SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models

链接https://arxiv.org/abs/2609.03544

作者:Caoyuan Ma,Tian Gu,Wenpu Liu,Weichu Xie,Shuai Dong,Yuqi Xu,Ji Zhao,Ziyue Wang,Wenzheng Chang,Taiqiang Wu,Yongfu Zhu,Wenqi Shao,Zheng Wang,Yinqiang Zheng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:model behavior globally, Existing safety alignment, Existing safety, behavior globally, vision-language models

备注: Preprint. 13 pages, 4 figures. Main paper with appendix

点击查看摘要

Abstract:Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade general multimodal capabilities. We argue that safety alignment should be an on-demand intervention rather than a permanent modification to every decoding trajectory. To this end, we propose a streaming recognition and gated LoRA framework for intrinsic VLM safety. During autoregressive generation, a lightweight recognizer estimates whether the current pre-token generation state is safe or unsafe. Its output updates the LoRA gate for the following decoding step; otherwise, generation follows the frozen-backbone policy. The LoRA module is trained from unsafe prefixes, transition statements, and safe continuations, so that it learns to redirect unsafe generations back to safe responses after activation. Experiments across multiple safety and general-purpose benchmarks demonstrate the effectiveness of our method in post-alignment settings.

72. 【2609.03534】runcGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates

链接https://arxiv.org/abs/2609.03534

作者:Theo Morales,Nhat-Quynh Le-Pham,Robin Atkins,Binh-Son Hua

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:input remains challenging, visual input remains, Gaussian Splatting, dynamic Gaussian Splatting, facto scene representation

备注: Accepted at Pacific Graphics 2026

点击查看摘要

Abstract:3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learning 3D Gaussian primitives from visual input remains challenging. Standard optimization relies on gradient-based updates, but a common issue is the gradient vanishing phenomenon: a pixel far from a Gaussian primitive often has diminishing gradient magnitudes to influence primitive attributes, resulting in suboptimal scene reconstruction. In this paper, we propose a method to address gradient vanishing with a piecewise truncated gradient formulation that improves the optimization stability and robustness to initializations. We show that our method consistently improves 3D Gaussian Splatting with random and COLMAP initializations while being generalizable across static and dynamic Gaussian Splatting. As a by-product, we also examine the limitations of current benchmarks for dynamic scenes, and introduce a novel dataset for benchmarking dynamic Gaussian Splatting using synthetic 3D scenes. We demonstrate the effectiveness of our method in both static and dynamic settings for the public benchmarks and our proposed dataset.

73. 【2609.03520】Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion

链接https://arxiv.org/abs/2609.03520

作者:Chuyue Shan,Songlin Sun,Wang Chenwei,Shen Zihan

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:conditional coding-based neural, coding-based neural video, neural video compression, directly affects compression, context directly affects

备注

点击查看摘要

Abstract:In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance. Existing methods mostly construct context from prop- agated reference features, but they are vulnerable to motion esti- mation and local alignment errors in regions with complex mo- tion, occlusion, and high-frequency textures, resulting in inaccu- rate temporal information. To address this issue, this paper pro- poses a method combining deformable temporal alignment and difference-aware spatial selective fusion. A Context-aware Tem- poral Alignment Module is used to generate complementary tem- poral context, while a Difference-aware Spatial Selective Fusion module adaptively selects reliable temporal information and sup- presses misalignment. Experiments show that the proposed method achieves certain rate-distortion performance improve- ment over DCVC-DC.

74. 【2609.03516】Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection

链接https://arxiv.org/abs/2609.03516

作者:Yue Zhao,Hua Yu,Yukun Zhao,Yuzhi Zhang,Maoguo Gong,Xin Mei,Zhuping Hu,Yanchi Li,A. K. Qin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Infrared-visible object detection, integrates complementary evidence, Infrared-visible object, object detection, integrates complementary

备注

点击查看摘要

Abstract:Infrared-visible object detection (IVOD) integrates complementary evidence from visible and infrared sensors for reliable perception in challenging scenes. In practice, sensors may fail or drop frames, leaving one modality unavailable or intermittent. Existing methods for IVOD assume both modalities are always present, and fixed fusion collapses when one stream is missing. Furthermore, it remains a critical challenge to reliably estimate semantic correlation across heterogeneous modalities, especially under spectral distribution discrepancy. We present FlexibleFusion, a unified and adaptive method that flexibly allocates integration pathways and fusion strength, operating seamlessly across complete and missing-modality regimes. At its core, the Modality-Aware Experts Collaboration (MAEC) mechanism selectively activates and aggregates cross-modal or intra-modal expert pathways. It allows cross-modal fusion when full modalities are available and falls back to self-fusion under missing conditions. Additionally, we design Residual Self-Paced Entropic Optimal Transport (RSPEOT) to align heterogeneous feature distributions from a transport perspective. Instead of relying on the fixed sparsity coefficient in standard entropic optimal transport (EOT), RSPEOT introduces a residual-driven self-paced update that prioritizes reliable matches and progressively refines harder ones. This design alleviates the additional optimization burden of standard EOT while preserving reliable semantic alignment. Comprehensive experiments under complete and missing-modality protocols show consistent performance across arbitrary modality configurations. Code will be released upon publication.

75. 【2609.03480】ree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings

链接https://arxiv.org/abs/2609.03480

作者:Alkiviadis Koukos,Spyros Kondylatos,Thomas Nord-Larsen,Lotte Nyborg,Christian Tøttrup,Kenneth Grogan

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:large-scale forest characterization, National Forest Inventory, Forest Inventory plots, Forest Inventory, Inventory plots

备注: Submitted to Remote Sensing of Environment. This preprint presents a national-scale tree species mapping framework for Denmark using Sentinel-1/2 time series, National Forest Inventory data, and EO foundation model embeddings. The resulted national map can be found here: [this https URL](https://zenodo.org/uploads/22108850)

点击查看摘要

Abstract:We map tree species across Denmark using National Forest Inventory plots and EO data, while evaluating the potential of foundation models for large-scale forest characterization. We compare two alternative input representations for tree species classification: (i) manually engineered spectral-temporal features (STF) derived from multi-temporal Sentinel-1 and Sentinel-2 observations, and (ii) embeddings generated by the EO FMs TESSERA and AlphaEarth. Both representations are complemented with canopy height information. Random forest, XGBoost, and Multi-Layer Perceptron (MLP) classifiers are evaluated for all input representations, with separate assessments for pure and mixed forest stands. The STF-based MLP achieves the highest classification performance, yielding macro F1 scores of 0.843 and 0.653 for pure and mixed stands, respectively. The MLP trained on TESSERA embeddings delivers competitive performance for pure stands, achieving results within 1.1 percentage points of the best-performing model. TESSERA consistently outperforms STF-based models when fewer than approximately 25% of training plots are available, demonstrating a substantial advantage under limited training data. Multi-year observations systematically improve classification accuracy relative to single-year inputs, while ablation experiments reveal the complementary contributions of Sentinel-1 backscatter, spectral indices, and canopy height data. The best-performing model is subsequently applied at the national scale to generate a 10 m tree species map of Denmark. Area-adjusted validation indicates an overall map accuracy of 79.9%. The resulting map, released as an open-access product, is the first high-resolution national tree species map of Denmark and provides a valuable resource for forest monitoring, ecological research, and land management applications.

76. 【2609.03475】SafeRestore: Detector-Relative Risk Certificates for Selective Industrial Image Restoration

链接https://arxiv.org/abs/2609.03475

作者:Shaoliang Yang,Jun Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Applications (stat.AP)

关键词:Industrial inspection pipelines, create clean-region activations, suppress detector-supported defect, detector-supported defect structure, Industrial inspection

备注: 27 pages, 10 figures, 6 main-text tables, and 17 supplementary tables. Shares the Carinthia-S image identities with [arXiv:2607.17401](https://arxiv.org/abs/2607.17401) by the same authors; the overlap and the distinct estimand are stated in Section 2.4 and Table 1

点击查看摘要

Abstract:Industrial inspection pipelines often restore a measured image before a detector acts on it, yet restoration can suppress detector-supported defect structure or create clean-region activations. We formulate restoration as a selective action problem over the measured display, five restored candidates, and review. SafeRestore ranks candidates with action-specific fitted scores, chooses a gate on threshold-tuning data, and evaluates the fixed gate on a disjoint certification sample with two one-sided exact binomial bounds: one for the positive-conditional evidence-loss incident rate and one for the all-accepted excess-activation incident rate. The guarantee is marginal for one policy fixed before its certification outcomes are observed, under an image-level i.i.d. working model. In a retrospective split-sample study of 4,591 public Carinthia-S images, the protocol yields auditable risk-coverage behavior. The primary all-action policy passes in one of five training repetitions (12.0% +/- 26.9% pass-gated test coverage when failures count as zero), whereas fixed bicubic and reduced-complexity variants pass more often. On reserved morphologies, evidence-loss incidence rises to 81.1-90.3%, and KolektorSDD lacks both detector competence and enough positive certification images for the stated target. The contribution is therefore an auditable, detector-relative framework for deciding when a transformed image may be returned automatically and when review remains necessary -- not a claim that adaptive routing outperforms simpler policies on the present evidence.

77. 【2609.03463】BMCTrack-d: Pig re-identification and tracking via back marks in challenging camera settings

链接https://arxiv.org/abs/2609.03463

作者:David Brunner,Maciej Oczak,Marie Bordes,Jean-Loup Rault,Stephan M. Winkler,Viktoria Dorfer

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Automated pig monitoring, Automated pig, assessing their health, essential for assessing, pig monitoring

备注

点击查看摘要

Abstract:Automated pig monitoring is essential for assessing their health, behaviour, and welfare. To date, most pig monitoring solutions operate on the group-level, because individual-level monitoring requires reliable long-term identification and tracking of each animal. For domesticated pigs this remains challenging because pigs of the same breed often have highly uniform appearances. Moreover, research on pig monitoring is almost exclusively reported in top-down view camera settings, which considerably ease tracking, but are not always an option in practice. In this work, BMCTrack-d is presented, a novel tracking-by-detection approach that leverages unique back marks to enable robust pig re-identification and tracking in a challenging side-view camera setting, afflicted by rapidly moving pigs, severe occlusions and low resolution. The method first predicts the detected pigs' identities using a neural network-based back mark classifier. To improve re-identification reliability over time, two dedicated post-processing stages are introduced: a temporal prediction consistency check, which validates the identity assignments against the recent prediction history, and deduplication, which resolves conflicting identity assignments in each time step. By explicitly prioritising accurate, appearance-based re-identification over continuous tracking, the proposed approach addresses a key limitation of existing trackers for individual-level monitoring scenarios. On a demanding test set BMCTrack-d outperforms two strong baselines, BoT-SORT-ReID and TrackTrack-ReID, by 9.11% and 1.03%, respectively, in higher-order tracking accuracy. These results demonstrate the effectiveness of back mark-based re-identification and tracking for robust individual-level pig monitoring in challenging settings.

78. 【2609.03453】Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems

链接https://arxiv.org/abs/2609.03453

作者:Jannatul Masruk Mukta,Rifa Sanjida,Adrita Rahman Tory,Md. Saifur Rahman,Khondokar Fida Hasan

类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)

关键词:standard first-line response, Preprocessing-based defenses, model-agnostic mitigations, first-line response, widely recommended

备注: 14 pages

点击查看摘要

Abstract:Preprocessing-based defenses are the standard first-line response to adversarial attacks on edge vision systems, requiring no retraining, no architectural changes, and widely recommended as model-agnostic mitigations. Yet the foundational evaluations of these defenses were conducted on residual or Inception-class architectures, not on the depthwise-separable CNNs that dominate edge deployments. This untested assumption leaves a gap in the security evaluation literature. This paper closes that gap by evaluating six preprocessing defenses against adversarial perturbations across both architecture families. Across all perturbation levels and defenses tested, the two depthwise-separable architectures show consistently poor recovery while the residual architecture shows partial recovery; ablation results are consistent with an architectural rather than parametric explanation, though only three architectures and one attack family are evaluated. Crucially, this failure is not merely a negative result. The same output divergence that disqualifies preprocessing as a recovery mechanism reveals a detection opportunity: preprocessing consistently disrupts clean predictions while leaving adversarial predictions largely unchanged, an asymmetry that is directly measurable without retraining or architectural modification. We further show that standard image quality metrics are unreliable proxies for defense effectiveness, a methodological gap in current evaluation practice. A practitioner decision framework is provided for adversarially resilient edge vision deployment.

79. 【2609.03447】STARS-GS: Structure-Aware Regularized Gaussian Splatting for Large-Scale Aerial Surface Reconstruction

链接https://arxiv.org/abs/2609.03447

作者:Bocheng Li,Wenjuan Zhang,Jie Pan.Dongxu Han,Xuesong Ma,Yiling Yao,Yaning Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:urban modeling, imagery is fundamental, fundamental to geospatial, geospatial mapping, mapping and urban

备注

点击查看摘要

Abstract:Large-scale 3D surface reconstruction from aerial imagery is fundamental to geospatial mapping and urban modeling. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated considerable potential for this task. However, existing methods still face three major challenges in large and complex scenes: scene partitioning may split continuous scene elements across independently optimized sub-regions; geometric constraints mainly focus on the attributes of individual Gaussians while overlooking their local organization; and uniform regularization struggles to accommodate heterogeneous geometric structures. To address these issues, we propose STARS-GS, a structure-aware 3DGS framework for large-scale surface reconstruction. First, we introduce a structure-aware scene partitioning strategy that better preserves continuous scene structures during partitioning and reduces cross-region geometric inconsistencies and stitching artifacts through boundary refinement. Second, we develop neighborhood-aware Gaussian organization that extends geometric constraints from individual primitives to their neighborhood organization, encouraging Gaussians to better conform to local surface geometry. Third, we introduce adaptive surface regularization that adjusts the regularization strength according to local geometric characteristics, promoting geometric consistency in structured regions while preserving plausible variations in unstructured regions. Extensive experiments on large-scale aerial photogrammetry benchmarks demonstrate that STARS-GS consistently outperforms the evaluated Gaussian-based methods in surface reconstruction. It increases the average F1-score from 0.640 for the second-best method to 0.698, corresponding to a relative improvement of approximately 9.1\%, demonstrating effective improvements in geometric accuracy and surface completeness.

80. 【2609.03446】Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

链接https://arxiv.org/abs/2609.03446

作者:Taehoon Kim,Jongwook Choi,Heejae Jo,Byungmin Park,Jongwon Choi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:deepfakes requires detectors, capture video-specific cues, forgery patterns, existing approaches, fail to capture

备注: Accepted by ECCV 2025. Code will be available at [this http URL](http://github.com/rama0126/MSFD)

点击查看摘要

Abstract:The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forgery patterns, yet existing approaches, which are designed for deepfake images, fail to capture video-specific cues. Unlike deepfake images that contain only spatial artifacts, deepfake videos leave distinct evidence along both spatial and temporal axes, necessitating the separate preservation of each modality during sequential model updates. To overcome this limitation, we introduce a continual deepfake video detection framework, Modality-Specific Frequency Distillation (MSFD), that explicitly decomposes video features into spatial, temporal, and spatiotemporal modalities in the frequency domain. This decomposition enables independent preservation of each modality, as different deepfake video types exhibit varying reliance on spatial and temporal cues across tasks. Furthermore, MSFD adopts a cross-modality decorrelation loss that encourages spatiotemporal representations to remain orthogonal to single-modality cues. Extensive experiments show that our framework achieves stronger adaptation and preserves performance more effectively than state-of-the-art methods across diverse continual deepfake video scenarios.

81. 【2609.03445】OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement

链接https://arxiv.org/abs/2609.03445

作者:Linnan Zhao,Kang Liu,Hao Yu,Jiabo Zhan,Chong Sun,Chen Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:formats remain error-prone, long-tail formats remain, document OCR systems, systems perform increasingly, OCR

备注

点击查看摘要

Abstract:Although document OCR systems perform increasingly well on routine documents, complex formulas, structured text, and long-tail formats remain error-prone. OCR predictions may omit fine-grained content or hallucinate unsupported outputs, while equivalent encodings of the same visible content must be accommodated. Existing OCR evaluation methods mostly report aggregate metrics, offering limited support for analyzing case-level errors and improving OCR performance. We propose OCR-EDR (OCR Error Diagnosis and Repair), a rendering-aware framework that advances from fine-grained diagnosis to iterative repair. Given a source image, an editable OCR prediction, and its rendered image, OCR-EDR first jointly assesses whether the prediction and its rendering are consistent with the source, preserving valid predictions, including rendering-equivalent ones, while diagnosing and localizing genuine errors. It then applies executable edits and may request an updated rendering for iterative reassessment. We construct OCRErrBench from diverse real OCR predictions, covering text and formulas, exact and rendering-equivalent positives, and genuine errors, and develop the DocEDR model to execute the diagnosis--repair loop. On OCRErrBench, DocEDR achieves 94.78% diagnostic accuracy. It repairs 86.23% of erroneous inputs to visual consistency, raises formula Case-F1 by 30.99 percentage points over DOCR-Inspector-7B on DOCRcaseBench, and improves formula CDM by up to 4.62 percentage points on the identified Bad subsets of four OCR systems on UniMER-Test. These results show that OCR-EDR turns fine-grained OCR analysis into verified corrections and performance gains.

82. 【2609.03429】When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

链接https://arxiv.org/abs/2609.03429

作者:Wonbin Son,Gyumun Choi,Junil Seo,Seungmin Rho,Mi Young Lee,Hyungjoon Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Answering what-if queries, Answering what-if, what-if queries, injecting the assumption, assumption as text

备注: 10 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Answering what-if queries about a scene with a VLM usually means injecting the assumption as text or repainting the scene with a generative model. We instead move the edit to the representation level, before the model input. The image is abstracted into a set of object-level tokens, and the original image never enters the VLM. This design rests on an open question: when do frozen VLMs actually respond to such token edits? We introduce an answer-key-free protocol: no post-edit answer is annotated. It scores edits whose answers are logically determined, and audits itself by reversing each scoreable choice. The protocol reveals three structures. The response is not free: explicit edit teaching, not ordinary VQA training, produces it in dense scenes and multiplies it in sparse ones, on all three operations. Once on, it is governed by token cleanliness and density, with deployable detector+segmenter tokens competitive with the oracle and outperforming it on VRSBench. And reading is a separable axis: the image-free token route preserves 92-96% of a matched patch-token baseline's free-text VQA, and the answers measurably depend on the tokens. The response, cleanliness, and reading structures are sign-preserved across two remote-sensing datasets (iSAID, VRSBench) and three frozen LM backbones. We release the probe generator, records, judge logs, and code.

83. 【2609.03415】Mudragen: Geometrically Supervised Generation of Interacting Two-Hand Mudras for Preserving Indian Classical Dance Heritage

链接https://arxiv.org/abs/2609.03415

作者:Jagadish Kashinath Kamble,Jayanta Mukhopadhyay,Debaditya Roy,Partha Pratim Das

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Indian classical dance, Indian classical, Samyukta Hasta Mudras, Automatic generation, classical dance

备注: Accepted for publication in ACM Journal on Computing and Cultural Heritage (JOCCH) Special Issue on Visual Heritage

点击查看摘要

Abstract:Automatic generation of hand gestures is essential for the transmission of Indian classical dance and critical for its preservation. Indian classical dance gesture datasets are inherently low-resource, and the canonical Sanskrit definitions of many mudras lack precise textual descriptions, limiting the effectiveness of conventional text-conditioned image generation models. We present \textbf{MudraGen}, a conditional diffusion framework that synthesizes realistic RGB images of \textit{Samyukta Hasta Mudras} -- interactive two-hand gestures from Bharatanatyam (an Indian classical dance form). Unlike prior work on simple hand signs or single-hand gestures, MudraGen introduces geometry-aware supervision to capture the precise coordination, anatomical validity, and cultural nuance of interacting hands. We formulate three geometry-aware objectives: Keypoint Loss for 3D joint alignment, Joint Offset Loss for inter-hand spatial coherence, and Shape Consistency, which serves as an anatomical regularizer by encouraging consistent hand morphology while allowing independent hand poses. Together, these objectives guide the diffusion model toward anatomically plausible and well-coordinated hand configurations, enabling the synthesis of photorealistic and pose-accurate gesture images. Experimental results show that MudraGen surpasses existing state-of-the-art generative approaches in visual realism, anatomical correctness, and preservation of fine hand-pose structure, enabling faithful reproduction of complex Samyukta Hasta mudras. Beyond quantitative gains, its ability to generate culturally grounded and structurally consistent gestures highlights practical applications in cultural preservation and dance education.

84. 【2609.03406】Neural-Collapse-guided Task-Free Continual Anomaly Detection

链接https://arxiv.org/abs/2609.03406

作者:Xiaotong Kong,Chaoyang Song,Ziai Zhou,Jinxia Zhang,Kanjian Zhang,Haikun Wei

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:witnessed growing interest, Recent years, industrial visual inspection, task-free continual learning, continual learning

备注

点击查看摘要

Abstract:Recent years have witnessed growing interest in continual anomaly detection for industrial visual inspection. However, real-world manufacturing environments exhibit unpredictable shifts in data distributions, rendering task-dependent continual learning assumptions impractical. To address this limitation, we formulate industrial anomaly detection as a task-free continual learning problem and propose NC-TFAD, a neural-collapse-inspired, geometry-driven framework for learning from non-stationary data streams without task boundaries. NC-TFAD freezes a pretrained backbone and aligns streaming features to a simplex Equiangular Tight Frame (ETF) prototype space to stabilize representation geometry under non-stationary streams. To satisfy the NC-inspired geometric construction in the absence of real anomalies, we generate synthetic anomaly samples as auxiliary anchors during training. Building on this geometry, we further introduce inter- and intra-class regularization together with a Focal Neural Collapse Contrastive (FNCC) loss to suppress representation drift and improve normal-anomaly separability. Finally, a normal-patch-prototype-guided localization branch constructs calibrated patch-wise deviation maps from normal training samples and fuses them with a weak self-attention prior, producing anomaly heatmaps without pixel-level annotations. Extensive experiments on MVTec AD and VisA show that NC-TFAD consistently outperforms representative task-free continual learning methods adapted from general vision, as well as unified anomaly detection baselines, in both image-level detection and pixel-level localization under the task-free continual learning protocol. These results highlight that geometry-driven modeling offers an effective and robust solution for task-free continual anomaly detection in real-world industrial applications.

85. 【2609.03391】Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data

链接https://arxiv.org/abs/2609.03391

作者:Xiangyang Miao,Kelu Yao,Yekai Huang,Xiaogang Xu,Junxiao Xue,Minjun Shen,Chenghui Lv,Shanji Liu,Yaying Chen,Chao Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:remote sensing, sensing vision-language understanding, remote sensing vision-language, Contrastive language-image learning, contrastive learning

备注: 9 pages, 4 figures, 5 tables. Submitted to AAAI 2027

点击查看摘要

Abstract:Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imaging (MSI), and hyperspectral imaging (HSI). To address this limitation, we propose OmniRSCLIP, an end-to-end contrastive learning framework that supports multi-source sensor inputs for remote sensing vision-language modeling. The key idea is to extend CLIP beyond its fixed RGB input interface without breaking the pretrained visual knowledge. To this end, OmniRSCLIP introduces Spectral-Spatial Basis Decomposition (SSBD), which formulates arbitrary-channel adaptation as a basis recomposition problem: pretrained CLIP patch embeddings provide transferable spatial bases, while wavelength-conditioned coefficients span sensor-specific embedding kernels within a constrained visual prior space. This design avoids forcing heterogeneous sensors into a fixed-channel input space, while aligning them in a unified image-text semantic space. We further introduce a spectral-context-aware mask-based contrastive learning scheme to suppress modality-specific redundant features and enhance fine-grained image-text alignment. Finally, to support multi-modal training, we construct OmniRS5M, the first large-scale remote sensing image-text corpus covering RGB, SAR, MSI, and HSI. Experiments on retrieval, zero-shot classification, and semantic localization show that OmniRSCLIP preserves strong RGB-domain performance while effectively extending CLIP to heterogeneous remote sensing modalities.

86. 【2609.03384】FoRIS: Progressive Foreground Refinement for Training-Free In-Context Segmentation

链接https://arxiv.org/abs/2609.03384

作者:Ming Hu,Jianfu Yin,Mingyu Dou,Miaomiao Zhang,Yao Wang,Cong Hu,Bingliang Hu,Quan Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:annotated visual exemplars, precisely segment arbitrary, arbitrary semantic concepts, segment arbitrary semantic, aims to precisely

备注

点击查看摘要

Abstract:In-Context Segmentation (ICS) aims to precisely segment arbitrary semantic concepts, such as objects or parts, given one or a few annotated visual exemplars. In this paper, we revisit ICS from a more classical segmentation perspective, viewing it as a coarse-to-fine progressive refinement process. Rather than directly predicting the final mask through reference-query matching, we progressively refine the segmentation from coarse and ambiguous foreground responses to precise and complete foreground structures. Building upon this perspective, we propose a training-free in-context segmentation framework, termed FoRIS. Specifically, FoRIS consists of three key stages: Foreground Purification, Foreground Localization, and Foreground Consolidation, which progressively suppress background distractions, localize discriminative target regions, and recover complete foreground structures through semantic aggregation. Experimental results demonstrate that FoRIS achieves SOTA performance across semantic and part segmentation tasks, with average improvements of 4.5 and 4.8 mIoU points over existing approaches in the 1-shot and 5-shot settings, respectively. Code: this https URL.

87. 【2609.03378】When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection

链接https://arxiv.org/abs/2609.03378

作者:Xuehao Wang,Jiaxin Hua,Runmei Li,Zhenyu Wu,Chenglizhao Chen,Ke Gu,Aimin Hao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:salient object detection, RGB-D salient object, resolve appearance ambiguity, object detection, uniformly reliable

备注

点击查看摘要

Abstract:Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Existing quality-aware approaches regulate observed depth but remain dependent on the same potentially defective modality. We propose \method, a reliability-aware geometry distillation framework developed for RGB-D SOD benchmarks without using dataset-provided depth during training or inference. A frozen Depth Anything V2 model serves only as a training-time teacher, transferring dense relative geometry, hierarchical spatial attention, and boundary structure to a compact edge-aware geometry branch. Pooled bidirectional interaction aligns geometry with appearance, and a pixel-wise reliability estimator selectively injects geometry that is compatible with the current RGB representation. The teacher is removed after training, leaving an RGB-only inference network. Trained on 2,985 RGB-mask pairs, \method{} achieves the best or tied-best result in 26 of 36 metric-dataset comparisons against ten recent RGB-D SOD methods, including a 13.4\% relative MAE reduction on ReDWeb-S. When retrained on DUTS-TR, it also improves the strongest prior $F$-measure by 4.2\% on PASCAL-S, showing that the distilled geometry transfers beyond a particular sensor or dataset domain. Code will be released upon publication.

88. 【2609.03349】P-CORE: Self-Supervised Surface Consistency for Point-Based Neural Editing

链接https://arxiv.org/abs/2609.03349

作者:Yanshu Zhang,Shichong Peng,Mehran Aghabozorgi,Alireza Moazeni,Ke Li

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:enabled high-fidelity multi-view, high-fidelity multi-view reconstruction, rendering have enabled, enabled high-fidelity, large deformations

备注: Accepted to ECCV 2026. Project Page: [this https URL](https://zvict.github.io/p-core/)

点击查看摘要

Abstract:Advances in neural rendering have enabled high-fidelity multi-view reconstruction of 3D scenes. However, free-form non-rigid shape editing remains a significant challenge. Point-based neural representations are highly desirable for multi-view reconstruction because they lack fixed connectivity, which does not constrain the learned surface topology to that of the initialization. Yet this same property causes point-based representations to struggle with holes and surface discontinuities under large deformations. To address this, we propose a novel self-supervised method to enable point-based representations to adapt to large deformations without requiring ground truth multi-view images of deformed geometry. The key idea is to generate random deformations and to ensure consistency in the predicted surface before and after deformation. In particular, the surface prediction from the deformed point cloud should be the same as the deformation applied to the surface prediction from the original point cloud. We incorporate our approach into attention-based point representations, which differ from splatting-based point representations in their use of a learned interpolation kernel between points as opposed to a Gaussian kernel around each point. This learned interpolation kernel can learn to adapt to large deformations, without requiring addition or removal of points. We show that our framework significantly enhances its robustness to large deformations. Experiments on synthetic geometry editing benchmarks (Neural Editor, Objaverse) demonstrate that our approach outperforms existing point-based methods in zero-shot editing and significantly reduces artifacts. Furthermore, qualitative results on the DTU and Mip-NeRF 360 datasets demonstrate our method's effectiveness on real-world scenes.

89. 【2609.03341】PointGT: Simultaneous Geometry and Texture Editing for Point-Based Representations

链接https://arxiv.org/abs/2609.03341

作者:Yanshu Zhang,George Shramko,Pratul P. Srinivasan,Ke Li

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Gaussian Splatting representations, Gaussian Splatting, enables simultaneous editing, editing, simultaneous editing

备注: Accepted to ECCV 2026. Project page: [this https URL](https://zvict.github.io/pointgt/)

点击查看摘要

Abstract:We present PointGT, a point-based 3D representation that enables simultaneous editing of object geometry and appearance. Existing reconstruction and view synthesis techniques produce volumetric 3D representations that are high-quality and photorealistic, but are difficult to edit. In particular, recent efforts to enable texture editing for 3D Gaussian Splatting representations are not compatible with geometry edits and deformations. Our method combines a point-based representation that is well-suited for geometry deformations with a learned UV mapping technique that enables high-resolution texture editing. We show that PointGT enables fine-grained editing of both geometry and texture in point-based neural representations with high rendering quality.

90. 【2609.03334】Laplacian Frequency Hierarchies for Efficient 3D Gaussian Splatting Training

链接https://arxiv.org/abs/2609.03334

作者:Yixiong Yang,Sisheng Zhang,Qingsong Yan,Shaohuai Shi,Qiang Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting training, Laplacian Frequency Hierarchies, increases optimization cost, Gaussian Splatting, propose Laplacian Frequency

备注: Accepted to Pacific Graphics 2026 (conference track). Project page: [this https URL](https://sorenzhang574.github.io/Laplacian-GS/)

点击查看摘要

Abstract:A key bottleneck in 3D Gaussian Splatting training is the continual growth of Gaussian primitives, which increases optimization cost and slows convergence, especially at high resolutions. We propose Laplacian Frequency Hierarchies, a simple yet efficient 3DGS scheme that combines Laplacian image decomposition with coarse-to-fine, frequency-staged training. After fitting lower-frequency structure, we archive the corresponding Gaussian field so that subsequent fields can optimize higher-frequency residuals without carrying the full primitive burden, and we compose the rendered components in the image domain via a Laplacian-style reconstruction at inference time. This design reduces the number of active Gaussians during training, thereby lowering optimization overhead and accelerating training. The proposed scheme is plug-and-play and orthogonal to prior 3DGS accelerations: it can be directly combined with strong backbones such as Taming-3DGS and FastGS to improve training speed with competitive reconstruction quality. It achieves average speedups of 1.73x and 1.21x at 1K setting, and 1.74x and 1.33x at 4K setting on Taming-3DGS and FastGS, with larger gains on more challenging scenes and increasingly pronounced benefits at higher resolutions.

91. 【2609.03302】nsor-based Brain Surface Modeling and Analysis

链接https://arxiv.org/abs/2609.03302

作者:Moo K. Chung,Keith J. Worsley,Steve Robbins,Alan C. Evans

类目:Computer Vision and Pattern Recognition (cs.CV); Neurons and Cognition (q-bio.NC)

关键词:unified computational approach, magnetic resonance images, magnetic resonance, surface shape differences, surface

备注

点击查看摘要

Abstract:We present a unified computational approach to tensor-based morphometry in detecting the brain surface shape differences between two clinical groups based on magnetic resonance images. Our approach is novel in a sense that we combined surface modeling, surface data smoothing and statistical analysis in a coherent unified mathematical framework. The cerebral cortex has the topology of a 2D highly convoluted sheet. Between two different clinical groups, the local surface area and curvature of the cortex may differ. It is highly likely that such surface shape differences are not uniform over the whole cortex. By computing how such surface metrics differ, the regions of the most rapid structural differences can be localized. To increase the signal to noise ratio, diffusion smoothing based on the explicit estimation of Laplace-Beltrami operator has been developed and applied to the surface metrics. As an illustration, we demonstrate how this new tensor-based surface morphometry can be applied in localizing the cortical regions of the gray matter tissue growth and loss in the brain images longitudinally collected in the group of children.

92. 【2609.03261】MedQA-MM: Shortcuts Behind Medical Visual Reasoning

链接https://arxiv.org/abs/2609.03261

作者:Benlu Wang,Yifan Zhang,Jiaqing Yu,Chin Siang Ong,Juncheng Huang,Zhuohao Li,Zhenyu Zhang,Arman Cohan,Hong Yu,Zonghai Yao

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:benchmark score credits, score credits final, credits final answers, benchmark score, score credits

备注

点击查看摘要

Abstract:A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

93. 【2609.03258】An Ensemble-Based Self-Taught Learning Approach for Parking Space Classification Under Limited Data

链接https://arxiv.org/abs/2609.03258

作者:Lucas de Oliveira Cunha,Joelton Deonei Gotz,Paulo Lisboa de Almeida,Andre Gustavo Hochuli

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Parking spot classification, intelligent transportation systems, deep learning approaches, learning approaches rely, Parking spot

备注

点击查看摘要

Abstract:Parking spot classification is a fundamental task in intelligent transportation systems, yet most deep learning approaches rely on large amounts of annotated data and exhibit limited generalization across heterogeneous environments. To address these limitations, we investigate a self-taught learning framework based on unsupervised representation learning with convolutional autoencoders. The proposed approach learns transferable visual representations from unlabeled data and reuses the learned encoders as fixed feature extractors for supervised classification with limited annotated samples in the target domain. To further enhance robustness and mitigate architectural bias, an ensemble of heterogeneous autoencoders is employed, with independent classifier heads and prediction fusion at inference time. Experiments conducted on the PKLot and CNRPark benchmarks under cross-dataset evaluation protocols show that the proposed ensemble-based strategy substantially reduces annotation requirements while improving robustness under significant domain shifts, achieving accuracies between 93\% and 96\% in data-constrained scenarios.

94. 【2609.03233】Counting Animals in Camera-Traps Image Sequences without Count Labels: Winning Solution to the iWildCam 2021 Challenge

链接https://arxiv.org/abs/2609.03233

作者:Fagner Cunha,Juan G. Colonna,Eulanda M. dos Santos

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:computer vision methods, wildlife monitoring, motivating the development, essential tool, tool for wildlife

备注

点击查看摘要

Abstract:Camera traps have become an essential tool for wildlife monitoring, motivating the development of computer vision methods for the automated extraction of information from these data. While most prior work has focused on species identification, many ecological applications also require estimating the number of unique individuals appearing across short image sequences. This task is particularly challenging because camera traps typically acquire bursts of images at approximately one frame per second, creating large temporal discontinuities that may make conventional multi-object tracking methods unreliable, and because manually collecting individual count annotations is prohibitively expensive. In this work, we describe the winning solution to the iWildCam 2021 Challenge, which introduced a benchmark for counting animals at the sequence level under realistic annotation constraints where count annotations are unavailable for training. Our approach, MaxBoxCount, combines a strong species classification pipeline with a simple yet effective counting heuristic based on MegaDetector detections to estimate the number of unique individuals without requiring count annotations. Code is available at this https URL.

95. 【2609.03216】ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers

链接https://arxiv.org/abs/2609.03216

作者:Ali Hojjat,Janek Haberer,Olaf Landsiedel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision Transformers, typically process, substantially less computation, classified with substantially, fixed input resolution

备注

点击查看摘要

Abstract:Vision Transformers (ViTs) typically process every image using a fixed input resolution and model width, even though many images can be classified with substantially less computation. We introduce ProgResViT, an input-adaptive ViT that performs inference progressively across multiple rounds. The first round processes a low-resolution image with a narrow subnetwork. Inference terminates when the prediction is sufficiently confident; otherwise, the model reuses the representations produced in the current round and proceeds with a higher-resolution input and a wider subnetwork to refine its prediction. As all rounds share a single backbone, we propose Progress-Conditioned Soft Gating (PSG), which conditions token fusion and layer outputs on the current round, block, and input resolution. On image classification, applying ProgResViT to DeiT yields better accuracy-compute trade-offs than adaptive-width, adaptive-depth, and dynamic-token baselines. With knowledge distillation, a DeiT-based ProgResViT achieves 84.9% top-1 accuracy, slightly exceeding the reported DeiT-III-S accuracy under a comparable evaluation setting. We show that the same design also provides favorable accuracy-compute trade-offs for self-supervised DINO representations and downstream semantic segmentation. Code is available at this https URL.

96. 【2609.03206】Learning to Zoom Efficiently with a Contrastive Curriculum

链接https://arxiv.org/abs/2609.03206

作者:Falko Helm,Iryna Gurevych

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:modern visual agents, important foundational part, involving high-resolution images, efficiently handle tasks, handle tasks involving

备注: EMNLP 2026

点击查看摘要

Abstract:Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images. Most previous methods need an extensive warm-start supervised fine-tuning phase for teaching models zoom-in. We show that this is not necessary by proposing a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT. Our InfoNCE-style reward uses a curriculum of increasingly hard negative tool calls as a contrastive training signal. Empirical experiments on $V^*$, HRBench and MME-RealWorld show that our approach is competitive while being more efficient. When used as a drop-in replacement for SFT, we even outperform all baselines. To directly measure the zoom-in ability of models, we further introduce the scalable synthetic MuffinChihuahua (MC) dataset. Each image consists of a grid with every cell either showing a muffin or chihuahua. Leveraging the MC dataset's unique region of interest labels, we find that recall is the metric that most strongly correlates the zoom-in region with final task performance. Our model and code for reproduction is publicly available under this https URL

97. 【2609.03199】RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning

链接https://arxiv.org/abs/2609.03199

作者:Howard Qian,Yiting Chen,Yunfei Xie,Kejia Ren,Podshara Chanrungmaneekul,Gaotian Wang,Bowen Wen,Chen Wei,Kaiyu Hang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:data remains expensive, learning increasingly depends, collecting robot data, robot data remains, increasingly depends

备注

点击查看摘要

Abstract:Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains expensive and poorly suited to covering the long tail of real-world tasks. To address this bottleneck, we introduce RoboTok, an internet-scale data engine that, given a query human manipulation video, retrieves manipulation-relevant human demonstrations from web videos for training dexterous robot policies. Specifically, we learn a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames. This representation enables manipulation behaviors to be compared across variations in camera viewpoint, scene appearance, and actor occlusions, while remaining compact enough for efficient search and continual indexing over internet-scale video collections. We evaluate RoboTok against existing robot-data retrieval approaches on retrieval benchmarks and downstream robot policy performance. Our results show that RoboTok retrieves more relevant manipulation demonstrations and improves downstream task success, establishing hand-pose trajectory-aware retrieval as a way to make web video a scalable and continuously growing source of supervision for robot learning.

98. 【2609.03181】Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

链接https://arxiv.org/abs/2609.03181

作者:Alejandro Barón García,Feng Wang,Emilia Garcia Casademont,Han Xiao

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:parsing model built, built to serve, document parsing model, document parsing, parsing model

备注: 15 pages, 5 figures, 8 tables. Model at [this https URL](https://huggingface.co/jinaai/jina-ocr-v1)

点击查看摘要

Abstract:We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at this https URL.

99. 【2609.03158】Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

链接https://arxiv.org/abs/2609.03158

作者:Qingchan Zhu,Weihang You,Hanqi Jiang,Changdi Yang,Tianming Liu,Geng Yuan

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Visual token pruning, Visual token, vision-language models, reduces the inference, inference cost

备注: Accepted to EMNLP 2026 main

点击查看摘要

Abstract:Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.

100. 【2609.03153】VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

链接https://arxiv.org/abs/2609.03153

作者:Wenzhuo Xu,Yuchen Zhu,Chongjian Ge,Xuan Shen,Jing Shi,Jason Kuen,Yongxin Chen,Molei Tao,Christopher McComb,Noelia Grande Gutiérrez,Jiuxiang Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scalar quality score, imply physical reliability, Visual fluency, typed physical obligations, moment it fails

备注

点击查看摘要

Abstract:Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed. During execution, observations gate and scope only declared calls to frozen low-level experts (e.g., segmentation and tracking, counting, eleven typed physical measurements over the resulting tracks, depth, OCR, and audio-event detection). Each action returns a provenance-carrying evidence record whose payload, when usable, is either a typed measurement or an explicitly tagged learned state. Typed resolvers and fixed composition map usable records to a three-valued state (supported, contradicted, or unknown, surfaced as plausible, implausible, or abstain) with full provenance, so that every verdict is traceable to the evidence that produced it. We anchor evaluation in a 1,500-clip corpus of human-annotated flaw records that localize real generation failures in prompt reference, space, and time. On a 149-clip core carrying 304 such records, VeriPhy accounts for 228, against 164 for a published question-decomposition evaluator given the same clips and the same claims. Recall alone does not separate it from prompting the same backbone monolithically, which reaches 222; what separates them is that each decision retains its evidence record and provenance, making the traces auditable one verdict at a time and usable as the interface through which a critic verdict could be written back into generation.

101. 【2609.03142】Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies

链接https://arxiv.org/abs/2609.03142

作者:Yue Yang,Diego Romeres,Chiori Hori,Gedas Bertasius,Daniel Szafir,Siddarth Jain

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:multimodal sensory inputs, policies fuse multimodal, fuse multimodal sensory, homogeneous robot demonstrations, robot demonstrations encourages

备注

点击查看摘要

Abstract:Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).

102. 【2609.03139】Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers

链接https://arxiv.org/abs/2609.03139

作者:Ahmed Abdelnaby,Mohamed Elmahallawy

类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)

关键词:real-world vision systems, Deep neural networks, neural networks, vision systems, increasingly deployed

备注

点击查看摘要

Abstract:Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted misclassification while preserving high clean accuracy. Existing defenses often rely on model internals, training data, or clean validation samples, making them difficult to deploy when only black-box access to a trained model is available. We propose TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples. The key insight behind TRIM is to identify image regions that are responsible for anomalous model behavior and purify only those regions while preserving benign content. TRIM innovates via three key components: (i) region-based segmentation with deep feature representations, (ii) adaptive trigger discovery through inpainting and diffusion-based reconstruction to isolate regions responsible for misclassification---without assumptions about trigger type, shape, or location, and (iii) selective region purification that cleans poisoned regions while retaining benign content. To support practical deployment, TRIM further caches feature embeddings of previously identified triggers, enabling efficient recognition and avoiding redundant detection and purification. Extensive experiments across diverse datasets and backdoor types, including blended, sparse, varying-size, and multiple triggers, show that TRIM consistently outperforms existing black-box defenses, reducing attack success rates (ASR) to as low as 1.16% while preserving clean accuracy of up to 87.87%. These results demonstrate that effective backdoor mitigation is possible at inference time even when the defender has no access to any auxiliary data.

103. 【2609.03117】Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields

链接https://arxiv.org/abs/2609.03117

作者:Amir Mallak,Alaa Maalouf,Lior Wolf,Daniela Rus,Dan Rosenbaum

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:map continuous coordinates, observations remains difficult, Classical Neural Tangent, Neural Tangent Kernel, fast high-quality reconstruction

备注: Published in IEEE TPAMI, vol. 48, no. 9, pp. 10940-10957, Sep. 2026. Author version adds related-work references and biography updates; Figures 12 and 13 were regenerated from the same locked hyperparameter sweep. Tabulated results, reported best points, scientific claims, and conclusions are unchanged

点击查看摘要

Abstract:Neural fields (NFs) map continuous coordinates to signals such as color or density, but fast high-quality reconstruction from sparse observations remains difficult. Classical Neural Tangent Kernel (NTK) regression gives closed-form fits, yet it is fundamentally linear and cannot accumulate reusable task priors. We develop three algorithms that address these gaps. NTK-KIP learns a distilled support set of coordinates (and optional labels) so that a finite NTK can inpaint large missing regions from little observed data, yielding a compact non-linear representation instead of a raw kernel solve. MetaQuill meta-learns a shared initialization for an INR so that new scenes can be adapted by updating only a small task-specific weight offset, which provides true feature learning and a reusable prior. Finally, MetaQuill-KIP fuses both ideas: it seeds the task with a KIP-style non-linear warm start, then refines only that small offset around the meta-learned initialization. MetaQuill-KIP achieves high-PSNR reconstructions and semantically plausible inpainting under very sparse observations, while requiring only lightweight per-instance adaptation, whereas diffusion-style baselines typically depend on large pretrained generative priors and costly per-image tuning. This shows that NTK-driven neural fields can be made both non-linear and meta-learnable, narrowing the gap between analytic kernels and practical few-shot reconstruction.

104. 【2609.03109】SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

链接https://arxiv.org/abs/2609.03109

作者:Haozhen Zheng,Fulin Wang,Tianhu Xiong,Yingjie Yu,Shengyi Qian,Hanchao Yu,Alex Schwing,Klara Nahrstedt,Mingyuan Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:compellingly describe slides, agents compellingly describe, AI-assisted slide editing, compellingly describe, Current

备注: EMNLP 2026 Main Conference. Haozhen and Fulin contributed equally

点击查看摘要

Abstract:Current AI agents compellingly describe slides. However, AI-assisted slide editing requires more than understanding: the output must retain layout, style, component structure, and native editability. Towards, AI-assisted slide editing, existing agents operate on screenshots or weak document representations and often fragment coherent visual units, rasterize editable content, or break layout. In contrast, for controllable slide editing, we introduce an agentic framework, SLIDEFORGE, which builds a Deck State Graph, an executable slide state that links visual decomposition, native pptx object structure, and perceptual organization. By recovering human-referable components while retaining fine-grained editable structure, SLIDEFORGE supports theme-preserving reconstruction through slide-native operations and rendered-state verification. We further introduce an evaluation paradigm for controllable slide transformation that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability. Experiments show that SLIDEFORGE outperforms direct prompting, screenshot-based agents, and generic code-agent baselines across these dimensions. Code is available at this https URL.

105. 【2609.03102】WireSeg-32K: A Physics-Grounded Synthetic Dataset for Wire Instance Segmentation

链接https://arxiv.org/abs/2609.03102

作者:Zilin Dai,Lehong Wang,Yi Yang,Xiang Fei

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Deformable linear objects, highly deformable, large-scale instance-level annotations, Deformable linear, frequently self-occluded

备注: Synthetic Data for Computer Vision @ CVPR 2026 Workshop

点击查看摘要

Abstract:Deformable linear objects such as wires and cables are difficult to segment because they are thin, highly deformable, and frequently self-occluded, while large-scale instance-level annotations are expensive to obtain in real scenes. Existing resources either focus on cable tracing or semantic segmentation under constrained settings, or generate visually plausible images without physically grounded wire deformation. We present WireSeg-32k, a synthetic dataset for wire instance segmentation with 32,000 RGB images, instance masks, depth maps, and a complementary real-world test set with annotations. To generate this dataset, we develop DeformX, a co-simulation pipeline that couples Cosserat-rod dynamics with photorealistic Isaac Sim rendering, enabling physically plausible, contact-consistent wire shapes, CAD-based wire assets, and diverse visually grounded scenes. As a simple baseline, LoRA fine-tuning SAM3 on WireSeg-32k alone improves real-world mAP@75 by 10.2% over the off-the-shelf model, showing that physically grounded synthetic data can transfer to real wire perception.

106. 【2609.03085】Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset Sampling

链接https://arxiv.org/abs/2609.03085

作者:Young Seok Jeon,Beatrice Brown-Mulry,Rohan Satya Isaac,Anjana Dissanayaka,Theo Dapamede,Mohammadreza Chavoshi,Judy Gichoya,Hari Trivedi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:adopting CLIP-style vision, growing interest, interest in adopting, adopting CLIP-style, language model

备注

点击查看摘要

Abstract:There is growing interest in adopting CLIP-style vision--language model (VLM) pretraining for mammography. However, models that directly employ the standard CLIP architecture and training objective exhibit limited zero-shot performance in clinically important tasks such as cancer, finding-type, and BI-RADS predictions. We argue that this underwhelming performance is due to neglecting two characteristics of mammography data: (1) its high-res nature, and (2) homogeneity of radiology reports, largely driven by a predominance of negative/benign findings on examinations. We propose TopKSigLIP, a VLM designed to address these two limitations through a novel architecture and learning objectives. Instead of downscaling high-res mammography images to satisfy GPU memory constraints, TopKSigLIP introduces TopK-Patch module that learns to sample a sparse set of high-res patches likely to contain lesions, sidestepping the resolution--batch size tradeoff of VLM training. The sampled patch locations additionally serve as a built-in localization tool. To address report homogeneity, we replace the contrastive loss, which falsely repels semantically similar pairs, with a Sup-sigmoid loss. Sup-sigmoid loss extends the sigmoid loss from SigLIP with soft labels derived from structured data. TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation. TopKSigLIP remains competitive under linear probing despite using a significantly smaller vision encoder and smaller training batches than baselines. The TopK-Patch module additionally achieves superior lesion localization over post-hoc Grad-CAM. Code and weights are made public:this https URL.

107. 【2609.03080】Exemplar: Classical Priors Complement Frozen Features for Few-Shot Microscopy Segmentation at Native Resolution

链接https://arxiv.org/abs/2609.03080

作者:Michal Průšek,Adam Novozámský,Filip Šroubek

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:foundation model steered, domain-specific model trained, domain-specific model, foundation model, model steered

备注: 5 pages, 3 figures, 2 tables. Code, configurations, and the per-image score records behind every reported number: [this https URL](https://github.com/michalprusek/Exemplar)

点击查看摘要

Abstract:Segmenting a new biomedical dataset usually means a domain-specific model trained on substantial annotation, or a foundation model steered at inference time. We present Exemplar, a few-shot segmenter that fuses a frozen DINOv3 backbone with a fixed bank of classical native-resolution filter responses in one lightweight head, fitted from the support masks alone. In the few-mask, native-resolution regime, classical priors and frozen self-supervised features are complementary: fused in one head, a single fixed configuration spans eleven biomedical imaging datasets. Under the same head, the classical bank alone reaches 0.693 on the eleven-dataset panel, scored by foreground intersection-over-union or centreline Dice, and the frozen features alone 0.672; the bank leads on seven of the eleven and the features on the rest, and fused they reach 0.782. Against five forward-pass few-shot methods, Exemplar leads in 54 of 55 method-dataset comparisons, 52 of them significant after Holm correction. From a single annotated mask it reaches 0.703 on the same panel, against 0.682 for a from-scratch nnU-Net trained on that same mask. At eight masks nnU-Net overtakes it on the panel mean, chiefly on centreline agreement, but takes 16-77x longer to fit.

108. 【2609.03077】Position: Unlabeled IS NOT Equal to No Human Supervision in Visual Learning

链接https://arxiv.org/abs/2609.03077

作者:Dong Lao

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:position paper argues, absence of labels, imply the absence, community to identify, identify sources

备注: ICML 2026

点击查看摘要

Abstract:This position paper argues that the absence of labels does not imply the absence of human supervision in visual learning, and urges the research community to identify sources of supervision more explicitly. Many recent methods in computer vision build upon representations learned from large-scale unlabeled data, and are therefore grouped under the same umbrella term ``unsupervised.'' However, different data curation schemes and training objectives embed substantially different human priors on which models rely, and we argue that one ``unsupervised'' umbrella term is no longer capturing these distinctions. This ambiguity makes it harder to compare unsupervised learning research conducted under different assumptions, coinciding with a sharp decline in papers titled with ``unsupervised'' in flagship computer vision conferences since 2021, despite continued growth of the field. While we fully embrace pre-training as a strong foundation for modern computer vision, we advocate for a community-level effort toward greater conceptual clarity: authors are encouraged to disclose priors in data selection and learning objectives, and to specify which components of a learning pipeline depend on which assumptions. Standardized disclosure practices can improve academic communication, ensure fairer comparisons, and preserve methodological diversity in unsupervised learning.

109. 【2609.03052】IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report]

链接https://arxiv.org/abs/2609.03052

作者:Lulu Xie,Yancheng Wang,Kanchan Chowdhury,Rolando Garcia,Yingzhen Yang,Jia Zou

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:services move online, move online, trust institutions, services move, governments must verify

备注

点击查看摘要

Abstract:As services move online, trust institutions such as banks, lenders, and governments must verify the identity of remote users. Fraud detection tools are widely available, but evaluating and fine-tuning them remains difficult because identity documents are sensitive and therefore scarce. Synthetic data generation offers a path forward, and demand is clear: our prior work in this area has been downloaded over $11{,}000$ times (aggregated from eight parts). We introduce IDSpace, extending this line of research in three directions. First, we propose model-guided Bayesian optimization, which tunes generation parameters to maximize both visual similarity and prediction consistency with target-domain models given only a few samples from a target domain. Second, we decouple user-specified metadata (demographics, fraud patterns, capture device) from automatically tuned control parameters (font styles, noise levels, image quality), allowing users to configure evaluations without low-level expertise. Third, we expand beyond template images to support scanned and mobile-captured documents. Experiments show IDSpace improves evaluation consistency by $15-45\%$ over baselines including CycleGAN, diffusion inpainting, and non-guided optimization, using only a few real samples, while improving training accuracy by up to $9\%$ and SSIM similarity with the target domain by $10\%$. We also released a new dataset consisting of $359{,}240$ high-quality synthetic documents across ten European ID types.

110. 【2609.03186】Improving Clinical Target Volume Segmentation Accuracy using Anatomical Priors and Active Learning for the AGITG TOPGEAR Clinical Trial

链接https://arxiv.org/abs/2609.03186

作者:Phillip Chlap,Mark Lee,Trevor Leong,Matthew Field,Jason Dowling,Hang Min,Julie Chu,Jennifer Tan,Phillip K. Tran,Tomas Kron,Annette Haworth,Martin A. Ebert,Shalini K. Vinod,Lois Holloway

类目:Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV)

关键词:Training deep learning-based, Clinical Target Volume, limited curated datasets, deep learning-based medical, learning-based medical image

备注

点击查看摘要

Abstract:Training deep learning-based medical image segmentation models is challenging with limited curated datasets. For AGITG TOPGEAR, a gastric cancer trial, the Clinical Target Volume (CTV) is complex and defined by multiple anatomical landmarks, making upfront training data preparation difficult for an automated contour QA segmentation model. We investigate anatomical priors, derived from surrounding organ segmentations, to provide spatial context and improve TOPGEAR CTV segmentation accuracy. We also evaluate active learning, iteratively expanding the training dataset by selecting cases expected to improve performance. One hundred TOPGEAR CT scans were retrospectively analyzed. An initial set of 10 expert-contoured cases was used to train an nnU-Net model. TotalSegmentator generated a voxel-wise anatomical prior map from surrounding structures as an additional input channel. Active learning was simulated over four iterations, selecting cases by model uncertainty and segmentation performance. All models used five-fold cross-validation for an ensemble uncertainty measure. Evaluation used a hold-out testing set of 50 cases. The anatomical prior improved CTV segmentation accuracy, increasing mean Dice Similarity Coefficient (DSC) from 0.84 to 0.86. Active learning similarly improved performance to 0.86, with greatest benefit in the final round. Combining the anatomical prior with active learning achieved the highest accuracy, with a DSC of 0.87. Model uncertainty correlated with DSC, supporting its use in identifying suboptimal predictions and guiding active learning. Anatomical priors and active learning each improved CTV segmentation accuracy and generalizability, with their combination achieving the best performance, supporting integration into segmentation model development for automated contour QA in radiotherapy clinical trials.

Subjects:

Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.03186 [physics.med-ph]

(or
arXiv:2609.03186v1 [physics.med-ph] for this version)

https://doi.org/10.48550/arXiv.2609.03186

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Phillip Chlap [view email] [v1]
Wed, 2 Sep 2026 22:03:11 UTC (1,816 KB)

111. 【2609.03095】Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief

链接https://arxiv.org/abs/2609.03095

作者:Robert Engel

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)

关键词:mobile care workflows, Smartphone skin photographs, Semantic Tri-view Pipeline, remains a critical, care workflows

备注

点击查看摘要

Abstract:Smartphone skin photographs are indispensable to teledermatology, yet assessing the diagnostic suitability of submitted cases (gradability) remains a critical bottleneck in mobile care workflows. Dermatologists routinely review multiple photographic views (regional, angled, and close-up) to identify consistent textural detail rather than relying on a single image. We present the Semantic Tri-view Pipeline, an interpretable architecture for automated teledermatology gradability screening that formalizes epidermal micro-relief as a computable biomarker of image quality. Using an expert-annotated subset of the public SCIN dataset, we train a lightweight DeepLabV3+ model to segment micro-relief fidelity. These spatial masks are then aggregated across up to three case views with a logistic regression classifier, leveraging viewpoint redundancy to support robustness under uncontrolled smartphone acquisition. This approach learns context-aware, clinically intelligible heuristics, such as penalizing high-fidelity texture in regional distance views. Evaluated at a predefined 90% sensitivity operating point, the system's apparent errors largely reflect subjective clinical variance on borderline cases where clinicians rely on non-visual metadata. On SCIN, performance improves from an AUC of 0.81 (80.6% PPV) on variance-heavy majority-consensus cases to 0.96 (97.7% PPV) on optically unambiguous unanimous cases. Overall, this work delivers an interpretable, privacy-by-design, edge-ready system that can provide real-time feedback during case submission to filter ungradable photo sets before review.

112. 【2609.02969】Learning from Scarce Labels: Multi-View Echocardiography for Ejection Fraction Prediction

链接https://arxiv.org/abs/2609.02969

作者:Zhiyuan Gao,Dominic Yurk,Yaser S. Abu-Mostafa

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:ventricular ejection fraction, predicting left ventricular, left ventricular ejection, ejection fraction, parasternal long-axis

备注: Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) [this https URL](https://melba-journal.org/2026:025)

点击查看摘要

Abstract:We present, to the best of our knowledge, the first publicly available resource for predicting left ventricular ejection fraction (EF) from parasternal long-axis (PLAX) echocardiography. Because no PLAX-EF datasets previously existed, our work focuses on an innovative data generation strategy to overcome this scarcity. By leveraging a time-based correlation between clinical notes and echocardiographic videos, combined with fine-tuning view classifiers and proxy labeling, we created a labeled dataset of over 25,000 PLAX videos. This enables us to train the first reproducible PLAX EF model, achieving a mean absolute error (MAE) of 6.86%. Given that apical four-chamber (A4C) methods, the clinical standard, report MAE values of 6%-7%, our results demonstrate that EF estimation from PLAX views is both feasible and clinically relevant. This surpasses the performance of existing methods and provides a clinically relevant solution for situations where apical views may not be feasible. Going further, we demonstrate that combining PLAX and A4C predictions via simple unweighted late fusion improves both single-view baselines to a 6.37% MAE, underscoring the value of multi-view integration. To promote continued research, we release the dataset labels, trained models, and runnable demos on GitHub, Hugging Face, and Google Colab: this https URL