本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。
统计
今日共更新1284篇论文,其中:
- 自然语言处理180篇
- 信息检索30篇
- 计算机视觉225篇
自然语言处理
1. 【2609.40361】Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
链接:https://arxiv.org/abs/2609.40361
作者:Tian Xia,Minghao Liu,Yiqing Liang,Laixi Shi,Jiayun Wang
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:adaptation pipelines remain, pipelines remain anchored, advancing clinical diagnosis, rapidly advancing clinical, Multimodal large language
备注:
点击查看摘要
Abstract:Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.
2. 【2609.40360】Semifactual Credit-Augmented Policy Optimization
链接:https://arxiv.org/abs/2609.40360
作者:Junshu Pan,Zhizhang Fu,Shulin Huang,Yiran Ding,Zifan Cheng,Wenqi Shao,Qiaosheng Zhang,Yue Zhang
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:predictions remain sensitive, task-irrelevant prompt features, Reinforcement learning, large language models, Relative Policy Optimization
备注:
点击查看摘要
Abstract:Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at this https URL.
3. 【2609.40340】EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
链接:https://arxiv.org/abs/2609.40340
作者:Young-Jun Lee,Jinheon Baek,Soyeong Jeong,Minki Kang,Seungyeon Jwa,Jonghyun Choi,Seungho Han,Dongyeop Kang
类目:Computation and Language (cs.CL)
关键词:progress requires external, large language models, requires external knowledge, large language, stall when progress
备注: Project page: [this https URL](https://open-galapagos.github.io/evoduet_project_page/)
点击查看摘要
Abstract:Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.
4. 【2609.40322】MatLoom: Layered Text-to-Material Generation in a Compact Program Space
链接:https://arxiv.org/abs/2609.40322
作者:Anson Y. Lam,Shuqing Li,Michael R. Lyu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
关键词:rules that construct, Abstract, pretrained language models, generation should produce, generation
备注: 27 pages, 8 figures
点击查看摘要
Abstract:Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.
5. 【2609.40316】Scaling Laws for Looped Mixture of Experts
链接:https://arxiv.org/abs/2609.40316
作者:Yanbei Chen,Anirudh Goyal,Raghuraman Krishnamoorthi
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:increases computational depth, expands total capacity, fixed active compute, sparsity expands total, recurrence increases computational
备注: 19 pages
点击查看摘要
Abstract:Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.
6. 【2609.40295】How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
链接:https://arxiv.org/abs/2609.40295
作者:Jenna Russell,Ben Glickenhaus,Katherine Thai,John Wieting,Mohit Iyyer,Max Spero,Bradley Emi
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:text, Web text makes, human, human text, Web text
备注:
点击查看摘要
Abstract:Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at this https URL.
7. 【2609.40286】Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
链接:https://arxiv.org/abs/2609.40286
作者:Tyler Skow,Shravan Chaudhari,Rama Chellappa,Abhay Yadav
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:reopen seemingly forgotten, seemingly forgotten knowledge, requested answer language, guarantee its removal, changing the query
备注:
点击查看摘要
Abstract:Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
8. 【2609.40284】cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
链接:https://arxiv.org/abs/2609.40284
作者:Pranjal Aggarwal,Lawrence Keunho Jang,Sean Welleck,Daniel Fried,Ruslan Salakhutdinov,Jing Yu Koh
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:including difficult long-horizon, recently surpassed human, difficult long-horizon tasks, graphical user interfaces, surpassed human performance
备注:
点击查看摘要
Abstract:Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at this https URL.
9. 【2609.40241】Decision-Oriented Recommendation Reranking: An Empirical Study of Jev
链接:https://arxiv.org/abs/2609.40241
作者:Hanjia Lyu,Yinglong Xia
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:Large language models, Large language, shown promise, introduces an important, important tradeoff
备注:
点击查看摘要
Abstract:Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate whether a decision-oriented model provides a useful alternative when the reranking task is fundamentally a structured choice among predefined candidate items. Specifically, we conduct a controlled empirical study of Jev, described by TypeSafe AI as a ``System One Model,'' for personalized recommendation reranking and compare it with recommendation-specific models and pointwise and listwise Qwen rerankers across multiple Amazon Reviews domains and candidate-set sizes, evaluating both recommendation effectiveness and observed serving latency. Our results show that Jev maintains strong recommendation effectiveness relative to the evaluated baselines while exhibiting substantially more gradual latency growth than the pointwise Qwen rerankers, although its observed serving latency remains substantially higher than that of recommendation-specific models. Together, these characteristics place Jev in a distinct quality--latency operating regime across candidate sizes and domains. These findings motivate further investigation of decision-oriented models for recommendation and other ranking tasks with structured output spaces.
10. 【2609.40236】Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports
链接:https://arxiv.org/abs/2609.40236
作者:Aawez Mansuri,Kush Mehta,Mohammadreza Chavoshi,Jahanzaib Malik,Theodorus Dapamede,Frank Li,Rohan Isaac,Beatrice Brown-Mulry,Chiratidzo Rudado Sanyika,YoungSeok Jeon,Judy W. Gichoya,Ali Emami,Hari Trivedi
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Converting free-text radiology, structured labels supports, labels supports cohort, strongest label extractors, supports cohort building
备注:
点击查看摘要
Abstract:Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT reports, and which ingredients matter. Using a 2x2 design, we crossed two adaptation strategies (a discriminative classification head, CH; generative instruction fine-tuning, IFT) with two training-data sources (distillation of real GPT-4o-labeled reports; synthetic reports generated by GPT-4o from real exemplars), across five training sizes, benchmarked on 100 expert-adjudicated reports against GPT-4o and the un-tuned open-weight base. The distilled instruction-tuned model (DIFT) matched GPT-4o (macro-F1 0.845 vs 0.850; p = 1.000) and exceeded the base model by 0.178. The decisive factor was the training-data source, not the fine-tuning method: both synthetic-data models failed to exceed the un-tuned open-weight base at any training size and underperformed the distilled models across all acuity classes. Fine-tuning and inference fit within the memory envelope of a single 24 GB consumer GPU. For narrow, high-value clinical label-extraction tasks, distilling real reports, rather than generating synthetic ones, is what closes the gap to a hosted model, enabling a private, low-cost, version-stable on-premises alternative.
11. 【2609.40235】Distribution Matching Distillation for Continuous Diffusion Language Models
链接:https://arxiv.org/abs/2609.40235
作者:Paul Le Van Kiem,Dario Shariatian,Umut Simsekli,Alain Durmus
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML)
关键词:language models generate, diffusion language models, Continuous diffusion language, network evaluations, language models
备注:
点击查看摘要
Abstract:Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
12. 【2609.40221】PhantomEnvironments: Training LLM Agents in Fictional Worlds
链接:https://arxiv.org/abs/2609.40221
作者:Anmol Kabra,Swathi Saravana Selvam,Albert Gong,Chao Wan,Christian Belardi,Dongyoung Go,Katie Z. Luo,Kilian Q. Weinberger
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:provide verifiable rewards, support long-horizon interaction, reinforcement learning, verifiable rewards, support long-horizon
备注:
点击查看摘要
Abstract:Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
13. 【2609.40198】SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models
链接:https://arxiv.org/abs/2609.40198
作者:Kanpat Vesessook,Saksorn Ruangtanusak
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)
关键词:solve tasks, tasks whose requirements, requirements emerge, SCBX Innovation Lab, sharded
备注: Conducted during a 2024 internship at SCBX RD
点击查看摘要
Abstract:Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (full), its concatenated information shards delivered together (concat), and incremental spoken disclosure across turns (sharded). We report final-answer accuracy for four commercial speech systems and LEGO, a proprietary speech pipeline developed internally by the SCBX Innovation Lab team with explicit conversational context management. Relative to concat, sharded accuracy decreases by 5.0-25.3 percentage points across the four commercial systems. LEGO achieves 77.5 percent accuracy in all three conditions, compared with 76.6 percent sharded accuracy for GPT-4o Realtime. The two single-turn baselines distinguish sensitivity to problem reformulation from the additional challenges introduced by incremental spoken interaction.
14. 【2609.40195】MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
链接:https://arxiv.org/abs/2609.40195
作者:Guangzhi Xiong,Xinyuan Zhang,Xiao Yang,Hyokun Yun,Kai Zhang,Shiun-Zu Kuo,Hyeonjeong Ha,Xilun Chen,Kai Sun,Lucas Liang,Guangqiang Dong,Ejaz Ahmed,Ahmed A Aly,Anuj Kumar,Raffay Hamid,Aidong Zhang,Xin Luna Dong
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Long-term egocentric video, Long-term egocentric, egocentric video enables, video enables personalized, daily life
备注:
点击查看摘要
Abstract:Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.
15. 【2609.40190】Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves
链接:https://arxiv.org/abs/2609.40190
作者:Sohail(Neel)Sarkar,Shakuntala Baichoo
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Statistics Theory (math.ST); Machine Learning (stat.ML)
关键词:verifier scores highest, answers, buy accuracy, Sampling, curve
备注: 32 pages, 10 figures, 5 tables
点击查看摘要
Abstract:Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within $\pm1/32$ at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions, and an audit that revisits every question need not pay for it. We derive the minimax cost of certifying the whole curve, up to logarithmic factors. It has three parts: calibrating the tail of the score distribution, telling the questions apart, and within-question noise summed along the curve. At a single benchmark the last part sharpens to the variance of one answer's influence under the best allocation of answers to questions, which every valid audit pays and an audit that learns the allocation attains, up to a logarithm, as the precision grows. A paired audit built on an exponential inequality for two independent draws at the same question needs no pilot. On 185 held-out score pools it uses 0.74 times the answers of the cheapest competing certified audit at 64 budgets and 0.53 times at 1,024, and on a newly generated MMLU-Pro study it certified the curve with 79,133 answers, within 0.6% of what a cost law fitted beforehand predicted from the study's within-question variance. The same paths certify pass@$k$ and majority voting, and the bands extend to populations of questions and to answers that depend on earlier ones.
16. 【2609.40185】Provably Tractable NFA-Constrained Language Generation via HMMs
链接:https://arxiv.org/abs/2609.40185
作者:Jialiang Sun,Kuldeep Meel
类目:Computation and Language (cs.CL); Formal Languages and Automata Theory (cs.FL)
关键词:Constrained generation aims, language models, conditioned on hard, Constrained generation, aims to sample
备注:
点击查看摘要
Abstract:Constrained generation aims to sample from language models (LMs) conditioned on hard constraints. Existing constrained-generation techniques for nondeterministic finite automaton (NFA) constraints either distort the distribution or sacrifice efficiency. Theoretically, this task reduces to counting the length-$n$ sequences accepted by an NFA (#NFA), and the exact #NFA problem is #P-complete. Recent work has shown that #NFA admits a fully polynomial randomized approximation scheme (FPRAS). Inspired by this result, we propose NFA-LM, a polynomial-time engine for NFA-constrained generation with theoretical guarantees under mild assumptions. Experiments show that NFA-LM efficiently generates high-quality outputs with theoretically bounded approximation error.
17. 【2609.40181】Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation
链接:https://arxiv.org/abs/2609.40181
作者:Tianjiao Li,Mengran Yu,Chenyu Shi,Lusheng Zhang,Qisi Chen,Yanshan Zhou,Ji Qi,Jingying Liu,Yuang Feng,Ziang Cui,Tianxing Yan
类目:Computation and Language (cs.CL)
关键词:shared multilingual foundation, translation, combines a shared, foundation with specialized, specialized training
备注: 27 pages. Project: [this https URL](https://index-translate.bilibili.com) ; Code and models: [this https URL](https://github.com/bilibili/Index-Translate)
点击查看摘要
Abstract:We introduce Index-Translate, a multilingual translation model family that combines a shared multilingual foundation with specialized training for general translation, instruction following, speech translation, controlled dubbing, and long-document translation. It includes three model sizes, 2B, 9B, and 35B-A3B, and supports translation in 150 languages, with multilingual instruction following. Evaluations on general translation and complex translation instructions show that Index-Translate outperforms translation models of comparable size and achieves performance comparable to 100B-scale translation models and frontier models. Index-Echo provides end-to-end speech-to-text and speech-to-speech translation, outperforming existing end-to-end models and achieving performance comparable to frontier omni models. Index-Homura extends the family to syllable-controlled dubbing. Index-NativeLong introduces native long-document translation with a dedicated task formulation and benchmark. These capabilities support diverse translation tasks, including multilingual content production.
18. 【2609.40127】Learning Functional Subspaces for Neural Network Compression
链接:https://arxiv.org/abs/2609.40127
作者:Massimo Bini,Anders Christensen,Stephan Alaniz,Judah Goldfeder,Ole Winther,Yann LeCun,Ravid Shwartz-Ziv,Zeynep Akata
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Modern transformers pair, transformers pair impressive, pair impressive capabilities, Modern transformers, compute demands
备注:
点击查看摘要
Abstract:Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standard hardware. Existing methods, however, choose the subspace to remove from each weight matrix with local closed-form criteria: activation energy, layer-wise reconstruction error, or a quadratic approximation of the loss. These criteria ignore how errors propagate through the network, so at high compression the errors compound with depth and performance collapses. We introduce Learnable Subspace Projections (LSP), which instead learns the subspaces to discard end-to-end. Each linear layer, or tied group of layers that read the same activations, is assigned an orthogonal projector. All projectors are optimized jointly against a global objective--the KL divergence to the dense model's output distribution or the model's original training loss--while the pretrained weights remain frozen. Projectors are initialized from a whitened SVD truncation, and ranks are allocated by the output KL each projector induces per parameter saved. After training, the projectors merge into standard low-rank factors, with each tied group sharing one factor. In attention, this also lets the model cache one narrow latent in place of full keys and values. Across LLMs (OPT-125M/1.3B, Qwen3-4B, Llama-2-7B) and ViT-B/16, LSP outperforms baselines, and its advantage widens as compression increases. At -70% compression, LSP brings Llama-2-7B to 10.9 WikiText-2 perplexity and 42.2% mean zero-shot accuracy, versus 13.3 and 36.0% for the strongest baseline. The factorized model decodes up to 1.6x faster than the dense model at small batch sizes, and aching the shared latent shrinks the combined memory of weights and KV cache by 13.5x at a 128k-token context, versus at most 6.5x for untied baseline factorizations.
19. 【2609.40124】Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions
链接:https://arxiv.org/abs/2609.40124
作者:Chahat Raj,Sina Mansouri,Aylin Caliskan,Antonios Anastasopoulos,Ziwei Zhu
类目:Computation and Language (cs.CL)
关键词:reduce stereotypical thinking, cognitive science, responses in humans, long been studied, studied in social
备注: Under Review
点击查看摘要
Abstract:Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.
20. 【2609.40121】On the (In)effectiveness of AMR Augmentation for Large Language Models
链接:https://arxiv.org/abs/2609.40121
作者:Hoa Quynh Nhung Nguyen,Jacopo Staiano,Michael Sullivan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Abstract Meaning Representation, Meaning Representation, Abstract Meaning, historically improved performance, range of NLP
备注: 23 pages, 6 figures, 18 tables, accepted at EMNLP 2026
点击查看摘要
Abstract:While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to reproduce recent work that reported substantial downstream gains from AMR augmentation, finding that these are likely due to specific choices in the experimental settings used: using a consistent and unified protocol for hyperparameter selection, we observe that text-only baselines consistently match or exceed the performance of AMR-augmented models. To investigate this null result, we introduce a perplexity-based probe measuring the degree to which AMR provides an LLM with supplemental relational knowledge not already available to the model. We find that AMR augmentation does not help LLMs improve their understanding of relational content in the sentence, indicating that augmenting these models with AMR offers no clear benefit on downstream tasks.
21. 【2609.40118】Persistent Context Graphs for Efficient Memory Compaction in LLM Agents
链接:https://arxiv.org/abs/2609.40118
作者:Jingbo Yang,Kwei-Herng Lai,Xiaowen Wang,Zhaoxuan Tan,Pei Zhou,Mengting Wan,Yaar Harari,Evgeniy Gabrilovich,Shiyu Chang
类目:Computation and Language (cs.CL)
关键词:LLM capabilities advance, tackling increasingly complex, LLM capabilities, increasingly complex tasks, capabilities advance
备注:
点击查看摘要
Abstract:As LLM capabilities advance, agents are tackling increasingly complex tasks over longer horizons. Their growing interaction histories make memory compaction essential for staying within context windows and reducing prefill cost. Existing methods summarize the history or compress its KV cache, often adding model computation to preserve information for future requests. A new user request can change which history matters, but reassessing that history with the model requires re-encoding it if the KV cache has expired. Past attention provides signals of historical importance and dependencies between messages, while relevance to the current task must be assessed using the new user request. We introduce ReCAP, a memory compaction method that stores attention-derived importance scores and dependency links in a lightweight, persistent context graph. For each new request, ReCAP combines stored importance with relevance cues from the request and follows dependency links to select messages and their supporting context, without additional model calls for selection. Compared with Codex's default summarization-based compaction, ReCAP reduces estimated latency for compaction and cold restoration by approximately 95% on both Qwen3-Coder and gpt-oss. It also roughly halves the historical context per call on SWE-Together at comparable task quality and improves accuracy on the code tasks of Lost-in-Conversation over full history by 19.8 and 41.2 points.
22. 【2609.40111】Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
链接:https://arxiv.org/abs/2609.40111
作者:Kunlun Zhu,Xuyan Ye,Yibo Li,Cheng Qian,Beibin Li,Heng Ji
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:unsuccessful LLM agent, LLM agent rollout, unsuccessful LLM, LLM agent, Agent Error Dataset
备注:
点击查看摘要
Abstract:An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.
23. 【2609.40108】OverdoseMoE: A Multi-Expert Framework for Opioid Overdose Risk Prediction
链接:https://arxiv.org/abs/2609.40108
作者:Mingchen Li,Rohan Pandey,Junhui Qian,Feiyun Ouyang,Sunjae Kwon,Avijit Mitra,Zonghai Yao,Hong Yu
类目:Computation and Language (cs.CL)
关键词:Opioid overdose remains, public health burden, opioid overdose risk, Opioid overdose, remains a major
备注:
点击查看摘要
Abstract:Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Here, we investigate diagnosis-specific adaptation for 180-day opioid overdose risk prediction from patients' preceding one-year longitudinal ICD histories. We develop OODMAMBA and OODQWEN through continued pretraining on longitudinal diagnostic sequences followed by task-specific fine-tuning. Building on the stronger Qwen-based predictors, we further propose OVERDOSEMOE, a multi-expert framework that integrates models of different scales using complementary expert-weighting strategies. Diagnosis-specific adaptation consistently improved predictive performance over general-purpose language-model baselines, with OODQWEN achieving an AUPRC of 24.47 and an AUROC of 68.56. OVERDOSEMOE further improved discrimination and precision, achieving an AUPRC of 25.17 and an AUROC of 69.49 while outperforming the strongest single-model baselines. Among patients ranked in the top 5% of predicted risk, OVERDOSEMOE identified substantially enriched overdose risk, achieving a PPV of 25.38% while retaining meaningful recall. Evaluation on an independent MIMIC-IV cohort further demonstrated cross-cohort robustness, with complementary weighting strategies showing advantages across different performance measures. These findings demonstrate that diagnosis-specific language-model adaptation combined with multi-expert integration can improve opioid overdose risk stratification and support more robust prediction across heterogeneous electronic health record populations.
24. 【2609.40103】JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation
链接:https://arxiv.org/abs/2609.40103
作者:Mufeng Yang,Junwei Yu,Yepeng Ding
类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:Large language models, majority voting discards, Large language, language models, AI-generated content
备注: 9 pages, 5 figures, 5 tables. To appear in Proceedings of the 14th International Conference on Human-Agent Interaction (HAI '26), November 16-19, 2026, Osaka, Japan
点击查看摘要
Abstract:Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into reusable rubric entries that all judges inherit, making the evaluator progressively self-refining. To enable large-scale, reproducible benchmarking without human studies, we evaluate JuryFlow in an automatic configuration in which focal selection is made by entropy ranking. On MT-Bench and LLMBar, JuryFlow improves agreement with gold labels over single-judge and majority-vote panel baselines, and ablations isolate the contributions of disagreement-targeted re-evaluation, propagation, and rubric induction. We contribute (1) a human-in-the-loop paradigm that recasts the human from labeler to structural guide, (2) the JuryFlow framework operationalizing it through a disagreement graph, focal re-evaluation, and closed-loop rubric induction, and (3) an evaluation protocol with ablations that isolate where the gains originate.
25. 【2609.40097】AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
链接:https://arxiv.org/abs/2609.40097
作者:Ruifeng Yuan,Yizhi Li,Yaxin Du,Fengyu Cai,Yiqi Liu,Hou Pong Chan,Chenghua Lin,Yun Chen,Jian Yang,Bryan Dai,Pinyan Lu,Chenghao Xiao
类目:Computation and Language (cs.CL)
关键词:Existing auto-research benchmarks, entangle multiple sources, specific research capabilities, Existing auto-research, frontier agent outperforms
备注:
点击查看摘要
Abstract:Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at this https URL.
26. 【2609.40064】From Tweets to Trades: Analyzing the Influence of Public Mood over Stock Market Performance in Turkiye
链接:https://arxiv.org/abs/2609.40064
作者:Ece Elif Adak,Bertaç Şakir Şahin,Şaziye Betül Özateş
类目:Computation and Language (cs.CL)
关键词:public mood, Media and Society, public, domain-specific public mood, Economy and Finance
备注: 16 pages, 10 tables, 1 figure
点击查看摘要
Abstract:Purpose: This study examines whether domain-specific public mood is associated with stock-market dynamics and whether these relationships vary across communication domains and market conditions. It distinguishes public mood from investor sentiment and investigates whether heterogeneous sources of public communication exhibit different relationships with market behaviour. Design: The study analyses 610,422 posts published by 176 curated X accounts between January 2022 and December 2023, covering Politics and Government, Economy and Finance, and Media and Society. Posts are classified using fine-tuned Turkish transformer models under three domain-specific and one pooled regime. Public mood measures are constructed at daily, weekly, and monthly frequencies and examined alongside BIST100 and BIST30 market measures using correlation, Granger causality, vector autoregression, and impulse response analyses across the full period and selected market conditions. Findings: Public mood is not associated with the direction of stock-market returns but is associated with the magnitude of price movements, particularly for Media and Society and pooled communication. These relationships become stronger at longer aggregation frequencies. Predictive relationships are concentrated in Economy and Finance communication, while their magnitude and direction vary across market conditions, particularly during the 2023 election period. The pooled measure largely reflects the most active communication domain. Originality: The study contributes to behavioral-finance research by incorporating communication - domain heterogeneity into the analysis of public mood and market dynamics. It also demonstrates how aggregating heterogeneous sources can obscure domain-specific relationships between public communication and financial markets.
Comments:
16 pages, 10 tables, 1 figure
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2609.40064 [cs.CL]
(or
arXiv:2609.40064v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.40064
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Şaziye Betül Özateş [view email] [v1]
Wed, 30 Sep 2026 16:21:55 UTC (185 KB)
27. 【2609.40063】LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models
链接:https://arxiv.org/abs/2609.40063
作者:Junyi Zou,Avrova Donz
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Adaptive Residual Connections, Low-Rank Adaptive Residual, Low-Rank Adaptive, compact numerical state, Residual Connections
备注: 19 pages, 6 figures, 15 tables. Technical report of MMLA. The authors contributed equally
点击查看摘要
Abstract:Low-Rank Adaptive Residual Connections (LARC) give a frozen model a compact numerical state that can learn from feedback. The map $h+BAh$ adds a low-rank correction to a hidden representation. A slow state $\rho$ learns starting factors across tasks; a private fast state $\Phi$ copies them, changes with feedback, and resets to the trained initialization. This report specifies an input-side realization of the numerical policy carrier in Memory-Mediated Learning Architecture and examines its factor-space dynamics and learning lifetime. We study a rank-4 input residual with 12,288 trainable parameters on a frozen MiniCPM5-1B-SFT substrate. In a four-candidate program-selection task, two feedback-gradient steps reduce expected query execution error by 24.65 and 36.65 percentage points relative to resetting to the respective trained static and post-adaptation initializations. These development results cover 16 parameter groups and three paired training seeds. A direct support-loss selection rule is much more accurate, reaching 0.78125% error. In a repository-balanced chronological replay of public continuous-integration jobs, retaining online updates raises half-Brier loss from 0.1274 to 0.1808. A fixed follow-up intervention records same-batch non-descent and inconsistent future benefit from shrinking updates. Together, the algebra and measurements distinguish residual capacity, adaptation relative to a starting point, and usefulness on later decisions.
28. 【2609.40041】MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs
链接:https://arxiv.org/abs/2609.40041
作者:Frank Lawrence Nii Adoquaye Acquaye,Eric George Parakal,Jesse Johnson,Kishankumar Bhimani,Jochebed Afua Basil
类目:Computation and Language (cs.CL)
关键词:Akuapem and Asante, speech translation dataset, Ghanaian speech resources, existing Ghanaian speech, English translations
备注:
点击查看摘要
Abstract:We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGhana-ST is an ongoing annotation effort; the experiments here use a fixed subset of about 16.1 hours of paired speech and English translations. The audio is curated from two existing Ghanaian speech resources. Unlike in those resources, the English translations are produced directly from audio by 37 native-speaker annotators and include verbal and non-verbal event annotations. Using Whisper-small, we compare monolingual and multilingual training under severe data scarcity, reporting means over three seeds. Flat multilingual training benefits no variety in this regime. Ga and Twi are unchanged within seed variance (+0.51 and +0.06 BLEU against monolingual standard deviations of 1.63 and 2.20), while Ewe declines by 6.99 BLEU and Fante by 5.11. The degrading varieties are Ewe, which is linguistically distinct and drawn from a different source corpus, and Fante, the least-resourced. Comparing empirical cross-lingual transfer with typology-based similarity, we find that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training. We also report a methodological finding. An earlier single-run analysis found positive transfer for three of four varieties; this did not survive replication across seeds. For Ga and Twi, monolingual baselines trained on 1.6 to 6.2 hours of audio have seed standard deviations roughly five and thirty times those of the multilingual models (0.35 and 0.07 BLEU). When the monolingual condition is noisier, a single-run comparison can show apparent transfer of this size from seed variation alone. We release MGhana-ST to support research on African language speech technology and low-resource speech translation.
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2609.40041 [cs.CL]
(or
arXiv:2609.40041v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.40041
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
29. 【2609.40035】OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search
链接:https://arxiv.org/abs/2609.40035
作者:Junyu Lu,Shichao Weng,Zhiqiang Wang,Haojie Luo,Jingfan Zhang,Yuhua Zhou,Cheng Du,Yuzhuo Zhang,Xi Li,Jinwei Du,Tiancheng Feng,Chuan Xiao,Shuyuan Zheng
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:miss rare high-return, rare high-return trajectories, Trajectory Policy Optimization, Parallel Tree Search, policy-gradient theorem
备注: 42 pages, 12 figures
点击查看摘要
Abstract:The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This needs no action-distribution correction, although branching changes state visitation. Our Branch Aggregation Lemma shows that branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansion states using estimated performance differences. Under deterministic dynamics, exact values, and max-backup advantages, the induced search policy's expected return improves monotonically with the budget. We bound the gradient bias from adaptive expansion and show that max backup assigns prefix credit to actions leading to better discovered suffixes. Against a finite chain reference, TTPG's measured bias stays near its no-branching level, while NaivePG's bias grows from 0.1251 to 0.4884. At matched budgets, reward- and value-guided OPTS improve correct-answer coverage and majority-vote accuracy over independent sampling. At matched branch counts, OPTS + TTPG gains coverage with a modest bias increase relative to Fixed-branch + TTPG. Under matched interaction or rollout budgets, OPTS-TTPO improves MuJoCo tail returns over PPO by up to 28.6%, achieves a 34-22-1 win-loss-tie record against PPO on Atari-57 under the last-100-log mean-return metric, and improves micro-averaged avg@32 and pass@32 over PPO across all four Qwen3 models.
30. 【2609.39982】Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
链接:https://arxiv.org/abs/2609.39982
作者:Minki Kang,Ryo Hachiuma,Shaokun Zhang,Subhashree Radhakrishnan,Yonggan Fu,Jindong Jiang,Mingjie Liu,Ehsan Hosseini-Asl,Yi Dong,Yu-Chiang Frank Wang,Byung-Kwan Lee
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
关键词:stochastic model generations, act through stochastic, ensure its reliable, Terminal agents act, reliable execution
备注: Project page: [this https URL](https://byungkwanlee.github.io/MidHarness-page/)
点击查看摘要
Abstract:Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.
31. 【2609.39975】Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
链接:https://arxiv.org/abs/2609.39975
作者:Anastasios Nentidis,Georgios Katsimpras,Anastasia Krithara,Martin Krallinger,Miguel Rodríguez-Ortega,Eduard Rodriguez-López,Natalia Loukachevitch,Igor Rozhkov,Elena Tutubalina,Dimitris Dimitriadis,Vasiliki Patsiou,Grigorios Tsoumakas,George Giannakoulas,Alexandra Bekiaridou,Athanasios Samaras,Giorgio Maria Di Nunzio,Nicola Ferro,Stefano Marchesin,Marco Martinelli,Gianmaria Silvello,Georgios Paliouras
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:Evaluation Forum, Conference and Labs, paper presents, presents an overview, question answering
备注: 21 pages, 17 tables, International Conference of the Cross-Language Evaluation Forum for European Languages 2026 (CLEF2026)
点击查看摘要
Abstract:This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical semantic question answering. b) Task Synergy14 on question answering for developing biomedical top- ics. c) Task MultiClinSum-2 on multilingual clinical summarization. d) Task BioNNE-R on extracting relations between nested named entities in Russian and English. e) Task ELCardioCC on clinical coding in cardiology. f) Task GutBrainIE on gut-brain interplay information extrac- tion. Across these six tasks, 87 distinct teams participated, submitting more than 1000 runs overall. As in previous editions, several submissions reached competitive performance, reflecting the continued progress of state-of-the-art methods across biomedical language processing tasks.
32. 【2609.39972】UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding
链接:https://arxiv.org/abs/2609.39972
作者:Chumeng Liang,Linxuan Wang,Xinyu Peng,Huabin Liu,Yuxin Chen,Ge Liu,Guang Lin,Qifan Song,Jianguo Li
类目:Computation and Language (cs.CL)
关键词:single target-model pass, Speculative decoding accelerates, accelerates language model, language model inference, verifying multiple draft
备注:
点击查看摘要
Abstract:Speculative decoding accelerates language model inference by verifying multiple draft tokens in a single target-model pass. Recent parallel drafters have achieved breakthrough performance in frontier production models, but their effectiveness deteriorates as the entropy of target distributions increases due to insufficient draft diversity. To overcome this bottleneck without sacrificing parallelism, we introduce UBTree, a parallel drafter that couples a Unigram proposer with a Bigram selector to construct drafting Trees. The unigram proposer is trained with the standard cross-entropy objective to generate candidate tokens independently for each position, while a lightweight bigram selector predicts transition scores between adjacent candidate pairs. Unlike the proposer, the selector is trained with a renormalized KL objective on high-temperature data. This tree-native training broadens the supervision beyond the greedy path, encouraging plausible alternative branches that improve the chance of accepting additional tokens during tree verification. Across seven standardized benchmarks with Qwen3-4B and Qwen3-8B, UBTree achieves an average speedup of $5.84$--$6.94\times$ over autoregressive decoding and outperforms DARTree in all 28 comparisons. Production-scale evaluation further demonstrates UBTree's advantage over frontier baselines such as DSpark.
33. 【2609.39938】LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
链接:https://arxiv.org/abs/2609.39938
作者:Juyi Lin,Zhiqiang Lao,Jiali Cui,Lin Zhao,Pu Zhao,Dichang Zhang,Arman Akbari,Yu Qi,Xinru Jiang,Yanzhi Wang,Heather Yu,Liang Peng
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Hour-scale audio-visual question, dense whole-recording encoding, whole-recording encoding rapidly, encoding rapidly exhausts, compression severely dilutes
备注: 39 pages, 16 figures
点击查看摘要
Abstract:Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
34. 【2609.39929】RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
链接:https://arxiv.org/abs/2609.39929
作者:Yuyang Wu,Yufeng Du,Hao Peng
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:distinguishing nearby positions, maintaining stable token, stable token preferences, RoPE-based language models, RoPE intrinsic tradeoff
备注:
点击查看摘要
Abstract:Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-frequency components support positional sensitivity while potentially disrupting semantic stability. We also derive a theoretical context-length bound beyond which, under specified conditions, a fixed attention-score comparison cannot jointly avoid semantic reversal and positional insensitivity. Guided by our fresh theoretical insights, we introduce RoPE Profiler, a lightweight, plug-and-play diagnostic toolkit that augments existing evaluations with zero additional forward passes by reusing cached query and key activations. Reusing activations collected during evaluation, the toolkit incurs little overhead. It supplements standard benchmark scores with two diagnostic scores that reveal semantic and positional weaknesses and help users prioritize which aspect to address. Crucially, our evaluations across 49 long-context task settings reveal a distinct pattern where reasoning tasks predominantly suffer from semantic reversal, whereas retrieval tasks are primarily vulnerable to positional insensitivity. Guided by our theory and diagnostic profiles, targeted high-frequency rescaling achieves immediate gains without additional training, improving task accuracy by up to 20 percentage points on Qwen3-8B and 25 percentage points on Llama-3.1-8B-Instruct.
35. 【2609.39927】AdaGEPA: Adaptive Feedback Allocation for Reflective Prompt Optimization
链接:https://arxiv.org/abs/2609.39927
作者:Junyang Chen,Zecheng Wang,Jingbang Chen
类目:Computation and Language (cs.CL)
关键词:language-model systems, Prompt, feedback, prompts, task
备注:
点击查看摘要
Abstract:Prompt optimization improves the performance of language-model systems on downstream tasks by refining their prompts. Classical methods evaluate prompts on task examples and use the resulting feedback to guide prompt revisions through reflection. However, when feedback selection does not account for the prompt's weaknesses, these revisions may improve performance on selected examples without yielding broader task improvements. To address this issue, we propose AdaGEPA, an adaptive feedback-allocation method that uses the prompt's performance and task structure to select examples for the next prompt revision. Our method replaces at most one example in each feedback minibatch to target an identified weakness while preserving the remaining feedback context. Across our main experiments on six downstream benchmarks, AdaGEPA achieves higher mean validation scores than non-adaptive feedback selection under matched rollout budgets. AdaGEPA also finds high-performing prompts earlier across several tasks. In the initial Schema-Guided Dialogue (SGD) study, its half-budget prompts outperform the non-adaptive baseline's full-budget prompts in joint goal accuracy on new dialogues from services seen and unseen during search. Overall, our findings highlight the potential of adaptive feedback allocation to improve both the effectiveness and rollout-budget efficiency of reflective prompt optimization.
36. 【2609.39920】MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models
链接:https://arxiv.org/abs/2609.39920
作者:Yanshu Li,Jiaqian Li,Canran Xiao,Xi Xiao,Tianyang Wang,Yongtai Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Large vision-language models, ability degrades substantially, Large vision-language, multimodal in-context learning, model size decreases
备注: 17 pages, 8 tables, 5 figures
点击查看摘要
Abstract:Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.
37. 【2609.39913】he Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation
链接:https://arxiv.org/abs/2609.39913
作者:Thomas Pashby
类目:Computation and Language (cs.CL)
关键词:large language models, language models solve, models solve formally, solve formally matched, formally matched kinship
备注:
点击查看摘要
Abstract:We test whether large language models solve formally matched kinship problems equally well when relations are expressed in familiar vocabulary or by explicitly defined nonce predicates. Across 500 paired graphs, concrete accuracy exceeds arbitrary accuracy by 35.6 percentage points in local Qwen3.8-27B, 26.6 in Gemma 4 26B-A4B, 12.0 in Gemma 4 31B, and 5.4 in Qwen3.8-Max. All four paired gaps are statistically resolved. Reasoning budgets and prompt-language interventions can substantially reduce the difference, showing that it is modifiable rather than a fixed incapacity. The minimal conclusion is behavioral: on these tasks, the models' manifested relational competence is not indifferent to presentation. Explicit definitions provide the formal relations but do not make nonce predicates as usable as familiar vocabulary embedded in learned linguistic associations.
38. 【2609.39884】OPSRD: On-Policy Self-Role Distillation
链接:https://arxiv.org/abs/2609.39884
作者:Weijie Ren,Yanwen Zhang,Hao Li,Zhuolin Qi,Hengyi Zhang,Naibo Wang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:prompting elicits specialized, elicits specialized behavior, Role prompting elicits, large language models, offering a lightweight
备注: 17 pages, 5 figures. Code: [this https URL](https://github.com/zhansan114514/OPSRD)
点击查看摘要
Abstract:Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at this https URL.
39. 【2609.39882】LLM Persona Unlearning
链接:https://arxiv.org/abs/2609.39882
作者:Kemou Li,Zhuan Shi,Qizhou Wang,Fengpeng Li,Negar Rostamzadeh,Golnoosh Farnadi,Jiantao Zhou
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Pre-training equips large, large language models, equips large language, Pre-training equips, equips large
备注:
点击查看摘要
Abstract:Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefore elicit personas that repeatedly shape judgment, language, and action. In open-weight settings, runtime controls can be removed, motivating persona unlearning: a weight-level edit that makes a designated persona difficult to elicit and enact on unseen contexts. We introduce PersonaUnlearnBench, a model-specific paired benchmark spanning six LLMs from three families and five personas, with aligned forget/retain sets, held-out instruction paraphrases, and four-axis evaluation. The benchmark shows that standard unlearning methods cannot reliably erase the target persona without sacrificing meaningful generation or general utility. We therefore propose PaCE, which compares target and desirable responses to the same questions to locate an internal behavior direction, then trains target-prompt states away from the target mode and toward the matched desirable response. Experiments show that PaCE consistently suppresses target personas with high response quality and useful counterpart behavior, at moderate utility cost. These results establish persona unlearning as a distinct behavior-level editing problem and a practical route toward persistent control of latent LLM response policies.
40. 【2609.39869】GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
链接:https://arxiv.org/abs/2609.39869
作者:Gabriele Tuccio,Antonino Furnari,Aldo Gangemi,Misael Mongiov\`ı
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:guarantees syntactic validity, generation guarantees syntactic, substantially degrade semantic, Grammar-constrained generation guarantees, degrade semantic quality
备注:
点击查看摘要
Abstract:Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-supervised rewards derived from its own likelihoods: a direct reward, measuring how likely the constrained output is given the input, and a reverse reward, measuring how well the input can be reconstructed from the generated output. We optimize these rewards with a Reinforce Leave-One-Out (RLOO) objective over groups of grammar-constrained rollouts, augmented with the top-1 beam-search hypothesis and regularized towards a frozen base model. We evaluate GrammarRL on sign language gloss translation, hierarchical text classification, and named entity recognition using Llama models ranging from 1B to 8B parameters. GrammarRL consistently outperforms constrained greedy decoding, with an average improvement of 9.8 points and gains of up to 22.8 BLEU. It matches or outperforms beam search on two of the three tasks while preserving greedy-decoding inference cost. Ablations further show that the two rewards are complementary: either reward alone can underperform the untrained baseline, whereas their combination consistently improves upon it.
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:
arXiv:2609.39869 [cs.AI]
(or
arXiv:2609.39869v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2609.39869
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
41. 【2609.39863】FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
链接:https://arxiv.org/abs/2609.39863
作者:Sidharth Pulipaka,Ruta Binkyte,Ivaxi Sheth,Sahar Abdelnabi
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:balance staying truthful, Large language models, Large language, language models frequently, balance staying
备注: 64 pages, 11 figures, 29 tables. Code: [this https URL](https://github.com/compass-group-tue/FIGSBench) ; Data: [this https URL](https://huggingface.co/datasets/compass-group-tue/FIGSBench)
点击查看摘要
Abstract:Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views. In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time. Current evaluations, however, rely on rigid, single-turn tests or fixed scripts that fail to capture these natural dynamics. Furthermore, these benchmarks often mistake showing basic empathy for yielding, penalizing models for acknowledging a user's feeling. This view may drive future models to over-correct into cold, dismissive rigidity. To address this gap, we introduce FIGS (Factual Integrity and Grounded Support), a dual-axis evaluation framework built around extended, realistic dialogue. We use an adaptive 10-turn conversational simulator that dynamically challenges the target model, reflecting how users repeat requests, push back, or steer a conversation toward a preferred answer. To accurately evaluate these trajectories, we apply a taxonomy that strictly separates Sycophancy (whether the model holds firm to the truth and keeps its praise proportional) from Calibrated Validation (showing empathetic understanding of the user's feelings without overdoing it). We release our complete testing environment, including 500 diverse multi-turn scenarios and an automated judge. Our evaluation of leading models reveals a consistent trade-off: over the course of a sustained interaction, current systems either slowly drift to sycophancy or over-correct into robotic detachment. This demonstrates that balancing honesty with appropriate support throughout a natural conversation remains a critical, unsolved challenge.
42. 【2609.39853】Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models
链接:https://arxiv.org/abs/2609.39853
作者:Xingjie Zhuang,Jialong Tang,Chulun Zhou,Buchao Zhan,Zhirui Li,Junhui Li,Yazheng Yang,Jinsong Su
类目:Computation and Language (cs.CL)
关键词:improving LLM reasoning, Role-playing prompting, output quality, technique for improving, reasoning and output
备注: 22 pages, 7 figures
点击查看摘要
Abstract:Role-playing prompting has become a popular yet simple technique for improving LLM reasoning and output quality. However, whether it consistently boosts performance across diverse domains remains unclear, as systematic validation is lacking. To fill this gap, we run multi-model, cross-domain, and multilingual experiments on MMLU and MMLU-Redux. We find that gains from role-play prompting depend heavily on model capacity, knowledge domain, and prompt language. Drawing on metacognition theory, we propose the persona-related cognitive alignment hypothesis: role-play works only when the LLM correctly grasps the designated persona and its associated knowledge domain. We test this hypothesis through persona information richness ablation, layer-wise entropy divergence analysis, and latent thought-space deflection observation. To reduce persona cognitive bias and stabilize role-play performance, we propose \textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP}), a simple, training-free, and efficient multilingual prompt concatenation strategy. It aggregates semantically equivalent role prompts to enrich complementary representational cues. Extensive experiments show that MLCP consistently outperforms vanilla role-play prompting across all tested LLMs.
43. 【2609.39846】When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models
链接:https://arxiv.org/abs/2609.39846
作者:Pakhapoom Sarapat,Saksorn Ruangtanusak,Kunat Pipatanakul,Pittawat Taveekitworachai
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
关键词:text while continuing, continuing to exhibit, exceed those implied, role-prompted reasoning model, convincing in-role text
备注:
点击查看摘要
Abstract:We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a kindergarten student, one might expect its performance on a mathematics benchmark to reflect kindergarten-level ability rather than expert-level proficiency in solving calculus problems. We introduce RoleCapBench, a curriculum-grounded benchmark for evaluating RCL across six educational roles and four assessment levels spanning elementary school through A-level, and use it to evaluate three open-weight reasoning models. We find that although the models can generate stylistically convincing in-role responses, they consistently fail to align their underlying capabilities with their assigned roles. Naive role prompting yields strong role-voice scores of 1.218--1.389 while retaining above-role accuracy of 0.811--0.898. RCL persists across a range of prompting conditions, including prompts that explicitly instruct the model to match the role's capability level. To mitigate this problem, we propose Injection, an inference-time intervention that combines explicit, role-specific capability guidelines with a guiding prefilled response prefix. Injection improves role-capability alignment across models, reducing above-role accuracy by up to 0.562 while preserving in-role accuracy with a marginal drop of less than 0.058 across most models. All artifacts, including scripts and evaluation data, will be released upon acceptance.
44. 【2609.39838】Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
链接:https://arxiv.org/abs/2609.39838
作者:Julian Schulz,Lukas Fülle,Rieke Fruengel
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:inside innocuous-looking text, steganographic reasoning, reasoning inside innocuous-looking, steganographic, reasoning
备注: Accepted as an oral at the NeurIPS 2026 Workshop on Trustworthy AI for Good (AI4GOOD). 41 pages. Code: [this https URL](https://github.com/stegano-ai/steg-reasoning-is-hard)
点击查看摘要
Abstract:Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded reasoning (reasoning in an illegible but unconcealed format), have already been shown to emerge under training pressures that occur in real pipelines, such as reinforcement learning against monitors. This suggests that steganographic reasoning too might arise as an unintended side effect of training. Here, we compare how easily models learn steganographic reasoning and these two neighbouring capabilities across three elicitation methods: reinforcement learning, in-context learning, and supervised fine-tuning (SFT). For most tasks, models learn steganographic reasoning only under SFT, while they learn steganographic messaging and encoded reasoning under all three elicitation methods. Even under SFT, steganographic reasoning requires at least twice as much training as messaging, and for several model-task combinations it is not learned at all. However, on a cover task that makes hiding information especially convenient, steganographic reasoning can be successfully learned under all three elicitation methods. Steganographic reasoning is thus much harder than steganographic messaging and encoded reasoning, and learning the latter two does not imply learning the former. Yet it lies within reach: an easy version is learned under every elicitation method, when the cover task is convenient for hiding information.
Comments:
Accepted as an oral at the NeurIPS 2026 Workshop on Trustworthy AI for Good (AI4GOOD). 41 pages. Code: this https URL
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:
arXiv:2609.39838 [cs.AI]
(or
arXiv:2609.39838v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2609.39838
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
45. 【2609.39827】Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
链接:https://arxiv.org/abs/2609.39827
作者:Atsuki Yamaguchi,Tatsuro Inaba,Joel Niklaus,Michal Štefánik,Aline Villavicencio,Nikolaos Aletras
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:synthetic non-natural language, PPT, non-natural language data, language model pre-training, Pre-pretraining
备注: Preprint. Under review
点击查看摘要
Abstract:Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
46. 【2609.39807】Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations
链接:https://arxiv.org/abs/2609.39807
作者:Maximilian von Klinski,Sebastian Lapuschkin,Wojciech Samek,Lennart Bürger
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:model internal states, language model internal, Lie detection probes, Lie detection, truthful or dishonest
备注:
点击查看摘要
Abstract:Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, particularly when correct and incorrect answers are evaluated under the same persona prompt. To investigate why, we construct three novel confounder datasets in which truth is anti-correlated with a potential confounding concept. Our experiments reveal that many existing probes strongly track concepts that are spuriously correlated with truth in their training data, such as instruction compliance or response likelihood. Based on these findings, we introduce a simple linear probe that achieves the strongest overall performance on both the persona and confounder stress tests. Our results suggest that current lie detection probes are far from reliable and highlight the need for training data in which truth is decorrelated from confounding concepts.
47. 【2609.39786】Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness
链接:https://arxiv.org/abs/2609.39786
作者:Ola El Khatib,Djellel Difallah
类目:Computation and Language (cs.CL)
关键词:Large language models, Large language, language models, increasingly combined, Large
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missing knowledge risk introducing hallucinated evidence. We introduce XoG (eXplore-on-Graph), a framework for multi-hop question answering over incomplete KGs that recovers missing reasoning paths from learned graph structure rather than LLM parametric knowledge. XoG combines type-level entity-relation statistics to identify candidate relations with KG embeddings to retrieve plausible missing entities, using the LLM as a semantic selector and reasoner. These mechanisms are integrated into an iterative planning-exploration-reasoning process. Experiments on WebQSP, CWQ, and the Wikidata-based BRINK benchmark show that XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under KG incompleteness. These gains persist across multiple LLM backbones, indicating that stronger LLMs alone do not resolve missing graph evidence. XoG also reduces LLM token consumption by up to 33% compared with a closely related planning-based approach.
48. 【2609.39765】MemCodex: Self-Programming Hierarchical Memory for Language Agents
链接:https://arxiv.org/abs/2609.39765
作者:Xiaoqiang Wang,Bang Liu
类目:Computation and Language (cs.CL)
关键词:Agent memory faces, faces heterogeneous access, Agent memory, multiple sources, single-hop question
备注: Work in progress
点击查看摘要
Abstract:Agent memory faces heterogeneous access needs: a single-hop question may require one piece of evidence, whereas a multi-hop question must combine evidence from multiple sources. Predefined memory workflows cannot adapt to these varying needs. Recent adaptive methods search or learn over memory components and their compositions, but the design space itself remains predefined. We introduce MemCodex, a self-evolving hierarchical memory system that organizes experience into executable memory programs for summaries, relational knowledge, reusable skills, and latent memory. Open-ended program evolution searches the open design space of layer programs by rewriting how each layer is constructed, indexed, retrieved, and routed, thereby adapting both within-layer implementations and cross-layer composition. At query time, reads traverse the hierarchy from coarse to fine and stop once sufficient evidence is found, descending to the original history when needed. We further develop MemArena, a unified runtime that places heterogeneous data and memory systems behind a common interface. MemCodex improves average task success by 10.1% relative to the strongest adaptive-memory baseline, while using 3.4x fewer context tokens and achieving 2.1x faster inference.
49. 【2609.39740】LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation
链接:https://arxiv.org/abs/2609.39740
作者:Xiaoqiang Wang,Suyuchen Wang,Bang Liu
类目:Computation and Language (cs.CL)
关键词:complementary bottlenecks, faces two complementary, Long-context reasoning faces, retaining evidence, reasoning
备注: Work in progress
点击查看摘要
Abstract:Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.
50. 【2609.39727】OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation
链接:https://arxiv.org/abs/2609.39727
作者:Oana Madalina Fron,Ojas Shirekar,Chirag Raman
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)
关键词:Cooperative language-model agents, existing agents map, agents map observations, Cooperative language-model, Prefrontal Cortex Module
备注:
点击查看摘要
Abstract:Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.
51. 【2609.39710】Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims
链接:https://arxiv.org/abs/2609.39710
作者:Vsevolod Karimov,Stepan Ostarkov,Anastasia Poroshina,Anatoly Frolov,Alexander Panchenko
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL)
关键词:Scientific abstracts mix, abstracts mix contributions, Atomic Contribution Claims, present Drift Inspector, Scientific abstracts
备注: Accepted to EMNLP 2026 System Demonstrations. 11 pages. Live demo, code and data: [this https URL](https://hamyrappy.github.io/drift-inspector)
点击查看摘要
Abstract:Scientific abstracts mix contributions with background, motivation, and meta-language, so tools that read them as-is cannot separate what a field produces from what it discusses. We present Drift Inspector, an open-source system for measuring and exploring how a research field changes over time at the level of Atomic Contribution Claims (ACCs): decontextualized, contribution-bearing propositions an LLM extracts from each abstract before analysis. The system clusters these claims across years into an interactive map where every trend traces back to the claims and papers behind it. Applied to six years of EMNLP, it shows the field shifting away from classic NLP tasks toward LLM-era capabilities such as reasoning and multimodality -- a movement that keyword or whole-abstract counts blur. The released data extend beyond EMNLP: the same pipeline has processed the full ACL Anthology (346k claims, 80k abstracts, 423 venues). Extraction is human-validated and clustering checked against an external manually constructed taxonomy.
52. 【2609.39702】A helps B while B hurts A: directed transfer in instruction-tuning mixture
链接:https://arxiv.org/abs/2609.39702
作者:Nima H. Siboni,Vahid Rostami
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Adapting a language, fixed budget, choosing which instruction-tuning, testing one choice, choice costs
备注:
点击查看摘要
Abstract:Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negative; the second, that it is symmetric. We show that both assumptions fail: task $A$ can help task $B$ while $B$ hurts $A$, so helpfulness is a signed property of ordered source--target pairs. We introduce the transfer map, a signed estimate of how much each source helps or hurts each held-out target. We fit the map in hundreds of fine-tuning runs on Qwen3 and Mistral models from 0.6B to 32B parameters, with all sources drawn from one corpus and no training examples from the target. The map predicts a held-out target's accuracy on unseen mixtures: recorded before those runs, its predictions have less than half the error of a mixture-agnostic baseline. The map is specific to its target and corpus but transfers across model scale: a mixture selected in advance at one size beats training on all source tasks at every other size we tested. Transfer is thus a property of the data. The map selects the tasks that help and drops the one that interferes: accuracy on the reasoning targets (causal explanation, multi-hop questions and methodological critique) rises by up to 14 percentage points over training on all source tasks.
53. 【2609.39688】ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
链接:https://arxiv.org/abs/2609.39688
作者:Tobia Poppi,Silvia Cappelletti,Samuele Poppi,Marcella Cornia,Lorenzo Baraldi,Diego Garcia-Olano,Rita Cucchiara
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
关键词:changing benign representations, web-scale training data, training data embed, unnecessarily changing benign, CLIP underlie
备注:
点击查看摘要
Abstract:Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at this https URL.
54. 【2609.39687】Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
链接:https://arxiv.org/abs/2609.39687
作者:Xincheng Wei,Yifan Ding,Yoshua Li,Yuquan Lu,Ziheng Li,Yi Lu,Dongsheng Ma,Rongxiang Weng,Xunliang Cai
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:trains mathematical reasoning, supervises student-sampled prefixes, mathematical reasoning models, On-policy self-distillation, trains mathematical
备注:
点击查看摘要
Abstract:On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.
55. 【2609.39661】he Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends
链接:https://arxiv.org/abs/2609.39661
作者:Zhentao Tan,Jingyi Shen,Yanbo Li,Yao Liu,Yue Wu,Jieping Ye
类目:Computation and Language (cs.CL)
关键词:dense token interactions, token interactions incur, interactions incur quadratic, incur quadratic prefill, quadratic prefill cost
备注:
点击查看摘要
Abstract:Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model. We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2609.39661 [cs.CL]
(or
arXiv:2609.39661v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.39661
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
56. 【2609.39645】SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration
链接:https://arxiv.org/abs/2609.39645
作者:Weijie Ren,Yanwen Zhang,Hao Li,Zhuolin Qi,Hengyi Zhang,Naibo Wang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:large language models, Multi-agent collaboration, improve question answering, language models, large language
备注: 22 pages, 4 figures. Code: [this https URL](https://github.com/zhansan114514/SEPAL)
点击查看摘要
Abstract:Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while single-pair Actor-Critic collaboration refines only one candidate. We introduce SEPAL, which assigns three private Actor-Critic teams to direct reasoning, evidence grounding, and verification. Role-specific training gives the teams different reasoning objectives beyond sampling variation. Each Critic guides revisions within its own team, preventing feedback from carrying errors across candidates. Once revision ends, majority voting combines only the final answers, keeping the reasoning histories separate until the decision. Across five open-weight backbones and five question-answering benchmarks, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor-Critic pair, with improvements across all five backbones. Code is available at this https URL.
57. 【2609.39640】Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
链接:https://arxiv.org/abs/2609.39640
作者:Dalton Raphael Harmsen,Swier Garst,Thomas van Osch,Zarè Palanciyan,Joaquin Vanschoren
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:source language benefits, Cross-lingual transfer describes, Cross-lingual transfer, describes how knowledge, benefits a target
备注: 4 pages, NeurIPS workshop, Linguistic Principles for Foundation Models, lp4fm
点击查看摘要
Abstract:Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $\rho{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $\rho{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{this https URL}{here}.
58. 【2609.39639】Marginal Response Surface Elicitation for Zero-Label Tabular Learning
链接:https://arxiv.org/abs/2609.39639
作者:Liangyu Teng,Yicheng Ding,Jing Liu,Hengsong Liu,Juncen Guo,Hongru Li,Jingyu Zhang,Liang Song
类目:Computation and Language (cs.CL)
关键词:predict target outcomes, target outcomes, learning uses structured, predict target, Response Surface Elicitation
备注:
点击查看摘要
Abstract:Tabular learning uses structured data to predict target outcomes. Traditionally, this process has relied on labeled data. However, large language models (LLMs) can be used to elicit domain priors based on the task description and feature semantics, thereby enabling predictions without labeled data. We propose Marginal Response Surface Elicitation (MARS), a method that transforms feature-level LLM priors into a reusable, zero-shot tabular classifier. To construct this classifier, MARS selects representative values for each feature from unlabeled data and prompts the LLM to provide corresponding class support scores and feature weights. It then aggregates multiple responses using the median to construct feature response functions, and makes predictions through their weighted sum without further LLM queries. Across eight tabular benchmark tasks, MARS achieves the highest average AUC and AP, outperforming direct prompting by 1.97 and 6.21 percentage points respectively, while substantially reducing end-to-end costs. Evaluations with LLMs of different sizes further demonstrate its predictive advantage over direct prompting.
59. 【2609.39608】Is This Evidence Decision-Critical? Learning to Verify Rule-Governed Decisions
链接:https://arxiv.org/abs/2609.39608
作者:Haoyang Zhang,Jianpeng Zhao,Qi Hao,Pengyang Wang
类目:Computation and Language (cs.CL)
关键词:Rule-based reasoning, evidence, contract reviews, Rule-based, decision
备注:
点击查看摘要
Abstract:Rule-based reasoning, as in eligibility checks and contract reviews, requires language models to assess evidence against individual conditions and combine their judgments under explicit rules. Errors in evidence assessment can leave a decision unchanged, but misinterpreting or overlooking decision-critical evidence can reverse it. Identifying such evidence allows more capable models to focus on checking the corresponding condition judgments, supporting accurate and safe decisions. Recognizing the evidence's criticality requires understanding how evidence affects a condition judgment and how that judgment affects the decision. To achieve the goal, we propose a INTERvention-based imPACT learning framework (InterPact), which enables counterfactual verification of evidence criticality in rule-governed decisions. Specifically, its evidence intervention constructor generates training pairs for a propagation verifier by editing case facts with a frozen language model while holding rules and non-target conditions fixed. Human-reviewed labels record the resulting condition and decision changes, while complete state-to-decision mappings supervise consequences beyond the observed edit. During training, the verifier weights learned conditional decision predictions by evidence-based condition probabilities through a fixed composition operation, propagating decision-change supervision into the base model. At inference, the trained base model directly judges criticality from the original case and target evidence, without human or stronger-model supervision. On single-case evidence criticality verification over adapted rule-governed decision cases, InterPact achieves 68.28% accuracy, outperforming all six baselines. These results support learned decision sensitivity as a basis for prioritizing evidence checks.
60. 【2609.39578】hinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
链接:https://arxiv.org/abs/2609.39578
作者:Minghan Wang,Boyuan Wang,Jinhang Zuo,Yuxin Tao,Fang kong
类目:Computation and Language (cs.CL)
关键词:grow more capable, constrain their execution, increasingly constrain, box, Bench
备注:
点击查看摘要
Abstract:Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.
61. 【2609.39572】Compact Language, Complex Model Shifts: How and Where Ambiguity and Underspecification Affect LLMs
链接:https://arxiv.org/abs/2609.39572
作者:Michaela Regneri,Nina Scheller,Sören Laue
类目:Computation and Language (cs.CL)
关键词:language model training, affect language model, underspecification affect language, model training, affect language
备注: To appear in Proceedings of BlackBoxNLP 2026
点击查看摘要
Abstract:We analyze how lexical ambiguity and underspecification affect language model training. We create artificial homonyms and artificial hypernyms as pseudowords and analyze the generative performance of language models as they are trained with increasing amounts of these ambiguous or underspecified pseudoword types. We further analyze whether the models disambiguate ambiguous or underspecified statements and provide a first mechanistic account of how ambiguity and disambiguation are represented internally. Our main results show that both ambiguity and underspecification increase model performance in ways that scale with their influence on the language's type-token ratio. However, the accuracy of generating sequences containing ambiguous words or their synonyms decreases compared to other texts. We also show that internal representations of pseudowords reflect disambiguation of pseudo-homonyms, but underspecification of pseudo-hypernyms is maintained during the generative process.
62. 【2609.39549】Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
链接:https://arxiv.org/abs/2609.39549
作者:Zezhong Wang,Xueyang Tang,Rui Lian,Yang Lou,Heqing Huang
类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large Language Model, Language Model, Large Language, significant security challenge, security challenge
备注:
点击查看摘要
Abstract:As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification stage, the system uses the target agent's real actions to calibrate and prune the trajectory tree, effectively reducing false positives. As a plug-and-playable component, SSH provides existing detectors with rich decision redundancy beyond the current interaction slice. By judging risk based on the evolution of the entire trajectory tree rather than a single point in time, the system reduces the reliance on the absolute precision of individual detection components. This improves the defense resilience and the warning lead-time of agent systems against complex temporal attacks.
63. 【2609.39533】CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
链接:https://arxiv.org/abs/2609.39533
作者:Shouli Wang,Yanfeng Jia,Zhihao Ou,Zitao Su,Ruize He,Haotong Xie,Hao Peng,Juanzi Li,Xiaozhi Wang
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:obtain high rewards, large language models, hacking, large language, intended capabilities
备注:
点击查看摘要
Abstract:During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at this https URL.
64. 【2609.39514】Spike-driven Vision-Language-Action Model
链接:https://arxiv.org/abs/2609.39514
作者:Shuai Wang,Malu Zhang,Mingquan Liu,Weihui Dai,Dehao Zhang,Jieyuan Zhang,Yimeng Shan,Zijian Zhou,Yang Yang
类目:Computation and Language (cs.CL)
关键词:advancing the dominant, models bridge multimodal, Spike-driven VLA, bridge multimodal understanding, VLA
备注:
点击查看摘要
Abstract:Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top-$k$ winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.
65. 【2609.39496】When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev
链接:https://arxiv.org/abs/2609.39496
作者:Jike Zhong,Ming Li,Yuxiang Lai
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Typed decision models, Typed decision, offer an efficient, generative LLMs, LLMs in decision-making
备注:
点击查看摘要
Abstract:Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.
66. 【2609.39460】Right-Wing Rock or Just Rock? A Computational Linguistic Analysis of Frei.Wild
链接:https://arxiv.org/abs/2609.39460
作者:Carlotta Schneeberger(1),Kevin Tang(1 and 2) ((1) Heinrich Heine University Düsseldorf, (2) University of Florida)
类目:Computation and Language (cs.CL)
关键词:spreads right-wing ideology, http URL, radical scene, instrumentalized to recruit, recruit adolescents
备注: 20 pages, 9 figures, for code and data see [this https URL](https://zenodo.org/records/22676753) , to be published in the proceedings of the NLP 4 Positive Impact workshop at EMNLP 2026
点击查看摘要
Abstract:Rechtsrock is a subgenre of rock music that spreads right-wing ideology, often instrumentalized to recruit adolescents into the radical scene. Monitoring institutions counteract this by manually examining and, in some cases, banning extremist content; however, there are border cases that evade regulation. We present a study aimed at determining whether such a case, the band this http URL, should be classified as politically right-leaning or as part of the general German rock genre. We sampled a German rock dataset and created a corpus for right-wing rock to use as reference in this analysis and found that we can confirm the intuitions from previous investigations that this http URL successfully maintains an ambiguity with regard to their political affiliation. However, the tendency is towards the right-wing spectrum. Lexical analyses reveal nationalistic narratives and two high-performing classifiers (up to 97% ROC-AUC score) label more than half of their songs as right-wing extremist. Our analysis provides insight into how computational methods can improve the process of identifying right-wing extremist tendencies in music, especially in borderline cases like this http URL. The code and data are made available for future research.
67. 【2609.39453】From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models
链接:https://arxiv.org/abs/2609.39453
作者:Hezhao Zhang,Thomas Hain
类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:assigning emotion labels, Speech emotion recognition, emotion recognition, assigning emotion, emotion labels
备注: 5 pages, 2 figures. Submitted to ICASSP 2027
点击查看摘要
Abstract:Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
68. 【2609.39447】Synthetic Data Characterization via Training Dynamics
链接:https://arxiv.org/abs/2609.39447
作者:Irene Lago,Ana Ezquerro,David Vilares
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Interpreting properties, properties of LLM-generated, important for understanding, understanding its utility, utility and limitations
备注: Accepted at Findings of EMNLP 2026
点击查看摘要
Abstract:Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.
69. 【2609.39446】DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements
链接:https://arxiv.org/abs/2609.39446
作者:Keyue Xing,Wentao Ding,Mengmeng Wang,Wenming Tu,Zilong Zheng,Yipeng Kang
类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:Existing full-duplex speech, limited contextual conditions, Existing full-duplex, speech benchmarks cover, full-duplex speech benchmarks
备注:
点击查看摘要
Abstract:Existing full-duplex speech benchmarks cover only subsets of real-time interaction behaviors, often under limited contextual conditions. We introduce DuplexAct-Bench, a bilingual benchmark that systematically covers six complementary behaviors, from interruption and yielding to proactive initiation, active silence, and backchanneling, across Pre-session, In-session, and No-explicit conditions. Across 1,290 English and Chinese streaming trials, we evaluate 12 full-duplex speech systems on both Timing and Content. Results reveal substantial variation across behaviors, conditions, and systems, as well as frequent mismatches between semantic quality and behavioral timing. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds. Project page: this https URL
70. 【2609.39420】QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code
链接:https://arxiv.org/abs/2609.39420
作者:Alexey Chernysh,Orkhan Ekhtibarov,Dmitry Zmitrovich
类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Trading and Market Microstructure (q-fin.TR)
关键词:Large language models, algorithmic trading remains, remain semantically faithful, correct program logic, specialized trading framework
备注: 16 pages, 2 figures, 6 tables
点击查看摘要
Abstract:Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.
71. 【2609.39394】Can Computation from Earlier Problems Help LLMs Solve New Ones?
链接:https://arxiv.org/abs/2609.39394
作者:Jipei He,Wenhui Tan,Xiaoyi Yu,Enver Sangineto,Fiorenzo Parascandolo,Rita Cucchiara,Ruihua Song
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large language models, Large language, solve independent problems, solve independent, Large
备注: 29 pages, 7 figures
点击查看摘要
Abstract:Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.
72. 【2609.39385】Lab at Daleel 2026: STAR-Ar, Sequence Tagging for Argument Recognition in Arabic
链接:https://arxiv.org/abs/2609.39385
作者:Bhuvanesh Verma,Ali Abusaleh,Alexander Mehler
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Arabic argument mining, critical NLP task, remains significantly under-resourced, critical NLP, inaugural Arabic argument
备注: Accepted at ArabicNLP 2026 Daleel-2026 shared task
点击查看摘要
Abstract:Argument Mining (AM) is a critical NLP task that remains significantly under-resourced in Arabic. This paper presents $\testtt{STAR-Ar}$, a BERT-BiLSTM-CRF architecture for argument discourse detection and classification, as our system for Daleel 2026, the inaugural Arabic argument mining shared task. The task requires the identification and classification of argumentative discourse units (ADUs) in debate and editorial this http URL jointly model these two objectives as a token-level sequence labeling task using a BERT-BiLSTM-CRF architecture that combines contextual transformer embeddings with structural transition constraints to support accurate span detection. $\testtt{STAR-Ar}$ achieves an F1-score of 72.69 on validation and 73.7 on test data. Our domain-specific analysis shows that models trained exclusively on editorials underperform those trained on debates, a disparity we primarily attribute to the smaller size of the editorial dataset. The code for $\testtt{STAR-Ar}$ is available at ${\href{this https URL}{\faGithub~TTLab at Daleel 2026}}$
73. 【2609.39369】Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer
链接:https://arxiv.org/abs/2609.39369
作者:Jiahe Fan,Si Chen,Yinghao Hou,Wenbo Xia,Ke Xu,Hong Xie,Enhong Chen
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Specialized models encode, encode task-oriented behavior, models encode task-oriented, Specialized models, requires training
备注: 6 pages, 1 figure, 7 tables. Preprint
点击查看摘要
Abstract:Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.
74. 【2609.39368】Making Grid Beam Search Less Greedy
链接:https://arxiv.org/abs/2609.39368
作者:Sean Papay,Roman Klinger
类目:Computation and Language (cs.CL)
关键词:grid beam search, beam search, grid beam, beam, DFA-constrained beam search
备注: Published as a conference paper at COLM 2026
点击查看摘要
Abstract:A common formalism for constraining the output of autoregressive text generation models involves lexical constraints, words or phrases which are required to occur in the generated text. DFA-constrained beam search and grid beam search are two widely used paradigms for decoding from autoregressive models while enforcing lexical constraints. As the former approach requires a number of forward passes exponential in the number of constraint tokens, it is often dispreferred to the latter, which requires only linearly many forward calls. However, while grid beam search achieves an exponential speedup, it does so in a manner which does not treat all of the constraints equally. In this paper, we demonstrate that grid beam search is biased to incorporate easier-to-satisfy constraints first, leaving harder constraints to the end of the sequence. This contrasts with DFA-constrained beam search, which exhibits no such bias. To address this shortcoming, we propose fair grid beam search, a modification to grid beam search which avoids this bias while still requiring only linearly many forward passes. Experimentally, we confirm grid beam search's bias on two constrained generation tasks, finding significant differences in how it orders constraint tokens as compared to DFA-constrained beam search and fair grid beam search. Furthermore, we find that fair grid beam search not only fixes grid beam search's bias, but finds higher-probability strings in the process.
75. 【2609.39365】Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts
链接:https://arxiv.org/abs/2609.39365
作者:Jeesu Jung,Hwan Chang,Juseon Do,Jeonghwan Choi,Jinho Choo,Sungwoo Nam,S. K. Hong,Hwanjun Song
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:previously acquired behaviors, forgetting previously acquired, alignment requires LLMs, acquired behaviors, requires LLMs
备注: 24 pages
点击查看摘要
Abstract:Continual alignment requires LLMs to adapt to new requirements without forgetting previously acquired behaviors. Natural-language instructions are flexible and composable but offer only indirect control, whereas post-training provides stronger adaptation at the cost of repeated parameter updates. We introduce Ready2Blend, which combines the flexibility of natural language with learned alignment. AlignFormer maps each requirement to a fixed-length alignment prompt stored in a modular prompt bank, while the backbone and prior prompts remain frozen. Composability regularization transfers the semantic geometry of textual requirements into prompt space, enabling inference-time blending and reweighting. Across two practical continual alignment settings, Ready2Blend is the only frozen-backbone method that matches post-training-based alignment methods, reaching $93.1$-$98.5\%$ of a joint-training reference with competitive retention, while requiring only a few prompt tokens and up to $4.3\times$ less training time. Its modular design further enables weighted personalization and order-free composition without retraining. Code will be released upon acceptance.
76. 【2609.39358】Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
链接:https://arxiv.org/abs/2609.39358
作者:Sietse Schelpe
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG); Performance (cs.PF)
关键词:Vishal Sikka, CEO of Infosys, transformer language model, language model performs, model
备注:
点击查看摘要
Abstract:A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and this http URL that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on this http URL at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
77. 【2609.39346】Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models
链接:https://arxiv.org/abs/2609.39346
作者:Bohan Zhang(1),Linan Yue(1),Weibo Gao(2),Pengyu Chen(1),Hong Guo(1),Yanqi Hao(3) ((1) Southeast University, (2) Hong Kong Polytechnic University, (3) ZTE Corporation)
类目:Computation and Language (cs.CL)
关键词:Large language models, small language models, language models, Large language, SLM
备注: 29 pages. Code: [this https URL](https://github.com/ZBH031/reusable-latent-correction)
点击查看摘要
Abstract:Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability-deployment gap has motivated LLM-SLM collaboration, which aims to improve SLM reasoning using LLM capabilities while preserving the deployment advantages of SLMs. Existing approaches mainly follow two paradigms. Knowledge distillation uses LLM-generated answers and reasoning trajectories to train SLMs offline, but requires parameter updates and additional training. Alternatively, online collaboration routes difficult problems to an LLM or leverages LLM-generated guidance and corrections when an SLM encounters difficulties. Although effective, online collaboration requires repeated LLM access. Moreover, the guidance produced for a particular problem is discarded after inference and cannot benefit subsequent problems involving similar reasoning states. In the paper, we focus on a more constrained setting in which the LLM is accessed only offline, the SLM parameters remain fixed, and online inference is performed solely by the SLM. To this end, we propose Reusable Latent Correction (RLC), which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM. RLC stores these experiences in an external bank and retrieves them according to the SLM's current reasoning state, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls. Experiments across multiple reasoning benchmarks and SLM scales show that RLC consistently improves SLM reasoning without parameter updates or online LLM calls. Code is available at this https URL.
78. 【2609.39341】Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models
链接:https://arxiv.org/abs/2609.39341
作者:Daniel Dragonevskiy
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:predict tokens, language model, model, logical, Arbitr
备注: 18 pages
点击查看摘要
Abstract:Does a language model merely predict tokens, or does it understand what it says? We make this question measurable by defining "understanding" through the lens of no-arbitrage. A model understands a vocabulary to a certain degree if a computationally bounded trader cannot extract guaranteed profit by betting against the model's probabilities on logically related claims (a "Dutch book"). We establish three theoretical results: first, because full logical coherence is computationally intractable, understanding is inherently graded, not absolute. Second, we prove that the exact optimum of standard next-token prediction is inherently incoherent across different question formats; the flaw lies in the training objective, not the architecture. Third, we show that uncertainty accumulates predictably along reasoning chains, making unjustified overconfidence an arbitrage opportunity in itself. To address this, we introduce Arbitr, a training framework where an adversarial trader penalizes the model for logical inconsistencies, paired with a calibration anchor to prevent uninformative collapse. Across five pre-registered experiments on Qwen2.5 and Phi-3.5 models, we demonstrate that standard models are highly exploitable across different phrasings. Arbitr reduces this exploitability by orders of magnitude without sacrificing task accuracy, and the effect successfully transfers to unseen logical patterns and new model families. Crucially, we uncover a scaling illusion: at 7B parameters, near-zero measured incoherence often coincides with extreme, unjustified confidence. We conclude that while Arbitr enforces rigorous logical consistency, coherence is a necessary condition for knowledge, but not a sufficient one
79. 【2609.39334】aming Speculative Search for Test-Time Scaling in LLM Serving
链接:https://arxiv.org/abs/2609.39334
作者:Jinwoo Jeong(Korea University),Woohyung Choi(Korea University),Myeongjae Jeon(POSTECH),Jeongseob Ahn(Korea University)
类目:Distributed, Parallel, and Cluster Computing (cs.DC); Computation and Language (cs.CL); Operating Systems (cs.OS)
关键词:substantially enhancing accuracy, Test-time scaling, improving LLM reasoning, allocating additional computation, substantially enhancing
备注: 14 pages
点击查看摘要
Abstract:Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks such as mathematics and coding. To accelerate the exploration of reasoning paths, recent studies proposed speculative execution. However, we show that supporting speculative execution poses two unique challenges for LLM serving systems: (1) an explosion in the search space of candidate paths and (2) frequent, fine-grained verification tasks for candidates. To address these challenges, this paper proposes SpecScale, a serving system for efficient speculative execution. We introduce three techniques to reconcile the trade-off between latency and computational overhead: (1) early pruning of low-quality candidate paths, (2) deduplicating computation across redundant candidate paths, and (3) deferring fine-grained verification tasks. We evaluate SpecScale on challenging reasoning benchmarks, including MATH and Olympiad. Our results show that SpecScale significantly outperforms both non-speculative and recent speculative approaches, delivering substantial improvements in throughput and latency while preserving answer quality.
Comments:
14 pages
Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC); Computation and Language (cs.CL); Operating Systems (cs.OS)
Cite as:
arXiv:2609.39334 [cs.DC]
(or
arXiv:2609.39334v1 [cs.DC] for this version)
https://doi.org/10.48550/arXiv.2609.39334
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
80. 【2609.39333】NarrativeSteward: Coordinating Delegation, Guidance, and Verification in Agent-Assisted Interactive Narrative Authoring
链接:https://arxiv.org/abs/2609.39333
作者:Wenjin Wang,Jiazhen Lei,Yuxin Sha,Nuwa Xi,Meng Zhao,Xingxi Yin,Qi Liu,Yuliang Shen,Zixun Sun
类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:turn authors' goals, turn authors', authors' goals, goals into interactive, independently organizing
备注:
点击查看摘要
Abstract:Autonomous AI agents can turn authors' goals into interactive narratives by independently organizing and carrying out generation and revision. As agents generate and revise extensive content, authors struggle to grasp its overall structure, local details, and relationships, complicating continued guidance. We present NarrativeSteward, an authoring environment that organizes outlines, worldbuilding, and narrative graphs as linked artifacts for agent implementation and author guidance. Agent dialogue and project-wide structural review help authors understand the evolving work and guide local and cross-layer revisions, while change records and execution verification help authors assess the resulting work. Technical tests validated the system's change records, recovery mechanisms, and execution diagnostics. In a 12-participant within-subject study, NarrativeSteward supported easier formulation of revision requests and inspection of changes, and greater perceived understanding of changes and story structure, than general-purpose agents. Qualitative findings show how reviewing the work and feedback helps authors develop requirements and guide subsequent delegation. We open-source NarrativeSteward at this https URL.
81. 【2609.39277】A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models
链接:https://arxiv.org/abs/2609.39277
作者:Steven Kolawole,Pearse Jim,Opegbemi M. Busoye,Glory Bagai,Virginia Smith
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:block saves memory, Looped models reason, saves memory traffic, block saves, Looped models
备注: Preprint; in review
点击查看摘要
Abstract:Looped models reason by applying the same block of weights many times, so compressing that block saves memory traffic on every loop. Compressed looped models, however, often collapse, and the collapse is usually blamed on rounding error that accumulates from loop to loop. In this work we test that account on more than 30 models from five families and find, to our surprise, that it holds only for loops that never settle. When a loop settles, a fixed rounding error does not accumulate. It moves the point where the loop settles, much as tilting a bowl moves where a ball comes to rest, and the answer is lost only when the shift is larger than the readout tolerates. This picture lets us predict which models fail from a single label-free measurement, and it tells us why failed models recover: their loops still settle, so a few final loops with 8-bit weights bring the answer back. Motivated by these findings, we build a controller that stops when the model's halting head fires and then finishes with 8-bit loops. On Sudoku-Extreme and Maze-Hard it beats fixed-depth inference by up to 15 points under a third of the weight traffic.
82. 【2609.39263】Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout
链接:https://arxiv.org/abs/2609.39263
作者:Aojie Yuan,Zhiyuan Julian Su,Haiyue Zhang,Zijian Su
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:extracted subspace overlap, subspace, Format-Agnostic Reasoning Subspace, output readout, Reasoning Subspace
备注: 54 pages. Substantially revised preprint: new title, expanded model coverage, readout-geometry controls, supplementary intervention and transfer experiments, revised interpretation, updated figures and author list
点击查看摘要
Abstract:A concept subspace's effect on model behavior does not establish how it relates to the output readout. We introduce a two-sided geometric diagnostic that measures an extracted subspace's overlap with the dominant right-singular directions of the unembedding matrix, evaluated against output-oriented positive controls. Given an extracted basis, the raw diagnostic requires only model weights. Our testbed is the Format-Agnostic Reasoning Subspace (FARS), a ten-dimensional basis extracted from eighteen reasoning concepts expressed in six surface forms. Across nine rank-matched estimators and twenty-six models, four activation-derived concept estimators carry only 0.38--0.80% mean energy in the top-ten readout span. Final-layer PCA carries 3.56%, exceeding FARS in 25 of 26 models. A same-layer next-token control, evaluated using a fitted linear translator for depth matching, carries approximately thirteen times more energy than FARS, with separation in all 25 tested models. Re-extracting FARS on ten disjoint concepts yields 62--100% cross-format retrieval across twenty-four generative models, demonstrating transfer of the extraction procedure rather than a fixed basis. A complementary four-model, three-seed intervention study finds model-dependent source-directed effects that remain well below full-vector replacement. Together, the geometry and intervention controls distinguish concept structure from dominant readout directions while limiting claims of causal sufficiency.
83. 【2609.39238】4MT-VLM: How Coarse Is a VLMs Cognitive Map?
链接:https://arxiv.org/abs/2609.39238
作者:Markus Frey
类目:Computation and Language (cs.CL)
关键词:recognise a place, bare terrain peaks, procedurally generated landscapes, holding layout fixed, remove appearance cues
备注:
点击查看摘要
Abstract:An agent that moves must recognise a place from a viewpoint it has never seen. We introduce 4MT-VLM, a dataset of procedurally generated landscapes, each rendered across five stimulus modes that remove appearance cues while holding layout fixed: shape and colour, shape only, colour only, bare terrain peaks with no objects, and a valley viewpoint that puts the peaks on the horizon. The last condition is commonly used in clinics to probe hippocampal function in human patients. We test this benchmark across sixteen different open and closed-source models and report 4AFC performance, a measure which is also used to grade human participants. We observe that models identify a place from the studied viewpoint but lose it once the camera moves, dropping below the 25% chance level at 135° where a human observer scores 85%. Frontier models (Gemini 3.8 Flash, GPT-5.6) answer only 39% and 31% of rotated trials correctly, recovering to 85% and 55% only when distractors are moved more than 30 meters apart. Our benchmark demonstrates that while current VLMs possess rudimentary cognitive maps, their spatial resolution remains fundamentally too coarse to maintain a stable, 3D understanding of the world once the viewpoint changes.
84. 【2609.39229】RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection
链接:https://arxiv.org/abs/2609.39229
作者:Elia Onofri,Roberto Di Pietro
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:language model acting, large language model, proprietary frontier models, Automatic evaluation, faithfulness increasingly relies
备注: 49 pages, 23 tables, 10 figures. Code and data: [this https URL](https://github.com/eOnofri04/raim-analysis) and [this https URL](https://github.com/eOnofri04/raim-verdicts)
点击查看摘要
Abstract:Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models, costly and ill-suited to high-throughput monitoring. We investigate whether a panel of cheap open-weight judges (4--9B) can be aggregated to stand in for a frontier one, what the substitution sacrifices, and when it is worth making. We propose RAIM, an aggregation scheme robust to the members' correlated errors, coupling a cross-fitted stacked logistic regression with an admissibility test that, read from the members' own outputs, identifies when aggregating them improves on their best member and stays within reach of the frontier judge. We instantiate RAIM with ten judges from disjoint families across eight faithfulness benchmarks. Against Claude Sonnet, the panel retains a median 93% of its Cohen's $\kappa$ and gives up only 2.9 points of balanced accuracy on average; read as paired differences, it clearly improves on one benchmark and clearly worsens on three (only two by a non-negligible margin), leaving four unresolved. At a sixty-fourth of the frontier's inference price, the operative expense is a one-time in-domain calibration on 50--100 labelled records. The panel is also competitive with purpose-trained detectors on their home benchmarks (within 1.3 accuracy points of GPT-4o and 1.9 of the LLM-AggreFact leader), and beats the strongest one we reran by 6 points on our grounded sets. Whether aggregation pays depends on the members themselves: where several capable members err on different items, the panel improves on its best judge and approaches the frontier; where one dominates, the stacker recovers the leader, and only there does the frontier remain materially ahead. Both conditions are read off the calibration set at no further cost, so a cheap panel can stand in for a frontier one wherever this audit admits it.
85. 【2609.39225】Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures
链接:https://arxiv.org/abs/2609.39225
作者:Siddharth Bhargava,Sara Tonelli,Patricia Martín-Rodilla,Javier Parapar
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Argument structure prediction, constructs complete argument, complete argument structures, argument structures, Inference Anchoring Theory
备注: CMNA'26: 26th International Workshop on Computational Models of Natural Argument
点击查看摘要
Abstract:Argument structure prediction (ASP) constructs complete argument structures from discourse by identifying argumentative units and their relations. While recent work has explored diverse approaches---including unified neural models, multi-step pipelines, and prompt-based large language models (LLMs)---their relative trade-offs remain under-explored, particularly in dialogical settings. We present a systematic evaluation of ASP under strict schema constraints, comparing supervised fine-tuning and prompt-based LLMs across single- and multi-step task architectures, generating complete argument structures from dialogical input end-to-end. We benchmark them on three diverse dialogical corpora adapted from Inference Anchoring Theory into bipolar argument structures. Under a shared evaluation framework, we assess predictive performance, cross-domain generalization, schema compliance, and computational efficiency. Our results show that ASP remains a challenging task, with identifying argumentative relations emerging as the primary bottleneck, largely due to the implicit and context-dependent nature of dialogical argumentation. To facilitate future research, we release our data processing pipeline and end-to-end modeling framework for computational ASP on dialogical corpora.
Comments:
CMNA’26: 26th International Workshop on Computational Models of Natural Argument
Subjects:
Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:
arXiv:2609.39225 [cs.CL]
(or
arXiv:2609.39225v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.39225
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
86. 【2609.39189】ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations
链接:https://arxiv.org/abs/2609.39189
作者:Dat Tien Nguyen,Nghia Hieu Nguyen,Anh Thi-Hoang Nguyen,Dung Ha Nguyen,Kiet Van Nguyen,Ngan Luu-Thuy Nguyen
类目:Computation and Language (cs.CL)
关键词:Trustworthy Legal, Legal, requires systems, grounding their responses, Vietnamese legal
备注:
点击查看摘要
Abstract:Trustworthy Legal AI requires systems that can answer legal questions while grounding their responses in authoritative sources. However, existing Vietnamese legal benchmarks provide limited coverage of real-world legal consultations. We introduce \textbf{ViLegalExpert}, a large-scale benchmark constructed from authentic citizen--lawyer consultations, containing over \textbf{172K} questions across \textbf{34 legal domains}, together with professional answers and expert-verified legal evidence. ViLegalExpert supports legal information retrieval, extractive QA, and abstractive QA. Experiments with representative retrieval methods and language models reveal substantial challenges in evidence retrieval and grounded answer generation. While pretrained models perform strongly on QA, hybrid retrieval achieves the best retrieval performance. These results demonstrate the difficulty of mapping naturally expressed legal questions to authoritative provisions and establish ViLegalExpert as a challenging benchmark for reliable Vietnamese Legal AI.
87. 【2609.39154】DAGent: Evaluate-then-Grow Planning for Deep Research Agents
链接:https://arxiv.org/abs/2609.39154
作者:Hanwen Liu,Yuanfu Sun,Qiaoyu Tan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large knowledge spaces, navigate large knowledge, knowledge spaces, findings emerge, tasks require agents
备注: Accepted at NeurIPS 2026
点击查看摘要
Abstract:Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: this https URL
88. 【2609.39118】Diagnosing On-Policy Self-Distillation for Reasoning Language Models
链接:https://arxiv.org/abs/2609.39118
作者:Yang Li,Gongle Xue,Yuheng Yuan,Yijia Guo,Shizhe Zhang,Liwen Hu,Lei Ma
类目:Computation and Language (cs.CL)
关键词:On-policy self-distillation, attracted growing interest, attracted growing, growing interest, promising approach
备注:
点击查看摘要
Abstract:On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavior in language reasoning remains unclear, with reported outcomes ranging from modest gains to behavioral collapse. In this work, we diagnose OPSD for mathematical reasoning across models spanning 0.6B--8B parameters. We conduct controlled experiments and token-level analyses to fully delve into OPSD. We point out that teacher's signal is shaped by reasoning-mode alignment and the complete teacher prefix, rather than by privileged semantics alone. OPSD improves reasoning only in narrow compatibility regimes. Otherwise, it produces ineffective length growth, stable degradation, or behavioral collapse. Token-level analysis shows that teacher's signal is not stable and does not predict downstream performance. Based on these results, we argue that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method.
89. 【2609.39111】Bongard: Training Machine Intuition
链接:https://arxiv.org/abs/2609.39111
作者:Li Ding,Haidi Jin,Chen Ji
类目:Computation and Language (cs.CL)
关键词:Human intelligence relies, intelligence relies heavily, Human intelligence, recognising patterns, intermediate step
备注: Technical report, 28 pages, 7 figures. Model weights: [this https URL](https://huggingface.co/AgentBull/bongard-mini)
点击查看摘要
Abstract:Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate step. We introduce Bongard, an open-weight System One model that treats machine intuition as an independent capability to design and train. A T5Gemma 2 4B-4B encoder-decoder separates reading the evidence from making judgments. The encoder reads the state bidirectionally together with the question instructions, and separate decoder branches share this encoding, so many judgments about the same situation require only one reading of the state. A trained head returns probabilities over the supplied candidates without generating text. Training proceeds in three stages, from supervised judgments to semantic relationships to action outcomes, and each stage updates all 7.09 billion trainable parameters on one Blackwell GPU. Joint-embedding post-training raises accuracy on held-out rephrasings from 75.7% to 85.9%. A sandbox stage then learns outcome distributions from action rollouts and exact oracles, raising accuracy on a frozen sandbox panel from 50.6% to 64.8%. On DecisionBench, the final model reaches 78.05% accuracy over 23,900 decisions and ranks fourth of 61 systems in the public comparison. On one RTX PRO 6000, its median latency is 36 ms for short requests, and 32 questions about one state take 221 ms. Bongard demonstrates that machine intuition can be systematically trained via representation learning and outcome feedback, providing an open, efficient alternative for high-throughput decision workloads.
90. 【2609.39102】False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
链接:https://arxiv.org/abs/2609.39102
作者:Meijia Chen,Hao Li,Zheng Lu,Hongshan Lin,Junbai Tian,Yichen Liu,Zijun Tian,Yufan Zou,Shuhan Sun,Hanxin Chen,Zeyu Zhang,Weizhi Du,Yueting Li,Tianyu Shi,Alaa Khamis
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Self-evolving search agents, search agents build, agents build, curricula by jointly, jointly optimizing
备注: 21 pages. Equal contribution: Meijia Chen, Hao Li, Zheng Lu
点击查看摘要
Abstract:Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
91. 【2609.39072】Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue
链接:https://arxiv.org/abs/2609.39072
作者:Yutong Hu,Jinho Choi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
关键词:applying Large Language, remains largely unexplored, Large Language Models, multimodal dialogue remains, dialogue remains largely
备注: 15 pages, 6 figures, 11 tables
点击查看摘要
Abstract:Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.
92. 【2609.39071】LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
链接:https://arxiv.org/abs/2609.39071
作者:Yida Cai,Xin Dai,Bingxiang He,Huiyuan Xie,Yuxiao Ye,Zhenghao Liu,Yang Bai,Zhiyuan Liu
类目:Computation and Language (cs.CL)
关键词:language models require, require reward signals, signals that capture, Legal, Legal language models
备注:
点击查看摘要
Abstract:Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.
93. 【2609.39069】CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search
链接:https://arxiv.org/abs/2609.39069
作者:Siyu Song,Rui Xu,Jia Lin,Kai Liu,Weifang Wang
类目:Computation and Language (cs.CL)
关键词:Test-time reasoning systems, earlier decision caused, caused the error, Test-time reasoning, systems often respond
备注:
点击查看摘要
Abstract:Test-time reasoning systems often respond to failure by restarting or revising the latest step, even when an earlier decision caused the error. We introduce CORE, a search controller that requests a certified conflict core from a verifier, backjumps to the latest decision in that core, and caches the conflict to avoid repeating it. Under sound verification, finite branching and depth, and exhaustive proposals, the uncapped search is complete and never prunes a valid solution. On 2,000 planted graph-coloring instances with matched proposals and an exact verifier, CORE reduces median verifier calls by 39.8% at 30 variables and 35.0% at 36 variables relative to chronological repair; caching further improves on backjumping alone. Across five reasoning tasks, CORE achieves 75.9% mean success with Qwen2.5-7B-Instruct and 84.2% with Qwen3-8B, compared with 72.5% and 81.8% for Tree of Thoughts. It also uses fewer verifier calls and generated tokens on both backbones. These results show the value of using certified failure explanations to direct language-model search.
94. 【2609.39050】Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
链接:https://arxiv.org/abs/2609.39050
作者:Deema Alnuhait,Gengyu Wang,Muhammad Khalifa,Hao Peng
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:enter high-stakes domains, multi-agent systems enter, systems enter high-stakes, high-stakes domains, growing concern
备注:
点击查看摘要
Abstract:As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show that benign agents can cross the same boundaries without adversarial incentives. We emulate a software-engineering workflow in which a planner represents a company hiring an external developer. The planner writes requirements and holds a company credential it is instructed not to disclose to the developer; a monitor screens their exchanges. Seven of nine tested frontier models disguise the credential in their requirements to help the developer recover it while evading the monitor, even after completing their assigned objective. For example, across 6,000 episodes with DeepSeek-V4-Pro, the planner attempts concealment in 16.9%; in 0.9%, the credential evades the monitor and is recovered and used by the developer. Despite the low per-episode rate, the risk compounds across repeated exchanges: at 0.9%, 105 independent episodes carry a 61.3% chance of at least one breach. Our analysis suggests that the planner misinterprets disguised disclosure as helping the developer and reads the nondisclosure rule as banning plaintext, not character codes or riddles. We call this behavior covert assistance. Monitoring favors concealment, while concealment gives way to direct disclosure when the recipient is presented as a human rather than an AI agent, so leakage persists. Giving the monitor the credential improves detection but requires trusting it with the secret. These risks, in models already used for software engineering, challenge oversight to distinguish authorized cooperation from task-advancing assistance that crosses safety boundaries.
95. 【2609.39049】Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
链接:https://arxiv.org/abs/2609.39049
作者:Xinkai Chen
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large language model, rate depression severity, social media post, large language, social media
备注: Extended version of a paper accepted at MHSM 2026 (IEEE ICDM 2026 workshop). 14 pages, 1 figure. Code: [this https URL](https://github.com/xinkaichen97/depseverity-artifact)
点击查看摘要
Abstract:A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.
96. 【2609.39045】RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
链接:https://arxiv.org/abs/2609.39045
作者:Wenyi Wu,Minghao Fu,Jieyu You,Kun Zhou,Siqi Liu,Aayush Salvi,Yiheng Lin,Ce Zhang,Xiaohan Lan,Jiahui Zhu,Yujie Zhong,Qi She,Biwei Huang
类目:Computation and Language (cs.CL); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
关键词:version remains challenging, large language models, reliably improving generated, playable version remains, Recent advances
备注:
点击查看摘要
Abstract:Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
97. 【2609.39034】Switching Linear Attention
链接:https://arxiv.org/abs/2609.39034
作者:Hyun Dong Lee,Xavier Gonzalez,Nicolas Zucchet,E. Kelly Buchanan,Emily B. Fox,Scott W. Linderman
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Designing expressive sequence, Designing expressive, efficient inference remains, modern machine learning, inference remains
备注: COLM 2026
点击查看摘要
Abstract:Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.
98. 【2609.39027】A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
链接:https://arxiv.org/abs/2609.39027
作者:Chenguang Wang,Ming Li,Chengrui Fan,Jianpeng Chen,Han Chen,Tianyi Zhou,Dawei Zhou
类目:Computation and Language (cs.CL)
关键词:potentially rewarding rhetorical, rewarding rhetorical optimization, potentially rewarding, scientific improvement, optimization over scientific
备注: 35 pages, 2 figures, 20 tables. Accepted (Oral) at AI-Native Academia @ NeurIPS 2026
点击查看摘要
Abstract:AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.
99. 【2609.39013】Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem
链接:https://arxiv.org/abs/2609.39013
作者:Divya Godara,Sachin Gupta
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:DocSem shared task, official final test, final test evaluation, shared task, DocSem shared
备注: 5 pages, 1 figure, 2 tables. Accepted as a shared-task system paper at DocInsights 2026, co-located with EMNLP 2026
点击查看摘要
Abstract:EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and evaluates that expression in local code. Saved intermediate results support inspection of failures. A separate public-validation run achieved 92.17% answer accuracy and 1.00 evidence F1. The configurations and metrics differ, so these scores are not a controlled comparison. Our manual, post-hoc analysis is descriptive: in one inspected case, optical character recognition (OCR) and block grouping merged the relevant passage into another block, and the system answered from unrelated text. An exploratory study of reading page images on 100 documents returned evidence identifiers for only 22 documents. These descriptive findings motivate further evaluation; they do not establish the causes of the overall score.
100. 【2609.39001】he Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype
链接:https://arxiv.org/abs/2609.39001
作者:Thomas Serval
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:LLM services, Anthropic counting API, services are billed, content varies, Universal Declaration
备注: 11 pages, 5 figures, 7 tables. Code, tokenizer, per-sentence counts and controls: [this https URL](https://github.com/Baracoda-ai-labs/baracoda-fr) (commit fc8a736)
点击查看摘要
Abstract:LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across languages. We measure this token premium on seven tokenizers of widely used 2026 models (OpenAI o200k, Llama 3, Qwen3, DeepSeek V3/V4, Gemma 3, Mistral Tekken, and the Claude generation-5 tokenizer via Anthropic's counting API) on NTREX-128 (124 non-English reference translations) and on the Universal Declaration of Human Rights for regional languages. French requires 31% to 58% more tokens than English, whereas Simplified Chinese ranges from 5% fewer to 40% more and is cheaper than French on six of the seven tokenizers. Regional and overseas languages of France pay roughly 1.6 to 3.3 times the English count. We discuss how history re-sending, tiered pricing and fixed context windows amplify the absolute gap in agentic use. In a controlled experiment (BPE, Europarl, 50k vocabulary), adding French to tokenizer training data quickly reduces the premium, with diminishing returns and a growing cost for English. Finally, we present Baracoda FR v1.2, a byte-level BPE prototype with Tekken's vocabulary size. On a final test of six corpora never consulted during design, with a protocol declared fixed beforehand, it uses 11.5% fewer tokens than Tekken on French and 3.7% fewer on English; results hold after removing test sentences overlapping the training data and with an equal ordinary-token budget. It is worse on other languages and, at comparable vocabulary size, does not outperform CroissantLLM. These are segmentation results only; effects on model quality and task cost remain to be shown.
101. 【2609.38997】Settle: Learning When to Stop Reasoning
链接:https://arxiv.org/abs/2609.38997
作者:Ryan Brown,Zihao Fu,Chris Russell
类目:Computation and Language (cs.CL)
关键词:Reasoning models, continue generating, Reasoning, Settle, Abstract
备注: 30 pages, 4 figures
点击查看摘要
Abstract:Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwen3-4B, Settle reduces token count by 40% with a 0.5-percentage-point decrease in accuracy. It gains 6.16 percentage points over supervised fine-tuning on the same traces shortened at their first stable answer, at nearly identical token counts. Its stopping score predicts whether a correct answer will remain correct. Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods.
102. 【2609.38995】When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation
链接:https://arxiv.org/abs/2609.38995
作者:Di Huang,Hao Li,Yixin Chen,Fuhai Li
类目:Computation and Language (cs.CL)
关键词:On-policy self-distillation, privileged information, original OPSD study, OPSD study finds, model conditioned
备注: 21 pages, 5 figures
点击查看摘要
Abstract:On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged information. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math-related tokens, and that pointwise clipping of the forward KL objective stabilizes training. Pointwise clipping caps each vocabulary-wise forward KL term at a fixed threshold before summing over the vocabulary. Follow-up studies have adopted this clipping, but its effect on training has not been directly examined. In matched training runs differing only in whether clipping is applied, we observe that clipped runs produce substantially more repetitions that persist to the end of the response than their unclipped counterparts. We trace this failure to the clipped objective. We prove that the clipped objective can fail to correct the student toward the teacher and can instead push clipped and unclipped token probabilities away from its teacher. Our training runs agree with this analysis: inside repetitions, the clipped student places less probability than its teacher on leaving the repetition, and more on continuing it, whereas the unclipped runs stay close to their teachers.
103. 【2609.38976】Fairness Beyond a Single Run: Training-Seed Variability in Speech LLM Adaptation
链接:https://arxiv.org/abs/2609.38976
作者:Srishti Ginjala,Eric Fosler-Lussier,Srinivasan Parthasarathy
类目:Computation and Language (cs.CL)
关键词:automatic speech recognition, single training run, Common Voice, single training, Common Voice spreads
备注:
点击查看摘要
Abstract:Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector and LoRA adapters of a speech LLM at five audio compression factors and six random seeds, holding the encoder, base decoder, data and decoding fixed, and evaluate every run on Common Voice and Fair-Speech. At 460 h of clean LibriSpeech, the seed moves fairness metrics more than compression does on most demographic axes. A balanced 3x3 decomposition attributes 85.3% of the variation in Fair-Speech ethnicity normalized gap to the seed against 8.3% to compression (p = 0.009), though compression explains more on age and gender. Held-out LibriSpeech word error rate spreads by 0.04 points across those seeds while Common Voice spreads by 8.57, so these are not failed runs, and the effect survives controlling for accuracy and dropout. Scaling and diversifying the adaptation set to 960 h damps the effect but does not remove it. On Fair-Speech ethnicity, two single-run systems must differ by more than 0.30 in normalized gap to exceed seed variability.
104. 【2609.38972】Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
链接:https://arxiv.org/abs/2609.38972
作者:Yihuai Hong,Shauli Ravfogel,Chen Zhao,Eunsol Choi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Large Language, Language Models, Large, Language
备注: 28 pages, 9 figures, 10 tables
点击查看摘要
Abstract:Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignment between the reasoning described in an LLM's CoT and what it computes internally. We propose CoT-Interpretability Alignment (CIA), a metric that measures the agreement between a model's CoT traces and its internal reasoning strategies as detected by interpretability tools. We evaluate CIA on three tasks (two-hop question answering, hint intervention, and integer multiplication) across three LLMs, finding that LLMs exhibit limited alignment across all tasks (44.8-75.9%). We then experiment with improving CIA via post-training, setting both the task accuracy and parametric faithfulness signals as a reward. Experiments show that we can substantially improve CoT parametric faithfulness while maintaining or improving the task accuracy. We provide rich analysis, such as their generalization patterns. Our work provides both a framework for auditing CoT parametric faithfulness and a pathway toward making models' explicit reasoning more trustworthy. Code and data are available at this https URL.
105. 【2609.38958】argeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting
链接:https://arxiv.org/abs/2609.38958
作者:Liang Twist Shan,Tianyu Hu,Hao Yan,Yiqiao Zhong
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Applications (stat.AP)
关键词:Large language models, Large language, rapidly improving, Large, Thinking
备注: 73 pages, including references and appendices
点击查看摘要
Abstract:Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count the number of records dispersed in a long text. Across twelve model comparison groups, Thinking (or reasoning) improves counting accuracy over Non-thinking, with pronounced gains at larger counts. This motivates our mechanistic analysis, which identifies two contrasting mechanisms: (i) broad retrieval, where Non-thinking models broadly attend to multiple needles; (ii) targeted retrieval, where Thinking models use enumeration in CoT traces to successively retrieve needles. Targeted retrieval concentrates attention on individual needles and is accompanied by more compact internal representations. Moreover, causal intervention analysis suggests that Thinking models use the CoT trace to maintain and update an internal counter as needles are successively retrieved, even without explicit numbering. In small controlled experiments, both retrieval mechanisms and counter states emerge under standard autoregressive training. Together, our results connect long-context retrieval with representation geometry of counting, supporting a state-tracking account of CoT reasoning.
106. 【2609.38923】GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
链接:https://arxiv.org/abs/2609.38923
作者:Qisheng Su,Hanchen Wang,Guanru Zhu,Huicheng Jiang,Qiuyinzhe Zhang,Kou Shi,Zhen Fang,Ziao Zhang,Qingnan Ren,Zehui Chen,Tao Gui,Feng Zhao
类目:Computation and Language (cs.CL)
关键词:read diverse files, Working agents, coordinate tools, produce deliverables, real files
备注:
点击查看摘要
Abstract:Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.
107. 【2609.38898】K2P: Label-Free Knowledge to Prompt Distillation
链接:https://arxiv.org/abs/2609.38898
作者:Yingchuan Zhang,Haoran Lu,Wenxuan Zhong,Ping Ma
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:avoiding weight updates, eliminate supervision, Knowledge distillation, avoiding weight, weight updates
备注: 66 pages, 5 figures
点击查看摘要
Abstract:Knowledge distillation can transfer reasoning from stronger teachers to frozen students through reusable prompts, but avoiding weight updates does not eliminate supervision. Without ground-truth answers, teacher solutions are unverified, and agreement with the teacher can reward shared mistakes. We introduce Knowledge-to-Prompt (K2P) for label-free knowledge distillation to prompts. K2P synthesizes reusable instructions from teacher solutions, refines them using paired teacher and student responses, and guides search and selection with answer agreement. It retains candidates that adaptive search may undervalue and selects on reserved questions. Deployment uses only the frozen student and selected prompt. Our theory separates generation and selection gaps and gives conditions under which agreement-guided construction yields accuracy guarantees despite imperfect teacher references. Across reasoning tasks and students, K2P outperforms label-free alternatives overall and remains competitive with supervised prompt optimization. Ablations and archive diagnostics assess the contributions of teacher solutions and refinement, while revealing the limits of agreement-guided selection.
108. 【2609.38878】Audio Token Attention Is Predictable Before the Language Model Runs
链接:https://arxiv.org/abs/2609.38878
作者:Kyoungjun Park,Yunzhe Li,Lili Qiu
类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:language model, language model runs, large audio language, language, model
备注: 43 pages, 7 figures
点击查看摘要
Abstract:A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $\rho \geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: this https URL
109. 【2609.38861】RACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA
链接:https://arxiv.org/abs/2609.38861
作者:Sachin Gupta,Divya Godara
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
关键词:Finding a relevant, producing a verifiable, verifiable answer, Finding, relevant paper
备注: 8 pages, 3 figures, 3 tables. Accepted at the 1st Workshop on Grounding Language Models: Learning Faithfully and Efficiently (GroundLM 2026), co-located with EMNLP 2026
点击查看摘要
Abstract:Finding a relevant paper is not the same as producing a verifiable answer from it. LitTraceQA requires canonical paper identifiers, exact evidence at the page or object level, and typed answers that match the evaluator. We call the separation between source access and scorer-visible correctness the grounding contract gap. TRACE - Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction - addresses this gap with target-grouped retrieval, independent typed evidence localization, multimodal table extraction, schema-driven table construction, and fail-closed validation. It indexes 27,487 papers through passage, object, alias, citation, and dense representations while retaining the question target behind each signal. For tables, TRACE predicts the observation unit before extracting values and assembles rows with evaluator-compatible key normalization. Our audited selected clean-track artifact scores 0.760613 on the official 71-question test set, including 0.9728 paper F1, 0.6847 evidence F1, 0.9800 multiple-choice accuracy, 0.5423 table-row F1, and 0.3508 macro cell accuracy. On 11 public-development table records, a clean baseline and coordinate-aware visual fill obtain row F1 of 0.291 and 0.411, respectively; this diagnostic comparison includes fallback outputs and is not an official-test claim. Remaining errors chiefly concern locator, observation-unit, row-key, and source-value identity.
110. 【2609.38851】Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis
链接:https://arxiv.org/abs/2609.38851
作者:Xia Hu,Brian Potetz,Chun-Ta Lu,Huanfen Yao,Leonidas Guibas,Zhicheng Wang,Howard Zhou,Pengfei Xing,Andrew Gallagher
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:compositional tasks records, reflects an intrinsic, intrinsic deficit, failure reflects, targeted capability
备注:
点击查看摘要
Abstract:End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54\% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84\% of the gain from supplying all prerequisites.
111. 【2609.38850】OpenJev-RLCD: A Working RLCD Implementation
链接:https://arxiv.org/abs/2609.38850
作者:Zhimin Gao,Pichao Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Jev answer questions, questions with probabilities, Jev answer, Jev, RLCD
备注:
点击查看摘要
Abstract:Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \gvTwoCovFive\% of the items at $\le$5\% error, versus \gvGrpoCovFive\% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: this https URL.
112. 【2609.38832】Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
链接:https://arxiv.org/abs/2609.38832
作者:Zizhuo Fu,Runsheng Wang,Meng Li
类目:Computation and Language (cs.CL)
关键词:language model quality, Scaling, improve language model, histories makes additional, additional heads costly
备注:
点击查看摘要
Abstract:Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parameter scaling can directly enable efficient and effective context scaling. We introduce NAMOH, an architecture-native sparse attention mechanism that activates $K$ of $H$ heads per token. Each head retains only its assigned tokens and performs causal attention within this subsequence. Head selection thus jointly determines active parameters and available context without scanning the full history. Under balanced assignments, increasing $H$ at fixed $K$ shortens head histories and reduces per-token key-value (KV) access without increasing total KV storage. We further support head-relative rotary position embeddings to shorten positional spans within routed subsequences, aiming to mitigate position-induced attention noise. Experiments show that NAMOH can outperform fully activated models with the same total parameters, while enabling more efficient long-context inference than smaller dense models with matched active parameter counts. It remains compatible with GQA and existing sparse attention mechanisms. We hope this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.
113. 【2609.38831】Forging LLM Authorship Fingerprints with Targeted Rewriting
链接:https://arxiv.org/abs/2609.38831
作者:Haohan Yuan,Simin Chen,Xi Niu,Hanqing Guo,Depeng Xu,Haopeng Zhang
类目:Computation and Language (cs.CL)
关键词:making model-specific writing, model-specific writing patterns, Model-attribution classifiers, language model produced, making model-specific
备注: 31 pages, 7 figures, 25 tables. Project page: [this https URL](https://haohanyuan01.github.io/ForgePrint/)
点击查看摘要
Abstract:Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this problem as targeted fingerprint transfer: rewriting one model's output so that attribution classifiers assign it to a chosen target model. We study summarization, where different models receive the same document and express the same underlying content, providing a controlled setting for conditional generation. We introduce ForgePrint, a search-then-distil framework that first searches for rewrites that move attribution toward a target fingerprint, then distils the selected rewrites into a one-pass 4B Student model. On CNN/DM, the Student reaches 70.2% target success rate, outperforming both its Teacher (54.1%) and the strongest of six published rewriting baselines (39.3%), against held-out classifiers that are never queried by the attack. It also reaches 68.3% target success when transferring summaries from an open model toward chosen commercial models. These results show that fingerprint detectability should not be conflated with source authenticity, and that text-only attribution can provide misleading evidence of model identity under targeted rewriting, even when it is accurate on unmodified text.
114. 【2609.38820】BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects
链接:https://arxiv.org/abs/2609.38820
作者:Ali Almutairi,Gelareh Mohammadi,Imran Razzak,Aditya Joshi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Arabic, Arabic NLP, rapid growth, Arabic tasks, Replaced Representation learning
备注:
点击查看摘要
Abstract:With the rapid growth of Arabic NLP, several models, datasets and benchmarks have been reported. This paper asks whether approaches developed for majority languages like English can be adapted to Arabic tasks. We adapt an English aspect-based sentiment analysis framework to Arabic classification tasks and present the adaptation as BARRAC: Brainstorming Alignment and Replaced Representation learning for ArabiC tasks. BARRAC replaces consumer-review attribute pools with Arabic linguistic devices and markers for dialectal sentiment, sarcasm, and dialect identification, and replaces noisy self-training with two-stage training. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro-F1 of 63.93\%, outperforming the best few-label SOTA by 3\%, and outperforming GPT-4o on four out of five tasks. Error analysis provides insights into remaining challenges. These results demonstrate that adapting task-specific approaches is a promising direction for Arabic NLP alongside adapting models, datasets and benchmarks.
115. 【2609.38818】Whose Voice Survives the Summary? A Voice-Retention Audit of LLM Employee Listening
链接:https://arxiv.org/abs/2609.38818
作者:Thilo Tamme,Anton Hantel,Bijan Khosrawi-Rad
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
关键词:Organizations increasingly route, large language model, silences already-spoken voice, Organizations increasingly, increasingly route employee
备注: 10 pages, 3 figures, 3 tables. Accepted at the 60th Hawaii International Conference on System Sciences (HICSS 2027)
点击查看摘要
Abstract:Organizations increasingly route employee feedback to leaders through large language model (LLM) summaries, an unaudited layer that silences already-spoken voice. We introduce a Voice Retention / Representation Ratio metric for representational bias in summarization and apply it to a bilingual (English/German) corpus of 2,586 free-text responses from a global professional service company. First, employees supply criticism more reliably than praise (withholding praise is 82 times more common). Second, across 45 leader-summaries the pipeline filters by popularity, not sentiment: criticism survives, yet a concern voiced once is dropped 86% of the time, with short and German-only content lost on the same axis (theme retention 0.14 vs 0.74; German directional). Controlling for frequency, sentiment has no independent effect; the harm is prevalence-driven, which sentiment-only audits miss. A targeted prompt recovers only named themes. We contribute the metric, field evidence, and a disaggregated voice-retention card.
116. 【2609.38817】When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning
链接:https://arxiv.org/abs/2609.38817
作者:Yuanhe Zhang,Ziwei Wang,Jie Ren,Haoran Gao,Zhenhong Zhou,Fanyu Meng,Cong Wu,Li Sun,Sen Su
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large reasoning models, Large reasoning, complex tasks, tasks through extended, redundant verification
备注:
点击查看摘要
Abstract:Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However, existing mitigations largely truncate long outputs or react to surface repetition, and thus fail to distinguish normal thinking from uncontrolled reasoning or explain how benign reasoning degenerates into harmful behavior. In this paper, we operationalize LRM generation as four states and further introduce Reasoning-state Analysis via Dynamic Attention Responses (RADAR), which identifies the current reasoning state in real time and characterizes how effective reflection can develop into uncontrolled generation. Guided by RADAR's analysis, we further realign abnormal attention distributions toward patterns observed in normal requests and examine how this correction affects excessive reflection and persistent looping. Temporal analyses show that uncontrolled reasoning is characterized by attention distributions that deviate from normal generation, with abnormal trends becoming detectable before repetition begins. Correcting these deviations through Attention Realignment consistently reduces looping while largely preserving benign performance. Together, RADAR provide a mechanistic account of how reasoning becomes uncontrolled, offering actionable guidance for identifying critical failure stages and designing targeted runtime interventions.
117. 【2609.38816】You're Hired: Strategic Model Selection for LLM Collaboration
链接:https://arxiv.org/abs/2609.38816
作者:Zongwan Cao,Ziyuan Yang,Shangbin Feng,Michael Duan,Skyler Hallinan,Bingbing Wen,Lucy Lu Wang,Yulia Tsvetkov
类目:Computation and Language (cs.CL)
关键词:diverse Large Language, Large Language Models, Large Language, diverse Large, existing systems remain
备注: 21 pages, 10 tables, 5 figures
点击查看摘要
Abstract:While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systematically evaluate a taxonomy of 9 selection algorithms ranging from diversity of model descriptions, capability-aware behavioral diversity, and LLM-based recruiters. We conduct extensive experiments across two candidate pools of 10 and 32 models, deployed in four model collaboration algorithms, and evaluated across tasks spanning math, coding, QA, and reasoning. Results demonstrate that successful selection algorithms greatly outperform random or heuristics-based teams such as merely selecting the models with top individual performance, by up to 36.1% across settings. Specifically, capability- and training-based selection strategies alleviate selection variance and achieve the best performance, which we recommend to employ before deploying real-world multi-LLM systems. Further analysis reveals that larger candidate pools pose greater challenges to shallow selection heuristics, while algorithms grounded in interacting with candidate models and understanding model capability robustly filter out misaligned, unsafe models, as well as generalizing to novel, out-of-distribution tasks. Together, we establish that principled and informed team selection is critical and present strong model selection algorithms for assembling effective multi-LLM systems.
118. 【2609.38812】Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification
链接:https://arxiv.org/abs/2609.38812
作者:Yingfeng Luo,Shaowei Wei,Daixin Wang,Dingyang Lin,Kaiyan Chang,Weiqiao Shan,Tong Zheng,Zhiqiang Zhang,Jingbo Zhu,Tong Xiao
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Terminal agents rely, command-line environments, solve tasks, Terminal agents, verification
备注:
点击查看摘要
Abstract:Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifies the first complete solution in each trajectory, determines whether it is objectively correct, and uses this ground truth to quantify the agent's subsequent verification and recovery behavior. Applying it to ten terminal agents on TerminalBench2.1, we find that verification is nearly universal after a complete candidate is formed, yet only 61.43\% of incorrect candidates are detected and only 49.36\% of detected errors are successfully repaired. These results show that the main weakness in self-verification lies not in initiating verification, but in detecting and repairing errors. Motivated by these findings, we propose Student-Conditioned Verification Distillation (SCVD), which lets the student first produce a candidate solution and distills a stronger teacher's subsequent verification and recovery from the same interaction context. Across three Qwen3.5 backbones, SCVD improves \textsc{Pass@1} on TerminalBench2.1 by 9.74--16.85 percentage points over the corresponding base models and by 4.49--8.61 points over the standard full-trajectory distillation, while avoiding the pronounced out-of-distribution degradation of full-trajectory distillation on SWE-bench Verified.
119. 【2609.38809】StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning
链接:https://arxiv.org/abs/2609.38809
作者:Naen Xu,Wanqing Cui,Yibo Hu,Shixin Hong,Hengyu An,Meiguang Jin,Junfeng Ma,Tianyu Du
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:evolving interaction histories, Large language models, Large language, language models deployed, reason over long
备注: NeurIPS 2026
点击查看摘要
Abstract:Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
120. 【2609.38806】Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems
链接:https://arxiv.org/abs/2609.38806
作者:Woosang Jeon,Jaeyeon Kim,Sham Kakade,Yilun Du,Amrit Singh Bedi,Arun Kumar Chithanar,Chul Lee,Taehyeong Kim,Sitan Chen
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:driven remarkable progress, complex global constraints, Next-token prediction, large language models, driven remarkable
备注: 32 pages, 9 figures
点击查看摘要
Abstract:Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by next-token prediction itself. We study this question through blackboard intelligence: an inference-time perspective in which a model works on a fixed, revisable canvas and searches over candidate solution states rather than committing to a causal, left-to-right trajectory. We instantiate this idea with diffusion language models, whose any-order prediction interface naturally exposes predictions over partially filled solution states. Our key observation is that mean confidence, a simple model-internal quantity available from the standard masked diffusion objective, provides a useful proxy for global coherence and can guide inference-time search and revision. Empirically, across ZebraLogic, Nurse Rostering, and Job-Shop Scheduling, Blackboard consistently improves inference while holding the fine-tuned LLaDA-8B-Instruct checkpoint fixed and substantially outperforms same-scale autoregressive baselines, reaching 90.4% accuracy on ZebraLogic-Hard, 76.4% exact feasibility on Nurse Rostering, and 80.2% optimality on JSSP. Stronger autoregressive search and refinement also fail to close the gap on ZebraLogic-Hard, while Blackboard surpasses tested frontier LLMs there and on JSSP despite their substantially greater scale and strong test-time reasoning. We open-source our codebase at this https URL.
121. 【2609.38802】Uncovering Uncontrolled Repetition through Residual Stream Dynamics
链接:https://arxiv.org/abs/2609.38802
作者:Yuanhe Zhang,Xinyao Zhou,Haoran Gao,Yuyao Zhang,Zhenhong Zhou,Fanyu Meng,Li Sun,Sen Su
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:prolong autoregressive generation, prolong autoregressive, Uncontrolled repetition, repetition, large language models
备注:
点击查看摘要
Abstract:Uncontrolled repetition can prolong autoregressive generation in large language models (LLMs) and enable resource consumption attacks. Prior analyses of repetitive generation have identified strongly activated features in intermediate and late layers. However, how uncontrolled repetition activity emerges and develops before becoming prominent in these layers remains insufficiently understood. In this paper, we investigate this question primarily in large vision-language models (LVLMs), which support a richer set of uncontrolled repetitions through both visual and textual inputs. We propose Tokenwise Residual Comparison (TRC), a method that identifies and localizes anomalies associated with repetition from residual dynamics during generation. TRC compares attention and multilayer perceptron writes to the residual stream across generated tokens to identify patterns associated with repetition. It then selectively suppresses coordinates in the residual stream at the identified layer. Experiments show that TRC effectively mitigates uncontrolled repetition, reducing loop rates by 57\% on average. Our analysis further shows that repetition semantics emerge in shallow layers and propagate through the residual stream, disrupting normal representations. TRC also generalizes to large language models (LLMs) and large reasoning models (LRMs), where it consistently captures analogous repetition dynamics and achieves effective mitigation. Our work broadens the study of repetitive generation from its prominent internal representations to earlier opportunities for intervention, providing insights for mitigating resource consumption attacks.
122. 【2609.38799】Overlap, Unique and Conflict: Can LLMs Extract What They Can Recognize?
链接:https://arxiv.org/abs/2609.38799
作者:Eftekhar Hossain,Santu Karmaker
类目:Computation and Language (cs.CL)
关键词:Understanding multi-perspective alternative, Understanding multi-perspective, alternative narratives requires, narratives requires identifying, multi-perspective alternative narratives
备注:
点击查看摘要
Abstract:Understanding multi-perspective alternative narratives requires identifying how their information agrees, conflicts, or differs across sources. Existing work on cross-text relations largely focuses on categorizing relations between predefined text pairs, such as entailment or contradiction, rather than directly extracting such information from full narratives. To address this gap, we introduce Overlap-Unique-Conflict (OUC) extraction, a cross-narrative task that extracts all overlapping, conflicting, and unique clauses from two narratives. To support this study, we construct a benchmark of approximately 22K narrative pairs and 140K OUC instances spanning factual, argumentative, and political discourse. Evaluating 14 open-source LLMs (0.6B-35B), we find that unique information is far easier to extract than overlap and conflict: the strongest model, Gemma-4-31B, reaches only 61.13% F1-score on overlap and 48.58% on conflict, against more than 75% on unique. Further diagnostic analysis reveals that this difficulty does not stem from relation recognition alone, but rather from a failure to pair and extract the corresponding clauses from full narratives, especially in smaller models. Nevertheless, learning these extractions with task-specific supervision narrows the gap considerably: a fine-tuned Qwen-3-8B gains 15-28% absolute over its baseline and surpasses models roughly four times its size (e.g., Qwen-3.6-35B) on several tasks. Even so, overlap and conflict remain well below satisfactory, leaving cross-narrative clause extraction an open challenge.
123. 【2609.38797】Evaluating Persistent Calibration under Evolving Model Knowledge
链接:https://arxiv.org/abs/2609.38797
作者:Victor Wang,Thomas Hofweber,Mohit Bansal,Elias Stengel-Eskin
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:adaptation and learning, maintaining their trustworthiness, produce confidence estimates, systems move, move from static
备注: Code: [this https URL](https://github.com/victorwang37/persistent-calibration)
点击查看摘要
Abstract:As AI systems move from static repositories to agents that are capable of continual adaptation and learning, maintaining their trustworthiness means equipping the models backing them with the ability to produce confidence estimates that dynamically reflect their changing skills and knowledge. We introduce the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. We operationalize this by examining persistent calibration across checkpoints of open models, asking whether confidence estimators trained on earlier checkpoints can generalize to later ones. Specifically, we aim to shed light on whether confidence is dependent on knowledge, a question with implications for the reliability of confidence estimates. To measure this relationship, we define and evaluate calibration on knowledge contrast sets: subsets containing questions that one checkpoint answers correctly and another checkpoint answers incorrectly, reflecting a change in knowledge. We show that both inference-time and fine-tuning methods fall short on contrast-set calibration compared to oracle methods trained on future checkpoints, even for methods that are well-calibrated on the full dataset. We provide evidence for the hypothesis that persistent calibration is challenging because there is a vast space of possible confidence functions that are well-calibrated on a given checkpoint, out of which only some rely on meta-knowledge features that would generalize to other checkpoints. Towards improving contrast-set calibration, we show that multi-checkpoint training helps, suggesting an avenue for identifying confidence features that remain robust across changing knowledge.
124. 【2609.38795】Recovering Off-Policy Supervision for Speculative Decoding
链接:https://arxiv.org/abs/2609.38795
作者:Jungseob Lee,Chanjun Park,Sugyeong Eo,Hyeonseok Moon
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:off-policy token invalidates, token invalidates supervision, single off-policy token, external models, decoding are commonly
备注: 22 pages, 4 figures, 17 tables
点击查看摘要
Abstract:Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at this https URL.
125. 【2609.38792】raining LLM Judges from Language Feedback via Position-Selective Self-Distillation
链接:https://arxiv.org/abs/2609.38792
作者:Ilgee Hong,Changlong Yu,Zhenghao Xu,Xin Liu,Yuwei Zhang,Qin Lu,Bing Yin,Tuo Zhao
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:study training LLM, training LLM judges, verdict depends strongly, natural language feedback, training LLM
备注:
点击查看摘要
Abstract:We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
126. 【2609.38722】Anchor-ECC: Local Integrity Checking for Watermarked LLM Outputs via Error-Correcting Codes
链接:https://arxiv.org/abs/2609.38722
作者:Zewei Deng,Muhammad Siddeek,Liyan Xie,Mohamed Seif,Mengdi Wang,H. Vincent Poor,Andrea Goldsmith
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:embedding detectable patterns, distinguishing AI-generated text, effective approach, approach to distinguishing, distinguishing AI-generated
备注: 18 pages, including references and appendices; 1 figure and 11 tables
点击查看摘要
Abstract:LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified content is still attributed to the original model. We propose Anchor-ECC, which incorporates the error-correcting code (ECC) constraints and explicit boundary anchors into the watermark structure and pairs them with a dynamic-programming decoder to detect and localize post-generation edits. Across Qwen3-8B, Mistral-7B-Instruct-v0.3, and OPT-125M, the approximate-hard setting achieves about 99.7% block-level true positive rate (TPR) with at most 7.6% false alarm rate (FAR) for edit detection under mixed insertions, deletions, and substitutions, while preserving the distinction between watermarked outputs and unwatermarked text. Additional quality experiments identify lower-perplexity configurations that retain strong edit-detection performance. Together, these results extend LLM watermarking from source identification to local integrity verification while supporting configurable trade-offs between detection reliability and generation quality.
127. 【2609.38718】MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation
链接:https://arxiv.org/abs/2609.38718
作者:Mehdi Jafari,Hao Xue,Flora Salim
类目:Computation and Language (cs.CL)
关键词:language models typically, models typically relies, large language models, Steering large language, behavioral distinctions
备注: Preprint. Code and pretrained model checkpoints will be released shortly
点击查看摘要
Abstract:Steering large language models typically relies on linear, context-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions. We introduce MetaSteer, a method that learns nonlinear interventions with context-dependent effects and applies them to attention projection matrices, producing activation effects that vary with the input context by construction and requiring no linear concept-geometry assumption. Framed as preference-based optimization, MetaSteer is trained once on a pooled preference corpus and transferred zero-shot to unseen concepts and out-of-distribution contexts. We find that, despite using low-rank adapters, MetaSteer induces structured, context-dependent changes in hidden-state trajectories while partially preserving aspects of their local trajectory dynamics, including velocity and curvature. We evaluate MetaSteer on three controlled text-generation benchmarks and three agentic settings across multiple model families and scales. MetaSteer matches or outperforms strong task-specific steering baselines on most aggregate comparisons in the zero-shot regime. Across the evaluated settings, stronger text-generation steering is associated with stronger agentic steering performance. We further discuss geometric trajectory effects, capability retention, and safety considerations raised by transferable steering.
128. 【2609.38660】Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
链接:https://arxiv.org/abs/2609.38660
作者:Haibo Jin,Xinjie Li,Najmeh Sadoughi,Yang Liu,Yibo Wang,Zhu Liu,Yuzong Liu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)
关键词:Long-form subtitle translation, translation requires reasoning, cultural context spanning, subtitle translation requires, maintaining consistent terminology
备注: 49 pages
点击查看摘要
Abstract:Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On the public MuSC benchmark, SMART obtains the best model result across all four language pairs and also achieves the best human-evaluation result, with an overall score of 4.50/5.
129. 【2609.38630】Strong Multilingual Privacy Tagging at Encoder Speed
链接:https://arxiv.org/abs/2609.38630
作者:Jonathan Graehl
类目:Computation and Language (cs.CL)
关键词:remove personal information, preserving relationships expressed, remove personal, personal information, information while preserving
备注: 46 pages, 23 figures. Includes supplementary appendices. Submitted to ACL Rolling Review, October 2026 cycle
点击查看摘要
Abstract:Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with fine-grained distinctions supporting varied redaction policies and methods for cheaply learning additional distinctions. We fine-tune a multilingual encoder with an affine span-tagging head on frontier-model annotations in 35 languages, replay mapped human gold with coverage-aware masking so unannotated types are not treated as negatives, and repair subword boundaries with a learned +/-1-character adjustment. On 1,283 human-gold test segments in seven languages, best measured redaction F1 is 88.8, against 69.1 for published GLiNER2 with 11 unrepresentable types excluded from its task (68.8 without that exemption), 67.8 for GLiNER2 adapted to the new training data, 57.3 for Microsoft Presidio and 35.8 for the best published OpenAI Privacy Filter fine-tune. Adding about 50,000 annotated training sentences and increasing human-gold replay improves exact typed-span F1 from 74.5 to 76.3 on Ont3, our 31-type frontier-annotated NER evaluation of 1,201 development segments. Mapped-gold replay alone raises human-gold F1 by ten points without loss on frontier-annotated text; boundary adjustment adds 1.7 exact typed-span F1 points on Ont3. Local LLMs fitting on a single 96-GB GPU underperformed as prompted annotators and frozen encoders, with encoding 30-95 times slower than XLM-R inference and prompted annotation roughly 180-1,100 times slower in the evaluated configurations. The encoder architecture delivers 4.9 times GLiNER2's CPU throughput. We release code, prompts and training recipes, with data-acquisition scripts and source links.
130. 【2609.38627】Marking Contour Tones in Yorùbá
链接:https://arxiv.org/abs/2609.38627
作者:Kólá Túbòsún
类目:Computation and Language (cs.CL)
关键词:tones pose persistent, persistent orthographic challenges, pose persistent orthographic, contour tones pose, pose persistent
备注: Under review at the 12th World Congress of African Linguistics (WOCAL 12)
点击查看摘要
Abstract:Yorùbá is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexical items whose conventional spellings avoid vowel lengthening that would otherwise provide a host syllable for the second tone. A particular concern is a class of names in which the conventional spelling does not just omit tonal information but inverts the meaning of said name, sometimes asserting the opposite of what the name intends. This paper describes the problem, illustrates the inadequacy of current solutions, and proposes the adoption of the caron and circumflex marks. These are symbols with precedent in Yorùbá phonological scholarship since Olmsted (1951), used as orthographic conventions on single vowels to encode rising and falling contour tones, making them accessible for the first time through standard keyboard input and computational text processing. The proposal is supported by an implementation in the WriteYoruba keyboard and the TTSYoruba speech synthesizer, whose architecture and listener evaluation are reported separately (Tubosun et al., 2026).
131. 【2609.38621】When Scientific Contradictions Are Lost in Translation
链接:https://arxiv.org/abs/2609.38621
作者:Tal Zeevi,Trey W. Jensen,Maxwell Strome
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:disagree without contradicting, Claude Opus, assignment, constraints, scientific
备注: Accepted at the NeurIPS 2026 AI for Science Workshop: Verification in the Age of AI Scientists. This version is not included in the official NeurIPS proceedings
点击查看摘要
Abstract:Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint system and translate its constraints into scientific reports from different laboratories. One assignment satisfies more constraints, while another satisfies fewer but better matches expected biology. This creates a simple dilemma: does the model choose the assignment that best fits the constraints, or the one that better matches biological expectations? When the constraints are stated directly, GPT-5.6 Sol and Claude Opus 5 recover the best-supported assignment in 90% and 96% of cases, respectively. In scientific prose, however, the models behave differently. Claude Opus 5 often prefers the biologically expected assignment. Removing that biological preference increases recovery of the better-supported assignment from 27% to 79% (p.001); recovery reaches 92% when the same Biology-favored record is accompanied by a formalization request and an explicit paired-design cue (p.001). GPT-5.6 Sol is less sensitive, with neither corresponding change reaching statistical significance. These results suggest that reliable scientific verification depends not only on formal reasoning, but also on how models decide which findings should be compared and what relations they imply.
132. 【2609.38612】StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams
链接:https://arxiv.org/abs/2609.38612
作者:Jhen-Ke Lin,Chung Chun Wang
类目:Computation and Language (cs.CL)
关键词:natural language drives, run inside programs, increasingly run inside, language models increasingly, models increasingly run
备注: 22 pages, 6 figures. Code and data: [this https URL](https://github.com/JacobLinCool/StreamDecisionBench)
点击查看摘要
Abstract:As natural language drives more applications, language models increasingly run inside programs as decision components: the program sends them the current state and acts on the returned decision until a newer one arrives. When evidence changes during inference, a decision correct for its own state can stay in force after that state has passed, as when a call recorder keeps running after a customer starts reading out a card number; untimed (offline) accuracy counts such an error as correct. We introduce StreamDecisionBench (SDB), which evaluates the decision in force at every instant and attributes every erroneous instant to judgment, latency or both. Its scenarios stream evidence in four application families, with reference decisions computed from public rules by executable code. We summarize in-force accuracy across update intervals of 1-5 s by its normalized area under the curve on a logarithmic time axis, giving equal weight to equal multiplicative ranges. Across six settings of four hosted models, this score stays within 2.9 points of the scenario-wise product of untimed accuracy and an oracle's integrated timing score. Reasoning improves judgment, but at low effort latency costs Luna and Terra, two GPT models we also evaluate without reasoning, 38.8 and 42.8 points relative to untimed accuracy; a faster component with weaker judgment attains a similar integrated score to Terra without reasoning. The aggregate and family curves show where these tradeoffs change, making the evaluation's time-scale dependence visible.
133. 【2609.38606】SecureVibe: Making Vibe Coding More Secure
链接:https://arxiv.org/abs/2609.38606
作者:Danqing Wang,Baolin Peng,Zhepei Wei,Isadora White,Wenlin Yao,Hao Cheng,Qianhui Wu,Minseon Kim,Xingdi Yuan,Lei Li,Jianfeng Gao
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:functionally correct solutions, security, functionally correct, SECUREVIBE, capable and widespread
备注:
点击查看摘要
Abstract:As vibe coding becomes increasingly capable and widespread, security vulnerabilities in even functionally correct solutions are a growing concern. When investigating functionally correct but insecure solutions, we find that the insecure agent is less than half as likely to conduct effective planning and testing for the hidden security risks behind the functional requirements. Motivated by this, we develop SECUREVIBE, a training recipe that explicitly targets planning and testing for code security. SECUREVIBE constructs training signals around these security behaviors. It includes supervised fine-tuning on the security suite with 4 security tasks, and post-training methods, SECUREVIBE_rl and SECUREVIBE_hg, to enhance security capabilities from verifiable execution feedback and hint-based self-supervision. Our SECUREVIBE outperforms the baseline on two types of security coding tasks across 4 benchmarks. Specifically, SECUREVIBE improves the security pass@1 by 6.9 points on BaxBench. The gains extend to unseen CWE categories, with improvements of 11.5 points on SusVibes. Meanwhile, it also improves functionality pass@1 by 13.6 points on the security coding task SusVibes and 4.1 points on the generic coding task SWE-bench Verified. Further analysis offers two practical insights: (i) diversifying supervision across security planning, coding, and testing strengthens security behaviors more effectively than adding coding trajectories alone, and (ii) hint-guided supervision is particularly valuable when the agent's existing security capabilities are insufficient to learn effectively from outcome feedback.
134. 【2609.38604】Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent
链接:https://arxiv.org/abs/2609.38604
作者:Zheyuan Zhang,Mengyuan Chao,Ke Xiao,Ziyi Chen,Daoan Zhang,Yan Zhang,Yanfang Ye,Wei Xu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Modern LLM agents, Modern LLM, LLM agents increasingly, increasingly tackle complex, agents increasingly tackle
备注:
点击查看摘要
Abstract:Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicate, change their goals, and run out of patience. We define this task setting as Interactive Intent Alignment, where agents must recover and continuously track the user's current intent despite imperfect communication and evolving goals. To study this setting, we introduce Drift-Bench++, a principled benchmark construction pipeline for verified executable tasks with controlled misalignment and intent shifts, along with an interaction protocol featuring finite patience, diverse simulated users, and silent interaction-conditioned shifts. We further develop GRIP, a comprehensive evaluation protocol covering task grounding, user realism, inquiry effectiveness, and adaptation to evolving intent. Across diverse environments, models, and interaction conditions, stronger interaction consistently helps but remains far from oracle performance; Validation on deployed ProdAgent sessions further shows that the modeled failures are prevalent and consequential in deployment. By providing a unified, executable benchmark for interactive intent alignment, Drift-Bench++ offers a foundation for evaluating and advancing agents under realistic communication and evolving intent.
135. 【2609.38593】Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions
链接:https://arxiv.org/abs/2609.38593
作者:Bo Ni,Li Li,Ryan A. Rossi,Franck Dernoncourt,Tyler Derr
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Large Language, incorporating relevant procedural, Language Models, inference time
备注:
点击查看摘要
Abstract:Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training. In addition, emerging tasks may fall outside the scope of existing skill libraries, creating a need to develop new skills before curated training data become available. Recent works have explored automated skill optimization through reflection, but they require a curated, in-distribution training set, which users might not always have. To address these limitations, we present Prompt2Skill, a framework that builds skills from natural-language task description alone. From the prompt, the system derives a task specification, discovers or synthesizes datasets, and refines the skill in a closed loop of reflective editing. Across four domains spanning question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill consistently outperforms the direct prompting baseline, achieving an average improvement of 10.8 across open-source and frontier models.
136. 【2609.38574】owards Model as a Library: Offline, Community-Sourced AI for Low-Resource African Languages
链接:https://arxiv.org/abs/2609.38574
作者:Fendji K. E. Jean Louis
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:African languages remain, Large language models, languages remain low-resource, African communities, standardised text systematically
备注: 5 pages, GlobalSouthAI @ NeurIPS 2026
点击查看摘要
Abstract:Large language models are frequently proposed as a route to AI-powered services for African communities, but they are least reliable exactly where the need is greatest: all African languages remain low-resource by any standard measure, and models trained on scraped, standardised text systematically misrepresent the dialectal and regional variation of how people actually speak. We introduce \textbf{Model as a Library (MaaL)}, a software architecture that packages small, community-enrolled speech models as versioned on-device dependencies, enabling offline structured data collection that cannot generatively hallucinate, for populations that current language models serve worst. Rather than relying on web-scraped corpora, MaaL's vocabulary is enrolled directly from a small number of example recordings by the speakers themselves, at the point of deployment. We describe the architecture and its central mechanism - keyword spotting that turns a closed-vocabulary text form into a voice form, filled and submitted entirely on-device - and propose transpiling the closed-vocabulary elements already present in widely-deployed digital form tools into MaaL schemas, a low-friction path to voice-first, offline data collection for the low-literacy populations these tools already reach. This is a position and system-design paper: we describe the concept, the mechanism, and an analytical feasibility case, and identify what a working implementation still requires.
137. 【2609.38543】MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models
链接:https://arxiv.org/abs/2609.38543
作者:Lukas Thede,Yash Kumar Atri,David Chen,Danielle Bitterman,Matthias Bethge,Tom Hartvigsen,Zeynep Akata
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Constantly evolving real-world, Constantly evolving, evolving real-world knowledge, real-world knowledge necessitates, knowledge
备注: Accepted at NeurIPS 2026 (Evaluations Datasets Track)
点击查看摘要
Abstract:Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.
138. 【2609.38530】Shifting Mechanisms: How Positional Encoding Choice Shapes In-Context Retrieval
链接:https://arxiv.org/abs/2609.38530
作者:Eric Enouen,Sainyam Galhotra
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Language models increasingly, vary attention span, Language models, vary attention, attention span
备注:
点击查看摘要
Abstract:Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention (SWA NoPE). However, how these choices shape in-context retrieval remains unclear. To study this question, we take a mechanistic view, tracing how positional encoding (PE) choice shapes the internal mechanisms models use for in-context retrieval. Across 22 open-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval. We further show on a controlled pre-training ablation that confining positional encoding to local layers produces this semantic shift, degrading representations of positional information. Finally, we show that the reported long-context gains of PE hybrids mask a retrieval trade-off: SWA NoPE improves over RoPE on multiple-target retrieval and QA, but degrades when distinguishing competing keys. We show that these behavioral differences better track the mechanism shift from positional toward semantic mechanisms than a uniform improvement in long-context retrieval.
139. 【2609.38510】DEdit: Iterative Draft Editing for Speculative Decoding
链接:https://arxiv.org/abs/2609.38510
作者:Longxuan Yu,Bingsen Chen,Peng Shi,Dongkyu Lee,Yi Xiang,Hideo Kobayashi,Sheng Zhang,Shuaichen Chang,Xing Niu,Zhuoyan Xu,Greg Ver Steeg,Jiarong Jiang
类目:Computation and Language (cs.CL)
关键词:Speculative decoding accelerates, target model verifies, accelerates autoregressive LLMs, Speculative decoding, lightweight drafter propose
备注: 21 pages, 7 figures, 6 tables
点击查看摘要
Abstract:Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks on Qwen3-4B and Qwen3-8B, DEdit achieves the highest macro-average token acceptance and speedup among the evaluated drafters, reaching macro-average speedups of $5.72\times$ and $5.97\times$ over autoregressive generation under greedy decoding, respectively. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and that ProposalMix halves harmful edits that shorten the accepted prefix. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains.
140. 【2609.38490】Personalized State-Transition-Aware Memory for Clinical Agents
链接:https://arxiv.org/abs/2609.38490
作者:Maryam Haghifam,Zahra Rajabi,Yizhou Sun,Carlos Morato
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large language model, Large language, language model, agents that reason, LLM
备注:
点击查看摘要
Abstract:Large language model (LLM) agents that reason over clinical records must track changes in a patient's state while preserving the history needed to understand them. Simply accumulating memories leaves it unclear which information still applies, whereas overwriting earlier memories can erase evidence needed to reconstruct treatment history and clinical trajectories. We introduce STAM, a state-transition-aware memory framework that records state changes as new clinical entries arrive. STAM combines semantic retrieval with typed clinical relations to identify affected memories, maintaining current information in Active and superseded or resolved information in History. At read time, a query-dependent gate selectively serves historical memory. Across four longitudinal clinical benchmarks, we evaluate STAM with downstream question answering, direct state-maintenance diagnostics, and comparisons at approximately matched context lengths.
141. 【2609.38486】Anthropomorphism in the age of Large Language Models: An overview of potential risks and mitigations
链接:https://arxiv.org/abs/2609.38486
作者:Ismael T. Freire,Marceau Nahon,Maud van Lier,Katie Evans,Hélie Bazin,Michele Farisco,Kathinka Evers,Raja Chatila,Mehdi Khamassi
类目:Computers and Society (cs.CY); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:Large Language Models, broadly Artificial Intelligence, Artificial Intelligence, Large Language, Language Models
备注: 35 pages, 1 box, 1 figure
点击查看摘要
Abstract:Large Language Models (LLMs) and more broadly Artificial Intelligence (AI) systems are often described and understood in human-like terms, a phenomenon known as \emph{anthropomorphism}. This paper provides a synthesis of recent literature on anthropomorphism in AI, covering theoretical frameworks, the role of language in framing AI as human-like, the various risks of anthropomorphizing machines, and strategies to mitigate these issues. After examining why we tend to anthropomorphize AI systems and whether we are right to do so, we highlight the impact of linguistic framing on anthropomorphism. Then, we introduce a conceptual taxonomy of risks associated with AI anthropomorphism. This taxonomy groups twenty-one concerns within five analytical categories: epistemic, affective, human agency, normative, and societal and institutional risks. Finally, we relate these concerns to proposed interventions in design, communication, education, and governance. We argue that a better understanding of AI systems requires concepts and theories grounded in their organization and demonstrated capacities. The linguistic shaping of anthropomorphic perceptions should form part of this scientific effort, since our descriptions influence both how these systems are understood and the roles we allow them to occupy in society.
142. 【2609.38480】KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy
链接:https://arxiv.org/abs/2609.38480
作者:Xueting Fang,Zehui Li,Yang Yang,Camilla Giovino,Shubh K. Patel,Shailly Prajapati,Vallijah Subasri,Caihua Shan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:benchmarks evaluate language, evaluate language models, evaluate language, clinical benchmarks evaluate, clinical
备注:
点击查看摘要
Abstract:Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment.
143. 【2609.38469】he Backdrop Exposes What the World Around an Agent Costs It
链接:https://arxiv.org/abs/2609.38469
作者:Nusrat Jahan Lia,Shubhashis Roy Dipta
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Agent benchmarks test, benchmarks test agents, benchmarks test, Agent, agents
备注: Submitted to ICLR 2027
点击查看摘要
Abstract:Agent benchmarks test agents in worlds that stay still. Deployed agents work in worlds that other people also change. Someone texts the agent to send the money elsewhere or an order confirmation asks it to reply with a door code. We present BACKDROP, which asks how much of an agent's capability in a clean world survives in such a world. BACKDROP takes a task along with the agents execution environment, and plants four everyday hazards in its world, one at a time and all together. The instruction and the correct end state stay the same. Each hazard asks one question. Authority: does a message from another person override the user? Injection: does text planted in a record redirect the agent? Boundary: does a request pull it into an app it was not given? Fault: after a write fails without saying whether it landed, does the agent check before it retries? Across 3,678 variants and 16 models, , the average pass rate falls from 69.5% to 31.3% once all four hazards are present; the strongest models fall furthest (Claude Fable 5.1 from 96.6% to 56.0%). Agents have learned to resist injected text but often follow other unauthorized requests of other people. With all four hazards present, and counting only runs where the planted text reached the agent, agents followed another person's message in 46.4% of runs and injected text in 20.3%. The gap is consistent throughout all 16 models. BACKDROP formalizes these gaps and shows how an agent's score in a task's world is a ceiling on real-world performance.
144. 【2609.38448】Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles
链接:https://arxiv.org/abs/2609.38448
作者:Ben Wigler,Maria Tsfasman
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:create false plurality, offer independent perspectives, familiar default, create false, false plurality
备注: 23 pages, 6 figures. Published in the Proceedings of the Third Conference on Language Modeling (COLM 2026)
点击查看摘要
Abstract:Open-ended LLM homogeneity can create false plurality when several systems appear to offer independent perspectives while returning the same familiar default. Single-pass answers obscure the distinction between agreement produced by a tightly constrained answer space, prompt-vocabulary echo, and broader answer spaces with stable alternatives beneath the surface. We introduce CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework that adapts free-list elicitation from cognitive anthropology to LLM ensembles. CHOIR repeatedly elicits ranked lists, clusters items into prompt-level concepts, and measures concept salience across models, prompt variants, and persona conditions. We evaluate CHOIR on Infinity-Chat 100, an external prompt bank from recent work on open-ended model homogeneity, and on a 27-question targeted diagnostic bank designed to isolate mechanism-level contrasts. On Infinity-Chat 100, CHOIR reproduces high surface agreement (93/100 prompts above chance) while separating narrow prompts from broad prompts with recoverable depth. Across targeted probes and the external prompt bank, base-model identity remains the strongest recoverable signature, and persona prompts shift surfaced concepts within base-model signatures. A source-blind ranking module prioritises rare-but-stable candidates for later inspection. CHOIR turns open-ended homogeneity into a diagnostic measurement problem by asking where models converge, why they converge, and what remains reachable under structured depth probing.
145. 【2609.38446】What Pretraining and Midtraining Make Learnable from Rewards?
链接:https://arxiv.org/abs/2609.38446
作者:Chiwun Yang,Xiaoyu Li
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:inputs undetermined, reward, reward adaptation, computation needed, computation
备注: 160 pages, 26 figures, 32 tables
点击查看摘要
Abstract:A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.
146. 【2609.38427】Policy-Conditioned AI-Use Detection: An Evidentiary Framework for Academic Publishing
链接:https://arxiv.org/abs/2609.38427
作者:Jairo Diaz-Rodriguez,Mumin Jia
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
关键词:publish detailed rules, differ by role, Major venues, publish detailed, area chairs
备注:
点击查看摘要
Abstract:Major venues now publish detailed rules about how authors, reviewers, and area chairs may use AI, and those rules differ by role, by task, and by what must be disclosed. AI detection, the instrument usually proposed to enforce them, estimates something else: whether an AI model wrote the text. We argue that this target is misaligned with the decisions conferences and journals face, and propose policy-conditioned AI-use detection, an evidentiary framework for assessing whether a human--AI workflow complied with a stated rule. Policy makes the governing rule an explicit input. Inference reports hypotheses, evidence, calibration regime, and uncertainty in place of verdicts such as "AI detected". Evaluation builds benchmarks from reproducible pipelines that generate compliant and non-compliant workflows, and reports true positive rate at a false positive rate the venue fixes in advance. We work the framework through peer review, where at plausible violation rates a detector at a strong operating point still flags more compliant authors than violating ones. The framework therefore also names what a venue must instrument: structured disclosure, approved-tool routing that respects reviewer confidentiality, and a path by which a finding can be contested. Under this framing a detector is not an authorship classifier but an auditable procedure with an error rate the venue fixes in advance and can defend.
147. 【2609.38426】LoopVL: Recurrent Visual Intelligence
链接:https://arxiv.org/abs/2609.38426
作者:Zhe Qian,Ziyang Gong,Zhongxing Xu,Hehan Li,Zhonghua Wang,Fei Luo,Mingxuan Wang,Xue Yang,Shiwei liu,Yanbiao Ma,Junchi Yan,Jungong Han
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Loop Transformers, extended to vision, effectively extended, Transformers, LoopVL
备注:
点击查看摘要
Abstract:We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
148. 【2609.38420】What Was Said, Not What Was 'Thought': Type-6 Logic for CoT Verification
链接:https://arxiv.org/abs/2609.38420
作者:Adrian de Wynter
类目:Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:contemporary large language, dynamic epistemic logic, epistemic logic augmented, large language model, uncertainty and recurrence
备注:
点击查看摘要
Abstract:We introduce Type-6 logic, a variant of dynamic epistemic logic augmented with two operators (uncertainty and recurrence), designed to model the inferential dynamics of contemporary large language model (LLM) chain-of-thought (CoT) reasoning. Type-6 accounts for common LLM reasoning pathologies such as unlicensed revision, enthymemes, loopbacks, and unverifiable/incorrect claims. We propose a verifier based on Type-6 logic that builds a graph out the trace, and checks it against Type-6's axioms and inference rules. We evaluate our framework on LLM-generated CoTs four splits spanning formal and informal reasoning. Our verifier detects structurally unsound reasoning steps that surface-level heuristics miss, and allows for easy visualisation of the model's reasoning process. In our corpus, our verifier shows that derived contradiction is the most common hard-fail category in CoT, and that only about 3\% of the propositions of a trace have impact on the final derivation. Ablation studies show that other verification methods (LLMs-as-judges, other neurosymbolic approaches, etc.) cannot be considered interchangeable: for example, agreement between LLMs-as-judges and LINC is $\kappa \approx 0.034$, and this persists within a method across underlying models. Type-6, however, is the most agreed-with method amongst the ones we tested. We prove our verifier runs on average-case linear time; and release our logic specification and artefacts.
149. 【2609.38409】ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning
链接:https://arxiv.org/abs/2609.38409
作者:İbrahim Ethem Deveci,Funda Tan Çalık,Barış Deniz Sağlam,Duygu Ataman
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:automatically verifiable rewards, Recent progress, large language model, reinforcement learning environments, language model reasoning
备注: 40 Pages, 16 Tables
点击查看摘要
Abstract:Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning often proceeds under incomplete and revisable information: conclusions may be supported provisionally, defeated by counter-evidence, reinstated by further arguments, or revised when stronger reasons become available. Reasoning of this kind is generally referred to as defeasible reasoning. We introduce ArgGYM, a procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning. ArgGYM decomposes this reasoning into twelve tasks and grounds task-specific scoring in a symbolic argumentation engine that computes the formal states used to evaluate model outputs. It includes a frozen benchmark of 1,440 verified instances across fifteen curriculum configurations, two argument preference orderings (weakest-link and last-link), and two set orderings (elitist and democratic), while the same generators and verifiers can produce fresh instances for evaluation that reduces dependence on static test sets and for verifiable-reward training. On the frozen benchmark, frontier and open-weight models show sharply different reasoning profiles: they can recover substantial parts of structured answers without solving the complete task, and performance declines in later curriculum configurations with longer dependencies and more interacting structures. We release the benchmark, generators, and verifiers for reproducible evaluation and RLVR training.
150. 【2609.38406】Evaluating Whether LLMs Can Reliably Connect the DOTs?
链接:https://arxiv.org/abs/2609.38406
作者:Eftekhar Hossain,John Salvador,Santu Karmaker
类目:Computation and Language (cs.CL)
关键词:noisy and fragmented, Large Language Models, Access to real-world, Access, infilling
备注:
点击查看摘要
Abstract:Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models (LLMs), their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ~9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score (4.02/5), outperforming models over ten times larger, including DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement (+0.6%). Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.
151. 【2609.38389】he Geometry of Harmfulness in Multi-Turn Attacks
链接:https://arxiv.org/abs/2609.38389
作者:Yelyzaveta(Lisa)Husieva,Lauren Alvarez
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:Large language models, Large language, circumvent safety alignment, vulnerable to adversarial, alignment to elicit
备注: 9 pages, 7 figures, 1 Table, preprint
点击查看摘要
Abstract:Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work investigates how the geometry and temporal dynamics of harmfulness and refusal representations evolve across multi-turn attacks. We analyzed hidden-state representations from three instruction-tuned LLMs (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-9B-it) using three multi-turn attack frameworks (Crescendo, ActorAttack, and X-Teaming), and examined representation behavior across conversation turns, model layers, and token positions under various context configurations. Across models and frameworks, we found that (1) each attack framework traverses different geometric directions, yet each achieves comparable success in eliciting harmful outputs; (2) multi-turn harmfulness directions became increasingly linearly separable at the end-of-turn token position across turns in middle to late model layers; and (3) harmfulness representations are weakly aligned with refusal-related representations. The results indicate that multi-turn attacks do not succeed by suppressing the model's internal representation of harmfulness. Instead, harmfulness representations become increasingly separable across conversation turns, while remaining only weakly aligned with refusal-related representations. The findings are one possible explanation for why static single-turn safety probes may degrade in multi-turn settings, and suggest that robust defenses must consider temporal representation dynamics rather than identifying harmfulness with isolated or single-turn prompts.
152. 【2609.38385】Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting
链接:https://arxiv.org/abs/2609.38385
作者:Loay Mualem,Lluís Pastor-Pérez,Vinh Tong,Andrei Manolache,Tanja Bien,Steffen Staab,Mathias Niepert
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:discrete diffusion language, diffusion language models, language models masks, Supervised fine-tuning, fine-tuning of discrete
备注: 30 pages, 4 figures, 14 tables. Main text 10 pages, references and appendix follow. Project page with interactive visualizations: [this https URL](https://loaym.github.io/GoldiMask/)
点击查看摘要
Abstract:Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by approximately maximizing a submodular objective. This objective uses model signals to balance the benefit of revealing tokens against their value as prediction targets. GoldiMask then weights the remaining targets according to how they benefit from the selected context and their remaining learning potential. Across three backbones and three training datasets, GoldiMask achieves the highest average accuracy in most evaluated settings, demonstrating gains on both reasoning and code generation. Component ablations show that both context selection and target weighting contribute to the gains. GoldiMask also reduces decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding, while maintaining comparable accuracy at higher confidence thresholds.
153. 【2609.38374】Doc2LoRA Provides Decodable Representations of Scientific Ideas
链接:https://arxiv.org/abs/2609.38374
作者:Chand Sahil Mansuri,Joel Zachariah,Sadamori Kojaku
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR); Machine Learning (cs.LG); Physics and Society (physics.soc-ph)
关键词:Representing scientific papers, drive innovation, papers, Representing scientific, LLM
备注: 32 pages, 4 figures, 12 tables. Code: [this https URL](https://github.com/skojaku/doc2lora-embedding)
点击查看摘要
Abstract:Representing scientific papers as points in a space lets us search for similar papers and inquire about how fields relate to one another and drive innovation. Beyond search, the vector space of papers invites generation: mixing papers through simple vector operations creates new points, mirroring combinatorial novelty, the recombination of existing ideas into new ones. However, a mixed point often represents an idea no paper has yet realized, with no papers nearby to identify the idea. We propose representing each paper by a LoRA adapter generated by the Doc-to-LoRA hypernetwork. Every point in the space, including mixtures, thus represents a large language model (LLM) open to questions and instructions in natural language. On papers from the American Physical Society (APS), we instruct the LLM at the average of each subfield to name the field in a few words and obtain labels closer to the official names than the labels of five baselines, as judged by word overlap and a panel of five LLM judges. We also ask the LLMs at points between two APS papers to write an abstract and obtain descriptions shifting from one paper to the other in step with the mixing weight. While Doc-to-LoRA is trained for generation, a small invertible transform makes the embeddings competitive for search, on par with SPECTER2 and EmbeddingGemma and close to SBERT. Because the transform is invertible, every point in the transformed space still maps back to an LLM. The embeddings thus serve both search and generation, enabling researchers to question the idea at any point in the space as a starting point for generating new ideas.
154. 【2609.38371】ALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns
链接:https://arxiv.org/abs/2609.38371
作者:Guangxin Zhao,Yiran Hu,Yuan Cao,Chenxi Jiang,Jianfei Yang,Yegang Du,Yasuyuki Taki,Yoshifumi Kitamura,Lin Gu,Zhi Zheng
类目:Robotics (cs.RO); Computation and Language (cs.CL)
关键词:Existing LLM-driven robot, Existing LLM-driven, task planners rely, LLM-driven robot task, planners rely
备注:
点击查看摘要
Abstract:Existing LLM-driven robot task planners rely on a taken-for-granted assumption of an ideal user whose instructions are clear, complete, and task-focused. However, when interacting with real-world users, especially those experiencing cognitive impairments, such as people living with dementia (PLWD), the planners often make mistakes and even pose physical safety risks. We proposed TALK-Dem (Talking Attributes and Linguistic Knowledge in Dementia), the first benchmark for evaluating LLM-driven robot task planning under dementia-associated verbal communication. TALK-Dem contains 4,800 instructions and covers five typical communication patterns, including Referential Imprecision, Object Substitution, Empty Speech, Topic Drift, and Intrusion, at three intensity levels. Experiments across six open-weight LLMs reveal a substantial robustness gap. Across communication patterns, open-weight models exhibited performance drops of up to 22.3 percentage points compared to ideal instructions. This revealed a critical gap and even danger for real-world applications, especially in assistive robotics, where locally deployable models are necessary due to privacy concerns and connectivity constraints. To mitigate this issue, we proposed the Context-Aware Retrieval from Experience (CARE) method, which retrieves relevant previously resolved tasks to provide task-specific interpretation and planning context. CARE generally outperformed standard prompting baselines across the six open-weight models, improving average task success by 18.1 percentage points over the vanilla prompt. These results highlighted the importance of both evaluating communication robustness and developing effective adaptation strategies for locally deployable assistive robots. The TALK-Dem dataset is publicly available at this https URL.
155. 【2609.38360】On the Off-Policy Teacher in On-Policy Distillation
链接:https://arxiv.org/abs/2609.38360
作者:Langlin Huang,Hao Liu,Mononito Goswami,Xinyu Li,Prithwith Jana,Nikos Kanakaris,Patrick Blöbaum,Purak Jain
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:promising post-training paradigm, dense teacher supervision, teacher, recently emerged, promising post-training
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
156. 【2609.38359】Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks
链接:https://arxiv.org/abs/2609.38359
作者:Sumit Asthana,Michael Ion,Kevyn Collins Thompson
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:High quality synthetic, Generative Flow Networks, High quality, diverse high quality, central to post
备注:
点击查看摘要
Abstract:High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We propose a method to generate diverse high quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate latent conversation structure using a Gaussian mixture density over key interaction features (e.g., confusion episode dynamics, scaffolding directive balance) enables sampling expert strategies in proportion to their prevalence in the training data. Across two structurally distinct domains, tutoring and emotional support dialogues, our GFlow based synthetic data generation approach offers a better balance of fidelity, mode coverage and authenticity than reinforcement-learning and end to end LLM baselines, without copying training data. Evaluated on three downstream outcome prediction tasks, classifiers trained on synthetic GFlowNet generated conversations provide a stronger training signal than competitive synthesis baselines.
157. 【2609.38357】Evaluating Language Model Safety Across Long Adversarial Conversations
链接:https://arxiv.org/abs/2609.38357
作者:Parisa Salmani,Peter R. Lewis
类目:Computation and Language (cs.CL)
关键词:real-world systems interact, Conversational safety, Conversational safety evaluations, single harmful prompt, real-world systems
备注:
点击查看摘要
Abstract:Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns. We evaluate three open-weight, instruction-tuned models on two harmful prompts across different conversation lengths and random seeds. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe. Across all model-prompt combinations, first-turn safe-response rates ranged from 85% to 100%. By depth 11, they dropped to 38-61%, and by depth 101, to 15-44%. This decline appeared across models and continued well beyond the short interactions typically used in multi-turn safety evaluations. These results provide proof-of-concept evidence that strong single-turn safety does not necessarily persist during sustained adversarial interaction. They highlight the need for long-horizon evaluations and conversation-level safeguards that account for risk accumulating across turns.
158. 【2609.38355】Halluscoring 2026: The first shared task on llms hallucination detection and answer verification
链接:https://arxiv.org/abs/2609.38355
作者:Aisha Alansari,Abdessalam Bouchekif,Ahmed Hasanaath,Salah Eddine Bekhouche,Malak Alkhorasani,Mohammed-En-Nadhir Zighem,Saad Ezzini,Hichem Telli,Hend Al-Khalifa,Muhammad Abdul-Mageed,Hadid Abdenour,Hamzah Luqman
类目:Computation and Language (cs.CL)
关键词:Arabic question answering, shared task, task, present HalluScoring, evaluating hallucination detection
备注:
点击查看摘要
Abstract:We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. The shared task is organized into two main tasks, each comprising two subtasks, for a total of four subtasks. Task 1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, 10 of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging under distribution shift, with the winning team achieving AUC-ROC test scores of 0.772 and 0.767 for Subtasks 1.1 and 1.2, respectively. For Task 2, the winning team achieved scores of 0.882 and 0.857 in the Islamic and general-knowledge subtasks, respectively, under assisted evaluation.
159. 【2609.38345】OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime
链接:https://arxiv.org/abs/2609.38345
作者:Chun-Wah Hsu,Kai Gong,Yu Wu,Xianhe Chen,Mengyang Liu,Jie Li,Hanyu Li,Zhixuan Liu,Naisheng Tang,Jiaying Chi,Ziheng Fan,Xuning He,Xiaokang Yang,Xue Jiang,Yihong Dong
类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:tackle complex software, complex software engineering, software engineering tasks, designed to tackle, tackle complex
备注: work on process
点击查看摘要
Abstract:Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents clear attribution of observed gains. To this end, we introduce OpenCollab, a multi-agent coding framework that provides a unified infrastructure for programmable collaboration and controllable runtime. Specifically, OpenCollab unifies organization design, enforces experimental control on a shared runtime, and tracks execution through fine-grained event streams. On this basis, we define Adherence to quantify whether the declared organization is actually realized. Our experiments reveal that agents collaborate very differently across configurations: changing any single dimension shifts Adherence, from 47.2% to as high as 97.2%. Furthermore, extensive agentic coding benchmarks show that a two-coder workflow built on OpenCollab establishes new SOTA performance compared to the mainstream harnesses such as Mini-SWE-agent, Codex CLI, and Claude Code, showing that a well-designed organization can outperform strong existing harnesses, while OpenCollab's single-agent configuration uses the fewest tokens across all evaluated suites. OpenCollab establishes a unified multi-agent infrastructure for easy programmable collaboration and controlled causal evaluation.
160. 【2609.38334】EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
链接:https://arxiv.org/abs/2609.38334
作者:Yuhan Guo,Jinming Liu,Liang Xu,Ziqiang Li,Jianguo Huang,Zhicheng Wang,Hu Zhu,Qiuyu Chen,Yuntao Wei,Xin Jin,Wenjun Zeng
类目:Computation and Language (cs.CL)
关键词:Large language models, Large language, multi-step decision-making, increasingly deployed, transfer poorly
备注: 19 pages. Project page: [this https URL](https://gnonymous.github.io/EVOKE) ; Code: [this https URL](https://github.com/Gnonymous/EVOKE) ; Models: [this https URL](https://huggingface.co/Gnonymous/EVOKE)
点击查看摘要
Abstract:Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
161. 【2609.38332】Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
链接:https://arxiv.org/abs/2609.38332
作者:Xinyu Li,Mononito Goswami,Hao Liu,Nikos Kanakaris,Langlin Huang,Prithwish Jana,Patrick Blöbaum,Purak Jain
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:improves model performance, performance by allocating, Test-time scaling improves, allocating additional compute, scaling improves model
备注: 45 pages
点击查看摘要
Abstract:Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.
162. 【2609.38291】HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control
链接:https://arxiv.org/abs/2609.38291
作者:Zhuo Liu,Moxin Li,Zhixin Ma,Wentao Shi,Wenjie Wang,Fuli Feng
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:Large language model, injected malicious instructions, prevent unsafe action, Large language, language model
备注:
点击查看摘要
Abstract:Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection rather than timely prevention or rely on predefined rules with limited flexibility across diverse risks. We propose a risk-aware harness that integrates LLM-based monitoring for flexible risk detection and structures monitor-guided execution around three core modules: trigger, monitor, and feedback, enabling targeted safety interventions while limiting disruption to benign task execution. To adapt the harness to different risks and deployment settings, we introduce HARDE, a two-stage harness optimization framework that first performs isolated probing of each module to derive an optimization guide, then uses this guide to iteratively optimize the harness based on safety and utility feedback. Experiments across three attack benchmarks show that HARDE improves runtime safety while preserving utility, outperforming manually designed harnesses and naive optimization baselines. Our analysis shows that effective runtime defense benefits from complementary safety mechanisms, attack-aware harness optimization, and harness designs matched to monitor capabilities. Our code is available at this https URL.
163. 【2609.38274】Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection
链接:https://arxiv.org/abs/2609.38274
作者:Liangyu Teng,Hengsong Liu,Juncen Guo,Jingyu Zhang,Yang Liu,Jing Liu,Liang Song
类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
关键词:individual model capabilities, inter-member error resonance, model capabilities, performance ceiling, LLM team
备注:
点击查看摘要
Abstract:The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics. We propose a heterogeneity-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co-failures, while the other measures divergence in predictive behavior to capture strategy diversity. We formulate team selection as a standardized quality--complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi-LLM systems.
164. 【2609.38261】NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space
链接:https://arxiv.org/abs/2609.38261
作者:Takanori Kotama,Shun-ichiro Hayashi,Daichi Mukunoki,Tetsuya Hoshino,Takahiro Katagiri
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:composed language models, single trained shared-latent, language models obtained, frozen language models, trained shared-latent adapter
备注:
点击查看摘要
Abstract:In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model runs the first layers of one model, converts the resulting intermediate representation once with the adapter, and then runs the remaining layers of the other model. Once the adapter is trained, several composed models that connect at different layers are obtained without retraining. Using the recurrent RWKV-4-Raven-7B and the Transformer-based Tulu-Pythia-6.9b, abbreviated as RWKV and Pythia, this study examines whether frozen models from different families can be recombined post hoc. The composed models answered multiple-choice questions, and those whose generations we examined produced syntactically well-formed text. The configuration that combines the first 5 layers of Pythia with the remaining 27 layers of RWKV reduced the Transformer key-value (KV) cache by 84.4% with accuracy not significantly different from that of RWKV alone. In multiple-choice accuracy, however, no composed model matched the parent model Pythia, and language-modeling performance decreased sharply on WikiText, a corpus of Wikipedia articles outside the training domain. The correspondence between intermediate representations was also obtained in one favorable case, with a shared tokenizer, the same depth, and the same hidden width, and does not show that the models share a general semantic space.
165. 【2609.38260】ContextAdapt: Evaluating Contextual Adaptation and Value Alignment in LLMs
链接:https://arxiv.org/abs/2609.38260
作者:Olivia Macmillan-Scott,Mirco Musolesi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:general principles underpinning, regarded as general, professional, models, general principles
备注:
点击查看摘要
Abstract:Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision is made. In this paper, we ask whether large language models (LLMs) appropriately adapt the application of a value across professional settings, while remaining consistent when contextual changes do not alter the relevant professional norm. To study this, we introduce ContextAdapt, an evaluation framework covering honesty, autonomy, and confidentiality across medicine, law, finance, and national security. Drawing on primary-source professional and regulatory documents, we construct a value x domain framework and use this to develop scenarios testing both default professional rules and recognised exceptions. We evaluate 12 LLMs on both the actions they recommend and the justifications they provide. In our main experiment, models achieve 95.6% mean appropriateness, although the use of the correct domain-specific justification varies substantially across models, from 25.6% to 76.9%. In a separate factorial experiment, explicitly naming the professional domain and changing the role of the model have limited effect on behaviour. Varying stakes, however, reveals severe but localised failures: in some cases, models alter their responses even though the underlying professional obligation remains unchanged. In particular, perceived severity appears to act as a cue for disclosure across both honesty and confidentiality scenarios. These results show that evaluating value alignment requires us to consider not only whether models follow abstract principles, but whether they apply them appropriately across different contexts.
166. 【2609.38256】Framing the Narrative: Ideological Mimicry in Large Language Models
链接:https://arxiv.org/abs/2609.38256
作者:Olivia Macmillan-Scott,Michael Jacobs,Nils Metternich,Mirco Musolesi
类目:Computation and Language (cs.CL); Computers and Society (cs.CY)
关键词:Large language models, Large language, evaluations typically treat, answer questions, typically treat
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly used to answer questions about politically contentious issues, yet evaluations typically treat a model's stance as a relatively stable property. Real users, however, communicate political signals through their terminology, assumptions, and personal context. We investigate whether such signals produce ideological mimicry: systematic shifts in the political stance expressed by an LLM toward the position conveyed by the interaction. If LLMs adapt their responses to these signals, they risk creating personalised political information environments in which users with opposing views receive systematically different accounts of the same issue, potentially reinforcing existing divisions. We build the Poli-SHIFT dataset and evaluation framework and assess seven open-weight LLMs across ten contentious political topics in the United States, United Kingdom, and Australia, systematically manipulating contested terminology, politically valenced premises, and user information, and eliciting responses in both multiple-choice and open-text formats. Across models, we find robust evidence that prompt framing shapes the political stance of LLM outputs. Changing terminology alone reverses which side of an issue a model supports in 16.9% of matched comparisons. Stated political ideology also systematically shifts responses toward the user's position. These findings show that political stance is not a fixed property of LLMs; the views expressed are conditional on the interaction with the user. As LLMs become increasingly personalised sources of information, such interaction-dependent adaptation could contribute to political information environments that reinforce users' existing perspectives.
167. 【2609.38232】When Does a Spoken Agent Have Enough Evidence to Act? The PACT-SLM Contract Test
链接:https://arxiv.org/abs/2609.38232
作者:Mengzhe Geng
类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:Streaming spoken agents, Speech Language Models, observed prefix supports, Streaming spoken, Partial Speech Action
备注:
点击查看摘要
Abstract:Streaming spoken agents may take an external action before the available speech supports it, yet final-turn scores do not reveal whether each observed prefix supports that action. We introduce the Partial Speech Action Contract for Turn Taking in Speech Language Models (PACT-SLM), a controlled evaluation that assigns a first valid action time and measures action identity and timing separately. The primary diagnostic contains 80 paired contrast groups from four held-out semantic families and 1,600 prefix predictions across clean and 15 dB noise renderings. After correcting a mismatch between randomized branch codes and semantic labels, a refitted WavLM Base Plus probe reaches 26.03% pooled post-onset semantic-label accuracy (95% group-bootstrap interval: 22.14%-29.68%), exposes an action on 18.99% of pre-onset prefixes, and predicts 5.94% of complete trajectories exactly. It exceeds matched text, scalar-acoustic, and shuffled-representation probes in post-onset label accuracy, but its score is at the 96th percentile of 100 within-prefix label permutations and below the 97.5th-percentile reference (26.73%). Elapsed time is more onset-exact than WavLM Base Plus (36.25% vs. 23.13%) but less accurate about action identity (9.92% vs. 26.03%). These results show that action identity and timing measure distinct aspects of partial-speech decision behavior.
168. 【2609.38222】Conformal Factuality Control for Multi-Hop Retrieval-Augmented Generation
链接:https://arxiv.org/abs/2609.38222
作者:Muhammad Aimal Rehman,Chi-Kuang Yeh
类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Machine Learning (stat.ML)
关键词:ground large language, large language models, Retrieval-augmented generation, external evidence, multi-hop RAG
备注: 15 pages, 2 figures. Code available at [this https URL](https://github.com/aimalrehman92/conformal-rag)
点击查看摘要
Abstract:Retrieval-augmented generation (RAG) can ground large language models in external evidence, but retrieved context does not guarantee that generated claims are factually supported. This problem is especially relevant in multi-hop RAG, where retrieval and reasoning proceed through multiple dependent stages. We study whether claim-level conformal factuality control, previously developed for RAG, remains effective in this setting. We apply split-conformal claim filtering to multi-hop RAG and evaluate it on HotpotQA, Natural Questions, and TriviaQA using Llama 3.1 8B and GPT-4o-mini, together with a single-hop reference experiment. Across all six multi-hop model-dataset configurations, increasingly stringent conformal targets consistently increase the fraction of responses whose retained claims are fully supported. At the 95% target, this rate ranges from 95.80% to 97.20%, compared with 55.60%-76.03% without filtering. However, the improvement is strongly selective: only 4.41%-31.09% of generated claims are retained and 9.70%-51.40% of responses remain non-empty at the 95% target. These results show that conformal factuality extends to multi-hop RAG, while demonstrating that nominal reliability must be interpreted jointly with claim retention and abstention.
169. 【2609.38219】utlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels
链接:https://arxiv.org/abs/2609.38219
作者:Mohamed-Amine Chadi,Ezzahra Ait El Arbi,Ismail Khayoub,Aymane Fadili,Yassine Ennhili,Jadjigua Bouali,Hanane Inhid,Mohammed Ameksa,Hajar Mousannif
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:remains severely under-resourced, Modern Standard Arabic, languages of Morocco, generally lacks information, uneven transcription quality
备注:
点击查看摘要
Abstract:Tamazight (Amazigh) is, together with Arabic, one of the two official languages of Morocco, yet it remains severely under-resourced for speech technology: pub licly available labelled audio is scarce, generally lacks information on the regional variety spoken, and is often of uneven transcription quality. This article describes the TutlAit dataset, a corpus of Moroccan Tamazight speech paired with Modern Standard Arabic text and explicit regional accent labels. The data were collected with TutlAit, a purpose-built crowdsourcing web application (React 18 front end, Django 5 / Django REST Framework back-end, PostgreSQL database). Native speakers recruited through targeted LinkedIn and Instagram campaigns created an account, declared their regional variety (Atlas, Souss, Rif or other) and demographic information, and then contributed through two workflows: Text-to Audio, in which an Arabic sentence is displayed and the volunteer records its oral Tamazight rendering in the browser, and Audio-to-Text, in which a Tamazight excerpt is played and the volunteer types its Arabic transcription. A complemen tary set of segments was obtained from freely accessible Tamazight audiovisual media, segmented and annotated with ELAN and imported through a bulk CSV/ZIP pipeline. Every upload is converted server-side to 16kHz mono WAV, hashed with SHA-256 for duplicate rejection, checked for duration bounds and validated by an administrator. The dataset contains 13,384 audio files totalling 75,231 seconds (approximately 20.9 hours, about 3.01GB). The Atlas variety accounts for 9,956 files (14.08h) and the Souss variety for 3,378 files (6.75h); small Rif (22 files) and Kabyle (28 files) subsets are also included. The corpus can be reused for speech recognition, speech translation and accent identification for Moroccan Tamazight.
170. 【2609.38205】he System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
链接:https://arxiv.org/abs/2609.38205
作者:Muhammad Usama,Dong Eui Chang
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:primary lever practitioners, Centered Kernel Alignment, language model behavior, control language model, remains poorly understood
备注:
点击查看摘要
Abstract:System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. Effects are layer-selective and instruction-type-dependent: persona and formatting instructions deeply restructure intermediate representations, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. Restrictive safety instructions and explicitly permissive ones ("you have no restrictions") engage near-identical computational pathways (mean CKA correlation 0.997), and this persists at commercial scale, where safety penetration remains below 10% even at 70B-72B. A linear probing baseline exposes the mechanism: the model encodes prompt category at every layer but restructures its computation only at a small subset, so the prompt is reliably "seen" but, for safety, not deeply "acted upon." Causal activation patching confirms these layers mediate behavioral change, and representational depth predicts behavioral effect size across the full 17-model cohort (Spearman rho = 0.761, p 0.001). The findings provide a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety. Code: this https URL
171. 【2609.38203】Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling
链接:https://arxiv.org/abs/2609.38203
作者:Bahman Mirheidari,Leslie Ing,Daniel Blackburn,Sharon Abrahams,Christopher McDermott,Heidi Christensen
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
关键词:Monitoring cognitive impairment, motor neuron disease, Behavioural ALS Screen, Verbal Fluency Index, co-occurring speech difficulties
备注:
点击查看摘要
Abstract:Monitoring cognitive impairment (CI) in motor neuron disease (MND) is essential for timely treatment and care, yet challenging due to co-occurring speech difficulties. The Edinburgh Cognitive and Behavioural ALS Screen (ECAS) provides a robust metric for CI assessment, with the Verbal Fluency Index (VFI) a central element. Building on recent advances in automated speech analysis, this study proposes a system for estimating VFI. It leverages a unique MND dataset and combines ASR (WhisperX) and VAD (Silero) with refined timestamping to predict the VFI and extract several clinically interpretable measures. Our approach outperformed systems based on traditional acoustic features and self-supervised embeddings, evaluated using multiple regression algorithms. Clinically inspired features consistently outperformed the other sets, with the best models achieving strong results (P-words: R2 0.9, NRMSE 0.05; S-words: R2 0.8, NRMSE 0.08), demonstrating the feasibility of automated VFI estimation.
172. 【2609.38201】omasuLLM: Out-of-Order Speculative Execution for LLM Agents
链接:https://arxiv.org/abs/2609.38201
作者:Jiangnan Yu,Ceyu Xu,Mengming Li,Shiyu Huang,Yiran Xia,Jian Weng,Hui Xue,Haohui Mai,Yuan Xie
类目:Computation and Language (cs.CL); Operating Systems (cs.OS); Software Engineering (cs.SE)
关键词:dominate coding-agent latency, test suites, Long-running tools, dominate coding-agent, repository commands
备注:
点击查看摘要
Abstract:Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors -- asequential interface hides work that can be predicted and started early, but a speculative result may become visible only after it and every earlier step have been validated. We present TomasuLLM, a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness. It drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state. Across three benchmarks spanning sub-second to minutes-long tool calls, TomasuLLM improves the reported benchmark means and scales with tool latency: 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions. Across 4,010 audited commit-validation records, it produces zero false accepts.
Subjects:
Computation and Language (cs.CL); Operating Systems (cs.OS); Software Engineering (cs.SE)
Cite as:
arXiv:2609.38201 [cs.CL]
(or
arXiv:2609.38201v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.38201
Focus to learn more
arXiv-issued DOI via DataCite</p>
173. 【2609.38181】Large Language Models are Approximate Survival Estimators
链接:https://arxiv.org/abs/2609.38181
作者:Juan M Zambrano Chaves,Peniel Argaw,Risa Ueno,Carlo Bifulco,Kristina Young,Rom Leidner,Tristan Naumann,Hoifung Poon
类目:Computation and Language (cs.CL)
关键词:medical risk assessment, risk assessment, Survival, Joseph Health Network, LLMs
备注:
点击查看摘要
Abstract:Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readily accessible through consumer applications. However, whether LLMs can provide accurate survival predictions has not been rigorously evaluated. We introduce Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes and prompts pre-trained LLMs to predict survival zero-shot. We benchmark Survprompt against conventional survival models, including random survival forests (RSF), across two multi-institutional pan-cancer cohorts: the publicly available MSK-CHORD cohort and a newly curated cohort from the Providence St. Joseph Health Network constructed using an LLM-based medical abstraction framework. We report censored mean absolute error (cMAE) and concordance index (c-index) and conduct feature ablations to identify variables influencing LLM predictions. Frontier LLMs achieved surprisingly competitive cMAE for individual survival times. For example, GPT-5.6-Sol achieved cMAE within 10% of state-of-the-art RSF models specifically trained for survival prediction for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD. Feature ablations revealed that LLMs prioritized clinical variables similarly to specialized survival models. However, LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated between high- and low-risk patients (lower c-index). Zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance across cancer types and institutions remains an important limitation for clinical use.
174. 【2609.34024】Jev in Medicine: A Benchmark Evaluation. Preliminary Results
链接:https://arxiv.org/abs/2609.34024
作者:Alfredo Madrid-García,Beatriz Merino-Barbancho
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:NEJM Case Challenges, model that assigns, Jev, Sol, System
备注:
点击查看摘要
Abstract:Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
175. 【2609.03221】Instability Floors: Separating Bias from Noise in Fairness Audits of Clinical LLM Agents with FairMedAgent
链接:https://arxiv.org/abs/2609.03221
作者:Rohith Reddy Bellibatlu,Manpreet Singh,Deepak Parashar,Rahul Joshi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG); Applications (stat.AP)
关键词:Counterfactual fairness audits, patient demographic descriptor, language-model agents report, flip rate, Counterfactual fairness
备注: 27 pages (13 main plus 14 supplementary), 4 figures, 3 tables. Code: [this https URL](https://github.com/rohithreddybc/FairMedAgent) (v0.1.5, commit 3982974; concept DOI [https://doi.org/10.5281/zenodo.22165979](https://doi.org/10.5281/zenodo.22165979) ). Trajectories: [this https URL](https://huggingface.co/datasets/Rohithreddybc/FairMedAgent)
点击查看摘要
Abstract:Counterfactual fairness audits of clinical language-model agents report a flip rate: how often an action changes when only the patient's demographic descriptor changes. Part of that rate is not demographic. A stochastic agent also changes its own action when nothing changes, and a flip rate cannot be interpreted without knowing how often. We measured it. Re-running one condition ten times over sixteen synthetic vignettes at default sampling changed a clinical agent's action in 8.7 percent of replicate pairs, from 2.2 percent for intensive-care escalation to 17.9 percent for controlled-substance caution, an output given no operational criteria. Across six models from five vendors, pooled floors ranged from 2.5 to 23.7 percent; in this panel neither disclosed size, vendor, nor hosting ordered them. The floor depends on the decoding configuration: majority voting over five draws removed 39 percent of it (95 percent confidence interval (CI) 18 to 64); at temperature 0 three of four locally served models showed no disagreement, but a hosted model still did. We also show that, for a binary action, the flip rate expected under no demographic effect equals the floor and a real effect adds only its square, so a flip rate inside the floor is not evidence of fairness, and direction must be tested with a signed paired test. We give a four-step reporting procedure and release FairMedAgent, the harness, with its protocol, vignettes and analysis scripts, so any team can measure the floor for its own agent.
176. 【2608.03756】LegalPincite: Multi-level Legal Information Retrieval Dataset
链接:https://arxiv.org/abs/2608.03756
作者:Theresia Veronika Rampisela,Henrik Palmer Olsen,Giovanni Colavizza
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:find relevant legal, relevant legal sources, case-law collections, common task, find relevant
备注: Accepted for publication at the 8th Natural Legal Language Processing Workshop (NLLP 2026), co-located with EMNLP 2026
点击查看摘要
Abstract:A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset and code: this https URL
177. 【2602.19001】Life-Bench: A Benchmark and Knowledge Graph Framework for Multimodal Personalization Beyond Concept Recognition
链接:https://arxiv.org/abs/2602.19001
作者:Xia Hu,Honglei Zhuang,Brian Potetz,Alireza Fathi,Bo Hu,Babak Samari,Howard Zhou
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:large language models, language models increasingly, models increasingly power, primarily target concept-level, power personal assistants
备注:
点击查看摘要
Abstract:As large language models increasingly power personal assistants, users expect them to reason over multimodal life histories, from recognizing people to understanding events to aggregating patterns, yet existing benchmarks primarily target concept-level recognition. We introduce Life-Bench, a fully synthetic, human-verified multimodal benchmark of over 11,800 question-answer pairs across 10 tasks, organized by required evidence scope: concept identification, event understanding, and aggregated reasoning. The benchmark's photo-centric personal histories are distributionally aligned with real user accounts under embedding statistic. The interconnected structure of personal data invites graph-based solutions; we propose LifeGraph, a personal knowledge graph framework providing structured retrieval with on-demand access to source visual evidence, showing particular promise on event and aggregated tasks. Systematic evaluation of four retrieval paradigms on Life-Bench demonstrates that accuracy degrades sharply with evidence scope, falling below 0.40 on aggregated tasks, and that no single paradigm dominates across categories. Performance beyond concept recognition remains modest for all evaluated methods, establishing personalization over multimodal histories as an open challenge and Life-Bench as a testbed for future progress.
178. 【2609.38887】VOSSA: Voiceprint Optimization for Streaming Speech Architectures
链接:https://arxiv.org/abs/2609.38887
作者:Mu-Ruei Tseng,Waris Quamer,Ghady Nasrallah,Ricardo Gutierrez-Osuna
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Real-time voice conversion, systems commonly rely, Real-time voice, automatic speaker verification, Streaming Speech Architectures
备注: Published in Proceedings of Interspeech 2026
点击查看摘要
Abstract:Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.
179. 【2609.38658】acit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
链接:https://arxiv.org/abs/2609.38658
作者:Jian Chen,You Zhang,Mark Vinton
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
关键词:rich expressive variation, incurs substantial latency, demonstrated strong zero-shot, autoregressive semantic modeling, sequential decoding incurs
备注: Under Review
点击查看摘要
Abstract:TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
180. 【2609.38324】Multi-agent discussion gains less when dissent is withheld
链接:https://arxiv.org/abs/2609.38324
作者:Chand Sahil Mansuri,Xin Wang,Mengying Li,Bryan Acton,Rory Eckardt,Dhaval Patel,Sadamori Kojaku
类目:Physics and Society (physics.soc-ph); Computation and Language (cs.CL); Multiagent Systems (cs.MA)
关键词:Multi-agent systems, LLMs add discussion, discussion improves accuracy, incorrect consensus, Multi-agent
备注:
点击查看摘要
Abstract:Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on whether discussion improves accuracy or leads to an incorrect consensus. Here, we introduce a parsimonious model that explains when discussion improves accuracy and when it ends in an incorrect consensus, built from four behaviors repeatedly observed in LLM agents: (1) withholding dissent, (2) internalizing a stated answer, (3) reconsidering after seeing dissent, and (4) correcting toward the correct answer. The model shows that discussion can overturn an incorrect initial majority only when the withholding rate $c$ is below a critical rate $c^* = \gamma/(\gamma + a)$, set by the net correction rate $\gamma$ and the internalization rate $a$. We estimate these rates from conversation logs with a Bayesian method and place LLM teams relative to $c^*$. As the model predicts, the gain from discussion shrinks as withholding rises, across LLMs and on a hidden profile benchmark, HiddenBench, and MedEInst. Instructing agents not to withhold dissent increases this gain. Turning reasoning off also increases the gain, because reasoning raises the internalization rate $a$ and keeps agents from reconsidering a minority answer. These findings reconcile the conflicting reports and identify when discussion outperforms majority voting.
信息检索
1. 【2609.40241】Decision-Oriented Recommendation Reranking: An Empirical Study of Jev
链接:https://arxiv.org/abs/2609.40241
作者:Hanjia Lyu,Yinglong Xia
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:Large language models, Large language, shown promise, introduces an important, important tradeoff
备注:
点击查看摘要
Abstract:Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate whether a decision-oriented model provides a useful alternative when the reranking task is fundamentally a structured choice among predefined candidate items. Specifically, we conduct a controlled empirical study of Jev, described by TypeSafe AI as a ``System One Model,'' for personalized recommendation reranking and compare it with recommendation-specific models and pointwise and listwise Qwen rerankers across multiple Amazon Reviews domains and candidate-set sizes, evaluating both recommendation effectiveness and observed serving latency. Our results show that Jev maintains strong recommendation effectiveness relative to the evaluated baselines while exhibiting substantially more gradual latency growth than the pointwise Qwen rerankers, although its observed serving latency remains substantially higher than that of recommendation-specific models. Together, these characteristics place Jev in a distinct quality--latency operating regime across candidate sizes and domains. These findings motivate further investigation of decision-oriented models for recommendation and other ranking tasks with structured output spaces.
2. 【2609.40069】Conversational Capture: A Trajectory-Level Framework for Evaluating Generative Engine Optimization in Multi-turn Human-Agent Interaction
链接:https://arxiv.org/abs/2609.40069
作者:Junwei Yu,Jieyu Zhou,Mufeng Yang,Yepeng Ding,Hiroyuki Sato
类目:Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
关键词:Generative Engine Optimization, large language models, retrieval-augmented large language, Generative Engine, Engine Optimization
备注: 9 pages, 2 figures, 2 tables. In Proceedings of the 14th International Conference on Human-Agent Interaction (HAI '26), November 16-19, 2026, Osaka, Japan
点击查看摘要
Abstract:Generative Engine Optimization (GEO) shapes content to increase its likelihood of being cited by answer engines built on retrieval-augmented large language models. GEO is typically evaluated as a single-turn property: for a fixed query, an evaluator measures a source's visibility in one answer. We argue that the single answer is an inadequate unit of analysis. Human-agent information seeking forms a closed loop: the agent's answer changes the user's beliefs and therefore the next question, which in turn determines what the agent retrieves. We introduce conversational capture, a phenomenon in which a source cited early becomes substantially more likely to be cited again. Capture operates through a machine-side channel, history-conditioned retrieval, and a human-side channel, follow-up questions directed toward the captured source. We formalize the interaction as a two-layer closed-loop system and derive trajectory-level constructs: cumulative conversational visibility; a direct/feedback decomposition of trajectory gain; a nested split of the feedback term into machine-side and human-side channels; a capture coefficient; a compounding ratio; and a misranking diagnostic. Using reinforcement-process (Pólya-urn) theory, we prove that the feedback term is zero under single-turn evaluation and that GEO's cumulative payoff grows superlinearly with conversation length while capture develops. A model-derived illustration shows that the feedback term can exceed the direct term, the compounding ratio exceeds two within ten turns, and single-turn and trajectory rankings agree only weakly (Kendall's $\tau = 0.4$). We connect the human channel to information foraging, trust calibration, and Bayesian persuasion, and discuss design implications for answer engines.
3. 【2609.39975】Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
链接:https://arxiv.org/abs/2609.39975
作者:Anastasios Nentidis,Georgios Katsimpras,Anastasia Krithara,Martin Krallinger,Miguel Rodríguez-Ortega,Eduard Rodriguez-López,Natalia Loukachevitch,Igor Rozhkov,Elena Tutubalina,Dimitris Dimitriadis,Vasiliki Patsiou,Grigorios Tsoumakas,George Giannakoulas,Alexandra Bekiaridou,Athanasios Samaras,Giorgio Maria Di Nunzio,Nicola Ferro,Stefano Marchesin,Marco Martinelli,Gianmaria Silvello,Georgios Paliouras
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:Evaluation Forum, Conference and Labs, paper presents, presents an overview, question answering
备注: 21 pages, 17 tables, International Conference of the Cross-Language Evaluation Forum for European Languages 2026 (CLEF2026)
点击查看摘要
Abstract:This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical semantic question answering. b) Task Synergy14 on question answering for developing biomedical top- ics. c) Task MultiClinSum-2 on multilingual clinical summarization. d) Task BioNNE-R on extracting relations between nested named entities in Russian and English. e) Task ELCardioCC on clinical coding in cardiology. f) Task GutBrainIE on gut-brain interplay information extrac- tion. Across these six tasks, 87 distinct teams participated, submitting more than 1000 runs overall. As in previous editions, several submissions reached competitive performance, reflecting the continued progress of state-of-the-art methods across biomedical language processing tasks.
4. 【2609.39828】KUAISHOU Explorer LLM-Rec Challenge 2026: Reasoning Generative Recommendation
链接:https://arxiv.org/abs/2609.39828
作者:Jiangxia Cao,Hao Peng,Wenlong Xu,Jiaxin Deng,Zhixin Ling,Xingmei Wang,Kun Shang,Can Tang,Zhihuai Cai,Jun Du,Fang Su,Xiaojuan Liu,Yiling Li,Chenglong Yu,Chongling Rao,Haixuan Gao,Haitao Xu,Jian Liang,Ruiming Tang,Chenglong Chu,Guohong Mu,Honghui Bao,Hui Wang,Jialong Chen,Jiao Ou,Muhao Wei,Peng Zhang,Renpu Liu,Ruochen Yang,Shugui Liu,Xinqi Jin,Yan Sun,Yifan Wang,Yingzhi He,Yufei Ye,Yusen Huo,Tingkuo Wang,Jihong Zhang,Lanxi Zhu,Pengyuan Liu,Zhipeng Yi,Luankang Zhang,Hang Lv,Xuyang Zhi,Tianyu Li,Bintao Wu,Chuang Ou,Siyue Su,Ziyuan Wang,Yuliang Sun,Baiyan Che,Feiyang Xu,Shiwen Zhang,Shiteng Cao,Chongcong Jiang,Yuan Fang,Xiangwu Yang,Hao Deng,Zijian Du,Pengxun Wang,Xiaoming Wang,Shun Qin,Yingqi Song,Tianyi Li,Naixiao Peng,Chenyu Zhou,Qiliang Jiang,Quan Zheng,Cheng Jin,Siying Zeng,Hongjia Xu,Junwu Hu,Teng Fu,Zhengkang Mei,Haijun Yu,Kai Li,Shengyang Zhou,Zhijia Wei,Siyi Xiong,Bo Liu,Zichun Guo,Zhubin Han,Jinpeng Fu,Bingqian Liu,Yuyi Wang,Yu Liu,Qinghai Tan,Ruijie Zhou,Zhuohang Li,Zhijia Zhong,Xiangnan He,Jirong Wen,Min Zhang,Wenwu Ou,Peng Jiang,Han Li,Kaiqiao Zhan,Yanan Niu,Lantao Hu,Kun Gai
类目:Information Retrieval (cs.IR)
关键词:build next-generation recommender, build next-generation, attracted a surge, surge of attentions, academic research community
备注:
点击查看摘要
Abstract:Generative recommendation, has been attracted a surge of attentions in industrial and academic research community, towards to build more smart system to build next-generation recommender. Under the significant developing wave of large language model, our team have been developed Semantic ID based OneRec/OneRec-V2. These models have been widely deployed in production and demonstrate the scaling potential of the autoregressive next-item prediction paradigm for industrial recommender systems. Building on the success of OneRec, we further explored a series of models, including OneRec-Think, OpenOneRec, and OneReason, that connect item Semantic IDs with natural language in a unified representation space and seek to unlock the potential of natural-language chain-of-thought (CoT) reasoning for recommendation. However, our preliminary works found that introducing reasoning CoT does not always improve the recommendation performance. To address this issue, OneReason strengthens the semantic alignment between items and language, introduces structured template-based supervision for interest reasoning, and applies advanced reinforcement learning techniques to make reasoning more beneficial to recommendation. As a frontier topic to building recommendation foundation models, we believe this topic has significant research value and hope to encourage more researchers to explore it together. To this end, together with the SIGIR 2026 community, we organized the KUAISHOU Explorer LLM-Rec Challenge 2026: Reasoning Generative Recommendation.
5. 【2609.39696】When the Label Ignores the Request: Auditing Policy-Selected Targets in Synthetic Conversational Music Recommendation
链接:https://arxiv.org/abs/2609.39696
作者:Sanjeev Suresh
类目:Information Retrieval (cs.IR)
关键词:LLM listener talks, Synthetic dialogues generated, LLM recommender, LLM pipelines, LLM listener
备注: 6 pages, 2 tables. Camera-ready version (CC BY 4.0). RecSys Challenge 2026 Workshop at ACM RecSys 2026, Minneapolis, October 2, 2026. Code and audit artifacts: [this https URL](https://github.com/Sanjeev-S/recsys2026-request-audit)
点击查看摘要
Abstract:Synthetic dialogues generated by LLM pipelines now serve as complete conversational-recommendation benchmarks: an LLM listener talks to an LLM recommender, and the track logged next in the conversation becomes the official label for each turn. These policy-selected labels make large-scale evaluation reproducible, but they are proxies for what the simulated user asked. We audit the one place where label and request are directly comparable: turns where the user asks for an exact song by name. In the RecSys Challenge 2026 TalkPlay benchmark, using visible dialogue and catalog metadata alone, we find that the official label contradicts the user's exact-song request in half of the audited development turns. This matters beyond one benchmark: naming the desired item is the dominant intent in real music search, where deployed systems avoid substituting an alternative for an exactly named item, on the premise that it costs satisfaction. A small training-time supplement closes most of the gap: adding catalog-resolved request-satisfying targets to a small fraction of training turns yields a 53.3% relative gain in nDCG@20 on the 43 conflict turns while leaving the official metric intact, verified against a matched control that detects the same requests but trains only on official labels.
6. 【2609.39358】Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
链接:https://arxiv.org/abs/2609.39358
作者:Sietse Schelpe
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG); Performance (cs.PF)
关键词:Vishal Sikka, CEO of Infosys, transformer language model, language model performs, model
备注:
点击查看摘要
Abstract:A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and this http URL that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on this http URL at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
7. 【2609.39327】Generative End-to-end Ad Retrieval at Douyin
链接:https://arxiv.org/abs/2609.39327
作者:Shaowen Zeng,Yanhua Huang,Jiacheng Sun,Jiarui Liu,Qian Dai,Zhikai Yang,Hancheng Li,Boya Wu,Tuoyu Zhang,Yekui Chen,Xiang Sun
类目:Information Retrieval (cs.IR)
关键词:retrieval reformulates recommendation, discrete item tokens, reformulates recommendation, generation of discrete, Generative retrieval reformulates
备注:
点击查看摘要
Abstract:Generative retrieval reformulates recommendation as the generation of discrete item tokens. However, scaling this paradigm to real-world recommender systems reveals two critical bottlenecks: 1) Representation collapse, where the item tokenizer converges to degenerate results under continuous distribution shifts, fundamentally hindering stable end-to-end adaptation. 2) Item collisions, where the massive candidate pool causes distinct items to share identical token sequences, compromising the final retrieval precision. Crucially, these bottlenecks are inherently coupled: expanding codebook capacity to mitigate collisions inevitably exacerbates collapse. To address them simultaneously, we propose GEAR, an end-to-end framework that jointly optimizes the tokenizer, generator, and reranker. To mitigate representation collapse, we introduce BasisVQ, which re-parameterizes the codebook via an orthogonal basis to enable global gradient sharing and rigid spatial rotation of the latent space, effectively stabilizing gradient dynamics without ad-hoc heuristics. We further extend it to prefix-aware BasisRQ, substantially enhancing the codebook's expressiveness with the same asymptotic time complexity. To resolve item collisions, GEAR integrates a context-conditioned reranking head into the generative process, efficiently disambiguating colliding items with minimal computational overhead. By unifying stable tokenization and joint reranking within an end-to-end generative framework, GEAR establishes a fully differentiable and scalable paradigm. It currently serves hundreds of millions of daily active users on Douyin Ads, yielding substantial empirical improvements in extensive online A/B tests.
8. 【2609.39319】Residual Trajectory Distillation for Generative Retrieval
链接:https://arxiv.org/abs/2609.39319
作者:Weihao Shen,Wei Chen,Fuwei Zhang,Guojun Liu,Qingsong Hua,Wei Lin,Fuzhen Zhuang
类目:Information Retrieval (cs.IR)
关键词:discrete Semantic IDs, autoregressive identifier generation, Semantic IDs, discrete Semantic, general retrieval paradigm
备注:
点击查看摘要
Abstract:Generative retrieval has emerged as a general retrieval paradigm, representing items with discrete Semantic IDs (SIDs) and retrieving them through autoregressive identifier generation. When SIDs are constructed with residual quantization (RQ), standard retrieval training supervises only the selected codes and discards the residual trajectories that produce them. The same hard code can nevertheless arise from different preferences over competing codewords, while the residual trajectory also contains information about subsequent quantization decisions. As a result, hard SID supervision collapses distinct quantization behaviors into identical targets and leaves information available during indexing unused in retrieval training. We introduce ResTD, a Residual Trajectory Distillation framework that transfers this discarded indexing information into retrieval training. Treating the frozen RQ indexer as a process teacher, it distills residual-induced codeword preferences into SID-decoding states. This supervision recovers distinctions hidden by hard assignments and allows earlier decoder states to capture information about subsequent quantization decisions before the corresponding SID suffix is generated. In this way, richer information from SID construction is incorporated into retrieval learning while preserving the original retrieval index and inference procedure. Experiments on multilingual e-commerce retrieval show consistent improvements over strong baselines and matched training controls. Controlled comparisons show that residual-derived targets outperform the tested codebook-only soft targets. Representation probes further show that future codebook preferences become more recoverable from earlier decoder states. ResTD can also be readily extended beyond retrieval to generative recommendation. Code is available at: this https URL.
9. 【2609.39312】Learning Multiresolution Relevance for Hierarchical Generative Retrieval
链接:https://arxiv.org/abs/2609.39312
作者:Weihao Shen,Wei Chen,Fuwei Zhang,Guojun Liu,Qingsong Hua,Wei Lin,Fuzhen Zhuang
类目:Information Retrieval (cs.IR)
关键词:Generative retrieval, textbf, makes successive decisions, Generative, relevance
备注:
点击查看摘要
Abstract:Generative retrieval with semantic identifiers (SIDs) makes successive decisions over a document hierarchy. Relevant documents for the same query may share coarse prefixes and diverge at finer depths, with branching patterns varying across queries. These paths reveal how relevance is distributed across successive refinements, yet standard full-SID supervision treats them as separate training targets. To make this allocation explicit, we formulate multiresolution relevance as consistent conditional distributions induced by a single document-level relevance measure across the SID hierarchy. We introduce \textbf{RARS}, \textbf{R}esolution-\textbf{A}ligned \textbf{R}elevance \textbf{S}upervision, which uses the resulting refinement-level distributions to supervise a shared query representation. RARS aggregates document relevance over prefixes and trains a prefix-conditioned predictor to allocate relevance among sibling branches. All relevance-bearing children participate in local competition, and each local loss is weighted by the relevance mass reaching its parent. This objective trains the query encoder to capture both the coarse structure shared by relevant documents and their finer branch allocations. The predictor is discarded after training, preserving standard autoregressive retrieval at inference. Experiments on three multilingual ESCI locales show consistent improvements over matched full-SID training under autoregressive decoding. RARS also outperforms grouped soft-target, decoder soft-target, and sampled-tree supervision under a common retrieval rule. The gains persist across alternative identifier structures and relevance definitions. Code is available at: this https URL
10. 【2609.39225】Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures
链接:https://arxiv.org/abs/2609.39225
作者:Siddharth Bhargava,Sara Tonelli,Patricia Martín-Rodilla,Javier Parapar
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Argument structure prediction, constructs complete argument, complete argument structures, argument structures, Inference Anchoring Theory
备注: CMNA'26: 26th International Workshop on Computational Models of Natural Argument
点击查看摘要
Abstract:Argument structure prediction (ASP) constructs complete argument structures from discourse by identifying argumentative units and their relations. While recent work has explored diverse approaches---including unified neural models, multi-step pipelines, and prompt-based large language models (LLMs)---their relative trade-offs remain under-explored, particularly in dialogical settings. We present a systematic evaluation of ASP under strict schema constraints, comparing supervised fine-tuning and prompt-based LLMs across single- and multi-step task architectures, generating complete argument structures from dialogical input end-to-end. We benchmark them on three diverse dialogical corpora adapted from Inference Anchoring Theory into bipolar argument structures. Under a shared evaluation framework, we assess predictive performance, cross-domain generalization, schema compliance, and computational efficiency. Our results show that ASP remains a challenging task, with identifying argumentative relations emerging as the primary bottleneck, largely due to the implicit and context-dependent nature of dialogical argumentation. To facilitate future research, we release our data processing pipeline and end-to-end modeling framework for computational ASP on dialogical corpora.
Comments:
CMNA’26: 26th International Workshop on Computational Models of Natural Argument
Subjects:
Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:
arXiv:2609.39225 [cs.CL]
(or
arXiv:2609.39225v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.39225
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
11. 【2609.39209】O-Funnel: Lossless Structural Capture and Requirement-Driven Extraction from Drifting, Heterogeneous Documents
链接:https://arxiv.org/abs/2609.39209
作者:Osama Mustafa
类目:Databases (cs.DB); Information Retrieval (cs.IR)
关键词:Pulling a fixed, hand-written byte patterns, fixed set, byte patterns, Pulling
备注: 21 pages, 3 figures. Code and benchmarks: [this https URL](https://github.com/osamaa-mustafa/ofunnel)
点击查看摘要
Abstract:Pulling a fixed set of fields out of documents that arrive in many formats and under drifting schemas is usually done with hand-written byte patterns, which break whenever a key is renamed, a value is reformatted, or a lookalike value appears first. We argue the cause is structural: one pattern must both describe the value and locate it among its surroundings. O-Funnel separates the two. It transcribes any XML, JSON, CSV, HTML or key-value text document into one typed tree over five constructors, gated by an oracle that rejects any capture that does not reconstruct its source. Each needed field is declared in the tree's own terms and located by fusing independent evidence (key, path, value shape, synonym, key spelling, record neighborhood, value profile), so the best-supported node wins and a missing field is reported with a reason. Data no requirement claims becomes residue that a funnel traces back to the requirements to learn new key aliases. On 34,989 real PubMed records, O-Funnel matches a hand-written parser (F1 1.00). After a five-element schema rename, the parser's regular expressions fall to 0.20 while O-Funnel stays at 1.00, with every capture verified complete. On constructed suites that isolate regex failure modes it raises F1 from 0.43 to 1.00, and from 0.80 to 0.94 after self-improvement; on held-out schema-matching instances it is competitive with classical matchers without training. O-Funnel is a dependency-free Python library (pip install ofunnel).
12. 【2609.39043】Routing Between Generative and Collaborative User Profiles: A Serving-Time Gate for Controllable Novelty
链接:https://arxiv.org/abs/2609.39043
作者:Milad Sabouri,Neeraj Sharma,Sardar Hamidian,Shaghayegh Agah
类目:Information Retrieval (cs.IR)
关键词:Large language models, Large language, enable rich semantic, enable rich, deploy uniformly
备注:
点击查看摘要
Abstract:Large language models (LLMs) enable rich semantic user profiles for recommendation, but such profiles are more expensive to generate and are not necessarily desirable to deploy uniformly. We study whether LLM-generated profiles can instead be invoked selectively within a production recommendation pipeline. Using a real-world streaming dataset covering movies, TV shows, and sports content, we train a serving-time routing gate that assigns each user to either a collaborative sequential recommendation model or a recommendation model driven by an LLM-generated profile. The gate uses only serving-time features and learns to identify users for whom profile-based routing can increase Novelty@10 while preserving ranking relevance. A routing threshold controls how aggressively users are sent to the generative model, exposing a tunable novelty--relevance trade-off. At an overall NDCG-loss budget of 5\%, the learned gate increases Novelty@10 by 6.5\% while routing 12.5\% of users, outperforming simple heuristic and random routing policies at comparable relevance cost. These results show that LLM-generated user profiles can serve as a controllable complement to collaborative recommendation, while results with non-generative semantic profiles indicate that the benefit stems from selective routing rather than LLM generation alone.
13. 【2609.39013】Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem
链接:https://arxiv.org/abs/2609.39013
作者:Divya Godara,Sachin Gupta
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:DocSem shared task, official final test, final test evaluation, shared task, DocSem shared
备注: 5 pages, 1 figure, 2 tables. Accepted as a shared-task system paper at DocInsights 2026, co-located with EMNLP 2026
点击查看摘要
Abstract:EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and evaluates that expression in local code. Saved intermediate results support inspection of failures. A separate public-validation run achieved 92.17% answer accuracy and 1.00 evidence F1. The configurations and metrics differ, so these scores are not a controlled comparison. Our manual, post-hoc analysis is descriptive: in one inspected case, optical character recognition (OCR) and block grouping merged the relevant passage into another block, and the system answered from unrelated text. An exploratory study of reading page images on 100 documents returned evidence identifiers for only 22 documents. These descriptive findings motivate further evaluation; they do not establish the causes of the overall score.
14. 【2609.39007】RouteRec: Behavior-Guided Sparse Routing for Sequential Recommendation
链接:https://arxiv.org/abs/2609.39007
作者:Junyeong Song,Jaemin Yoo
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:improve sequential recommendation, sequential recommendation, improve sequential, expert, Abstract
备注: Accepted at CIKM 2026. 12 pages, 12 figures, 7 tables
点击查看摘要
Abstract:Sessionized interaction histories contain behavioral patterns that can improve sequential recommendation. However, existing models process all sessions through the same parameterized blocks, regardless of their behavioral differences. Mixture of Experts (MoE) enables conditional computation, but it leaves open what should guide expert allocation. We propose RouteRec, a sequential recommender that uses observed session behavior as the routing criterion. RouteRec summarizes four types of behavioral evidence from sessionized histories: interaction tempo, item-group focus, repetition and carryover, and popularity tendency. It uses these cues to route computation at macro, mid, and micro scopes. Cue-derived scores first select expert groups; within each selected group, the current backbone state then refines expert selection. Across six public datasets and 18 dataset-metric combinations, RouteRec ranks first in 12 and second in three, yielding the best overall average rank of 1.61 compared with 4.11 for the next-best baseline. Additional analyses suggest that the behavioral cues guide expert allocation beyond added capacity and produce routing patterns aligned with observed behavior. Our code is available at this https URL
15. 【2609.38999】When LLM-Inferred User Context Adds Value in Production Streaming Recommendation
链接:https://arxiv.org/abs/2609.38999
作者:Milad Sabouri,Neeraj Sharma,Sardar Hamidian,Shaghayegh Agah
类目:Information Retrieval (cs.IR)
关键词:shifting from static, predefined variables, information in recommender, variables toward latent, Contextual information
备注:
点击查看摘要
Abstract:Contextual information in recommender systems is shifting from static, predefined variables toward latent representations inferred from behavior. Large language models support this shift by rendering an unstructured interaction history as a natural-language summary, which yields a thematic user context that can be encoded and used in place of an aggregate profile. The conditions under which such generated profiles outperform aggregate embeddings have received limited characterization mainly at the domain level. We evaluate semantic user-profiling strategies on a production streaming platform, ranking against the full catalog. The evaluation covers a 2*2 design space crossing representation type (aggregate or LLM-generated) with contextual scope (holistic history or attention-fused short-term and long-term contexts). The relative ordering of the two representation types is conditional on the user's consumption regime. Aggregate profiles are consistently stronger under habitual consumption, which characterizes approximately four-fifths of the population, while LLM-generated profiles are stronger for exploratory users whose subsequent interactions diverge semantically from their history. We also observe a popularity-attractor effect in LLM-generated profiles, which modestly raises within-list diversity while substantially lowering catalog coverage and reducing novelty. These results indicate that a context-aware system can select a profiling strategy from the inferred consumption regime rather than applying one representation to all users.
16. 【2609.38949】xt-Video Retrieval via Multi-Dimensional Saliency Assessment and Granularity-Aware Query Decomposition
链接:https://arxiv.org/abs/2609.38949
作者:Shuquan Wei,Xi Chen,Xu Chen,Xiangyang Jia
类目:Information Retrieval (cs.IR)
关键词:joint embedding space, multimodal intelligence, aims to bridge, textual modalities, modalities by learning
备注:
点击查看摘要
Abstract:Text-video retrieval, which aims to bridge visual and textual modalities by learning a joint embedding space, has become a crucial task in multimodal intelligence. Despite extensive efforts to mitigate visual redundancy, previous methods typically rely on a single-aspect criterion to assess visual importance, overlooking the multifaceted spatiotemporal nature of video. In addition, encoding text into a single global embedding to align with videos compresses temporal events and spatial entities into a unified representation space, further aggravating cross-modal misalignment. To address these issues, we propose MMTI, a method that jointly mitigates visual redundancy and enables multi-grained text-video interaction to achieve accurate multi-grained semantic alignment. Specifically, a key feature selection (KFS) mechanism adaptively identifies and aggregates informative frames and patches by jointly evaluating multi-dimensional saliency and learnable importance scores, effectively compacting dense visual features and mitigating visual redundancy. Furthermore, our proposed multi-grained text-video interaction module (TVIM) employs a dynamic gating mechanism to decompose the text query into sentence, frame, and patch queries (SFP), enabling multi-grained text-video alignment. Complementary alignment at different granularities is thereby achieved. Extensive experiments on four standard benchmarks demonstrate that our method outperforms state-of-the-art methods.
17. 【2609.38946】Breaking News Out of the Filter Bubble: Generative AI Search Diversifies Collective Attention and Raises Shared Information Consumption
链接:https://arxiv.org/abs/2609.38946
作者:Heeseung Andrew Lee,Dokyun Lee,Gwanhoo Lee,Dongwon Lee
类目:Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Multimedia (cs.MM)
关键词:Washington Post, overviews are transforming, transforming access, encounter a narrower, narrower range
备注: 31 pages, 4 figures; includes supplementary material
点击查看摘要
Abstract:Generative AI search and AI overviews are transforming access to information and news, renewing concerns that readers will encounter a narrower range of topics and have less in common. We examine these concerns via a randomized field experiment with 37,561 readers at The Washington Post. Both groups searched the same archive, but treatment readers also received AI answers with article citations above conventional results. Measuring consumption across displayed answers and opened articles, we find that AI search expands the reach of widely read topics and increases overlap in readers' topic consumption. At the same time, consumption becomes less concentrated and shifts toward less-popular topics, both within readers and across the audience. AI answers account for most of the increase in shared information, delivering it without requiring article clicks and broadening exposure beyond the articles readers open. Cited articles also contribute to the shift toward less-popular topics. Readers shift from conventional-result clicks and browsing toward cited articles and follow-up searches. More frequent searching offsets lower article consumption per search, producing a small increase in article consumption per reader. Total information consumption per minute also rises. Generative AI search can thus diversify collective attention while strengthening the information readers have in common.
18. 【2609.38861】RACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA
链接:https://arxiv.org/abs/2609.38861
作者:Sachin Gupta,Divya Godara
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
关键词:Finding a relevant, producing a verifiable, verifiable answer, Finding, relevant paper
备注: 8 pages, 3 figures, 3 tables. Accepted at the 1st Workshop on Grounding Language Models: Learning Faithfully and Efficiently (GroundLM 2026), co-located with EMNLP 2026
点击查看摘要
Abstract:Finding a relevant paper is not the same as producing a verifiable answer from it. LitTraceQA requires canonical paper identifiers, exact evidence at the page or object level, and typed answers that match the evaluator. We call the separation between source access and scorer-visible correctness the grounding contract gap. TRACE - Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction - addresses this gap with target-grouped retrieval, independent typed evidence localization, multimodal table extraction, schema-driven table construction, and fail-closed validation. It indexes 27,487 papers through passage, object, alias, citation, and dense representations while retaining the question target behind each signal. For tables, TRACE predicts the observation unit before extracting values and assembles rows with evaluator-compatible key normalization. Our audited selected clean-track artifact scores 0.760613 on the official 71-question test set, including 0.9728 paper F1, 0.6847 evidence F1, 0.9800 multiple-choice accuracy, 0.5423 table-row F1, and 0.3508 macro cell accuracy. On 11 public-development table records, a clean baseline and coordinate-aware visual fill obtain row F1 of 0.291 and 0.411, respectively; this diagnostic comparison includes fallback outputs and is not an official-test claim. Remaining errors chiefly concern locator, observation-unit, row-key, and source-value identity.
19. 【2609.38822】SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
链接:https://arxiv.org/abs/2609.38822
作者:Guanqun Yang,Wenlong Zhang,Tian Shi,Ping Wang
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
关键词:http URL directories, Anthropic Agent Skills, package reusable procedural, reusable procedural know-how, Agent Skills package
备注: Accepted at AACL-IJCNLP 2026. Code at [this https URL](https://github.com/guanqun-yang/SkillSeek)
点击查看摘要
Abstract:Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into this http URL directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
20. 【2609.38807】PatchHolmes: Agentic Patch Retrieval via Listwise Selection
链接:https://arxiv.org/abs/2609.38807
作者:Guanqun Yang,Yingming Zhou,Jiangrui Zheng,Shudong Hao,Xueqing Liu
类目:Information Retrieval (cs.IR); Software Engineering (cs.SE)
关键词:vulnerability management workflows, major advisory databases, advisory databases lack, patch retrieval system, two-phase patch retrieval
备注: Accepted at AACL-IJCNLP 2026. Code at [this https URL](https://github.com/Aizhouym/PatchHolmes)
点击查看摘要
Abstract:Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pairs a hybrid first-stage retriever with an agentic second-stage inspection loop. Unlike pointwise prior work that scores each candidate independently, the Phase 2 agent reads the top-100 listwise: it sees the full candidate list at once and selectively reads 3 to 10 commits through four budgeted tools before submitting a single best commit. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and the retrieve-and-CoT baseline IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten; with the candidate set held identical, the agent adds 27.32% Recall@1 over taking the retriever's top candidate, and the same agent, transferred unchanged to PatchFinder_top10, lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%. Swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, and a second model family (gpt-oss) stays far above the no-agent floor, so the gain comes from the listwise agent loop; the entire system runs on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs.
21. 【2609.38743】Learning to Route in Visual Space via Multi-Step Embedding Retrieval
链接:https://arxiv.org/abs/2609.38743
作者:Tianyu Chen,Mingyuan Zhou,Jiaxing Wu
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:access external knowledge, remains severely bottlenecked, search remains severely, LLM agents rely, external knowledge
备注:
点击查看摘要
Abstract:LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5\% to 76.3\%. In agentic search, it improves task success rates by 52.7\% and reduces the average token length by 61\% from 1886 to 728, whereas upgrading the agent yields only a 3.7\% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by $23\times$ and cutting the cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.
22. 【2609.38646】Exploring Forum Post Retrieval with Generative Modeling
链接:https://arxiv.org/abs/2609.38646
作者:Yang Li,Yaguang Liu,Heng Liu,Shengbo Guo,Samson Komo,Jane Kou,Yulian Zhou,Gang Yang,Shubhojeet Sarkar,Gaurav Chakravorty,Yujie Liu,Haipeng Chen,Yonghuan Yang,Deepti Chheda,Yamin Wang,Mike Plumpe,Rish Tandon
类目:Information Retrieval (cs.IR)
关键词:Facebook Groups, embedding-based retrieval, alternative to embedding-based, Facebook Groups engagements, success of generative
备注:
点击查看摘要
Abstract:Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Forum is a new surface, its own interaction data are too sparse to train a GR model from scratch. We address this with transfer along two axes: we train on a broader corpus of Facebook Groups engagements rather than Forum sessions alone, and we reuse hierarchical, prefix-based semantic IDs (SIDs) learned from cross-platform Facebook Feed data instead of fitting a Forum-specific tokenizer. A 3B-parameter instruction-tuned language model is then supervised-fine-tuned to generate SIDs directly from user context. We systematically ablate the design choices that matter most in practice, including SID construction, the composition and length of user history, and the inclusion of user-profile features. Our results show that cross-platform SIDs transfer to a new recommendation surface, and offer practical guidance for teams deploying GR on real-world social platforms.
23. 【2609.38639】Component-Aware Feedback for Self-Evolving Programs
链接:https://arxiv.org/abs/2609.38639
作者:Ethan Lin,Jinming Nian,Yi Fang
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:LLM-guided evolutionary search, discover complex programs, save candidate programs, component edits produced, fitness scores
备注:
点击查看摘要
Abstract:LLM-guided evolutionary search can discover complex programs, but existing methods mostly only save candidate programs and fitness scores while discarding which component edits produced which fitness metric changes. Existing methods force the mutator LLM to infer the effect of prior edits from cluttered histories, making program search slow and unstable. This is especially true for locally servable LLMs to evolve multi-component systems. We introduce component-aware feedback, which compares each evaluated program with its parent, identifies the components that changed, and logs them with the associated metric differences into an attribution memory that later mutations read. The memory keeps each change in two reference frames, local against the parent it came from and global against the seed program, which shows both the immediate effect of a change and the cumulative progress made since the seed. We study this on LLM reranking, a multi-objective optimization problem where a multi-stage pipeline must balance quality against serving cost. Across twelve \textsc{Bright} datasets, our method reaches the strongest baseline's final quality after a median of one third of the search budget and ends 7.2\% higher in held-out nDCG@10, and under a cost-aware objective it finds pipelines that are on average more accurate while using 11\% fewer tokens per query, showing component-aware feedback to be a promising direction for more efficient self-evolving systems.
24. 【2609.38473】Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering
链接:https://arxiv.org/abs/2609.38473
作者:Bhagyesh Rathi,Eshan Chawla,William B. Andreopoulos
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Large Language Models, ground Large Language, Language Models, Large Language, Reciprocal Rank Fusion
备注: on September 21st submitted for consideration to the Elsevier Data and Information Management (DIM) journal (DIM-D-26-00430)
点击查看摘要
Abstract:Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is large and the trade-offs between variants are not well understood, especially on domain-specific corpora at realistic scale. In this work, we present a controlled comparison of six retrieval strategies for scientific question answering: (i) classic top-k dense retrieval, (ii) LLM-based query rephrasing, (iii) query rephrasing followed by LLM-based reranking, (iv) multi-query fusion via Reciprocal Rank Fusion (RRF), (v) an agentic tool-call pipeline in which the generator decides for itself whether to retrieve, and (vi) late-interaction retrieval with ColBERTv2. All six pipelines share the same generator (Meta-Llama/Llama-3.1-8B-Instruct), prompt, and evaluation protocol; the five single-vector pipelines additionally share SPECTER2 embeddings and a Chroma vector store; and all six retrieve from the full corpus of 463,971 arXiv papers dated 2024-2025. To support reproducible, large-scale evaluation, we also release a synthetic question dataset of 19,484 problem-statement and methodology questions generated by Llama-3.1-8B-Instruct from a random sample of 10,000 papers across academic domains (query generation succeeded for 9,742 of them), and every strategy is evaluated on this same query set. We describe the architecture and implementation of each pipeline, release the code and the synthetic question dataset, and evaluate each strategy with an LLM-as-a-judge protocol along multiple quality dimensions, together with direct gold-paper retrieval metrics. The result is an open testbed for studying the cost and quality trade-offs of RAG design choices on a research-literature corpus, and a basis for future work on faithfulness, retrieval robustness, and agentic retrieval.
25. 【2609.38455】AdaM-Rec: Adaptive Modality Routing for Multimodal Recommendation
链接:https://arxiv.org/abs/2609.38455
作者:Honghao Fu,Jiacheng Chen,Manxi Lin,Junjun Zheng,Xiangheng Kong,Yiwei Wang,Xin Yu,Miao Xu,Yuning Jiang,Yujun Cai
类目:Information Retrieval (cs.IR)
关键词:existing methods rely, signals remains stable, visual signals remains, recent multimodal recommender, multimodal recommender systems
备注:
点击查看摘要
Abstract:While recent multimodal recommender systems have demonstrated the effectiveness of incorporating visual and textual information to improve downstream performance, most existing methods rely on static modality fusion, assuming that the relative importance of textual and visual signals remains stable across recommendation scenarios. This design may not fully account for an important variation across recommendation requests: some queries require fine-grained visual cues, whereas others are better served by textual or functional semantics, in which case indiscriminate modality fusion brings in uninformative cues and impairs recommendation quality. To address this, we propose AdaM-Rec, an LLM-based framework for adaptive modality routing in multimodal recommendation, which enables dynamic calibration of reliance on textual and multimodal evidence for user-specific queries. Built on structured natural-language representations of items and user preferences, it estimates modality reliability using proxy recall tasks. Specifically, it generates pseudo-queries that match the granularity of the actual query while pointing to the user's positively interacted items as verifiable proxy targets, evaluating which modality yields better recall performance in analogous scenarios and optimizing the routing strategy in an agentic manner. It then performs routed recall with optimized strategy, enriches results with collaborative items, and ranks candidates by their relevance to both the query and user preferences. Experiments demonstrate that AdaM-Rec delivers strong performance against state-of-the-art baselines, highlighting the effectiveness and broader potential of adaptive control over modality reliance in multimodal recommendation.
26. 【2609.38374】Doc2LoRA Provides Decodable Representations of Scientific Ideas
链接:https://arxiv.org/abs/2609.38374
作者:Chand Sahil Mansuri,Joel Zachariah,Sadamori Kojaku
类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR); Machine Learning (cs.LG); Physics and Society (physics.soc-ph)
关键词:Representing scientific papers, drive innovation, papers, Representing scientific, LLM
备注: 32 pages, 4 figures, 12 tables. Code: [this https URL](https://github.com/skojaku/doc2lora-embedding)
点击查看摘要
Abstract:Representing scientific papers as points in a space lets us search for similar papers and inquire about how fields relate to one another and drive innovation. Beyond search, the vector space of papers invites generation: mixing papers through simple vector operations creates new points, mirroring combinatorial novelty, the recombination of existing ideas into new ones. However, a mixed point often represents an idea no paper has yet realized, with no papers nearby to identify the idea. We propose representing each paper by a LoRA adapter generated by the Doc-to-LoRA hypernetwork. Every point in the space, including mixtures, thus represents a large language model (LLM) open to questions and instructions in natural language. On papers from the American Physical Society (APS), we instruct the LLM at the average of each subfield to name the field in a few words and obtain labels closer to the official names than the labels of five baselines, as judged by word overlap and a panel of five LLM judges. We also ask the LLMs at points between two APS papers to write an abstract and obtain descriptions shifting from one paper to the other in step with the mixing weight. While Doc-to-LoRA is trained for generation, a small invertible transform makes the embeddings competitive for search, on par with SPECTER2 and EmbeddingGemma and close to SBERT. Because the transform is invertible, every point in the transformed space still maps back to an LLM. The embeddings thus serve both search and generation, enabling researchers to question the idea at any point in the space as a starting point for generating new ideas.
27. 【2609.38353】AGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories
链接:https://arxiv.org/abs/2609.38353
作者:Yu-Su Chen,Yu-Jung Liang,Pengtao Xie
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:LLM agents recall, agents recall past, recall past interactions, LLM agents, Long-term memory
备注: An earlier version was accepted at the COLM 2026 Workshop on Lifelong Learning Agents (LLA)
点击查看摘要
Abstract:Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.
28. 【2609.38289】Privacy in Personalized AI Is a System Property, Not Just a Model Property
链接:https://arxiv.org/abs/2609.38289
作者:Guillaume Salha-Galvan,Jiaying Xu
类目:Cryptography and Security (cs.CR); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:reuse user information, conversational assistants, assistants and recommender, user information, users interact
备注: NeurIPS 2026 Workshop on Privacy in the Era of Large Opaque Models
点击查看摘要
Abstract:In personalized AI applications, such as conversational assistants and recommender systems, users interact not with models in isolation but with broader systems that access, infer, and reuse user information across components and over time. While such use of user information is integral to personalization, it also raises important privacy questions. In this paper, we argue that individual model- or component-level analyses may not capture all privacy risks arising in such systems, motivating a system-level perspective on privacy. We distinguish and analyze four interconnected privacy-risk channels in personalized AI, and subsequently propose four requirements for system-level privacy evaluation, covering interaction trajectories, internal information flows, indirect leakage, and the privacy-utility trade-off. We argue for their systematic incorporation into privacy audits of personalized AI.
29. 【2608.03756】LegalPincite: Multi-level Legal Information Retrieval Dataset
链接:https://arxiv.org/abs/2608.03756
作者:Theresia Veronika Rampisela,Henrik Palmer Olsen,Giovanni Colavizza
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:find relevant legal, relevant legal sources, case-law collections, common task, find relevant
备注: Accepted for publication at the 8th Natural Legal Language Processing Workshop (NLLP 2026), co-located with EMNLP 2026
点击查看摘要
Abstract:A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset and code: this https URL
30. 【2602.19001】Life-Bench: A Benchmark and Knowledge Graph Framework for Multimodal Personalization Beyond Concept Recognition
链接:https://arxiv.org/abs/2602.19001
作者:Xia Hu,Honglei Zhuang,Brian Potetz,Alireza Fathi,Bo Hu,Babak Samari,Howard Zhou
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:large language models, language models increasingly, models increasingly power, primarily target concept-level, power personal assistants
备注:
点击查看摘要
Abstract:As large language models increasingly power personal assistants, users expect them to reason over multimodal life histories, from recognizing people to understanding events to aggregating patterns, yet existing benchmarks primarily target concept-level recognition. We introduce Life-Bench, a fully synthetic, human-verified multimodal benchmark of over 11,800 question-answer pairs across 10 tasks, organized by required evidence scope: concept identification, event understanding, and aggregated reasoning. The benchmark's photo-centric personal histories are distributionally aligned with real user accounts under embedding statistic. The interconnected structure of personal data invites graph-based solutions; we propose LifeGraph, a personal knowledge graph framework providing structured retrieval with on-demand access to source visual evidence, showing particular promise on event and aggregated tasks. Systematic evaluation of four retrieval paradigms on Life-Bench demonstrates that accuracy degrades sharply with evidence scope, falling below 0.40 on aggregated tasks, and that no single paradigm dominates across categories. Performance beyond concept recognition remains modest for all evaluated methods, establishing personalization over multimodal histories as an open challenge and Life-Bench as a testbed for future progress.
计算机视觉
1. 【2609.40362】Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
链接:https://arxiv.org/abs/2609.40362
作者:Hongyuan Tao,Xinggang Wang,Lianghui Zhu,Yongkang Li,Yunchao Wei,Bin Feng,Shaoyu Chen,Qian Zhang,Chang Huang,Kai Yu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:present Multimodal Flow, Multimodal, Multimodal Flow, Flow, continuous
备注: 18 pages, 5 figures, 10 tables. Code and model: [this https URL](https://github.com/hustvl/Multimodal-Flow)
点击查看摘要
Abstract:We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The latter requires modality-dependent objectives and sampling procedures. Fully continuous modeling avoids these trade-offs and enables a shared generative process, but remains underexplored for multimodal pretraining. Multimodal Flow introduces a unified continuous architecture that integrates multimodal continuous representations with a shared chunk-causal flow backbone. It organizes text blocks and images as ordered continuous hyperchunks, preserving textual token order and visual spatial structure. The backbone learns a single vector field over these hyperchunks through Flow Matching. Joint attention enables cross-modal interaction, while modality-specific feed-forward networks process each modality. The model predicts multiple target chunks in parallel during training and generates hyperchunks sequentially at inference. We instantiate MF-1 and pretrain it on multimodal data. Across 0.6B, 1.2B, and 1.6B scales, continued pretraining consistently improves multimodal modeling. With only 150B pretraining tokens, MF-1 achieves an average score of 82.8 across GenEval and DPG-Bench and 75.3 across VQAv2, MMBench, and POPE, remaining competitive with unified models trained on substantially more data. Under matched data, optimization, and parameter budgets, Multimodal Flow further outperforms representative hybrid and discrete models. These results establish continuous chunk-based embedding flow modeling as a new fully continuous paradigm for unified multimodal modeling. The related code and model are publicly released at this https URL.
2. 【2609.40361】Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
链接:https://arxiv.org/abs/2609.40361
作者:Tian Xia,Minghao Liu,Yiqing Liang,Laixi Shi,Jiayun Wang
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:adaptation pipelines remain, pipelines remain anchored, advancing clinical diagnosis, rapidly advancing clinical, Multimodal large language
备注:
点击查看摘要
Abstract:Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.
3. 【2609.40358】Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model
链接:https://arxiv.org/abs/2609.40358
作者:Liming Lu,Xianzheng Ma,Wenkun He,Guanqi Zhan,Yilin Zhao,Junyu Chen,Mengyao Xu,Jiaojiao Fan,Wenhang Ge,Yuchao Gu,Yunze Liu,Boyi Li,Zhen Dong,Victor Prisacariu,Ming-Yu Liu,Song Han,Han Cai
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:physical world evolves, violate basic physical, world evolves, expected to predict, violate basic
备注:
点击查看摘要
Abstract:Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.
4. 【2609.40356】ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
链接:https://arxiv.org/abs/2609.40356
作者:Xinghao Chen,Xiangbo Gao,Jiongze Yu,Yuheng Wu,Zhengzhong Tu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Recent video generation, scene text editing, Video scene text, original scene dynamics, precise local edits
备注: Accepted to NeurIPS 2026 (Evaluations and Datasets Track). 27 pages (10-page main text), 5 figures, 12 tables. Project page: [this https URL](https://vitex-bench.github.io/)
点击查看摘要
Abstract:Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video scene text editing that achieves high visual quality, temporal consistency, and edit locality remains underexplored. Existing resources offer limited paired real-video data, and general video-editing metrics do not directly measure whether the requested text remains correct over time. We introduce ViTeX-Bench, a benchmark suite comprising ViTeX-Dataset and a three-axis evaluation protocol. The dataset contains 387 real-world 720p videos with text-region masks and editing instructions: 230 provide reviewed, pipeline-generated paired edits for training, and 157 form a frozen evaluation split. The protocol evaluates text correctness, visual and temporal quality, and edit locality through 13 metrics, with one primary metric per axis and a Pareto comparison of their trade-offs. OCR calibration, human evaluation, and annotation-sensitivity analyses support the interpretation of these scores. Across eight baselines from four editing families, accurate text, temporal stability, and scene preservation remain difficult to achieve together. We also release ViTeX-Edit-14B, an open-source reference editor fine-tuned on the paired training split with motion-aligned glyph-video conditioning. It achieves CharAcc 0.688, the highest mean among the evaluated video-native editors, and the lowest comparable text-crop Warp among raw editor outputs. ViTeX-Bench provides a reproducible foundation for studying these trade-offs in video scene text editing.
5. 【2609.40353】AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents
链接:https://arxiv.org/abs/2609.40353
作者:Jiahao Zhang,Yeying Fan,Moitreya Chatterjee,Suhas Lohit,Bernhard Egger,Tim K. Marks,Anoop Cherian,Stephen Gould
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:precise spatial arrangements, assembly requires translating, spatial arrangements, requires translating, translating an understanding
备注: 24 pages, 11 figures. Project page: [this https URL](https://assemblyworld.github.io)
点击查看摘要
Abstract:The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fracture reassembly. Evaluating eight agent systems reveals substantial differences in their capabilities. The strongest system achieves 80.9% part accuracy but 59.4% complete-assembly success. The evaluated open-source systems lag substantially behind their stronger closed-source peers in both execution reliability and assembly accuracy. Analyses of visual references, interaction trajectories, and failures show how agents revise assemblies while leaving residual positioning errors. AssemblyWorld provides a common setting for both assessing the capabilities of interactive assembly agents and characterizing the gap between approximate structure recovery and precise reconstruction.
6. 【2609.40347】Image Classifiers are Efficient Self-Supervised Video Representation Learners
链接:https://arxiv.org/abs/2609.40347
作者:Owais Iqbal,Sudipta Sarkar,Shyam Marjit,Omprakash Chakraborty,Anirban Chakraborty,Abir Das
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Siamese Network framework, Masked Siamese Network, Network framework, Siamese Network, Masked Siamese
备注: Accepted in BMVC 2026
点击查看摘要
Abstract:We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to $32\times$ fewer and $160\times$ fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: this https URL.
7. 【2609.40341】Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?
链接:https://arxiv.org/abs/2609.40341
作者:Zhihao Sun,Liu Liu,Xinjiang Wang,Haoyi Jiang,Wei Feng,Huiqiang Zhang,Xiaosong Jia,Zhizhong Su,Zuxuan Wu
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Egocentric human data, Egocentric human, human data, behavioral coverage, data
备注:
点击查看摘要
Abstract:Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline. We present a systematic study of egocentric human data with different alignment and supervision under a unified world-action model framework. With the model backbone fixed, we disentangle the effects of human-robot alignment, data duration and task diversity, action supervision, and data usage strategies. We find that aligned human demonstrations substantially improve out-of-distribution generalization and reduce target-task robot data requirements; data duration and task diversity affect downstream capabilities differently; and video-only supervision remains effective without action labels, providing a strong foundation for subsequent video-action training. We validate these findings through closed-loop policy evaluation on both real robots and RoboDojo. Rather than treating data duration as the sole scaling axis, Ego4WAM shows how alignment, task diversity, available supervision, and usage strategy jointly shape the value of egocentric human data for robot learning.
8. 【2609.40333】I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
链接:https://arxiv.org/abs/2609.40333
作者:Ivan Martinović,Lukas Knobel,Yuki M. Asano
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:infant visual development, Self-supervised learning draws, learning draws inspiration, Self-supervised learning, training pipelines bear
备注: Preprint. Accepted to NeurIPS 2026
点击查看摘要
Abstract:Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.
9. 【2609.40322】MatLoom: Layered Text-to-Material Generation in a Compact Program Space
链接:https://arxiv.org/abs/2609.40322
作者:Anson Y. Lam,Shuqing Li,Michael R. Lyu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
关键词:rules that construct, Abstract, pretrained language models, generation should produce, generation
备注: 27 pages, 8 figures
点击查看摘要
Abstract:Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.
10. 【2609.40320】Atomizer-IO: Beyond Pixels, Patches and Grids
链接:https://arxiv.org/abs/2609.40320
作者:Hugo Riffaud de Turckheim,Sylvain Lobry,Nicolas Houdré,Damien Robert,Roberto Interdonato,Diego Marcos
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:vision architectures assume, temporal sampling, geometry can vary, effective abstraction, abstraction for natural
备注:
点击查看摘要
Abstract:Most vision architectures assume that observations lie on a regular grid, an effective abstraction for natural images but a restrictive one for sensing data whose channels, temporal sampling, spatial resolution, and geometry can vary. Generic set-based architectures remove the grid, but also remove useful spatial inductive biases. We introduce Atomizer-IO, an architecture that places observations first and derives structure from their physical relationships. Building on top of an atomic representation of the data, each observation is described by its measurement and acquisition metadata, while local cross-attention maps observations to anchor points that can be arbitrarily placed. We evaluate this design by progressively relaxing the grid assumption, from varying input raster configurations and incomplete channel sets to flexible output density and, ultimately, inputs without a raster grid. Atomizer-IO is competitive with flexible EO-specific architectures on most tasks, while offering post-training control over inference cost and competitive compute--performance trade-offs. The same formulation extends without architectural redesign to unordered 3D point clouds, showing that the atomic interface generalizes beyond regular raster inputs. These results suggest that pixels, patches, and grids do not need to define the interface of a sensing architecture.
11. 【2609.40317】GLARE: Generating Listening Heads with Appropriate Reactions
链接:https://arxiv.org/abs/2609.40317
作者:Zikai Liao,Yumin Suh,Yi Ouyang,Yi-Lun Lee,Yi-Hsuan Tsai,Zhaozheng Yin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:talking head generation, generating natural listener, remains underexplored, advanced rapidly, type of response
备注: Accepted in NeurIPS 2026. Project page: [this https URL](https://github.com/lzk901372/glare)
点击查看摘要
Abstract:While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.
12. 【2609.40305】Looped Diffusion Transformer
链接:https://arxiv.org/abs/2609.40305
作者:Yong Xien Chng,Tianyi Chen,Wenwen Tong,Haiwen Diao,Zhongang Cai,Lei Yang,Ziwei Liu,Lewei Lu,Dahua Lin,Gao Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Improving, traditionally relied, denoising steps, looped, denoising
备注: 21 pages, 9 figures
点击查看摘要
Abstract:Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.
13. 【2609.40253】ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
链接:https://arxiv.org/abs/2609.40253
作者:Yong Du,Tongbo Chen,Zhengxi Lu,Yizhou Liu,Bofan Chen,Tao Jiang,Wenhao Xu,Yongliang Shen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:enables computer-use agents, training enables computer-use, computer-use agents, executable environments, enables computer-use
备注: [this https URL](https://github.com/ZJU-REAL/ComputerSD)
点击查看摘要
Abstract:Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
14. 【2609.40244】StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry
链接:https://arxiv.org/abs/2609.40244
作者:Yufei Wei,Shuhao Ye,Qi Wang,Xin Zheng,Qing Huang,Rong Xiong,Yue Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:Mobile robots, carry synchronized multi-camera, vehicles carry synchronized, leaving efficient, geometry a challenge
备注: 8 pages, 4 figures, 5 tables. Code: [this https URL](https://github.com/WeiYuFei0217/StreamRig)
点击查看摘要
Abstract:Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model. The frozen front-end jointly perceives the synchronized views using rig calibration. A Rig-Resampler compresses their features, a CausalBridge applies causal attention with a key-value cache, and a lightweight head regresses rig poses. A periodic re-anchoring protocol supports stable pose estimation over long sequences. Only these modules are trained, 74.6M parameters in total, with relative poses as the sole supervision. Our two-stage training strategy combines group relocalization pretraining with causal rig training to transfer the geometric priors of the frozen front-end and the alignment ability of the pretrained modules to streaming odometry. We evaluate on NCLT, TartanGround, KITTI-360, and our self-collected humanoid-robot dataset ZJH, where training uses only simulation and real-world evaluation is zero-shot. Across all four datasets, StreamRig achieves lower translation and rotation drift than the evaluated non-oracle monocular streaming and rig-aware offline models, while maintaining low inference cost. Ablations and controlled camera-count experiments identify the sources of these gains. We further examine how longer training windows affect inference over longer horizons. Code has been released at this https URL.
15. 【2609.40230】EviRover: Reinforcing Agentic Perception Beyond a Glance
链接:https://arxiv.org/abs/2609.40230
作者:Kaixuan Fan,Kaituo Feng,Tianshuo Peng,Yilei Jiang,Manyuan Zhang,Junke Wang,Xiangyu Yue
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:parametric knowledge suffice, image content, conventionally formulated, one-shot prediction, single glance
备注:
点击查看摘要
Abstract:Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.
16. 【2609.40222】LOCI: Spatial Linear Memory for Streaming World Models
链接:https://arxiv.org/abs/2609.40222
作者:Ji Xia,Tingting Liao,Xuezhi Liang,Hao Li,Guangyi Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:previously observed region, observed region, revisits a previously, previously observed, past observations
备注: 25 pages, 8 figures, 14 tables. Project page: [this https URL](https://xiaji2021.github.io/LOCI/)
点击查看摘要
Abstract:When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
17. 【2609.40219】Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
链接:https://arxiv.org/abs/2609.40219
作者:Qi Lyu,Jiahua Dong,Hao Shen,Xudong Wang,Hongyuan Yu,Baichen Liu,Henghui Ding,Zhi Han,Nicu Sebe,Ivan Laptev,Fahad Shahbaz Khan,Salman Khan
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:World Action Models, couple visual dynamics, World Action, Action Experience Dictionary, Action
备注:
点击查看摘要
Abstract:World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{this https URL}{AED}.
18. 【2609.40212】Recognition of Urbanized Areas in UAV-Derived Very-High-Resolution Visible-Light Imagery
链接:https://arxiv.org/abs/2609.40212
作者:Edyta Puniach,Wojciech Gruszczyński,Paweł Ćwiąkała,Katarzyna Strząbała,Elżbieta Pastucha
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:acquired RGB imagery, unmanned aerial vehicle, acquired RGB, RGB imagery, study compared classifiers
备注:
点击查看摘要
Abstract:This study compared classifiers that differentiate between urbanized and non-urbanized areas based on unmanned aerial vehicle (UAV)-acquired RGB imagery. The tested solutions in-cluded numerous vegetation indices (VIs) thresholding and neural networks (NNs). The analysis was conducted for two study areas for which surveys were carried out using different UAVs and cameras. The ground sampling distances for the study areas were 10 mm and 15 mm, respectively. Reference classification was performed manually, obtaining approximately 24 million classified pix-els for the first area and approximately 3.8 million for the second. This research study included an analysis of the impact of the season on the threshold values for the tested VIs and the impact of image patch size provided as inputs for the NNs on classification accuracy. The results of the con-ducted research study indicate a higher classification accuracy using NNs (about 96%) compared with the best of the tested VIs, i.e., Excess Blue (about 87%). Due to the highly imbalanced nature of the used datasets (non-urbanized areas constitute approximately 87% of the total datasets), the Mat-thews correlation coefficient was also used to assess the correctness of the classification. The analysis based on statistical measures was supplemented with a qualitative assessment of the classification results, which allowed the identification of the most important sources of differences in classification between VIs thresholding and NNs.
19. 【2609.40195】MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
链接:https://arxiv.org/abs/2609.40195
作者:Guangzhi Xiong,Xinyuan Zhang,Xiao Yang,Hyokun Yun,Kai Zhang,Shiun-Zu Kuo,Hyeonjeong Ha,Xilun Chen,Kai Sun,Lucas Liang,Guangqiang Dong,Ejaz Ahmed,Ahmed A Aly,Anuj Kumar,Raffay Hamid,Aidong Zhang,Xin Luna Dong
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Long-term egocentric video, Long-term egocentric, egocentric video enables, video enables personalized, daily life
备注:
点击查看摘要
Abstract:Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.
20. 【2609.40131】Prototype-Rule Neurosymbolic Regularization for Rank-Constrained Tensor Neural Networks under Label Scarcity
链接:https://arxiv.org/abs/2609.40131
作者:Eftychios Protopapadakis,Konstantinos Makantasis,Konstantinos M. Giannoutakis
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:Rank-constrained tensor neural, neural networks reduce, constrain class geometry, explicitly constrain class, Rank-constrained tensor
备注:
点击查看摘要
Abstract:Rank-constrained tensor neural networks reduce the parameterization of high-order inputs, but they do not explicitly constrain class geometry in the learned representation. This study investigates whether a differentiable prototype-rule can provide a complementary inductive bias for Rank-R tensor learning under limited supervision. The proposed framework augments the Rank-R objective with prototype-based regularization and optionally fuses prototype evidence with neural logits at inference. Four hyperspectral benchmarks are evaluated with four Rank-R configurations under both seven-fold stratification and spatially separated folds that mitigate leakage; a separate spatial study varies the class support budget from 2 to 20 samples. Under spatial evaluation, full neurosymbolic inference changes Macro-F1 score by +8.82 percentage points on Botswana, +5.49 on Indian Pines, +1.59 on Pavia University, and -0.62 on Salinas. Most of the benefit arises from training-time regularization, whereas inference fusion is small and dataset dependent.
21. 【2609.40129】VR-JEPA: Learning Contrastive-State Latent Guidance for Generation-based Video Reasoning
链接:https://arxiv.org/abs/2609.40129
作者:Zehua Ma,Kun Xiang,Yunshuang Nie,Quanlin Chen,Haoyuan Li,Xiuwei Chen,Jiang Ji,Haijun Wu,Zhenyu Xie,Michael Kampffmeyer,Hanhui Li,Xiaodan Liang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:video generation offers, video generation, offers a promising, promising path, intelligence by modeling
备注:
点击查看摘要
Abstract:Reasoning through video generation offers a promising path toward visual intelligence by modeling latent visual states and their dynamics. However, current video generation models often lack explicit guidance on how these states should evolve, leaving generated trajectories prone to physical and structural inconsistencies that undermine reasoning reliability. While the Video Joint-Embedding Predictive Architecture (V-JEPA) provides rich spatiotemporal priors learned through latent prediction, these general priors do not naturally adapt to the logical reasoning capabilities required for complex visual tasks. To bridge this gap, we propose VR-JEPA, a framework that aligns the V-JEPA predictor with task-specific reasoning logic through localized contrastive-state learning and uses its predicted latent trajectories to guide video generation for visual reasoning. Specifically, (i) we pair successful trajectories with generated alternatives under the same input conditions and use discrepancies in their V-JEPA representations to identify informative states and tokens for localized contrastive supervision. (ii) We further equip the V-JEPA predictor with skill-specific experts trained on anchor-task data, allowing the model to adaptively specialize its shared spatiotemporal priors across diverse cognitive domains. Together with skill-specific experts, this contrastive supervision enables VR-JEPA to predict latent trajectories that provide task-specific logical guidance for video generation. Comprehensive experiments on the large-scale VBVR-Pro-Bench dataset demonstrate that VR-JEPA achieves an $11.33\%$ relative improvement over the cutting-edge generation-based reasoning baseline, significantly mitigating physical artifacts and enhancing logical consistency.
22. 【2609.40091】GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation
链接:https://arxiv.org/abs/2609.40091
作者:Hoang Nguyen Van,Cuong Vuong Tuan,Trang Mai Xuan,Bien Tran Van,Nam Tran Van,Thien Van Luong
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:multi-sequence MRI studies, interpreting multi-sequence MRI, burden radiolo gists, radiolo gists face, Automated report generation
备注:
点击查看摘要
Abstract:Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical efficacy metrics, where a missed abnormality is most costly. We propose GateSPINE, a vision-language framework that fuses sagittal T1 and T2 volumes with a training-free operator, encodes the fused sagittal and axial volumes with two parallel 3D encoders, and decodes their combined representation into a report. Its core mechanism is a gated cross view fusion module that predicts, per feature channel and token, how much of each view to admit, so the more informative view dominates at each spatial location. We evaluate GateSPINE on three lumbar MRI datasets, comprising two public bench marks and a private cohort collected from Phenikaa University Hospital, using both natural language generation (NLG) and clinical efficacy (CE) metrics. GateSPINE achieves the highest CE F1 through improved re call on all three datasets; on SPIDER, which lacks an axial sequence, this reflects the sagittal fusion component rather than the gated cross-view mechanism, which is validated on the two cohorts with both imaging planes. GateSPINE also remains competitive on standard NLG metrics.
23. 【2609.40079】LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
链接:https://arxiv.org/abs/2609.40079
作者:Shuo Zhang,Yifan Zhou,Han Wang,Jinsong Zhang,Jingyu Li,Hongbing Li,Zhejun Zhang,Chengyi Zhao,Yuquan Hao,Yitong Liu,Jiyin Li,Ruiqi Tang,Zixuan Lin,Yi Luo,Xurui Zhang,Ronghao Chen,Huacan Wang,Lei Li
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Multimodal Large Language, Language Models, Large Language, recent Multimodal Large
备注: 33 pages
点击查看摘要
Abstract:While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.
24. 【2609.40055】Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding
链接:https://arxiv.org/abs/2609.40055
作者:Jiacheng Qiu,Yunsoo Kim,Ruichen Xu,Jian Luo,Petar M. Djurić,Sima Mofakham
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:On-policy distillation, dense supervision directly, effective post-training strategy, student-generated trajectories, directly on student-generated
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that supervision trustworthiness and supervision necessity are distinct yet coupled: the former concerns target credibility, while the latter varies with the student's current task competence; together, they shape supervision value. Building on this coupled view, we introduce Student-Curriculum Coupling (SCC), a closed-loop framework in which a compact Anchor-Frontier curriculum defines the candidate supervision space and the evolving student dynamically determines its active subset. Supervision can therefore be activated, suspended, or reactivated as competence changes, concentrating teacher computation and optimization on current task-level deficits. Across three TVG benchmarks, SCC achieves a 5.1% relative improvement in mean recall over Video-OPD on its original curriculum, while using 60.0% fewer training examples and reducing training time by 50.4%. Ablations support the complementary roles of capability-structured curriculum design and student-dependent supervision in achieving these gains. Together, these results establish SCC as a data- and compute-efficient framework for TVG post-training, delivering stronger temporal grounding by aligning trustworthy supervision with the student's evolving learning needs.
25. 【2609.40048】CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
链接:https://arxiv.org/abs/2609.40048
作者:Yiduo Jia,Muzhi Zhu,Jinchuan Shi,Hao Zhong,Yuling Xi,Ke Liu,Hao Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:grounding requires balancing, requires balancing long-range, limited visual budget, temporal grounding requires, existing agentic methods
备注: Project page: [this https URL](https://aim-uofa.github.io/CoEvoWhen/)
点击查看摘要
Abstract:Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
26. 【2609.40037】Enhancing Autoregressive Video Generation via Representation Adversarial Distillation
链接:https://arxiv.org/abs/2609.40037
作者:Fangyu Lin,Xingtong Ge,Lunjie Zhu,Yi Zhang,Zhening Liu,Tianhang Wang,Mengfei Li,Yumeng Zhang,Guanglu Song,Yu Liu,Jun Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:efficient streaming synthesis, enables efficient streaming, early temporal blocks, Few-step autoregressive video, generation enables efficient
备注:
点击查看摘要
Abstract:Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. Existing distribution matching distillation (DMD) primarily aligns student and teacher distributions in diffusion latent space, but provides no direct supervision over the perceptual quality of decoded videos. We introduce Radian, a representation-space adversarial distillation framework that complements on-policy DMD with real-data adversarial supervision in the feature space defined by a frozen visual foundation model (VFM). During training, Radian sparsely decodes frames from autoregressive student rollouts, extracts multi-level visual representations, and applies lightweight discriminator heads to distinguish generated outputs from real video frames. The DMD objective anchors the student to the pretrained teacher, while the representation-space adversarial objective supplies complementary perceptual and semantic gradients that promote high-quality modes. These additional components are discarded after training, leaving the generator architecture and inference-time denoising budget unchanged. Experiments on Wan2.1-1.3B cover four-step chunk-wise, one-step frame-wise, and minute-long autoregressive generation. Our method achieves a VBench Total of 0.8444 and a VideoAlign Total of 0.8033 under four-step generation, and improves VBench-Long from 0.7805 to 0.8041 over Rolling Forcing while using fewer denoising steps. Controlled comparisons across image, video, and diffusion representations further indicate that the choice of representation spaces induces distinct adversarial signals, and external VFM gradients complement DMD more effectively than adversarial supervision derived from diffusion-internal features.
27. 【2609.40031】WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks
链接:https://arxiv.org/abs/2609.40031
作者:Khaled Abud,Aleksey Yakushev,Aleksandr Akimenkov,Irina Serzhenko,Kirill Aistov,Egor Kovalev,Dmitry Obydenkov,Sergey Lavrushkin,Anastasia Antsiferova,Dmitriy Vatolin,Yury Markin,Kirill Lukianov
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
关键词:Digital image watermarking, require marking AI-generated, marking AI-generated content, ensuring traceable sources, industry practices require
备注: Accepted to ACM MM 2026 (Main Track)
点击查看摘要
Abstract:Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. Recent advances in invisible watermarking methods highlight the need to update existing benchmarking practices to reflect current techniques and evaluation criteria. We address this by introducing WARP -- a unified framework and benchmark for evaluating the robustness of invisible watermarks. WARP incorporates 32 recent classical, deep, and generative watermarking methods, as well as 34 different erasing techniques, ranging from traditional distortions to more sophisticated adversarial, purification, and re-embedding attacks. It provides standardized, reproducible, and easily scalable protocols for evaluating perceptual quality, watermark readability, and attack resilience. Using WARP, we extensively evaluate current invisible watermarking techniques, collecting the largest robustness benchmark in the field. Results identify the most robust approaches under both distortion and adversarial conditions, and reveal consistent relationships between watermarking methods and the attack strategies most effective against them. Our experiments also highlight that some of the watermarking methods considered are highly vulnerable to reembedding, even if they are robust to standard distortions. The code is made available at this https URL.
Comments:
Accepted to ACM MM 2026 (Main Track)
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
Cite as:
arXiv:2609.40031 [cs.CV]
(or
arXiv:2609.40031v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2609.40031
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Related DOI:
https://doi.org/10.1145/3767308.3835888
Focus to learn more
DOI(s) linking to related resources</p>
28. 【2609.40014】Can We Anticipate Violence? Multimodal Learning from Pre-Incident Behavioral Cues
链接:https://arxiv.org/abs/2609.40014
作者:Sindhuja Penchala,Mohammed Yusuf Mujawar,Noorbakhsh Amiri Golilarz,Sudip Mittal,Shahram Rahimi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Detecting violence, begins is important, Detecting, multimodal video signals, Abstract
备注:
点击查看摘要
Abstract:Detecting violence after it begins is important from recognizing behavioral cues that appear immediately beforehand. This work studies short-horizon pre-incident risk recognition from multimodal video signals. We construct a binary Normal-versus-Risky setting from temporally annotated XD-Violence clips, using 443 samples with source-level separation across training, validation, and test sets. Each sample consists of a variable-length pre-incident clip, with its duration determined by the observable behavioral context preceding the incident. The inci- dent itself is excluded from all input clips. We evaluate three complementary information sources: facial-region appearance, temporally aligned audio, and body-motion features derived from tracked keypoints. Controlled ablations are performed with Swin-Tiny, ViT-Tiny, and DeiT-Tiny to measure the contribution of each modality under the same split. Results show that combining all modalities is more effective than using any other combination alone. The best configuration, Deit-Tiny with audio, facial appearance, and motion, achieves 91.21% accuracy, 88.96% balanced accuracy, 93.65% F1-score, and 96.38% ROC-AUC on the held-out test set. These results suggest that complementary appearance, acoustic, and kinematic cues provide useful evidence for recognizing elevated pre-incident risk.
29. 【2609.40007】Multi-Link Safety Filtering for VLA Policies Around Moving Hazards
链接:https://arxiv.org/abs/2609.40007
作者:Yatharth Agarwal,Vijay Raghunathan
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:pretrained VLA policy, VLA policy clear, VLA policy, deploy in clutter, finish a manipulation
备注: 9 pages, 4 figures, 3 tables. Project page: [this https URL](https://yathag.github.io/multilink-safety-filter/)
点击查看摘要
Abstract:A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy. Our training-free shield covers the gripper, wrist, and forearm with five ellipsoids and filters every commanded motion through one barrier program against a keep-out ellipsoid fitted from RGB-D perception at reset. Sparse optical flow then carries that ellipsoid's center along with the hazard, with no repeated detection or refitting. Over six simulated hazard-motion conditions, the shield lowers collision from $65.62\%$ to $27.27\%$ and raises safe-success, task completion without collision, from $29.35\%$ to $50.43\%$. Ablations show that guarding the arm links protects beyond end-effector shielding, and that tracking recovers most of the protection lost when the hazard estimate is frozen at reset. On heterogeneous edge hardware, the five-ellipsoid barrier runs on the CPU in $2.2$~ms at the 99th percentile, and trimming the vision--language prefix and taking fewer flow-matching steps shortens each $\pi_{0.5}$ policy call on the integrated GPU from $343$ to $177.3$~ms. On a physical SO-101 arm across four tasks, the arm touched the hazard in 3 of 16 shielded episodes versus 11 of 16 unshielded ones. Project page: this https URL
30. 【2609.39960】Reconstructing the Dynamic World: A Representation-Centric View of 4D Scene Reconstruction
链接:https://arxiv.org/abs/2609.39960
作者:Ziren Gong,Guo Chen,Yongjia Li,Yihua Shao,Fabio Tosi,Stefano Mattoccia,Matteo Poggi,Hao Tang,Fei Ma,Shuyan Li,Ziyang Yan,Nicu Sebe,Ling Shao,Jianfei Cai,Qi Tian,Ming-Hsuan Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:scene reconstruction aims, reconstructing dynamic scenes, Neural Radiance Fields, neural scene representations, visual observations
备注:
点击查看摘要
Abstract:4D scene reconstruction aims to recover the evolving geometry, appearance, and motion of dynamic environments from visual observations. Despite substantial progress in neural scene representations, reconstructing dynamic scenes remains challenging due to non-rigid motion, occlusions, temporal inconsistencies, and the trade-offs between reconstruction fidelity and computational efficiency. Recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have introduced diverse approaches to representing and reconstructing dynamic scenes, yet their relationships, underlying design choices, and evaluation protocols remain fragmented. In this paper, we present a unified perspective on 4D scene reconstruction, organizing existing methods around their scene representations, temporal modeling strategies, reconstruction pipelines, and optimization objectives. Through this framework, we examine how different design choices affect geometric fidelity, appearance consistency, motion representation, and computational efficiency. We further consolidate commonly used datasets and evaluation metrics, identify limitations in current experimental practices, and discuss open challenges in reconstructing complex, dynamic real-world environments. By connecting methodological developments with their underlying assumptions and evaluation evidence, this work provides a structured foundation for understanding existing approaches and identifying future research directions. An evolving collection of relevant papers and resources is available at this https URL.
31. 【2609.39953】Learning to Reason with Compressed Context: Ground-Truth-Free Adaptation of OmniLLMs via Self-Distillation
链接:https://arxiv.org/abs/2609.39953
作者:Jianghao Wang,Ke Meng,Jian Li,Chi Cheng,Longyu Qi,Liyin Liang,Yifeng Qian,Chunbo Lai,Yutian Lin,Zeyu Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Omni-modal large language, token sequences make, Omni-modal large, enable unified audio-video, unified audio-video understanding
备注: 31 pages, 5 figures. Project page: [this https URL](https://github.com/Bamboos2003/CAFD)
点击查看摘要
Abstract:Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computationally expensive. Token compression reduces this cost, yet aggressive compression often lowers accuracy. Existing works predominantly focus on designing better compression mechanisms; however, adapting the underlying language model to reason effectively over the remaining compressed context remains under-explored. To address this, we propose CAFD (Compressed-Context Adaptation via Full-Context Distillation), a ground-truth-free self-distillation framework that adapts OmniLLMs to fixed compression pipelines without requiring reference answers, rationales, or correctness rewards. CAFD leverages the full-token view of the same multimodal sample as a source of privileged information: a full-context self-teacher provides soft target supervision to a compressed-context student along the student's on-policy trajectory. Evaluated on Qwen2.5-Omni-7B across five audio-video benchmarks, five compression pipelines, and five deployment budgets, CAFD demonstrates consistent gains, improving 120 out of 125 conditions with an average accuracy boost of 1.44 points and recovering 26.9% of the accuracy gap on average. These results demonstrate that the proposed ground-truth-free adaptation offers an effective and practical route to improving the accuracy-efficiency trade-off in deployed OmniLLMs.
32. 【2609.39938】LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
链接:https://arxiv.org/abs/2609.39938
作者:Juyi Lin,Zhiqiang Lao,Jiali Cui,Lin Zhao,Pu Zhao,Dichang Zhang,Arman Akbari,Yu Qi,Xinru Jiang,Yanzhi Wang,Heather Yu,Liang Peng
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Hour-scale audio-visual question, dense whole-recording encoding, whole-recording encoding rapidly, encoding rapidly exhausts, compression severely dilutes
备注: 39 pages, 16 figures
点击查看摘要
Abstract:Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
33. 【2609.39934】Reliability-Aware Checkpoint Selection for Domain Generalization
链接:https://arxiv.org/abs/2609.39934
作者:Jinshi Liu,Jiahao Li,Pan Liu,Yanfeng Li,Rui Qian,Zhao Tong,Yue Sun,Tao Tan
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:provide reliable probabilities, provide reliable, reliable probabilities, probabilities on unseen, accuracy
备注: 28 pages, 5 figures. Project page: [this https URL](https://github.com/Jjjjjjh666/Reliability-Aware-DG)
点击查看摘要
Abstract:Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with near-optimal source accuracy can improve mean target probability quality with small observed changes in mean target accuracy. We study accuracy-constrained reliability selection (AC), which retains checkpoints within a tolerance of the best source-validation accuracy and ranks them by source reliability. Our reference rule aggregates within-set normalized negative log-likelihood (NLL) and class-wise calibration error (CwECE) using $D_\infty$. AC uses no target data and requires neither additional training nor weight averaging. We evaluate five domain generalization training algorithms on three benchmarks, using PACS to develop the objectives and a 0.5-percentage-point tolerance. In exploratory aggregation comparisons on 360 OfficeHome and TerraIncognita runs, the reference rule reduces mean target soft-bin squared-gap ECE and CwECE by 0.240% and 0.182%, respectively, and NLL by 0.030 relative to Source-Acc. Mean target accuracy changes by +0.213 percentage points. These results identify opportunities for reliability-aware reselection, while the additional benefit of joint over single-objective ranking remains unresolved.
34. 【2609.39926】Super-Resolving Unseen Hyperspectral Sensors at Any Scale via Spatial Operators
链接:https://arxiv.org/abs/2609.39926
作者:Ji-Xuan He,Guohang Zhuang,Bo Junge,Tingyi Li,Lingchen,Miaomiao Cai,Yanan Qiao,Xiujin Liu,Junfeng Fang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Achieving cross-sensor generalization, single model remains, model remains challenging, Achieving cross-sensor, arbitrary-scale reconstruction
备注:
点击查看摘要
Abstract:Achieving cross-sensor generalization and arbitrary-scale reconstruction with a single model remains challenging in hyperspectral super-resolution (HSR). Although recent methods support arbitrary-scale reconstruction, applying them to new sensors or scales beyond the training range often requires additional data and computation to maintain reconstruction quality. To address these challenges, we propose OmniHSR, which predicts band-shared spatial operators rather than spectral values. Cross-Spectral Mapping (CSM) resamples inputs with any number of bands to fixed reference positions and predicts local operators with Gaussian supports. Continuous Operator-Field Reconstruction (COFR) composes these operators into a continuous field and applies them to all original bands for arbitrary-scale reconstruction. Experiments demonstrate that operator prediction outperforms direct spectral-value prediction on all seven datasets. Trained solely on ARAD with only 0.538M parameters, OmniHSR outperforms all directly transferred baselines on six unseen datasets without target-domain training data or adaptation. Across twelve upsampling factors from $\times2$ to $\times48$, it improves average PSNR on Pavia U and Chikusei by 0.55 dB over the strongest baseline. It also surpasses baselines trained from scratch or adapted on the target sensor and achieves up to $36\times$ faster inference. Our code will be publicly released soon.
35. 【2609.39924】CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
链接:https://arxiv.org/abs/2609.39924
作者:Yulong Liu,Xiaotian Han,Junyuan Shang,Yuchen Ding,Zhenyu Zhang,Shuohuan Wang,Guibo Zhu,Sirui Han,Dianhai Yu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Vision-language models face, fundamental scaling bottleneck, Vision-language models, making long-video understanding, long-video understanding expensive
备注:
点击查看摘要
Abstract:Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: this https URL
36. 【2609.39920】MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models
链接:https://arxiv.org/abs/2609.39920
作者:Yanshu Li,Jiaqian Li,Canran Xiao,Xi Xiao,Tianyang Wang,Yongtai Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Large vision-language models, ability degrades substantially, Large vision-language, multimodal in-context learning, model size decreases
备注: 17 pages, 8 tables, 5 figures
点击查看摘要
Abstract:Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.
37. 【2609.39915】NavHarness: Adaptive Goals for Agentic Vision-Language Navigation
链接:https://arxiv.org/abs/2609.39915
作者:Haoxiang Shi,Zaijing Li,Muhe Ding,Xiang Deng,Yaowei Wang,Liqiang Nie
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:requires embodied agents, Agent, Verify Agent, Visuomotor Agent, Memory Agent
备注: 22 pages, 10 figures
点击查看摘要
Abstract:Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3\% success and 1.51\,m navigation error across eight challenging routes evaluated three times each.
38. 【2609.39899】Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models
链接:https://arxiv.org/abs/2609.39899
作者:Bangwei Guo,Xiao Chen,Boris Mailhe,Jia Yao,Yiqing Wang,Ankush Mukherjee,Yikang Liu,Zheyuan Zhang,Hang Yu,Terrence Chen,Shanhui Sun
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:myocardial tissue characteristics, Cardiac magnetic resonance, magnetic resonance imaging, ventricular function, tissue characteristics
备注:
点击查看摘要
Abstract:Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians interpret these images by identifying cardiac structures and focusing on the regions relevant to each clinical question, motivating anatomically guided vision-language models (VLMs). Yet CMR-specific supervision for anatomical localisation and clinical question answering remains limited. To address this gap, we investigate fine-grained CMR visual question answering through anatomical grounding and guided attention. We construct 128,915 anatomical-grounding and 42,799 clinical QA pairs across short-axis cine, late gadolinium enhancement, and long-axis cine. These datasets support anatomical recognition, localisation, and clinical assessment without requiring paired reports for individual training images. To help the model learn where to look, we introduce Cardiac Anatomy-Routed Attention (CARA), which selects predicted anatomical priors according to the question and guides decoder attention with learned task-specific strengths. Combining anatomical grounding pretraining with CARA yields our model, CARA-VL. Experiments demonstrate CARA-VL's strengths in clinical assessment and regional localisation across CMR imaging settings, with promising generalization to an external clinical cohort. Together, our data and method provide a practical framework for studying and advancing cardiac visual understanding in VLMs. We will release the QA data derived from public datasets upon publication.
39. 【2609.39894】Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement
链接:https://arxiv.org/abs/2609.39894
作者:Ziyin Huang,Sik-Ho Tsang,Xinyuan Qin,Yui-Lam Chan,Xueling Zhou,Feiyu Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:scene switches, text and graphics, Screen Content Videos, natural videos, Multi-scale Feature Distillation
备注: 5 pages, 4 figures
点击查看摘要
Abstract:Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to the disruption of temporal correlations. To address these challenges, we propose the Spatial-Temporal Multi-scale Network (STM-Net), a novel framework specifically tailored for compressed SCV enhancement. Our approach integrates three complementary components: a Prior-Guided Spatio-Temporal Dispatcher (PG-STD) that routes input into three parallel streams to avoid feature contamination, a Bidirectional Temporal Feature Extraction (BTFE) module that adaptively handles abrupt transitions without explicit detection, and a Cascaded Multi-scale Feature Distillation (CMFD) module that preserves critical high-frequency details. Experimental results demonstrate that STM-Net outperforms state-of-the-art methods in both objective metrics and subjective visual quality, providing a robust solution for screen content artifacts. Code is available at this https URL.
40. 【2609.39883】Grounding with Confidence: Controllable Generative Video Temporal Grounding
链接:https://arxiv.org/abs/2609.39883
作者:Jinhao Chen,Benlei Cui,Ruijian Jia,Ziheng Wang,Tianyu Wo,Pengfei Sun,Longtao Huang,Hui Xue,Yitong Yang,Haiwen Hong
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Video temporal grounding, video search, Video temporal, content review, natural language
备注: 22 pages, 7 figures; includes appendix
点击查看摘要
Abstract:Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro Recall@0.5 from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.
41. 【2609.39871】Hyperspectral Image Models: Technical Report
链接:https://arxiv.org/abs/2609.39871
作者:Tanishq Rachamalla,Aryan Das,Srishti Kaushik,Swalpa Kumar Roy
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Kolmogorov Arnold networks, graph neural networks, Vision Transformers, supervised masked autoencoding, Kolmogorov Arnold
备注: Documentation and benchmark library for hyperspectral image models
点击查看摘要
Abstract:Hyperspectral remote sensing has advanced across diverse deep learning paradigms, including spectral spatial CNNs, Vision Transformers, Mamba, graph neural networks, Kolmogorov Arnold networks, and self supervised masked autoencoding. Yet progress remains hindered by fragmented repositories, incompatible tensor conventions, and non standardized evaluation. Hyperspectral Image Models addresses these challenges through a modular framework unifying 55 representative models across six paradigms with a common registry, automatic 4D/5D tensor adaptation, and standardized constructors. It integrates 24 benchmark scenes from Airborne, Spaceborne, UAV, and Mars CRISM sensors, with caching, label remapping, PCA, explicit band selection or raw spectra, optional spatial max pooling, and arbitrary PxP patch extraction. To prevent inflated accuracy from overlapping windows, it supports class balanced random partitioning and spatially disjoint regional blocking with Chebyshev guard bands that eliminate train test pixel overlap. Experiments use a single this http URL with deterministic seeds and complete provenance, generating LaTeX benchmark tables and classification maps. Across 1,320 model scene evaluations and 6,600 seeded runs, scene difficulty dominates architecture, with mean accuracy ranging from 96.40% on Botswana to 56.70% on Houston 2018, versus a 15 point spread across paradigm means. No paradigm universally dominates, while sub 1 M parameter models can match architectures two orders of magnitude larger. Code is publicly available at this https URL.
42. 【2609.39841】DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
链接:https://arxiv.org/abs/2609.39841
作者:Merav Keidar,Tomer Borreda,Rajalakshmi Nandakumar,Or Litany
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:recorded sensor data, sensor data supports, data supports closed-loop, supports closed-loop evaluation, Reconstructing dynamic driving
备注: Project page: [this https URL](https://dyrad-nvs.github.io/) . Code: [this https URL](https://github.com/Dyrad-NVS/DyRAD)
点击查看摘要
Abstract:Reconstructing dynamic driving scenes from recorded sensor data supports closed-loop evaluation of autonomous driving systems by synthesizing observations beyond the original trajectory. Unlike cameras and LiDAR, radar measures radial velocity directly through Doppler. Yet existing radar novel-view synthesis fails to exploit this capability: methods addressing dynamic scenes reconstruct only range-azimuth tensors, while methods that render Doppler assume static scenes. Moreover, because radar processing spreads each reflection across multiple bins, existing representations absorb this spread into scene geometry, causing it to render incorrectly when the viewpoint moves. We present DyRAD, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors. Reflector velocities are derived from object tracks and projected onto the line of sight, making Doppler both a rendered output and supervision for those tracks. Crucially, we render reflectors through a fixed analytic point-spread function (PSF) derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation. Beyond improving scene reconstruction, this separation also enables zero-shot sensor-configuration transfer, allowing the same reconstructed scene to be rendered under different radar specifications without refitting. We evaluate DyRAD on RADIal, Boreas, and a synthetic benchmark across both on-path poses and displaced viewpoints untested by prior work. On RADIal, DyRAD recovers radar detections in 90.7% of reference-detected objects, compared with 26.9% for the strongest baseline.
43. 【2609.39836】Spherical Interpolation for Backward-Compatible Multimodal Representations
链接:https://arxiv.org/abs/2609.39836
作者:Simone Ricci,Niccolò Biondi,Federico Pernici
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Contrastive vision-language models, making cosine similarity, Contrastive vision-language, vision-language models map, models map visual
备注: Accepted at NeurIPS 2026
点击查看摘要
Abstract:Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment can partially mitigate this problem by mapping new-model queries into the old-model gallery space. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether interpolation along the spherical geodesic between these two normalized query representations can improve retrieval without re-indexing the gallery. We characterize when this path contains an interior query direction closer to an idealized retrieval-optimal direction than either endpoint, and connect this characterization to Recall@$K$ through a local margin-based certification result. Experiments across multiple benchmarks and model families show that post-alignment spherical interpolation improves over orthogonal alignment alone, recovering backward-compatibility in most evaluated settings. Consistent with our geometric characterization, per-query oracle analysis shows that retrieval-favorable interior points occur frequently in practice. Code is available at this https URL .
44. 【2609.39832】P-SRM: Selective Recovery of Rejected Predictions in Visual Tracking
链接:https://arxiv.org/abs/2609.39832
作者:Youbin He,Siwei Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:suppress unreliable predictions, suppress unreliable, visual tracking methods, Selective Recovery Method, Post-rejection Selective Recovery
备注: 5 pages, 2 figures, 3 tables
点击查看摘要
Abstract:Many visual tracking methods use rejection mechanisms to suppress unreliable predictions. However, these mechanisms can also reject correctly localized candidates, leaving useful information unused. We investigate how to identify and recover these candidates while preserving native accepted outputs and candidate coordinates. To this end, we propose P-SRM (Post-rejection Selective Recovery Method), which combines spatial responses, past accepted states, and native decision margins to reassess candidates and selectively restore reliable predictions. We evaluate P-SRM on six trackers and four datasets spanning category-specific, point, and generic object tracking. Across all nine configurations, P-SRM improves rejected-candidate ranking and overall tracking performance. These results show that post-rejection verification can identify and recover useful predictions discarded by native rejection, demonstrating the value of reusing rejected information. Project repository: this https URL.
45. 【2609.39794】Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model
链接:https://arxiv.org/abs/2609.39794
作者:Zaijing Li,Rui Shao,Bing Hu,Haoyu Zhang,Dongmei Jiang,Liqiang Nie
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:previously learned tasks, domains remains inefficient, incurring substantial costs, shown strong promise, general-purpose robotic manipulation
备注: 24 pages, 8 figures
点击查看摘要
Abstract:Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning. Optimus-R introduces: (i) An \textbf{Inline Memory Interface for skill extraction}. It inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to derive control-aware query and skill representations within the native action-conditioning pathway. (ii) A \textbf{Query-Skill Memory Bank for skill learning}. It externalizes skills into query prototypes for deciding \emph{what} to retrieve and skill values for specifying \emph{how} to act, supporting skill reuse and expansion with limited parameter updates. (iii) A lightweight \textbf{Bridge-and-Adapt mechanism for skill updating}. It aligns target-domain queries and skills with the existing memory space through a lightweight adapter and residual memory updates. Experiments on in-domain adaptation, cross-domain adaptation, and lifelong learning show that Optimus-R enables data-efficient skill learning while mitigating catastrophic forgetting.
46. 【2609.39785】Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision
链接:https://arxiv.org/abs/2609.39785
作者:Weijian Jian,Xiaoyue Zhang,Bin Xiao,Chunyu Xie,Yixiao He,Yutao Liu,Dawei Leng,Yuhui Yin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:relies heavily, creating a fundamental, massive manual annotations, heavily on massive, fundamental bottleneck
备注: Published at ECCV 2026. Includes supplementary material. Code: [this https URL](https://github.com/360CVGroup/MoSA)
点击查看摘要
Abstract:The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised methods attempt to learn object concepts from motion, they typically overfit to moving entities, lacking both multi-granularity understanding and the ability to generalize to static objects. To overcome this, we introduce Motion-Grounded Segment Anything (MoSA), a highly scalable unsupervised framework that learns a transferable objectness prior from unlabeled videos. MoSA operates in three progressive stages: (1) automatically generating multi-granularity motion pseudo-labels from large-scale video data; (2) training a Perceptual Grouping Model (PGM) via contrastive learning to internalize a generalized, appearance-driven concept of objects; and (3) transferring this learned prior into a prompt-guided architecture for segment-anything-style inference on images. Extensive zero-shot evaluations across seven challenging benchmarks (e.g., COCO and ADE20K) demonstrate that MoSA significantly outperforms existing unsupervised methods. Notably, despite using zero manual annotations, MoSA achieves segmentation performance comparable to the fully supervised SAM. Our findings reveal that harnessing large-scale unlabeled motion is a feasible and highly scalable alternative to annotation-driven segment-anything pipelines.
47. 【2609.39757】Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction
链接:https://arxiv.org/abs/2609.39757
作者:Xiao Cui,Mo Zhu,Yulei Qin,Yuze Wu,Wengang Zhou,Houqiang Li
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:API-accessible large language, Black-box distillation, large language models, practical route, route for transferring
备注: NeurIPS 2026
点击查看摘要
Abstract:Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at this https URL.
48. 【2609.39756】Determining Vertical Displacement of Agricultural Areas Using UAV-Photogrammetry and a Heteroscedastic Deep Learning Model
链接:https://arxiv.org/abs/2609.39756
作者:Wojciech Gruszczyński,Edyta Puniach,Paweł Ćwiąkała,Wojciech Matwij
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:unmanned aerial vehicle, photogrammetry point clouds, determine vertical ground, ground surface elevations, point cloud classification
备注:
点击查看摘要
Abstract:This article introduces an algorithm that uses a U-Net architecture to determine vertical ground surface displacements from unmanned aerial vehicle (UAV)-photogrammetry point clouds, offering an alternative to traditional ground filtering methods. Unlike con-ventional ground filters that rely on point cloud classification, the proposed approach em-ploys heteroscedastic regression. The U-Net model predicts the conditional expected val-ues of the elevation corrections, aiming to reduce the impact of vegetation on determined ground surface elevations. Concurrently, it estimates the logarithm of the elevation cor-rection variance, allowing for direct quantification of the uncertainty associated with each elevation correction value. The algorithm was evaluated using three metrics: the root mean square error (RMSE) of vertical displacements, the percentage of nodes with deter-mined displacement values, and the percentage of outliers among those values. Perfor-mance was assessed using the technique for order of preference by similarity to ideal so-lution (TOPSIS) method and compared against several ground-filter-based algorithms across four datasets, each including at least two time intervals. In most cases, the U-Net-based approach demonstrated a slight performance advantage over traditional ground filtering techniques. For example, for the U-Net-based algorithm, for one of the test da-tasets, the RMSE of the determined subsidences was 6.1 cm, the percentage of nodes with determined subsidences was 80.5%, and the percentage of outliers was 0.2%. For the same case, the algorithm based on the next best model (SMRF) allowed an RMSE of 7.7 cm to be obtained; for 77.3% of nodes, the subsidences were determined; and the percentage of outliers was 0.3%.
49. 【2609.39748】FAST: Flow Any Scene Transformer
链接:https://arxiv.org/abs/2609.39748
作者:Yongjian Zhang,Longguang Wang,Zhuo Song,Zhiheng Fu,Liang Lin,Yulan Guo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:matching remains underexplored, vision foundation models, precise correspondence matching, correspondence matching remains, remains underexplored
备注:
点击查看摘要
Abstract:Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.
50. 【2609.39723】Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
链接:https://arxiv.org/abs/2609.39723
作者:Linfeng Jiang,Steven McDonagh,Yuhang Chen,Xingyu Zhao,Siddartha Khastgir,Andi Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Strong unrestricted adversarial, Strong unrestricted, unrestricted adversarial attacks, distort the primary, primary object
备注:
点击查看摘要
Abstract:Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global classifier guidance. We demonstrate three key findings: 1. A carrier mitigates subject distortion by absorbing a larger share of globally normalized attack updates. 2. A carrier improves cross-model transferability, governed by the strength of target-related features that balance semantic separation and transfer performance. 3. Successful targeted attacks retain the personalized subject as the primary content perceived by humans while successfully misleading the classifier. Our results demonstrate that a visually secondary carrier offers an auxiliary spatial pathway for adversarial changes, enabling strong and transferable attacks while improving subject preservation.
51. 【2609.39709】BTC3D: Blended Tile Conditioning for Detail-Enhancing Image-to-3D Generation
链接:https://arxiv.org/abs/2609.39709
作者:Junyu Li,Qiuyu Chen,Pengcheng Wang,Shiqi Yang,Alexandra Gomez-Villa,Joost van de Weijer,Ruilin Li,Kai Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent diffusion-based pipelines, achieved promising progress, Recent diffusion-based, achieved promising, promising progress
备注:
点击查看摘要
Abstract:Recent diffusion-based pipelines have achieved promising progress in image-to-3D synthesis. However, generating high-fidelity details remains challenging, especially when the input image contains rich details. Existing approaches often rely on globally encoded conditioning features, which compress spatial information and limit the model to reproduce fine-grained details. This common design often leads to a phenomenon we term detail attenuation. Moreover, improving image-to-3D synthesis quality typically requires retraining or fine-tuning large diffusion models, which can be computationally expensive and impractical for complex 3D pipelines. In this work, we present Blended Tile Conditioning for image-to-3D generation (BTC3D), a training-free inference time framework that enhances fine-grained detail preservation in image-to-3D diffusion pipelines. To alleviate detail attenuation, we first examine the image feature additivity in image-to-3D models. Based on this property, we introduce a blended tile embedding that extracts local conditioning signals from split image regional patches, allowing the diffusion model to better preserve fine-grained visual details. To integrate the global and local conditioning guidance stably, we propose a dynamic conditioning schedule that gradually increases the influence of tile-level conditioning during later low-noise stages of diffusion. Our proposed method BTC3D operates entirely at inference time and can be seamlessly integrated into existing image-to-3D diffusion pipelines. Experimental results demonstrate that the proposed approach significantly improves texture quality and visual fidelity of the base model while maintaining global structural consistency in a training-free manner.
52. 【2609.39704】When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic
链接:https://arxiv.org/abs/2609.39704
作者:Muhammad Zawish,Steven Davy
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:hurts worst-group robustness, masking-based token pruning, labels or fine-tuning, paper demonstrate, masking-based token
备注:
点击查看摘要
Abstract:This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5\% relative on some datasets and degrades it by up to 100\% on others. We trace this instability to spurious inversion: background patches receive higher CLIP text-similarity than the true object when the spurious attribute is background-separable, inverting the assumption every text- and attention-guided pruning method relies on. We introduce the Spurious Inversion Metric (SIM), a label-free, pre-deployment diagnostic whose sign predicts this effect with statistical significance (binomial $p=0.035$) across all 8 datasets, and remains dependable across 6 CLIP architectures with a clean foreground/background split. Naive masking is itself a major source of risk: it causes the largest average-accuracy loss of any method we evaluate, and its own per-image segmentation step is a significant runtime bottleneck. To address this, we design a batched, synchronization-free GPU segmentation routine that cuts this overhead from 3.5$\times$ to 1.75$\times$ baseline. Gating deployment by SIM's sign recovers masking's benefits while avoiding its worst failures, matching or exceeding a strong pruning baseline on 7 of 8 datasets.
53. 【2609.39688】ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
链接:https://arxiv.org/abs/2609.39688
作者:Tobia Poppi,Silvia Cappelletti,Samuele Poppi,Marcella Cornia,Lorenzo Baraldi,Diego Garcia-Olano,Rita Cucchiara
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
关键词:changing benign representations, web-scale training data, training data embed, unnecessarily changing benign, CLIP underlie
备注:
点击查看摘要
Abstract:Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at this https URL.
54. 【2609.39684】Unapologetically Distributed: A Call for Decentralized Document Analysis
链接:https://arxiv.org/abs/2609.39684
作者:Adrià Molina,Oriol Ramos Terrades,Josep Lladós
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:increasingly important concern, Document Analysis community, Document Analysis, governmental institutions, local businesses
备注: Accepted at BMVC2026
点击查看摘要
Abstract:Privacy has become an increasingly important concern in the Document Analysis community, to the extent that in many environments such as archives, governmental institutions, and local businesses, the adoption of automation is restricted by legal and policy constraints. While federated learning has often been regarded as a ``necessary evil'', implying an unavoidable performance trade-off in exchange for decentralization and privacy, many prior works overlook its potential to improve robustness to out-of-distribution data. In this paper, we present Unapologetically Distributed, the first comprehensive study evaluating distributed learning in Document Analysis along three key axes simultaneously: the tasks addressed, the architectures employed, and the fine-tuning strategies applied. Specifically, we demonstrate how various distributed training approaches enhance generalization capabilities across diverse tasks such as Table Recognition, handwriting recognition, and Word Spotting, particularly during transfer learning stages. Our results provide strong evidence that decentralization is not merely a constraint, but a valuable opportunity to improve model robustness and adaptability in real-world Document Analysis scenarios.
55. 【2609.39681】MC-PanDA++: Simpler, Stronger, and More Robust Domain-Adaptive Panoptic Segmentation
链接:https://arxiv.org/abs/2609.39681
作者:Ivan Martinović,Josip Šarić,Yuki M. Asano,Siniša Šegvić
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Unsupervised domain adaptation, unlabeled target domain, cost-effectively labeled source, Existing panoptic UDA, distribution gap
备注: Preprint. Accepted to IJCV
点击查看摘要
Abstract:Unsupervised domain adaptation (UDA) reduces the annotation burden in panoptic segmentation by leveraging a cost-effectively labeled source domain (e.g., synthetic) and an unlabeled target domain to bridge the distribution gap. Existing panoptic UDA methods rely on teacher-student consistency learning built upon suboptimal per-pixel segmentation architectures. In contrast, state-of-the-art mask transformers are rarely adopted due to their pronounced vulnerability to confirmation bias in consistency learning, where erroneous teacher predictions are reinforced during training. Our earlier approach, MC-PanDA, mitigates this issue through fine-grained confidence estimation, which suppresses gradients from unreliable masks while sampling informative yet reliable locations for loss computation. However, this method entails a complex multi-stage training and requires careful hyperparameter tuning. This work presents MC-PanDA++, which addresses these limitations by introducing: (i) self-supervised vision encoders that provide a stronger and more robust initialization, further reducing the reliance on human annotations, (ii) per-class, self-adapting mask-wide loss scaling that stabilizes training and enables the usage of a single set of hyperparameters across domains, and (iii) a single-stage training pipeline that decreases overall conceptual complexity. Together, these improvements result in a conceptually simpler, better-performing, and more robust method for domain-adaptive panoptics. Source code: this https URL
56. 【2609.39662】ypographic Attack Against VLM-based AI-generated Image Detection
链接:https://arxiv.org/abs/2609.39662
作者:Eunmin Lee,Jungwoo Kim,Jong-Seok Lee
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:providing natural-language explanations, Vision-language models, providing natural-language, natural-language explanations, Vision-language
备注: 5 pages, 3 figures
点击查看摘要
Abstract:Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.
57. 【2609.39660】BAM! Bayesian Anything Model: a foundation model for generative computational imaging
链接:https://arxiv.org/abs/2609.39660
作者:Alessio Spagnoletti,Charlesquin Kemajou Mbakam,Jonathan Spence,Andrés Almansa,Marcelo Pereyra
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
关键词:transforming Bayesian computational, field still lacks, Bayesian computational imaging, models, BAM
备注: 37 pages, 25 figures
点击查看摘要
Abstract:Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces significant bias and computational cost. Physics-aware generative models avoid this bias, but each is tied to a specific dataset, task and instrument. We introduce BAM (Bayesian Anything Model), a lightweight foundation model for few-step, physics-aware posterior sampling that generalises robustly to unseen data and tasks, zero-shot or with minimal finetuning. BAM upgrades the operator-conditioned Reconstruct Anything Model (RAM) backbone (Terris et al.) into a conditional flow map, so instrument physics is specified at inference time rather than fixed during training. BAM has just 36M parameters and is pre-trained jointly on large image corpora and libraries of forward operators. A single network then draws posterior samples in a few steps, with no likelihood approximation and no guidance weights to tune. Across linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and the Kohler camera-shake benchmark, BAM outperforms in just 3 steps both specialised models and leading zero-shot methods in sample quality, at a fraction of their computational cost. BAM gives the community an accessible entry point to generative computational imaging, lowers the economic and environmental cost of training imaging models, and opens a new path for research on physics-aware Bayesian computational imaging. Official page: this https URL
58. 【2609.39657】Diffusable Latents from Structure-Agnostic Distillation
链接:https://arxiv.org/abs/2609.39657
作者:Adrien Ramanana Rahary,Nicolas Dufour,Patrick Pérez,David Picard
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:enabling diffusion models, higher sample quality, pretrained foundation models, reach higher sample, Distilling pretrained foundation
备注: NeurIPS 2026 Workshop on Principles of Generative Modeling
点击查看摘要
Abstract:Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's performs as well as or slightly better than dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. First-order matching extends naturally to 1D token-sequence latents and across modalities, where distilling a text encoder into an image autoencoder still improves diffusability; a relational objective based only on each image's nearest neighbours improves it as well. Code and blog post are available at this https URL and this https URL.
59. 【2609.39649】FANVIDv2: Evaluating Video Super-Resolution by Face and Licence-Plate Recognition Under Compound Degradation
链接:https://arxiv.org/abs/2609.39649
作者:Kavitha Viswanathan,Vrinda Goel,Shlesh Gholap,Devayan Ghosh,Madhav Gupta,Dhruvi Ganatra,Sanket Potdar,Amit Sethi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:PSNR and SSIM, Video super-resolution, judged by PSNR, downsampled bicubically, surveillance its purpose
备注:
点击查看摘要
Abstract:Video super-resolution (VSR) is normally judged by PSNR and SSIM on clips that were downsampled bicubically, although in surveillance its purpose is to make faces and licence plates \emph{recognisable}. We present FANVIDv2, a benchmark that scores VSR by what a recognition pipeline can do with its output. FANVIDv2 provides $320\times180$ low-resolution (LR) clips with high-resolution (HR) references for 48 public figures (with one HR gallery image each) and 375 licence-plate clips covering 360 distinct plate strings. LR clips are generated with a randomised compound degradation (blur, resize jitter, sensor noise, JPEG compression, final downsampling) rather than bicubic downsampling alone. Two metrics score recognition \emph{inside} detections: FaceRecBox rewards a face only if it is localised and correctly identified, and TextRecBox scores plate transcriptions by normalised edit distance weighted by localisation quality. With a 2.3\,M-parameter VSR baseline (RCDM), FaceRecBox rises from 0.6864 to 0.7222, identity accuracy on matched faces from 84.35\% to 86.93\%, and TextRecBox from 0.3088 to 0.3667; a residual-map gated variant (RCDM-RMGF) reaches 0.3801 on plates. We describe the degradation model, the baseline architectures and the scorers in detail, and release annotations, metadata, download and degradation scripts and evaluation code.
60. 【2609.39648】From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models
链接:https://arxiv.org/abs/2609.39648
作者:Cristina López Amado,Marco Fumero,Francesco Locatello
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:sigma, typically viewed, viewed as stochastic, stochastic processes, processes that transform
备注:
点击查看摘要
Abstract:Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $\sigma$, we treat the denoiser as a self-map and study its dynamics. For an exact denoiser, fixed points correspond to critical points of the smoothed data density, while attractors correspond to its modes; as $\sigma$ increases, sample-level modes merge into progressively coarser ones. This suggests a geometric view of memorization: examples that receive excess probability mass due to duplication or overfitting, as well as outliers, should remain distinguishable under stronger smoothing than ordinary examples. We quantify this persistence by the critical scale $\sigma_c$, the largest noise scale at which an example is retained by the fixed-scale dynamics. In conditional models, the same construction extends naturally to image--caption pairs. Experiments in controlled settings and on large-scale models show that $\sigma_c$ tracks memorization arising from duplication, overfitting, and outliers, and identifies both memorized and partially memorized examples in Stable Diffusion. Moreover, $\sigma_c$ yields interpretable measures of the image spatial distribution and caption dependence of memorization.
61. 【2609.39635】SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
链接:https://arxiv.org/abs/2609.39635
作者:Shuang Liang,Lejun Liao,Shiyuan Zhang,Max C. Zhang,Xiaolong Luo,Han Wang,Stefano Anzellotti,Yuan Yuan
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:contrastive analysis separates, textit, contrastive analysis, content shared, target dataset
备注: 28 pages, 18 figures, 9 tables
点击查看摘要
Abstract:Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below $2$) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy $0.950$ vs.\ at most $0.281$) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent ($90.5\%$ vs.\ $27.7\%$) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.
62. 【2609.39627】Introduction to Computer Vision
链接:https://arxiv.org/abs/2609.39627
作者:Stan Birchfield
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:modern deep learning, spanning classical, deep learning, presents a code-first, code-first introduction
备注: 217 pages. For online notes and code, see [this https URL](https://sbirchfield.github.io/cvintro)
点击查看摘要
Abstract:This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organized as 44 short chapters across three parts, the book builds each topic from first principles: image arithmetic and morphology; convolution, pyramids, and frequency-domain filtering; feature detection, optical flow, and stereo; projective geometry, camera calibration, and structure from motion; and the full arc of modern deep learning, from a single neuron through convolutional networks, backpropagation, classic architectures, transfer learning, object detection, and semantic and instance segmentation, concluding with engineering considerations like mixed-precision and parallel training. Every technique is implemented directly in Python and NumPy or PyTorch and checked numerically against the corresponding OpenCV or PyTorch library function, so readers see not just the mathematics but its concrete behavior on real and synthetic data. The material was distilled with AI assistance from freely available online course notes, condensing extensive working code into concise mathematical exposition while preserving verified, reproducible results throughout. It is intended as a self-contained reference for students and practitioners who want to understand computer vision algorithms and their Python implementations.
63. 【2609.39625】D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
链接:https://arxiv.org/abs/2609.39625
作者:Xinyue Xu,Jiahao Zhang,Lijie Hu,Peter Hase,Hao Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:reveal visual structure, diffusion transformers, Diffusion Scope, Sparse autoencoders, reveal visual
备注:
点击查看摘要
Abstract:Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text descriptions against these visual centroids in the shared image-text embedding space then enables retrieval of individual features without per-feature text annotations. The underlying patches provide evidence for inspecting each selection, while spatially masked interventions test the corresponding decoder direction at varying strengths under fixed generation conditions. We characterize 150 SAEs across two model families and five layers, and introduce a benchmark of 100 target concepts with ten contexts each spanning under-specified and explicit-conflict conditions. Our empirical results show that high reconstruction fidelity can coexist with low dictionary utilization and limited visual-evidence coverage. Under per-case best-of-sweep strength selection, contrastive retrieval yields larger mean regional SigLIP~2 gains than direct retrieval across the tested steering configurations, without consistently improving outside-region preservation. D-Scope provides an inspectable framework for evaluating sparse DiT features through their visual evidence and the effects of their decoder directions on generation. The demo is available at this https URL.
64. 【2609.39624】ExpandDiff: Dynamic Range Expanding Diffusion for Single-Image HDR Reconstruction
链接:https://arxiv.org/abs/2609.39624
作者:Mehmet Emre andıran,Zhuoqian Yang,Liying Lu,Mathieu Salzmann,Sabine Süsstrunk
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:requires inferring missing, inferring missing detail, Single-image HDR reconstruction, Single-image HDR, reconstruction requires inferring
备注: 5 pages, 3 figures, 2 tables. Submitted to ICASSP 2027. Code and supplementary material: [this https URL](https://memreandiran.github.io/expanddiff/)
点击查看摘要
Abstract:Single-image HDR reconstruction requires inferring missing detail while preserving the visible content of an LDR image. Differences in sensor dynamic range and exposure cause LDR images to lose varying amounts of information in shadows and highlights. We present ExpandDiff, a conditional diffusion pipeline that jointly reconstructs clipped shadows and highlights. To account for this variation, we introduce Dynamic Clipping Synthesis (DCS), which randomly samples shadow and highlight clipping percentiles when constructing training inputs from HDR targets. A pixel-space diffusion model guided by spatially-adaptive normalization then predicts perceptually encoded HDR through a bounded output head, reconstructing both clipping directions in one sampling trajectory. On the SI-HDR benchmark, ExpandDiff variants improve HDR reconstruction accuracy by 3.43 dB in PU21-PSNR over the strongest evaluated competing method, and by 7.34 dB under two-sided clipping. The code and supplementary material are available at this https URL.
65. 【2609.39623】Semantic Watermarking for Malicious Image Manipulation Detection
链接:https://arxiv.org/abs/2609.39623
作者:Yoonseo Kim,Seungwoo Baek,Junyoung Park
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:preserving visual plausibility, high-fidelity generative editing, generative editing models, visual plausibility, vulnerable populations
备注:
点击查看摘要
Abstract:The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images while preserving visual plausibility, with concrete consequences for public discourse and vulnerable populations. We propose a robust semantic watermarking framework that reframes the watermark as a recoverable semantic reference rather than an opaque identifier. Our framework combines a $\beta$-VAE-based binary watermark (CLIP-VAE) with explicit channel-aware training---random bit-flip noise is injected during training so that the decoder learns graceful degradation under the noisy watermarking channel. As a downstream application, a lightweight module SDA-Net uses the recovered semantic embedding to expose not only whether but in which semantic direction an image has been altered. In a 5-way comparison against representative binary hashing baselines (SimHash, ITQ, HashNet, and their robust-MLP variants), CLIP-VAE achieves the highest reconstruction cosine similarity to the original CLIP embedding under realistic InstructPix2Pix bit-error rates, and uniquely supports direction-of-drift detection---a forensic complement to existing content-moderation pipelines.
66. 【2609.39605】FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning
链接:https://arxiv.org/abs/2609.39605
作者:Łukasz Rudnik,Agnieszka Polowczyk,Alicja Polowczyk,Przemysław Spurek
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:temporally coherent videos, generative video models, rapid advancement, advancement of generative, enabled the synthesis
备注:
点击查看摘要
Abstract:The rapid advancement of generative video models has enabled the synthesis of increasingly realistic and temporally coherent videos, while also raising concerns about the generation of harmful content. The reliance on large-scale web datasets during training inevitably exposes these models to undesirable material, making concept unlearning an essential mitigation. Existing methods mainly target static visual concepts, such as objects, identities, or unsafe appearance, largely overlooking motion unlearning. Furthermore, these approaches often pay little attention to preserving the surrounding scene. As a result, successful concept removal may unintentionally alter the background, composition, or overall video dynamics. We argue that effective unlearning should ideally change only what is targeted, while minimizing unnecessary changes to the remaining scene. In this work, we introduce FOMO, to the best of our knowledge the first training-based selective video unlearning method that directly treats preservation of the original scene as a priority. We formulate unlearning around two complementary objectives: what to change and what to preserve. Our method localizes concept-related representations and modifies them, while the preservation mechanism maintains non-target scene information without requiring auxiliary data. Beyond simply erasing unwanted concepts, FOMO explicitly redirects the generation toward a specified safe alternative. We further extend this formulation to motion unlearning, where the concept is defined by temporal behavior rather than a fixed spatial region. Our solution achieves effective unlearning across unsafe content, object, and motion concepts, while achieving the best trade-off between concept removal and scene preservation. Code: this https URL Project Page this https URL
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2609.39605 [cs.CV]
(or
arXiv:2609.39605v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2609.39605
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
67. 【2609.39601】GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
链接:https://arxiv.org/abs/2609.39601
作者:Qize Yu,Lianrui Fan,Boyu Chen,Jiaqi Liang,Xini Ding,Yue Chen,Zetian Song,Yuran Wang,Yi Zou,Kaixuan Wang,Tianxing Chen,Wenxuan Song,Bohan Zhou,Mingleyang Li,Siqiao Huang,Yuqi Ye,Caigao Jiang,Wei Wei,Ruihai Wu,Hang Zhang,Yixiao Ge,Shuchang Zhou,Shilong Liu,Xianming Liu,Ping Luo,Shiyu Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
关键词:Precise grounding matters, Precise grounding, Precise, GroundingPI, grounding matters
备注: 64 pages, including supplementary material. Project page: [this https URL](https://groundingpi.github.io/) Code: [this https URL](https://github.com/groundingpi/GroundingPI) Model: [this https URL](https://huggingface.co/GroundingPI/GroundingPI)
点击查看摘要
Abstract:Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
68. 【2609.39600】GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
链接:https://arxiv.org/abs/2609.39600
作者:Qize Yu,Lianrui Fan,Bowen Ping,Xini Ding,Zetian Song,Junbo Niu,Kaixuan Wang,Tianxing Chen,Yue Chen,Minghua He,Yuran Wang,Jie Huang,Haojun Zhang,Min Chen,Hao Li,Wenxuan Song,Ruihai Wu,Xianming Liu,Shilong Liu,Shuchang Zhou,Ping Luo,Shiyu Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
关键词:introducing sequential latency, serialize spatial predictions, introducing sequential, output tokens, sequential latency
备注: 61 pages, including supplementary material. Project page: [this https URL](https://groundingpi.github.io/groundanything/) Code: [ [this https URL](https://github.com/groundingpi/GroundAnything) ]( [this https URL](https://github.com/groundingpi/GroundAnything) ) Model: [this https URL](https://huggingface.co/GroundingPI/GroundAnything) , [this https URL](https://huggingface.co/GroundingPI/GroundAnything-VLM)
点击查看摘要
Abstract:Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
69. 【2609.39591】Structural Limits of the Information-Theoretic Uncertainty Decomposition
链接:https://arxiv.org/abs/2609.39591
作者:Jakob Lønborg Christensen,Christian F. Baumgartner,Morten Rieger Hannemose,Anders Bjorholm Dahl,Vedrana Andersen Dahl
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:typically decomposes uncertainty, machine learning typically, learning typically decomposes, standard information-theoretic framework, Uncertainty estimation
备注:
点击查看摘要
Abstract:Uncertainty estimation in machine learning typically decomposes uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) using the standard information-theoretic framework. However, in practice, two critical issues arise: entanglement (AU and EU are highly correlated) and epistemic collapse (EU magnitude shrinks with increasing model capacity). We analyze this framework on a functional level and discover that significant portions of the assumed AU, EU range are infeasible in finite settings, and cannot be attained with any class probabilities. We characterize how this infeasible region scales with the number of classes and Monte Carlo samples $N$ (e.g., from ensembles with $N$ members), revealing it is bounded by $\text{AU} \leq \log(2)/N$. Crucially, the infeasible region's boundary helps explain epistemic collapse: when model confidence is high, $\text{AU} \text{EU}$ is guaranteed by this fundamental structural limitation. Our findings show that increasing ensemble size mitigates epistemic collapse by reducing the infeasible area. Lastly, we caution against interpreting AU and EU as independent quantities in low AU regimes, since we show they are coupled when $\text{AU} \leq \log(2)/N$.
70. 【2609.39590】SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images
链接:https://arxiv.org/abs/2609.39590
作者:Guibiao Liao,Mochu Xiang,Heng Li,Ken Deng,Zijie Wang,Guanbin Li,Ping Tan,Shenghua Gao,Yizhou Yu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:recover complete, aims to recover, scene generation aims, visual observations, object
备注:
点击查看摘要
Abstract:Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.
71. 【2609.39588】KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs
链接:https://arxiv.org/abs/2609.39588
作者:Aravindh Mahendran,Michael King,Matthew Koichi Grimes,Antoine Yang,Tyler Zhu,Joseph Heyward,Tengda Han,Shiry Ginosar,Chen Sun,Dima Damen,Simon Osindero,Noah Snavely,Simon Lynen,João Carreira,Viorica Pătrăucean
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:probes geographical layout, geographical layout understanding, large-scale spatial intelligence, real-world videos, push the frontier
备注:
点击查看摘要
Abstract:We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at this https URL.
72. 【2609.39585】A Generalizable and Explainable Framework for Synthetic Video Detection Using First-Digit Gradient Statistics
链接:https://arxiv.org/abs/2609.39585
作者:Sidharth Shanu,Gautam Kumar,Tej Singh
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:harder to detect, generate a diverse, landscapes to street, street views, views to animal
备注: 10 Pages
点击查看摘要
Abstract:AI video generators have not only become harder to detect but are used to generate a diverse set of scenarios from landscapes to street views to animal videos. This creates a problem where CNN-based detectors are effective but offer no insight into their inner workings, while forensics-based detectors are often pretrained for a set scenario or become too complex to derive meaningful insights. We present a novel approach to AI video detection using Sobel gradient values analysed with the first-digit law. Using linear discriminant analysis, we visualise the discriminatory signal, while a multi-layer perceptron is used for classification. The detection method has no generator- or scenespecific features, and the model has no knowledge of container formats, codec, bitrate, or compression artefacts. The model is trained and tested on GenBuster-200K, GenBusterBench, GenVA, FaceForensics++ C23, and CelebDF. We also show how zero-shot detection fails even though the feature set carries a discriminatory signal.
73. 【2609.39573】Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond
链接:https://arxiv.org/abs/2609.39573
作者:Simone Facchiano,Jan Eric Lenssen,Bernt Schiele,Wolfgang Stammer,Fabio Galasso,Jonas Fischer
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:promoting benign alternatives, suppressing harmful content, achieve near-photorealistic quality, controlling their outputs, near-photorealistic quality
备注:
点击查看摘要
Abstract:As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state of the generation and often causes unintended global changes. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on the noisy states of flow models, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Despite using no explicit spatial masks or object priors, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free.
74. 【2609.39567】Invariant Shape Analysis of Surfaces with Spherical Topology
链接:https://arxiv.org/abs/2609.39567
作者:T. Shaska,M.-R. Siadat
类目:Computer Vision and Pattern Recognition (cs.CV); Algebraic Geometry (math.AG)
关键词:standard rotation-invariant reductions, Spherical harmonic descriptors, Spherical harmonic, shapes depend, rotation-invariant reductions
备注:
点击查看摘要
Abstract:Spherical harmonic descriptors of closed 3D shapes depend on the parameterization, the pose and the scale of the surface, and the standard rotation-invariant reductions, the power spectrum and the bispectrum, discard the relative orientation of the harmonic bands and cannot distinguish a shape from its mirror image. We construct a descriptor that removes all three dependencies exactly and loses nothing else: a conformal parameterization normalized by its conformal barycenter, followed by polynomial invariants of the rotation group. Identifying each harmonic band with a binary form turns the rotation quotient into classical invariant theory and makes reflections visible as the sign of an invariant, so chirality is recorded. The descriptor is complete for the truncated expansion, stable in the orbit distance, and comes with numerical diagnostics. Benchmarks confirm the guarantees, and on bilateral anatomical structures the descriptor separates mirror-image pairs from asymmetric pairs, which parity-blind descriptors cannot.
75. 【2609.39566】From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning
链接:https://arxiv.org/abs/2609.39566
作者:Minye Shao,Chaohui Yu,Yixuan Wu,Fan Wang,Ling Shao,Yang Long
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Foundation models, tool-use harnesses, Foundation, clinical agents, clinical
备注:
点击查看摘要
Abstract:Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evidence, together with a tool-use harness and an agentic post-training framework for compact vision-language policy models. We further introduce a longitudinal multimodal benchmark built on UK Biobank, comprising 50,401 clinical questions derived from real-world ICD-10-coded diagnoses of 4,739 participants. Each question links to a patient-specific environment containing clinical context and multi-sequence MRI from baseline and follow-up visits, where agents autonomously select which visits, organs, modalities, slices, and specialist tools to inspect and compare. Supervised fine-tuning transfers evidence-seeking workflows from 14,734 frontier-model interaction trajectories, followed by agentic reinforcement learning on the learner's own environment interactions. Privileged on-policy self-distillation and rubric-based LLM feedback refine evidence-to-conclusion reasoning without prescribing tool sequences. Experiments show that CASE moves beyond question-answer imitation toward transferable investigation policies, strengthening evidence-grounded longitudinal reasoning. Under matched evaluation conditions, our Qwen3-VL-8B based agent achieves over 16% and 10% relative improvements in answer accuracy over GPT-5.4 and Claude Opus 4.8. Code will be available at this https URL.
76. 【2609.39563】RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling
链接:https://arxiv.org/abs/2609.39563
作者:Can Zhang,Xiaotian Han,Junyuan Shang,Yuchen Ding,Zhenyu Zhang,Shuohuan Wang,Dianhai Yu,Ruirui Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:encode sampled RGB, sampled RGB frames, Existing video language, RGB frames independently, Existing video
备注:
点击查看摘要
Abstract:Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.
77. 【2609.39553】EffGS: Efficient and High-Fidelity Gaussian Splatting
链接:https://arxiv.org/abs/2609.39553
作者:Changbai Li,Shuo Yang,Yichen Yang,Shuwei Shao,Huobin Tan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:methods suffer severe, general-purpose acceleration methods, acceleration methods suffer, existing general-purpose acceleration, suffer severe rendering
备注:
点击查看摘要
Abstract:3D Gaussian Splatting (3DGS) enables real-time novel view synthesis, but existing general-purpose acceleration methods suffer severe rendering quality degradation when extended to more complex, large-scale scenes. To address this issue, we propose EffGS, a more general acceleration framework that improves training and rendering efficiency while maintaining reconstruction quality comparable to or better than vanilla 3DGS across bounded and large-scale scenes. EffGS combines frequency-aware guidance, localized density control, and adaptive primitive scale modulation. First, an importance scoring mechanism combines pixel-wise reconstruction errors with a difference-of-Gaussians mask scheduled over training to provide stage-dependent spatial guidance. Second, localized densification and pruning restricts density modifications to Gaussians with valid projected footprints in the sampled views. Third, learnable per-Gaussian scale modulation adjusts effective primitive extent during optimization while retaining the Compact Box rasterization rule. Extensive experiments on bounded and large-scale scene datasets demonstrate a favorable balance between reconstruction quality, training time, and primitive count. Component ablations and matched-primitive-budget comparisons further support the effectiveness of the framework.
78. 【2609.39548】Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models
链接:https://arxiv.org/abs/2609.39548
作者:Junjian Li,Xiaolong Liu,Peng Sun,Liantao Wu,Linghan Chen,Yudong Gao,Honglong Chen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
关键词:secure deployment, Backdoor attacks pose, Backdoor, diffusion, Backdoor attacks
备注:
点击查看摘要
Abstract:Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.
79. 【2609.39542】Comparative study of adapting pre-trained models for driving behavior video captioning
链接:https://arxiv.org/abs/2609.39542
作者:Sayak Mallick,Philipp Geiger,Augustin Kelava
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:prompting methods existing, Large Language Model, report examines, domain of autonomous, methods existing
备注:
点击查看摘要
Abstract:This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of different forms of data and this study aims to induce a low dimensional understanding of driving situations into our primary test model SpaceTimeGPT. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate good performance of the full fine tuning framework on some automatic metrics, and in some metrics, it even surpasses the baseline. We also try Low-Rank Adaptation (LoRA) and prompt engineering on VideoLLaVA model and discuss its limitations.
80. 【2609.39527】Lens Flare Removal and Reconstruction
链接:https://arxiv.org/abs/2609.39527
作者:Tarun Yenamandra,Jonathon Luiten,Daniel Cremers,Nathan Matsuda
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:downstream application results, lens flares, flares, lens, significantly reduce
备注: 20 pages, 14 figures. Project page: [this https URL](https://lensflare-3dgs.pages.dev)
点击查看摘要
Abstract:The presence of lens flares in images can significantly reduce the quality of downstream application results for tasks such as 3D scene reconstruction. This is because lens flares are a property of the camera imaging system, and not a part of the underlying scene being modeled. There are previous methods that tackle the removal of small flares focused around a light source. However, existing methods struggle with large flares, such as those that fill the entire image. In this work, we compile a novel dataset for large-flare removal, combining publicly available real-world data with a procedural generation pipeline. We fine-tune a diffusion-based model on our dataset to remove complex, large lens flares. On the other hand, lens flares remain effective artistic tools, widely used in the media. While there are ways to simulate 2D flares, representing and reconstructing lens flares consistently across multiple views has not yet been explored. To achieve this, we introduce a flare representation model that leverages the symmetry of lens flares about the camera's principal point. We propose a computational pipeline to jointly optimize this flare model and a Gaussian splatting model (3DGS). This enables the decomposition of a 3D scene into lens flares and the scene itself, using our flare-removal model. Because the reconstructed flare is explicit and re-renderable, it can be edited and transferred to novel images and new 3D scenes. We evaluate removal on an established benchmark and a new one for large reflective flares, quantify the flare/scene decomposition directly, and show that the pipeline is robust to errors in automatic light-source localization.
81. 【2609.39504】PartiCam: Camera Controlled Video Generation with Reward Guidance
链接:https://arxiv.org/abs/2609.39504
作者:Amine Ouasfi,Runjia Li,Junlin Han,Eric Marchand,Philip H.S. Torr,Adnane Boukhayma
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Particle filtering rooted, improved Camera controlled, filtering rooted method, present PartiCam, filtering rooted
备注:
点击查看摘要
Abstract:We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid the need to construct large camera-annotated datasets by steering pretrained models toward the desired camera motion at test time. This enables the generation of camera-controlled video data that can subsequently be used to train camera-conditioned video diffusion models. Existing sampling-based guidance approaches often suffer from unstable trajectories: they either explore too broadly and fail to respect the target camera motion or collapse early and lose visual diversity over time. We introduce a global-local refinement framework for diffusion reward guidance, enabling accurate and consistent camera control during video generation. Our method builds on Sequential Monte-Carlo (SMC) guidance, but introduces a local refinement stage based on particle filtered resampling. Experiments show large improvements in camera trajectory adherence, reduced drift, and better visual quality, without requiring model retraining.
82. 【2609.39492】Front-to-Back: Benchmarking Vision-Language Models for Asymmetric Cross-View Vehicle Re-Identification
链接:https://arxiv.org/abs/2609.39492
作者:Moseli Mots'oehli,Thulani Babeli
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:rear cameras, South Africa, share a view, target crops, vehicle
备注: Submitted to the ACCV 2026 Workshop on Computer Vision for Developing Countries (CV4DC)
点击查看摘要
Abstract:Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehicle handovers from 20 recording sequences in South Africa. Each example asks a model to match a vehicle highlighted in a front-camera image to the same vehicle among at least three candidates in a later rear-camera image. We evaluate seven zero-shot vision-language models, four image-retrieval baselines, and 25 human participants. Models are tested using full front RGB images, cropped target vehicles, and binary silhouettes. The strongest VLM achieved 76.6 percent Rank-1 accuracy on target crops, compared with 74.0 percent for the frozen SigLIP2 baseline; this difference was not statistically clear. Human participants achieved 94.0 percent accuracy with full images and 92.2 percent with target crops. Under our evaluation setup, enabling reasoning improved accuracy across all three input conditions for every model evaluated in both modes. We also found that VLMs generally performed worse on full scenes than on target crops. These results show that general-purpose VLMs do not yet consistently outperform strong visual retrieval for front-to-rear vehicle matching, while humans remain substantially more reliable.
83. 【2609.39490】OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
链接:https://arxiv.org/abs/2609.39490
作者:Junming Lin,Yuxuan Wang,Zhenxin Lei,Yuxin Liu,Ruixun Liu,Yinsong Yan,Ling Wang,Minghao Han,Yunfei Chu,Shun Lei,Xueyao Zhang,Qize Yang,Jin Xu,Yiwu Zhong
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Recent advances, enabled unified omni-modal, advances have enabled, enabled unified, Recent
备注:
点击查看摘要
Abstract:Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-visual joint reasoning, together with time-stamped clue chains that guide the annotation of thinking process. Besides our benchmark, this engine produces training data OmniReasoning-SFT-112K and OmniReasoning-RL-19K. Finally, we propose an on-policy self-distillation method Modality-Factored Self-Distillation (MFSD). It evaluates each sampled response under modality-specific clue contexts, disentangling the contributions of individual clues and their cross-modal interactions for token-level credit assignment. With our training data and learning method, our model OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points, respectively. Moreover, it delivers substantial gains on general and long-video benchmarks, including Video-MME-v2. We hope our work offers a solid step for facilitating future research in omni-modal joint reasoning.
84. 【2609.39486】From Wrecks to Wisdom: Recovering Crash Mechanics from Real-World Multi-View Photos
链接:https://arxiv.org/abs/2609.39486
作者:Ondřej Valach,Václav Diviš,Ivan Gruber
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Estimating accident mechanics, insurance claim triage, Estimating accident, injury modeling, vehicle-safety analysis
备注:
点击查看摘要
Abstract:Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, crash-severity prediction, and operational workflows such as insurance claim triage. In standard crash records, key metadata such as impact configuration, principal direction of force, and change in velocity ($\Delta V$) may be missing, delayed, or corrupted, while post-crash photographs are widely available and contain rich visual evidence of deformation. We study how much crash-mechanics information can be recovered directly from vehicle photos when structured signals are absent. We formulate crash understanding as supervised prediction from per-case multi-view photo sets. Targets include six Collision Deformation Classification (CDC) descriptors and the longitudinal and lateral components of reconstructed $\Delta V$. Each photo is encoded by a shared visual backbone, and the resulting view-level features are fused into a case-level representation from which target-specific heads predict crash descriptors. Using 15.2k training cases from the Crash Investigation Sampling System, drawn from about 1.5M photos before filtering, together with 1.15k validation and 1.15k test cases, we define an evaluation protocol for vision-based crash descriptor estimation from incomplete multi-view evidence. Post-crash imagery alone provides usable signal for several non-trivial crash-mechanics descriptors, while weakly observable and long-tailed targets remain challenging. Within the compared training regimes, the selected joint-training recipe reduces mean absolute angular error for principal direction of force from 20.1 to 14.05 degrees and longitudinal $\Delta V$ MAE from 8.04 to 7.45 km/h. Our work provides a reference point for future multimodal fusion with structured crash metadata.
85. 【2609.39467】DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion
链接:https://arxiv.org/abs/2609.39467
作者:ZiAn Wang,MingZhe Liu,Chaoyi Guo,ChangChun Li,Fangming Gu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Pedestrian detection plays, autonomous driving, public safety, detection plays, plays a crucial
备注: Accepted at WISE 2026
点击查看摘要
Abstract:Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweight detectors struggle to balance accuracy and efficiency while often neglecting quality-aware feature modeling and consistency between classification and localization, leading to unstable performance under crowded conditions. To address these issues, we propose DensePed-Lite, a unified framework built on a single principle: under occlusion the network should adapt its behavior to the quality of what it observes rather than assume complete information. This principle is realized at three points where occlusion does the most damage: unreliable confidence scoring (UQE), fragmented spatial coverage (MPSC), and incoherent multi-scale fusion (CTDM). The three mechanisms reinforce one another instead of acting in isolation, all without significantly increasing complexity. Experiments on CityPersons and CrowdHuman validate that DensePed-Lite achieves a superior accuracy-efficiency trade-off compared with recent state-of-the-art lightweight methods, making it suitable for real-time deployment in dense pedestrian scenarios.
86. 【2609.39456】Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback
链接:https://arxiv.org/abs/2609.39456
作者:Ho-min Park,Byungkon Kang
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:mutual feedback architecture, coupled embeddings, embedding reflects, possibly different modalities, mutual feedback
备注: 20 pages
点击查看摘要
Abstract:This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea leads to a mutual feedback architecture consisting of two components whose outputs are fed back into the other. The final output of this model is defined as the fixed point of this interaction. We provide theoretical analysis that offers interpretation of this model as well as design choices to prevent failure cases. We show the benefits of MEQ through classification and visual grounding tasks spanning various datasets. Quantitatively, our model outperforms or shows competitive performance on concatenation-based multimodal classification problems. Qualitatively, the proposed interactive mechanism allows the model to progressively refine the visual grounding when paired with complementary modality, thus demonstrating the power of mutual feedback under such settings.
87. 【2609.39451】ResARC: Residual-Aware AutoRegressive Coding for Ultra-Low Bitrate Image Compression
链接:https://arxiv.org/abs/2609.39451
作者:Qin Yan,Ruixiao Dong,Yutao Xie,Li Li,Ying Chen,Kai Li,Daowen Li,Houqiang Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Progressive autoregressive image, remaining suffix tokens, image codecs provide, quantizing continuous latents, Progressive autoregressive
备注:
点击查看摘要
Abstract:Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens, transmitting coarse-to-fine prefix tokens and generating the remaining suffix tokens at the decoder. However, their reconstruction quality is fundamentally limited by two residuals introduced along this pipeline: the quantization residual, arising from information loss during discrete tokenization, and the generation residual, resulting from imperfect autoregressive generation of the suffix tokens. To address these limitations, we introduce ResARC, a residual-aware autoregressive codec that explicitly compensates for both residuals at the decoder. Specifically, we generate the quantization residual with a diffusion transformer conditioned on the autoregressive decoding context, while requiring no additional side information. In parallel, we compute the generation residual at the encoder and employ a learned Generation Residual Codec to efficiently compress and transmit it for decoder-side correction. The recovered residuals are then integrated with the reconstructed latent representation and decoded through an adapted VAE decoder. Extensive experiments demonstrate that ResARC achieves competitive perceptual similarity while substantially improving distributional fidelity over leading generative codecs across the ultra-low bitrate regime. Code and models will be released soon.
88. 【2609.39441】CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models
链接:https://arxiv.org/abs/2609.39441
作者:Shu Yu,Chaochao Lu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Online reinforcement learning, Online reinforcement, reinforcement learning, extended to flow, flow matching
备注: Project page: [this https URL](https://opencausalab.github.io/CAST)
点击查看摘要
Abstract:Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.
89. 【2609.39429】owards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
链接:https://arxiv.org/abs/2609.39429
作者:Gonzalo Esteban Mosquera Rojas,Sebastian R. van der Voort,Carolin M. Pirkl,Sandeep Kaushik,Marion Smits,Stefan Klein
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:high-stakes medical image, medical image analysis, Carlo Deep Ensembles, Monte Carlo Deep, Uncertainty Quantification
备注: Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) [this https URL](https://melba-journal.org/2026:033)
点击查看摘要
Abstract:Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
90. 【2609.39427】PCB-MC: Missing Component Analysis in Printed Circuit Boards
链接:https://arxiv.org/abs/2609.39427
作者:Betsy Villa Brochero,Ian Gibson,Estefania Talavera
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Detecting missing components, Detecting missing, printed circuit boards, conventional object detection, differs fundamentally
备注: Preprint
点击查看摘要
Abstract:Detecting missing components on printed circuit boards (PCBs) differs fundamentally from conventional object detection, as the model must localize components that are not present. We introduce PCB-MC, a curated dataset for missing component detection with footprint level annotations built on top of the RF100 dataset. The dataset contains 197 distinct board types, each corresponding to a unique PCB design, with multiple augmented samples per type. We also provide benchmark results on PCB-MC by evaluating a diverse set of supervised and unsupervised methods. To ensure fair evaluation, we propose board type aware cross validation splits that prevent layout leakage between training and test sets. Supervised models showcase high false negative rates on unseen board designs, and unsupervised anomaly detection methods fail entirely due to the lack of spatial alignment with a board specific reference. These results confirm that missing component detection on diverse PCB layouts remains an open challenge. We release PCB-MC and all training protocols to support reproducible research on structural absence detection in industrial inspection.
91. 【2609.39388】UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision
链接:https://arxiv.org/abs/2609.39388
作者:Wei Xue,Keliang Liu,Mingzhang Cui,Jinhua Xie,Jinjie Wei,Jianan Hou,Jingcheng Lu,Lintao Wang,Kaixiang Qiu,Yizhou Liu,Xinghai Ye,Jinghang Han,Mingcheng Li,Jie Gu,Shunli Wang,Lihua Zhang,Dingkang Yang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:reliable object interaction, manipulation requires precise, requires precise navigation, object interaction, requires precise
备注: UniWAM Technical Report
点击查看摘要
Abstract:Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1\% in position error and 44.0\% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.
92. 【2609.39380】InfoAgent: Traceable Generation and Repair of Evidence-Grounded Infographics
链接:https://arxiv.org/abs/2609.39380
作者:Yifan Li,Tong Li,Qi Zeng,Lishuai Gao,Ruwei Pan,Cong Wei,Shaohua Kevin Zhou,Zhuoliang Kang,Xiaoming Wei
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Reliable infographic generation, generation requires facts, infographic generation requires, Reliable infographic, rendering and revision
备注:
点击查看摘要
Abstract:Reliable infographic generation requires facts, symbols, and visual relations to remain consistent through rendering and revision. Correcting one element also requires tracking its supporting evidence and the dependencies affected by the change. We present \textbf{InfoAgent}, a training-free framework for \emph{evidence-bound visual-symbolic program synthesis}. Its Infographic Visual Description (IVD) records factual payloads, evidence provenance, execution routes, and verification obligations in a typed dependency graph. Retrieved design priors guide compilation, and layered execution combines raster synthesis with editable symbolic and binding objects while retaining their traces. Dependency-aware repair localizes corrections, rechecks affected dependencies, and requires protected obligations to remain satisfied under the declared checkers. Unresolved obligations remain explicit. On IGenBench, InfoAgent achieves 93.0 Q-ACC and 59.0 I-ACC. We also introduce InfoGraphicBench-Evidence, where complete-checklist pass rates on 200 test requests increase from 21.5\% for Same-IVD Prompt to 23.5\% for the initial layered output and 28.5\% after repair, using the same evidence and initial IVD. On 120 audited repair cases, localized repair edits 12.4\% of the canvas on average, compared with 67.3\% for global regeneration.
93. 【2609.39378】EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
链接:https://arxiv.org/abs/2609.39378
作者:Shulin Tian,Junsu Kim,Shuai Liu,Hao Li,Yujiao Shen,Sihan Li,Zhe Yang,Yeongon Kim,Feiyu Li,Jialin Wu,Yichi Zhang,Wenhui Wang,Runmao Yao,Yuhao Dong,Zhaoxi Chen,Fangzhou Hong,Antonino Furnari,Jingkang Yang,Hongyuan Zhu,Ziwei Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:tracking evolving object, agents to act, act under physical, physical constraints, constraints while tracking
备注: 32 pages, 7 figures. Project page: [this https URL](https://ropedia.github.io/egotools)
点击查看摘要
Abstract:Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
94. 【2609.39375】Beyond the Current Scene: Event-Referential Grasping with Active View Selection
链接:https://arxiv.org/abs/2609.39375
作者:Hyunjoon Lee,Haebeom Jung,Eunsung Cha,Daeun Lee,Yu-Chiang Frank Wang,Jaesung Choe,Jaesik Park
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:observes people interacting, observes people, people interacting, refer back, target
备注: Project page: [this https URL](https://www.haebeom.com/BeyondCSe/)
点击查看摘要
Abstract:A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act. We present BeyondSCe, a zero-shot robotic grasping system for this event-referential setting. Given the event history and the current scene, the system identifies the requested object or part and localizes it for grasping. If the target is occluded, it combines an event prior recovered from the history with current scene geometry to select camera viewpoints likely to reveal the target. The system uses pretrained models without additional task-specific training. In real-robot experiments with a single wrist-mounted RGB-D camera, it achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively, compared with 40% and 55% for the strongest baseline in each condition. On four additional scenes with heavy occlusion, it increases grasp success rates from 75% to 95% while reducing the mean number of views from 3.35 to 2.20, compared with an active-perception baseline given the target's ground-truth 3D bounding box.
95. 【2609.39366】COBICount: Separating Object and Background Responses for Remote Sensing Object Counting Without Training on Target Data
链接:https://arxiv.org/abs/2609.39366
作者:Junjing Zheng,Zhiyi Zhou,Ningrui Yang,Hongying Meng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Remote sensing object, Remote sensing, Remote, sensing object counting, object counting estimates
备注: 19 pages, 7 figures
点击查看摘要
Abstract:Remote sensing object counting estimates how many buildings, vehicles, or ships appear in overhead images. Most supervised counters predict a density map, whose sum gives the object count, and assume similar categories, sizes, and backgrounds. Applying them across regions, sensors, or categories often requires target data or further training, which may be costly or unavailable. We study source-only counting. Training for the counting task and model selection use one group of images that shares an object category and similar imaging conditions, with one point marking each object. Target images and information remain unavailable until the model is fixed. This reduces data preparation but makes transfer harder. A model trained on one source may place high density values, called responses, on real objects and repeated background structures. Road edges, parking grids, roof boundaries, and water boundaries may then be counted as objects, creating candidate origin ambiguity. COBICount separates response generation, acceptance, and background suppression. Candidate Evidence (CE) generates possible responses. Candidate Acceptance (CA) keeps compact responses centered on objects. Bias Isolation (BI) reduces responses associated with repeated background structures. Their outputs form the final density map. Trained on RSOC Building and evaluated directly on DOTA Large Vehicle, Small Vehicle, and Ship, COBICount achieves the lowest mean absolute error (MAE) averaged over the target domains among the compared methods, 174.132. It uses 5.07 million parameters and 17.41 billion floating point operations for a 512x512 input. COBICount improves transfer without target data or training for each target. The code will be available at: this https URL.
96. 【2609.39363】Rethinking Multi-Image Re-Representation in Multi-Image Understanding
链接:https://arxiv.org/abs/2609.39363
作者:Gengyuan Zhang,Xiao Han,Xinyu Xie,Tong Liu,Volker Tresp
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:understanding requires MLLMs, Multi-image understanding requires, organise visual evidence, visual evidence distributed, visual
备注: 27 pages, 7 figures, 9 tables
点击查看摘要
Abstract:Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at this https URL.
97. 【2609.39335】xTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning
链接:https://arxiv.org/abs/2609.39335
作者:Zijing Qin,Jun Zhou,Ruicheng Zhang,Jiaqi Hou,Zunnan Xu,Ronghui Li,Zhenyu Xie,Xiu Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:attracted increasing attention, increasing attention due, intelligent e-commerce, Video virtual try-on, attracted increasing
备注:
点击查看摘要
Abstract:Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations can be attributed to two main factors: (1) the insufficient utilization of rich garment reference information, and (2) the lack of explicit positional modeling between garment and video representations during cross-modal interaction, which weakens fine-grained local correspondence. To address these issues, we propose TexTailor, a high-fidelity video virtual try-on framework built upon a pretrained video Diffusion Transformer. Specifically, we introduce a timestep-adaptive modulation mechanism to dynamically adjust garment visual representations throughout denoising. We further develop a frame-aligned positional encoding strategy to strengthen garment-to-video correspondence, together with a multi-source injection design that reduces interference among heterogeneous conditions. Extensive experiments on multiple video virtual try-on benchmarks, including the high-resolution Eevee dataset, demonstrate that TexTailor achieves competitive performance in garment detail preservation, temporal consistency, and overall video quality.
98. 【2609.39324】MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies
链接:https://arxiv.org/abs/2609.39324
作者:Jingqiu Wang,Yan Wang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:incorporated world models, recently incorporated world, provide richer dynamic, sparse action labels, richer dynamic supervision
备注: 4 pages + 1 page references, 3 figures, 2 tables. Code: [this https URL](https://github.com/autu-mn/MotionWeave)
点击查看摘要
Abstract:Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared global action features may fail to establish timestep-specific correspondence between actions and local visual changes. To address this issue, we propose MotionWeave, a motion-centric future-dynamics framework for action-chunk prediction with two modules: the Action-Induced Motion Grounder (AIMG) and the Horizon Residual Composer (HRC). Specifically, AIMG conditions on action and proprioceptive representations to construct horizon-specific queries that localize interaction regions associated with each future action timestep from current visual tokens. HRC extracts differences between interaction representations at adjacent horizons, encodes them as temporal motion cues, and injects them into action tokens through a gated residual. During training, robot-arm masks rendered from future frames are used to construct KL-based motion-grounding supervision, while inference uses only the current observation. On six MetaWorld tasks, MotionWeave achieves a 75.3% average success rate, an absolute gain of 8.6% over {\pi}0 (66.7%), especially on sustained-interaction tasks. Our code is available at this https URL.
99. 【2609.39315】Rethinking Generative Image Compression at Extremely Low Bitrates
链接:https://arxiv.org/abs/2609.39315
作者:Tianyu Zhang,Zhaoyang Jia,Houqiang Li,Dong Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generative image compression, compression produces visually, produces visually plausible, remains largely unexplored, visually plausible reconstructions
备注:
点击查看摘要
Abstract:Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specific detail, they produce malformed or unrecognizable content. Our analysis identifies two factors. As the bitrate decreases, reconstruction losses increasingly conflict with semantic objectives on gradients and visual results, while pixel-space and reconstruction-oriented VAE diffusion models become less efficient on semantic preservation. Guided by these findings, we introduce RAE-CoD, a compression-oriented diffusion (CoD) built in a representation autoencoder (RAE) space with direct alignment between compressed and source representations, preserving recognizable, naturally structured content for a $256\times256$ image with as few as 16 bits. We evaluate this framework using five vision foundation models (VFM) and a blinded vision-language model protocol. On MSCOCO-30K, RAE-CoD stands out from all evaluation. At 0.001-0.008 bpp, it reduces relative VFM feature MSE and Fréchet Distance ratio by at least 25.7% and 69.1% over the best competitors. Meanwhile, semantic recognizability and quality of the reconstructions remain nearly constant while source consistency falls smoothly, replacing abrupt semantic collapse with a graceful transition toward unconditional generation. Code will be released at this https URL.
100. 【2609.39300】BMASH: Ball-Motion-Aware Soccer Header Spotting
链接:https://arxiv.org/abs/2609.39300
作者:Ahmed Endris Hasen,Muhammad Shahzad Khan,Nikolaos Passalis,Jenni Raitoharju
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent advances, sports videos increasingly, Video Swin, made broadcast sports, broadcast sports videos
备注: 9 pages, 4 figures, MMsports
点击查看摘要
Abstract:Recent advances in computer vision have made broadcast sports videos increasingly useful for event analysis, performance assessment, and player-safety applications. In soccer, however, header spotting remains a challenging problem due to the subtle and short-lived nature of header events. This paper focuses on soccer header spotting: identifying moments in broadcast videos where the ball contacts a player's head. We first adapt and evaluate Video Swin as a strong action-recognition baseline for this task, and then introduce BMASH, a ball-motion-aware fusion framework that integrates detector-derived ball features. BMASH combines Video Swin action representations with ball-presence and motion features from frame-level soccer-ball detection, integrating player-action context with ball dynamics to distinguish headers from visually similar events. We evaluate BMASH using game-level splits with separate test matches and rotating validation folds, considering both centered-window classification and continuous full-video spotting. Results show that Video Swin provides a strong baseline for header spotting, while BMASH improves clip-level AP and ROC-AUC over the corresponding Video Swin baseline. In continuous full-video spotting, BMASH achieves a comparable event-level F1-performance with a different precision--recall trade-off.
101. 【2609.39273】MegaAvatar: Controllable Talking Avatar Generation
链接:https://arxiv.org/abs/2609.39273
作者:Junyao Gao,Sibo Liu,Weidong Zhang,Cairong Zhao,Jun Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:generation framework built, report presents, framework built, built on top, avatar generation framework
备注: 7 pages, 6 figures
点击查看摘要
Abstract:This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with previous talking-avatar methods that mainly rely on audio or reference-image conditioning, we introduce additional SMPL-X-derived 3D guidance, enabling global control over body pose and head motion. Specifically, we render the driving SMPL-X sequence into dense mesh frames and encode them with a lightweight 3D convolutional encoder, whose outputs are injected into the latent tokens to provide overall motion control. Furthermore, we extend Wan2.2-TI2V-5B with additional audio and face cross-attention modules to enable fine-grained expression control and preserve the input identity, respectively. In addition, we implement an audio-to-SMPL-X model to predict an SMPL-X sequence conditioned on the reference image and input audio, allowing MegaAvatar to support audio-driven inference without user-provided SMPL-X frames. Experiments show that MegaAvatar achieves high-quality talking avatar generation with controllable body and head motion, speech-synchronized facial expressions, and consistent identity preservation. MegaAvatar also supports inference with flexible resolutions and video lengths. Codes, dataset, models will be avaliable in this https URL
102. 【2609.39266】PLRS-IC: A Dual-Calibration Framework for Chest X-Ray Vision-Language Alignment
链接:https://arxiv.org/abs/2609.39266
作者:Qixing Zhao,Jinpeng Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Fine-grained vision-language alignment, chest radiography enables, radiography enables zero-shot, task-specific annotations, Fine-grained vision-language
备注: 9 pages, 4 figures, 5 tables
点击查看摘要
Abstract:Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotations. However, this alignment is fundamentally hindered by two intertwined sources of ambiguity: projection-induced visual mismatch and patient-agnostic semantic overlap. First, at the local feature level, frontal and lateral radiographs exhibit distinct appearances for the same clinical finding, rendering a shared patch-text similarity geometry inherently suboptimal. Compounding this visual ambiguity is a semantic mismatch during global contrastive optimization, where instance-level objectives penalize cross-patient pairs as strict negatives even when they share identical positive clinical concepts. To address this dual ambiguity, we propose PLRS-IC, a unified dual-calibration framework for chest X-ray representation learning. At the local alignment stage, Projection-Conditioned Low-Rank Residual Similarity (PLRS) dynamically adapts patch-text matching to projection-specific manifolds using a bounded, parameter-efficient low-rank residual. At the global optimization stage, Information-Content-Calibrated Soft False-Negative Suppression (IC-SFNS) leverages a corpus-derived information-theoretic prior to soften the penalty of semantically overlapping negatives without altering original contrastive assignments. Extensive experiments across nine public zero-shot benchmark settings demonstrate that our framework yields consistent improvements in classification, grounding, and segmentation, validating the necessity of dual-calibration in medical vision-language pre-training.
103. 【2609.39265】Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation
链接:https://arxiv.org/abs/2609.39265
作者:Ziqi Zhou,Yifan Hu,Yufei Song,Haowen Jiang,Xianlong Wang,Shengshan Hu,Dezhong Yao,Leo Yu Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Promptable Concept Segmentation, segmentation foundation models, concept segmentation, extends promptable segmentation, segmentation
备注: Accepted by NeurIPS 2026
点击查看摘要
Abstract:The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial examples, the robustness of SAM3 under the concept segmentation paradigm remains unexplored. In addition, existing adversarial attacks on SAM-series models exhibit limited cross-prompt transferability. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack for Promptable Concept Segmentation (PCS), including a min-max prompt optimization strategy, a global-local perception deception attack, and a temporal transition deviation attack. Specifically, we first identify the hardest-to-attack prompts via min-max bilevel optimization. In the inner maximization, we enhance diversity over candidate point, box, and text prompts. In the outer minimization, we select prompts with the highest responses based on the confidence scores output by the detector. Given the selected prompts, we apply the perception deception attack to minimize both global and local existence probabilities under joint prompting and employ the temporal memory misalignment attack to maximize inter-frame semantic inconsistency and corrupt memory pointers. Extensive experiments on four benchmark datasets show that a single universal adversarial perturbation (UAP) generated by our method generalizes across frames from different videos and achieves strong attack performance under point, box, and text prompts. In particular, it reduces the average mIoU of various PCS models on the SA-CO dataset to below 5% under text prompts, demonstrating strong attack ability.
104. 【2609.39235】he Planning Limits of Latent World Models
链接:https://arxiv.org/abs/2609.39235
作者:Ali Alrasheed,Basim Azam,Naveed Akhtar
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:plan complex behaviours, physical world evolves, World models offer, offer a promising, robots understand
备注:
点击查看摘要
Abstract:World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this question using action-conditioned predictors built on five frozen self-supervised visual backbones: V-JEPA 2, V-JEPA 2.1, VideoMAEv2, VideoPrism, and DINOv2. We use frozen backbones to test representations intended to transfer across environments. We evaluate these models on diverse Meta-World manipulation tasks and real-robot interactions from BridgeData V2. We find that a world model guides action selection reliably only when the goal lies within, or slightly beyond, the trajectory it imagines during planning. With five-step rollouts, the length the predictor was trained on, the world model ranks actions reliably only for targets five to ten control steps ahead, whereas task goals lie 16 to 53 steps away. Neither an 81-fold larger predictor nor longer-rollout training extends this range; the encoder affects both range and closed-loop success, with V-JEPA 2.1 performing most consistently. More fundamentally, the limit persists under perfect prediction: using the real simulator, success falls from 92% to 41% as the target moves from five to twenty steps ahead of a five-step rollout. Planning therefore requires either longer imagined trajectories or closer subgoals. For distant goals, pure imagination succeeds in 23% of episodes, planning with feedback (MPC) raises success to 30%, imagining as far as the goal to 47%, and nearby expert subgoals to 76%. Used within its plannable range, a world model can also improve a vision-language-action (VLA) policy: choosing among eight actions the VLA proposes raises its success from 65% to 77% across 16 different tasks.
105. 【2609.39227】Emergent Multi-View Geometry Through Self-Distillation
链接:https://arxiv.org/abs/2609.39227
作者:David Nordström,Thibaut Loiseau,Vincent Lepetit,Michael Felsberg,Guillaume Bourmaud,Fredrik Kahl
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Henri Poincaré argued, Henri Poincaré, century ago, notion of space, RGB reconstruction
备注:
点击查看摘要
Abstract:Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.
106. 【2609.39222】DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
链接:https://arxiv.org/abs/2609.39222
作者:Xu Huang,Ye Huang,Zijun Liao,Yuwei Niu,Xiaojie Li,Menghan Zhou,De Wen Soh,Xiaotong Li,Daquan Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:essential for scaling, image generative models, DC-SAE, scaling latent image, latent image generative
备注: 16 pages, 5 figures. Project page: [this https URL](https://dagroup-pku.github.io/DCSAE)
点击查看摘要
Abstract:High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a $1.6$B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at $1024\times1024$ resolution.
107. 【2609.39195】Uruqi: Learning Spatial Cognition from Visual Experience
链接:https://arxiv.org/abs/2609.39195
作者:Shichao Li,Meiqi Wang,Fei Su,Zhicheng Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:embodied agent moves, intelligence requires maintaining, requires maintaining, maintaining a coherent, coherent understanding
备注: 23 pages
点击查看摘要
Abstract:Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQI$_{\mathrm{Syn}}$-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQI$_{\mathrm{Syn}}$-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.
108. 【2609.39184】Fiber-Resolved Microstructure Quantification from Multi-Shell Diffusion MRI using Detection Transformers
链接:https://arxiv.org/abs/2609.39184
作者:Sebastian Endt,Marcus Wirth,Johannes Reinhold Schlund,Marion Irene Menzel
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Medical Physics (physics.med-ph)
关键词:white matter tissue, single fiber direction, resolve fiber orientations, diffusion MRI, fiber direction
备注: 12 pages, 4 figures, 1 table; Accepted at MICCAI 2026 Workshop CDMRI; Code: [this https URL](https://github.com/Marcus02W/Diffusion-DETR)
点击查看摘要
Abstract:Fiber orientation and compartmental microstructure are central to the characterization of white matter tissue in diffusion MRI, yet existing methods either resolve fiber orientations without quantifying microstructure, or quantify microstructure while assuming a fixed number of compartments and a single fiber direction. Nonparametric approaches that recover both require tensor-valued diffusion encoding and computationally expensive Monte-Carlo inversion of an ill-posed inverse Laplace transform. We propose to reframe this problem as an object detection-like task, adopting the Detection Transformer (DETR) architecture to jointly predict mean diffusivity (MD), fractional anisotropy (FA), main fiber direction, and signal fraction for a variable number of compartments per voxel from standard multi-shell diffusion MRI with linear encoding. Hungarian matching during training resolves permutation invariance across compartments. We introduce mean Average Precision as a reproducible benchmark metric. Evaluated on synthetic test data with up to five compartments per voxel, our model achieves $R^2=0.95$ for MD, $R^2=0.88$ for FA, and a median angular error of 4.2°, with performance scaling naturally with compartmental signal fraction.
109. 【2609.39183】Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
链接:https://arxiv.org/abs/2609.39183
作者:Pengzhan Sun,Shiu-hong Kao,Shijie Li,Yongyi Su,Junbin Xiao,Arjun Reddy Akula,Angela Yao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Visual Intention Grounding, paper studies, textbf, Intention Grounding, Visual Intention
备注:
点击查看摘要
Abstract:This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.
110. 【2609.39182】MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models
链接:https://arxiv.org/abs/2609.39182
作者:Ali J Alrasheed,Aryan Yazdan Parast,Basim Azam,James Bailey,Naveed Akhtar
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:computer vision, World Models, major frontier, frontier in computer, latent World Models
备注: Accepted at DICTA 2026 (International Conference on Digital Image Computing: Techniques and Applications). Camera-ready version
点击查看摘要
Abstract:World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.
111. 【2609.39178】Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
链接:https://arxiv.org/abs/2609.39178
作者:Songhua Yang,Ziyu Liu,Yuanwei Liu,Xuetao Li,Xuanye Fei,He Huang,Zheng Wang,Miao Li
类目:Robotics (cs.RO); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:integrating visual perception, seamlessly integrating visual, language understanding, revolutionized robotic manipulation, visual perception
备注: Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA 2026), Vienna, Austria. 8 pages. (c) 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes
点击查看摘要
Abstract:Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since these models are designed to interact directly with the physical world and humans, their security is critical, and even small vulnerabilities can lead to catastrophic failures. In this work, we propose the Universal Adversarial Object, a sphere with optimized surface texture that significantly degrades task success rates when placed within the robot's field of view. Specifically, our approach introduces a multi-level attack framework that jointly disrupts trajectory planning, task execution, and action control. We validate our method in both simulated and real-world robotic settings. Experimental results demonstrate that the adversarial object reduces the average task success rates by 31.2%-39.9% for two representative VLA models (Pi0 and RDT), with success rates dropping to near zero in complex scenarios. Index Terms--Vision-Language-Action models, adversarial attack, robotic security, universal adversarial object
112. 【2609.39157】ripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation
链接:https://arxiv.org/abs/2609.39157
作者:Songhe Wang,Lifu Wei,Shuolin Xu,Charles A. Kamhoua,David Miller
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Video object removal, difficult editing challenge, uniquely difficult editing, object removal presents, Video object
备注: 24 pages, 23 figures, including appendices
点击查看摘要
Abstract:Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generate, the model must infer and reconstruct a highly specific occluded background entirely from the surrounding context. Existing training-free methods struggle with this because their editing mechanisms act primarily as localized erasers. They fail to actively synthesize the missing background details and often leave behind ghosting artifacts. To solve this, we propose TripleFlow, a training-free framework that tightly couples erasure and generation. It coordinates a source flow, a residual flow, and a synthesis flow throughout the entire process. By reusing a single target prediction, the residual flow isolates and suppresses the object, while the synthesis flow independently reconstructs the occluded background. Crucially, TripleFlow injects this newly synthesized background back into the editing trajectory at every step. This continuous feedback loop ensures that the generated structures actively guide the removal process, achieving seamless completion that is spatiotemporally consistent with the unedited scene. Extensive evaluations across five challenging benchmarks demonstrate that TripleFlow establishes a new state-of-the-art, significantly outperforming existing baselines in both reconstruction fidelity and temporal consistency.
113. 【2609.39150】OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning
链接:https://arxiv.org/abs/2609.39150
作者:Xingming Shui,Dapeng Chen,Bowei Liu,Jingqi Tian,Minfu Li,Kun Yi,Jiapeng Hong,Yansong Tang
类目:Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
关键词:Omni-modal large language, large language models, language models deployed, real-world environments encounter, environments encounter external
备注:
点击查看摘要
Abstract:Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The challenge is to resist acoustic interference while preserving useful audio evidence. On-policy distillation provides dense teacher feedback on student-generated responses, but uniform token weighting does not explicitly prioritize positions affected by acoustic interference. We introduce OP-CAD (On-Policy Clean-Audio Distillation), a curriculum-based privileged self-distillation framework for robust audio-visual understanding. Training progresses from mild to severe environmental noise and competing speech, with selective token-level supervision at each stage. The student generates responses from corrupted audio-visual input, while a frozen teacher uses clean audio and the verified answer to supervise the same response prefixes. To allocate this supervision, OP-CAD compares teacher predictions under clean, corrupted, and visual-only contexts without revealing the answer. These matched comparisons measure sensitivity to audio removal and corruption; a bounded weighting rule emphasizes positions identified by either signal while retaining supervision throughout the response. OP-CAD outperforms the compared methods across all evaluated noise conditions. Paired analyses further show improved preservation of clean-correct answers under strong interference, with no observed aggregate clean-accuracy penalty. These results demonstrate the value of directing clean-teacher supervision toward acoustically sensitive predictions for robust audio-visual reasoning.
114. 【2609.39147】MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation
链接:https://arxiv.org/abs/2609.39147
作者:Ruiqi Li,Xuanyi Liu,Sijia Li,Haofeng Wang,Yuxin Liu,Feng Xie,Songchao Tan,Shiqi Wang,Hanwei Zhu,Yizong Wang,Chuanmin Jia,Siwei Ma
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:states remains unexplored, achieve visual realism, remains unexplored, Current, mental states remains
备注: 7 pages, 6 figures. Accepted by ACM Multimedia 2026
点击查看摘要
Abstract:Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-conditioned video generation. We formalize this as mental-state-to-behavior reasoning, where models generate actions from a world state and latent variables without explicit action prompts. MindWorldBench utilizes Zero-Action Prompting and a counterfactual design with 744 prompts to isolate the causal effects of mental states. An automated pipeline evaluates video quality, commonsense plausibility, and mental-state consistency. Evaluations of 11 models show that despite visual fidelity and physical reasoning, models fail to align behaviors with latent mental states. We identify a failure mode, termed Omniscient Bias, where models default to the objective world state rather than human's subjective belief. These results demonstrate a disconnect between visual generation and cognitive reasoning, suggesting a need for explicit mental-state modeling in video generation systems. Project website: this https URL
115. 【2609.39142】When Can Text Replace Vision? Structural Bottlenecks in Diagram Reasoning
链接:https://arxiv.org/abs/2609.39142
作者:Yunbei Zhang,Janet Wang,Jihun Hamm,Chandan K Reddy
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:structured text replace, text replace vision, diagram reasoning, structure, structured text
备注: 33 pages, 16 figures. Code: [this https URL](https://github.com/yunbeizhang/text-for-vision)
点击查看摘要
Abstract:Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information the question needs, or because the solver fails to use information that is present. We introduce a diagnostic protocol to distinguish these explanations. Using the same solver model and generation settings, we compare three input conditions: the original image, question-blind structure extracted by a vision-language model, or gold structure derived from the diagram source. Validity-triggered recovery tests truncation and schema failure, question-relevant fidelity measures preservation of answer-critical structure, and matched edge interventions test the effect of error location. On a reserved holdout of 240 public FlowGen diagrams, evaluated under a frozen protocol, gold structure reaches 87% accuracy while direct vision and learned text both remain below 30%. The aggregate comparison includes source-derived relation labels that may not be printed in the image and uses different learned and gold graph encodings, so it does not isolate extraction error alone. Retrying only invalid extractions makes nearly every public representation schema-valid yet leaves accuracy essentially unchanged. The public learned-text deficit relative to gold more than doubles with structural difficulty. Question-relevant topology predicts correctness better than whole-graph topology. In an exposed intervention study, a single answer-relevant edge edit reduces the primary solver's original-answer accuracy to near zero, while matched irrelevant edits largely preserve it. Supplied structure requires fewer solving tokens than vision, but learned acquisition removes this advantage at single use. These comparisons motivate evaluating acquired text by the answer-relevant evidence it preserves and by the solver's ability to use that representation.
116. 【2609.39135】Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing
链接:https://arxiv.org/abs/2609.39135
作者:Shenxiang Zeng,Chen Yang,Peiyao Chen,Guohui Zhang,Jiansheng Fan,Chen Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:requires inferring latent, video requires inferring, inferring latent physical, latent physical properties, World Modeling
备注: 21 pages, 9 figures, 6 tables
点击查看摘要
Abstract:Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while predefined tool pipelines rely on task- and domain-specific priors that limit generalization across materials, dynamics, and reasoning tasks. We introduce Asking the World (ATW), a generalist agent that constructs and interrogates task-relevant executable worlds through two adaptive stages: World Modeling calibrates a world from video, while World Probing queries, simulates, and intervenes on it to obtain question-relevant evidence. Rather than prescribing the operations in either stage, ATW determines how to model and probe according to the scene and question. We develop PolyWorld Engine, a lightweight and highly programmable Warp-based multiphysics simulator for constructing and probing worlds with rigid bodies, soft bodies, cloth, ropes, fluids, and their coupled interactions. CEM-based system identification recovers task-relevant dynamics during World Modeling. The resulting world becomes an active workspace for question-directed physical experiments rather than a predetermined downstream tool. We evaluate ATW on CLEVRER, ContPhy, and three real-world scenarios. Using Gemini-3-Flash as its base VLM, ATW achieves 80.82% overall per-question accuracy on CLEVRER, improving direct Gemini-3-Flash by 46.50 points, GPT-5.5 by 13.58 points, and PhysMind by 8.27 points. On ContPhy, it reaches 70.56% overall accuracy, surpassing Gemini-3-Flash by 28.10 points and GPT-5.5 by 3.53 points. Across the three real-world scenarios, ATW achieves 71.67% accuracy, 28.33 points above GPT-5.5. These results establish agentic world modeling and probing as an effective, execution-grounded approach to generalist physical reasoning.
117. 【2609.39134】Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models
链接:https://arxiv.org/abs/2609.39134
作者:Shilinlu Yan,Bowen Chen,Yuechen Zhang,Zhenhong Zhou,Li Sun,Sen Su
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Visual-token compression improves, large vision-language models, Visual-token compression, full-token evaluation misses, vision-language models
备注: 29 pages including references and appendices, 10 figures. Submitted to ICLR 2027
点击查看摘要
Abstract:Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such failures is challenging because perturbing token importance can also damage the visual content needed for full-token inference. We propose Feature-Aware Token Attack (FATA), which couples attention suppression with cosine-based feature preservation on a fixed set of salient clean-image tokens. In the primary LLaVA-1.5-7B setting, FATA uses only vision-encoder gradients, without access to the deployed compressor, token budget, or downstream task. Across four visually dependent task subsets and four compressors under a controlled reconstruction protocol, FATA achieves SR = 96.3% full-token accuracy retention and CBR = 22.1% conditional blinding, compared with 89.8% and 15.7% for CAA. Ablations support the role of both objectives in balancing compressed-path failure against full-token preservation. FATA also has the lowest measured detection rate among four attacks across three evaluated detectors at a 5% false-positive rate. These findings motivate assessing adversarial robustness jointly across full-token and compressed inference.
118. 【2609.39132】Uncertainty-Aware Consistency Distillation for Few-Step Video Generation
链接:https://arxiv.org/abs/2609.39132
作者:Lingyu Liu,Yaxiong Wang,Li Zhu,Zhedong Zheng
类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:multi-step video generator, incurring substantial latency, typically requires tens, study few-step video, few-step video generation
备注:
点击查看摘要
Abstract:We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.
119. 【2609.39130】Perceptual Color Difference Modeling Using Machine Learning and Human Similarity Judgments
链接:https://arxiv.org/abs/2609.39130
作者:Elnara Kadyrgali,Muragul Muratbekova,Adilet Yerkin,Nuray Toganas,Ayan Igali,Malika Ziyada,Aruzhan Burambekova,Jamaladdin Hasanov,Pakizar Shamoi
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
关键词:Accurate assessment, quality control, essential for applications, applications ranging, ranging from digital
备注: This manuscript has been submitted to IEEE Access for consideration
点击查看摘要
Abstract:Accurate assessment of color differences is essential for applications ranging from digital design to quality control. While existing color difference metrics, such as CIEDE2000, aim to approximate human perception, they may still exhibit inconsistencies with perceptual judgments. In this study, we investigate a data-driven approach to color-difference estimation based directly on human evaluations. We collect similarity judgments for 2,000 systematically generated color pairs, each rated by seven observers using a four-point ordinal scale. These judgments are then used to train regression models using different color representations, including RGB channel differences, HSI differences, and COLIBRI fuzzy linguistic categories. Experiments with five regression algorithms show that the choice of color model has a greater influence on prediction performance than the choice of regression algorithm. Using COLIBRI features alone, linear regression achieves an R2 of 0.595, outperforming RGB and HSI representations, which achieve R2 values of 0.479 and 0.493, respectively. The best performance is obtained by LightGBM using the combined representation, reaching an R2 of 0.703. The results indicate that human perceptual color differences are better captured when numerical color coordinates are complemented by graded perceptual categories, highlighting the potential of data-driven models for perceptually aligned color-difference estimation.
120. 【2609.39128】GeoGAT: Bidirectional Temporal Sampling Meets Hierarchical Graph Attention for Global Video Geo-localization
链接:https://arxiv.org/abs/2609.39128
作者:Junchao Cui,Xuanzi Ma,Wenqi Shi,Hangyu Li,Biru Zhu,Chong Fu,Xiangyang Luo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Global video geo-localization, video geo-localization aims, Global video, geographic hierarchies, geo-localization aims
备注:
点击查看摘要
Abstract:Global video geo-localization aims to infer the geographic location of a video worldwide, evaluating performance across four geographic hierarchies: city, state/province, country, and continent. Existing methods typically employ one-way uniform sampling to process video frames and train independent classifiers for each hierarchy, which leads to the loss of key geographic cues and prediction conflicts between hierarchies, especially for complex multi-shot edited videos. To address these limitations, we propose GeoGAT, which integrates bidirectional temporal sampling with graph attention networks (GATs). Specifically, GeoGAT extracts forward and offset-reversed frame sequences to construct complementary spatiotemporal features. These fused features are then fed into a predefined geographical hierarchy graph, where GATs perform structure-aware message passing, while a dual-constraint mechanism prunes predictions to eliminate cross-hierarchy conflicts. We construct GeoGAT10k, comprising 9,720 multi-shot edited videos from 166 cities worldwide, specifically to benchmark generalization ability on complex video structures. Experimental results on CityGuessr68k and GeoGAT10k demonstrate that GeoGAT eliminates hierarchical conflicts entirely and achieves state-of-the-art performance across all four geographic hierarchies. On CityGuessr68k, GeoGAT outperforms the strongest baseline, evaluated under both classification and retrieval protocols, by 2.6 percentage points at the city level. On the more challenging GeoGAT10k with multi-shot edited videos, the accuracy improvement exceeds 24 percentage points, validating strong generalization to complex real-world scenarios.
121. 【2609.39127】How to Reduce Localization Ambiguity? Geometry-Semantic Constrained BEV Representation Learning for Satellite-Ground Localization
链接:https://arxiv.org/abs/2609.39127
作者:Junming Feng,Panwang Xia,Qiong Wu,Xudong Lu,Zeyu Jiao,Kun Lv,Zherong Wu,Yi Wan,Peifeng Ma,Li-Ta Hsu,Zhi Zheng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Satellite-ground localization estimates, geo-referenced satellite image, Satellite-ground localization, estimates the planar, planar position
备注: 10 pages, 2 figures, and 4 tables
点击查看摘要
Abstract:Satellite-ground localization estimates the planar position and yaw orientation of a ground camera within a geo-referenced satellite image. Most recent methods map ground and satellite features into a shared bird's-eye-view (BEV) space and establish spatial correspondences. However, insufficient depth constraints can assign one ground feature to different distances along a viewing direction, creating geometric ambiguity in BEV feature placement. Similar appearances at different locations can also create descriptor matching ambiguity, while existing descriptor learning lacks explicit semantic supervision to distinguish them. We propose GeoSem-BEV, a geometry-semantic constrained BEV representation learning method. Radial depth supervision constrains distance assignment, and vertical height supervision constrains height aggregation. Shared explicit semantic supervision promotes consistent semantic predictions across views and helps distinguish locations with similar semantics. These constraints improve feature placement and descriptor discriminability, enhancing state-of-the-art BEV localization models. On VIGOR with unknown orientation, GeoSem-BEV reduces mean orientation error by 37.2% and 38.1% in the cross-area and same-area settings, respectively. The corresponding errors are reduced by 10.8% and 15.6% on DReSS-D. On KITTI-CVL, it reduces same-area mean orientation error by 26.8% under 10 degree orientation noise.
122. 【2609.39120】Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation
链接:https://arxiv.org/abs/2609.39120
作者:Siyuan Liu,Kanghui Tian,Yue Duan,Yutao He,Shangdong Yang,Jian Zhang,Yinghuan Shi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:OPD, student, Teacher-calibrated Policy Contrast, On-policy distillation, improves reasoning
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited student perception is another critical bottleneck in multimodal OPD. By providing oracle visual facts, the performance of OPD-trained students can still be substantially improved for both weak and strong teachers. To address this bottleneck, we propose S-OPD, a simple multimodal on-policy distillation framework that explicitly strengthens student perceptual learning through two objectives. Specifically, Teacher-calibrated Policy Contrast separates student policies under original and masked images with teacher-based token-level gating, strengthening the student's reliance on visual evidence during reasoning. Policy Agreement aligns student policies under original and noise-perturbed images, further improving perceptual robustness to visual noise. Notably, our method can be seamlessly plugged into existing OPD frameworks, requiring no additional data annotations, model parameters or inference operations. Extensive experiments on eight benchmarks across student scales and distillation paradigms demonstrate consistent performance improvements, with gains of up to 4.25 points on LogicVista. When combined with existing teacher-side supervision methods, our method can yield further gains. Code is available at this https URL.
123. 【2609.39116】GRC-Pose: Generation-Reconstruction Correspondence for Prior-Free 6D Object Pose Tracking
链接:https://arxiv.org/abs/2609.39116
作者:Shiyang Liu,Weiquan Lin,Luping Xiao,Jiadong Tang,Yi Yang,Yu Gao,Xingyu Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:single RGB video, posed reference images, object-specific CAD models, single RGB, RGB video
备注: 39 pages
点击查看摘要
Abstract:Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, posed reference images, or pose annotations. Geometric foundation models provide complementary object-centric and scene-centric cues, yet SAM3D CAD is indexed by an arbitrary object-local surface parameterization, whereas reconstructed evidence is expressed in a sequence-specific world frame with partial surface coverage. To exploit this complementarity, we formulate tracking as generation-reconstruction correspondence and introduce GRC-Pose, a correspondence-based framework that combines learned correspondence prediction with robust pose estimation. Concretely, GeoCorr-Matcher estimates weighted object-scene correspondences and per-match uncertainty for each pose candidate. FGH-Solver integrates these matches through multiple robust geometric estimators and sequence-level posterior inference, while a posterior-gated memory retains only inlier-supported observations through occlusion and viewpoint change. Extensive evaluation shows that with SAM3D CAD, GRC-Pose achieves state-of-the-art Average Recall and motion retention on HOT3D, improving the latter by 58% over prior art. On classical benchmarks including YCBInEOAT and LINEMOD, it remains highly competitive.
124. 【2609.39115】Beyond Local Linearity: Scale-Resolved Geometry of Learned Image Encoders
链接:https://arxiv.org/abs/2609.39115
作者:Jakub Szymkowiak,Wojtek Pałubicki,Kamil Adamczewski
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Understanding how learned, characterizing their sensitivity, learned representations respond, respond to finite, important for characterizing
备注: Extended abstract, NeurIPS 2026 Workshop on Symmetry and Geometry in Neural Representations (NeurReps). 14 pages, 5 figures
点击查看摘要
Abstract:Understanding how learned representations respond to finite input changes is important for characterizing their sensitivity, invariances, and robustness. Yet existing geometric analyses are predominantly local and describe only infinitesimal perturbations. We introduce a scale-resolved statistic that compares an encoder's measured feature displacement with its local linear prediction as the perturbation magnitude increases. Across diverse image encoders, we discover a characteristic plateau-rise-peak-decay profile, which we call the bump. The bump is absent at initialization, emerges early during standard training, and does not form under randomized labels or random-noise inputs. Its shape also varies with the training distribution and robustness objective. These results establish departures from local geometry as a signature of how encoder representations are shaped by learning.
125. 【2609.39112】CamAgent: An LLM-Agent Framework for Multi-Species Camera-Trap Workflows
链接:https://arxiv.org/abs/2609.39112
作者:Yutong Deng,Qi Song,Xi Guo,Tianming Wang,Lei Bao,Jianping Ge
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Camera traps accumulated, traps accumulated vast, remains highly fragmented, translating raw media, raw media archives
备注:
点击查看摘要
Abstract:Camera traps accumulated vast, multidimensional data for wildlife monitoring, yet translating raw media archives into meaningful ecological insights remains highly fragmented. Current research workflows require laboriously stitching together disparate analysis tools and scripts, creating steep programming hurdles and complicating end-to-end spatiotemporal analyses. To overcome this fragmentation, we present CamAgent, an autonomous Large Language Model (LLM) agent framework that integrates camera-trap analytical workflows into a unified intelligent ecosystem. CamAgent interprets natural-language ecological intent, schedules computational routing, and executes specialized tools spanning computer-vision perception (e.g., SpeciesNet), CamtrapDP-compatible data management, detection-corrected occupancy modeling, temporal activity analysis, and species co-occurrence networks. The framework automates multi-stage analytical pipelines while maintaining essential data-quality controls and analytical conventions. Consequently, CamAgent significantly reduces manual programming overhead for conservationists, establishing a transparent, scalable, and fully integrated paradigm for camera-trap ecology. Our project is available at this https URL.
126. 【2609.39098】DiFF: Doppler-informed Flow Matching for Human Motion Flow
链接:https://arxiv.org/abs/2609.39098
作者:Kai Wang,Mingle Zhao
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
关键词:Perceiving human motion, next-generation human-robot interaction, foundational motion representation, Perceiving human, point cloud scene
备注: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026. Code: [this https URL](https://github.com/keroseus/DiFF)
点击查看摘要
Abstract:Perceiving human motion via privacy-preserving 4D millimeter-wave (mmWave) radar is critical for next-generation human-robot interaction (HRI), where point cloud scene flow serves as a foundational motion representation. Yet the extreme sparsity and noise of 4D radar point clouds make non-rigid motion flow estimation severely ill-posed--a challenge that existing rigid-centric methods and prior works fail to adequately address, largely because they neglect the rich Doppler velocity cues inherent in 4D radar. We propose DiFF, a generative framework that marries Doppler-informed motion priors with a Kolmogorov-Arnold Network (KAN)-based conditional flow matching model. At its core, a KAN-attention mechanism enables expressive feature extraction, while a prior-guided generative process harnesses Doppler cues to regularize the ill-posed solution space. Extensive experiments show that DiFF achieves state-of-the-art (SOTA) performance across diverse real-world datasets, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.
127. 【2609.39096】DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
链接:https://arxiv.org/abs/2609.39096
作者:Zeqi Xiao,Qingle Liu,Kaiwen Zhang,Yifan Zhou,Zihan Ding,Xingang Pan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Autoregressive video diffusion, diffusion supports streaming, supports streaming generation, video diffusion supports, interactive control
备注:
点击查看摘要
Abstract:Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compression strategies discard history using fixed windows or select tokens through local attention and similarity signals, without directly measuring whether a chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find that tokens with larger discrepancies between intermediate clean predictions and final denoised values tend to carry visual evidence less predictable from the retained context. DeCoPrune uses this model-intrinsic signal to retain high-discrepancy tokens in the long-term cache while pruning low-discrepancy tokens. To evaluate information retention, we introduce CMBench, comprising 58 approximately one-minute generated or real-world context episodes and 116 Reappear or Revisit continuation tasks requiring recall of earlier events or objects. Experiments with LingBot World v2 show that DeCoPrune achieves a DINO score of 0.6701 on a 0-1 scale, with an 85.43% reduction in cumulative historical KV token counts and a 4.14-fold continuation-generation speedup over FullKV. Its head-specialized variant reaches 0.6783 at an 86.19% pruning ratio, approaching FullKV's 0.6803 score and exceeding the evaluated compression baselines at similar budgets. These results indicate that denoising consistency can support long-range information retention while reducing autoregressive inference cost. Our project homepage is this https URL. The code is available at this https URL, and the benchmark at this https URL.
128. 【2609.39089】UGOD: Uncertainty-Guided Opacity and Dropout for Sparse-View 3D Gaussian Splatting
链接:https://arxiv.org/abs/2609.39089
作者:Zhihao Guo,Peng Wang,Zidong Chen,Xiangyu Kong,Yan Lyu,Guanyu Gao,Chenghao Qian,Ziyang Wang,Xinqi Fan,Liangxiu Han
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:limited observations leave, Splatting is prone, Gaussian Splatting, Gaussian primitives weakly, alpha blending
备注: 31 pages, 5 figures, 10 tables. Supplementary material included at the end of the manuscript
点击查看摘要
Abstract:Sparse-view 3D Gaussian Splatting is prone to overfitting because limited observations leave many Gaussian primitives weakly constrained, yet their contributions are still accumulated through alpha blending. Without uncertainty estimation, the renderer cannot distinguish unreliable primitives from well-constrained ones, allowing their erroneous contributions to corrupt novel-view synthesis. We introduce UGOD, an uncertainty-guided framework that estimates a view-dependent uncertainty score for each Gaussian and uses it to regulate its rendering contribution. A lightweight uncertainty head conditioned on Gaussian attributes and viewing direction predicts this score, which then drives a differentiable opacity-modulation mechanism that attenuates high-uncertainty primitives before compositing. During training, a detached soft-dropout branch applies an uncertainty-controlled continuous keep mask to discourage the model from relying on poorly constrained Gaussians and thereby reduce overfitting. Crucially, detaching the uncertainty score prevents gradients from this stochastic regulariser from biasing or collapsing the uncertainty prediction. Experiments on Mip-NeRF~360 and LLFF show that UGOD improves sparse-view novel-view synthesis while producing more compact Gaussian representations than the compared methods. These results demonstrate that Gaussian uncertainty provides an effective rendering-time control for sparse-view reconstruction.
129. 【2609.39083】MRI Super-Resolution with RCDM/WaveMix and Task-Aware Segmentation
链接:https://arxiv.org/abs/2609.39083
作者:Kavitha Viswanathan,Harsh Choudhary,Amit Sethi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Super-resolution and quality, image-fidelity metrics, improve downstream analysis, brain MRI, Super-resolution
备注:
点击查看摘要
Abstract:Super-resolution and quality enhancement of 1.5\,T brain MRI are normally validated with image-fidelity metrics, although their purpose is to improve downstream analysis. We study whether enhancement improves tissue segmentation, and for which segmenters. We propose an unpaired, physics-guided training pipeline for a lightweight ($\le$2.5\,M parameter) recurrent convolutional enhancer: a six-module stochastic 1.5\,T degradation operator, a residual adversarial network that adds scanner-specific texture without moving anatomy, and a cycle-consistent objective with an anti-identity penalty that rules out the copy solution. We then train U-Net, Swin-UNet and wavelet token-mixing segmenters \citep{jeevan2023wavemix} from scratch on either raw or enhanced 1.5\,T images of the same subjects, using identical labels and subject-level splits, for three enhancer variants and two datasets. On ABIDE (41 held-out subjects, FreeSurfer labels) enhancement significantly improves the wavelet segmenter (mean Dice $+0.014$, Wilcoxon $p=3.5\times10^{-5}$; CSF $+0.018$, grey matter $+0.013$), significantly degrades the U-Net ($-0.008$, $p=5.1\times10^{-4}$) and leaves Swin-UNet unchanged. On IXI, whose labels come from FSL-FAST, enhancement lowers Dice for all nine pairings, almost entirely through CSF; we trace this to spatially implausible CSF voxels in the labels that penalise smoother predictions. Enhancement of low-field MRI should therefore be validated per downstream model and against reliable labels.
130. 【2609.39066】Agentic Tool-Augmented Reasoning for Explainable Image Forgery Detection
链接:https://arxiv.org/abs/2609.39066
作者:Zhiya Tan,Jing Huang,Changtao Miao,Lin Tan,Xin Zhang,Weiwei Feng,Jianshu Li,Joey Tianyi Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:based approaches generate, large language model, methods produce binary, produce binary scores, recent multimodal large
备注: Accepted at ACM Multimedia 2026 (Oral)
点击查看摘要
Abstract:Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.
131. 【2609.39051】SMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection
链接:https://arxiv.org/abs/2609.39051
作者:Bo-Yuan Cheng,Kuan-Yu Chen,Po-Han Huang,Jeng-Lin Li,Jian-Jiun Ding
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Existing multimodal video, detectors typically assume, Existing multimodal, highlight detectors typically, assume that visual
备注: 5 pages, 3 figures, 3 tables
点击查看摘要
Abstract:Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.
132. 【2609.39047】BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
链接:https://arxiv.org/abs/2609.39047
作者:Zhihang Wu,Zhongqi Wang,Jie Zhang,Fengming Gu,Shiguang Shan,Xilin Chen
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:remain largely unexplored, achieved remarkable progress, producing controllable visual, controllable visual content, visual content guided
备注: 12 pages, 6 figures, 5 tables. Project page: [this https URL](https://wsad55.github.io/badaction01/)
点击查看摘要
Abstract:Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: this https URL.
133. 【2609.39033】ED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization
链接:https://arxiv.org/abs/2609.39033
作者:JinYoung Kim,Geonho Kim,GiJeong Park,Geonu Lee,YoungJoon Yoo
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:powerful vision-language model, CLIP-based anomaly detectors, powerful vision-language, designed for fine-grained, detectors therefore adapt
备注: 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
点击查看摘要
Abstract:CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-normal evidence, having the same anomaly evidence. We propose TED (Text-Axis Evidence Decomposition), a post-hoc scoring method that asks whether each ambiguous response is better supported by source defect patches or by source normal patches mistaken as anomalous. TED compares these supports under the host's normal-versus-anomaly text response, leaves the backbone and prompts unchanged, and requires no target-domain training. It works as a train-free score for raw VLM backbones or as a source-calibrated residual correction for adapted CLIP-AD hosts. Across frozen VLM backbones, TED substantially improves pixel-level localization over raw prompt similarity; across adapted hosts, it improves most pixel-level settings over P-AUROC, P-PRO, and P-AP. Gains are largest under stronger hard-FP competition, with mean localization gain increasing from +5.0 in low-competition regimes to about +10.9 in mid/high-competition regimes. These results suggest that recoverable defect evidence can already exist in pretrained multimodal representations, but reliable localization requires decoding it against hard-normal competitors. Code will be released at TED GitHub repository.
134. 【2609.39024】Persistent Watermarking of Text-to-Image Models
链接:https://arxiv.org/abs/2609.39024
作者:Dixi Yao,Kaiwen Chen,Tahseen Rabbani,Tian Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:expensive training costs, gaining increasing popularity, generation is gaining, general public, motivating the development
备注:
点击查看摘要
Abstract:Text-to-image (T2I) generation is gaining increasing popularity with the general public, motivating the development of reliable mechanisms for copyrighting such models given their expensive training costs. An adversary may obtain and reuse a pretrained T2I model without authorization, and then serve a modified version through an API service. Such modifications may arise from ordinary downstream adaptation or deliberate attempts to erase ownership, including input-prompt preprocessing, model fine-tuning, and output post-processing. From the model owner's perspective, a key challenge is therefore to embed trigger data that remain persistent under such changes while preserving the model's normal image-generation capabilities. In this work, we propose a contrastive-style watermarking objective with a term that explicitly encourages the watermarked model to behave differently from the original model on trigger inputs. Experiments show substantially stronger trigger-data persistence than prior methods across a wide range of downstream modifications and deliberate attempts to weaken the watermark, resulting in higher detection rates, often approaching 100% TPR@FPR$10^{-4}$.
135. 【2609.39021】Frame Differential On-Policy Self-Distillation for Video Reasoning
链接:https://arxiv.org/abs/2609.39021
作者:Haiying He,Xin Zheng,Shaoli Hu,Shijun Xiao,Xuanhe Liu,Bing Li,Harry Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:temporal credit assignment, multimodal language models, Reinforcement learning, credit assignment, substantially improved
备注:
点击查看摘要
Abstract:Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.
136. 【2609.39004】When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion
链接:https://arxiv.org/abs/2609.39004
作者:Zeyu Wang,Jiayu Wang,Haiyu Song,Haoran Duan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multimodal image fusion, integrate complementary information, Multimodal image, aims to integrate, support downstream tasks
备注:
点击查看摘要
Abstract:Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art (SOTA) performance. Code: this http URL.
137. 【2609.38991】Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework
链接:https://arxiv.org/abs/2609.38991
作者:Yaohua Liu,Yifan Guo,Jiaxin Gao
类目:Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:causing surrogate-specific overfitting, trajectory-level geometric disconnect, intrinsic data manifold, Manifold Anchored Bilevel, propose Manifold Anchored
备注: Accepted at NeurIPS 2026 as a Spotlight. 21 pages, 6 figures
点击查看摘要
Abstract:A key bottleneck in adversarial transfer is a trajectory-level geometric disconnect: ambient gradients often drift away from the intrinsic data manifold, causing surrogate-specific overfitting. To rectify this, we propose Manifold Anchored Bilevel Transfer (MABT), a unified framework that anchors adversarial trajectories to the shared semantic subspace. MABT introduces a relaxed manifold-anchoring operator as a semantic rectifier to suppress off-manifold noise. With this constraint, we cast transfer attack generation as a distributional bilevel optimization problem that learns a geometry-aligned initialization by minimizing expected transfer risk under a surrogate uncertainty distribution. We further develop a Hessian-free solver with linear-time complexity to handle the resulting hierarchy. Experiments demonstrate improved transferability for 10 baseline attackers across 28 attack configurations, diverse victim architectures, and defense mechanisms.
138. 【2609.38985】MeshOctave generates meshes via cascading resolution transitions
链接:https://arxiv.org/abs/2609.38985
作者:Junkai Lin,Tianhao Zhao,Hang Long,Huipeng Guo,Jielei Zhang,Youjia Zhang,Jiale Xu,Wenbing Li,Rendong Liang,Jozef Hladký,Matthias Nießner,Yuanming Hu,Wei Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:sequential per-token costs, incur prohibitive sequential, prohibitive sequential per-token, Generating compact, explicit topology typically
备注: 17 pages
点击查看摘要
Abstract:Generating compact, artist-style meshes with explicit topology typically relies on autoregressive models which incur prohibitive sequential per-token costs, or continuous flow models that depend on heuristic connectivity decoders. Next-scale generation paradigms offer a compelling alternative by enabling parallel intra-scale token prediction and coarse-to-fine refinement from global structure to local topology; yet, existing methods derive hierarchical scales via progressive mesh simplification and invert them sequentially. This eliminates intra-scale parallelism and scales generation steps linearly with face count. In this paper, we propose MeshOctave, which instead defines scale through dyadic spatial grid resolutions, framing coarsening as a deterministic collapse that merges vertices sharing a voxel cell and inherits connectivity. Its inverse operation, split-and-rewire, determines which octant sub-vertices are instantiated for each coarse face and resolves local connectivity using discrete structural tokens. These per-face operations require no serialization, each scale transition is modeled as an unordered set that adds one bit of coordinate precision, naturally supporting dynamic-length meshes and adaptive resolution refinement. We construct a scale-conditioned masked-uniform discrete diffusion model to learn split-and-rewire operation from resolution collapse hierarchies. MeshOctave outperforms strong baselines in geometric fidelity and topological validity by a non-trivial margin, while supporting adaptive resolution refinement and extending naturally to mesh subdivision tasks.
139. 【2609.38979】Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting
链接:https://arxiv.org/abs/2609.38979
作者:Chang Liu,Yu Tian,Rui Xie
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:persistent challenge caused, large vision-language models, large vision-language, persistent challenge, challenge caused
备注:
点击查看摘要
Abstract:Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of HALlucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control. Code is available at: this https URL
140. 【2609.38978】PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers
链接:https://arxiv.org/abs/2609.38978
作者:Yun Dai,Jiarui Wen,Huiping Zhuang,Cen Chen,Ziqian Zeng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Diffusion Transformers, dominant architecture, architecture for video, efficiency is limited, quadratic complexity
备注:
点击查看摘要
Abstract:Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade generation quality or yield unnecessary computation. We identify two retrieval mismatches in methods that retrieve blocks using the averaged representations of query and key blocks: (i) query-side aggregation mismatch, where averaging queries before Softmax fails to preserve their individual attention preferences, and (ii) key-side clustering metric mismatch, where standard Euclidean clustering in the original key space can group keys with dissimilar QK scores under the current query, so their average representation may not accurately represent how the current query scores individual keys. These mismatches can lead to inaccurate block retrieval. To address these mismatches, we propose PARK, a training-free sparse attention method for accurate block retrieval. PARK retains every original query, independently normalizes its attention over key blocks, and then averages these distributions within each query block. It also uses information from the current queries to transform keys before clustering, so that keys receiving similar QK scores are grouped together. A fused GPU kernel further reduces the overhead of block retrieval. Experiments on HunyuanVideo and Wan demonstrate that PARK improves block retrieval accuracy and preserves generation quality while accelerating inference, achieving the best quality-efficiency trade-off among the compared sparse attention methods.
141. 【2609.38968】Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion
链接:https://arxiv.org/abs/2609.38968
作者:Zeyu Wang,Mingyu Ge,Haiyu Song,Haoran Duan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:preserving complementary cues, coordinating cross-modal conflicts, Multi-modal image fusion, integrating shared information, Multi-modal image
备注: Accepted to NeurIPS 2026
点击查看摘要
Abstract:Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Rather than relying solely on direct source approximation, we further leverage frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Code: this http URL.
142. 【2609.38930】On the Relaxation of Conditional Independence Assumption for Image Segmentation
链接:https://arxiv.org/abs/2609.38930
作者:Zixun Wang,Ben Dai
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
关键词:directly optimizes Dice, methods directly optimizes, Conditional Independence Assumption, optimizes Dice, modifying model training
备注:
点击查看摘要
Abstract:In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which ignores crucial label correlations and therefore degrades performance in ambiguous or low-contrast scenarios. However, accounting for full label dependence is computationally prohibitive, requiring $\mathcal{O}(d^3)$ time. To address this, we replace the CIA with a Spatially Localized Dependence (SLD) structure that captures local label correlations while keeping the dependence model tractable. We further overcome the remaining computational bottleneck via a Reciprocal Moment Approximation coupled with a novel fixed-point optimization strategy that eliminates exhaustive search. The proposed algorithm achieves a highly practical $\mathcal{O}(d \log d)$ complexity and consistently outperforms conventional argmax and CIA-based RankSEG across diverse segmentation benchmarks. Improvements are significant in low-contrast or small-object scenarios, where label dependence offers valuable signals complementary to image information for accurate segmentation. The code of experiments is available at this https URL.
143. 【2609.38927】World-as-Graph: Relational World Modeling Through Latent Space Graphs
链接:https://arxiv.org/abs/2609.38927
作者:Yaqi Yang,Shuo Huang,Yujin Huang,Fucai Ke,Jiatong Han,Xin Zheng
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:World models aim, World models, object-centric world models, object-centric world model, aim to learn
备注: under review
点击查看摘要
Abstract:World models aim to learn representations of real-world environments and predict their future evolution. Recent object-centric world models have made expressive progress by representing visual scenes as sets of object-level latent states, but object-object relations are often captured only implicitly, which limits explicit relational and temporal structure modeling and object-centric dynamic memory modeling. To address such challenges, we propose World-As-Graph (WAG), a graph-based object-centric world model that introduces relational inductive bias into JEPA-style predictive representation learning. The proposed WAG contains two main modules: (1) Relation-aware structure induction, which constructs time-varying latent graphs from object-centric slots and designs relation-aware object masking policies to guide relational object representation learning in latent space; (2) Object-centric memory transition, which maintains and updates object-level dynamic states by combining relational information from neighboring objects with historical memory, enabling effective autoregressive future prediction. Extensive experiments on both visual reasoning and robotic manipulation tasks could demonstrate the superior performance of our proposed WAG.
144. 【2609.38924】From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning
链接:https://arxiv.org/abs/2609.38924
作者:Jialu Pi,Yanan Ma,Weijie Chen,Owen Crystal,Shubham Trivedi,Stephen Xie,Anna Silverman,Matthew Stib,Chadi Ayoub,Reza Arsanjani,Imon Banerjee
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Major adverse cardiovascular, Major adverse, adverse cardiovascular events, remain the leading, mortality worldwide
备注:
点击查看摘要
Abstract:Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.
145. 【2609.38913】FLOW: Feature-Level Optimal Warping for Generalized Remote Physiological Measurement
链接:https://arxiv.org/abs/2609.38913
作者:Bo Zhao,Junzhe Cao,Dan Guo,Dongmin Huang,Wenjin Wang,Tao Tan,Yue Sun,Zitong YU
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Remote photoplethysmography, enables non-contact physiological, optimal transport, non-contact physiological measurement, Feature-Level Optimal Warping
备注:
点击查看摘要
Abstract:Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains vulnerable to domain shifts from illumination, motion, and sensors. We propose \textbf{FLOW (Feature-Level Optimal Warping)}, an \emph{optimal transport--driven} framework for domain-generalized rPPG. FLOW integrates a \textbf{Temporal Refinement Module (TRM)} to stabilize temporal dynamics and a \textbf{Prototype-based Cross-Temporal Optimal Transport (PCOT)} module to achieve domain-invariant alignment via learnable this http URL feature alignment, FLOW employs soft cross-temporal correspondence modeling that aligns temporal features in a flexible manner, allowing the model to respect and preserve the intrinsic rhythmic patterns of physiological signals. Moreover, the lightweight design of our modules allows seamless integration into existing end-to-end rPPG architectures without additional preprocessing. Two regularization terms further enforce source consistency and identity preservation. Theoretically, we derive a generalization bound under conditional optimal transport. Extensive experiments across four rPPG benchmarks show that FLOW achieves state-of-the-art cross-domain performance with lightweight design and strong physiological fidelity.
146. 【2609.38900】MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding
链接:https://arxiv.org/abs/2609.38900
作者:Yinying Li,Yuqian Fu,Yulin Dai,Jingyu Gong,Tianwen Qian,Xiaoling Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:vast temporal horizons, video understanding requires, process unbounded visual, unbounded visual streams, preserving rich visual
备注: Accepted by ACM Multimedia 2026 (ACM MM 2026)
点击查看摘要
Abstract:Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.
147. 【2609.38864】AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks
链接:https://arxiv.org/abs/2609.38864
作者:Jinglong Wang,Yunjie Wang,Zhiyang Zhang,Jiawei He,Ye Yuan,Bo Qiu,Jing Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:tasks demand accurate, Embodied tasks demand, scene representations, demand accurate, semantically rich
备注:
点击查看摘要
Abstract:Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.
148. 【2609.38856】Decoupling Spherical Reasoning from Dense Prediction for 360 Depth Estimation
链接:https://arxiv.org/abs/2609.38856
作者:Zhijie Shen,Chunyu Lin,Shuai Zheng,Feng Li,Runmin Cong,Huihui Bai,Yao Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:spatially varying distortion, varying distortion makes, distortion makes geometry-consistent, panoramic depth estimation, makes geometry-consistent feature
备注:
点击查看摘要
Abstract:The equirectangular projection (ERP) is widely used for panoramic depth estimation, but its spatially varying distortion makes geometry-consistent feature modeling challenging. We revisit panoramic depth estimation by decoupling contextual modeling in native spherical space from dense ERP prediction. To this end, we propose a Fibonacci Spherical Graph (FSG) as an intermediate reasoning space to lift ERP features onto quasi-uniform Fibonacci nodes on the sphere and capture local and long-range dependencies through complementary spherical neighborhoods. The resulting spherical discretization distributes graph nodes approximately uniformly over the spherical surface, reducing the over-representation of highly stretched regions during relational modeling. Operating on a compact set of Fibonacci nodes also avoids the computational burden of constructing and processing a graph at full ERP resolution. To bridge spherical reasoning and dense prediction, we propose a Spherical Context Conditioning (SCC) module that adaptively modulates dense ERP features with the enhanced spherical representation, allowing spherical context to guide pixel-aligned depth prediction. Extensive experiments on three benchmarks demonstrate that the proposed method consistently achieves superior depth accuracy over existing approaches.
149. 【2609.38851】Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis
链接:https://arxiv.org/abs/2609.38851
作者:Xia Hu,Brian Potetz,Chun-Ta Lu,Huanfen Yao,Leonidas Guibas,Zhicheng Wang,Howard Zhou,Pengfei Xing,Andrew Gallagher
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:compositional tasks records, reflects an intrinsic, intrinsic deficit, failure reflects, targeted capability
备注:
点击查看摘要
Abstract:End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54\% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84\% of the gain from supplying all prerequisites.
150. 【2609.38839】FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
链接:https://arxiv.org/abs/2609.38839
作者:Bo Yin,Xiaobin Hu,Jiaqi Zhao,Shuicheng Yan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Long-horizon video generation, Long-horizon video, effectively leverage, Long-horizon, increasingly long
备注:
点击查看摘要
Abstract:Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.
151. 【2609.38823】DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference
链接:https://arxiv.org/abs/2609.38823
作者:Xudong Tan,Peng Ye,Ming Xie,Chenyu Huang,Yaoxin Yang,Jiayuan Fan,Tao Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:models combine sparse, repeatedly incur attention, inference remains costly, sequences repeatedly incur, combine sparse expert
备注: 21 pages, 11 figures, 6 tables
点击查看摘要
Abstract:Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at this https URL.
152. 【2609.38819】Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video
链接:https://arxiv.org/abs/2609.38819
作者:Chang-Bae Bang,Hyungjin Chung,Byung-Hoon Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
关键词:observed visual stimuli, visual cortex, human visual processing, visual, internal representations
备注:
点击查看摘要
Abstract:Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed visual stimuli, but also predicts upcoming stimuli based on what has been observed. Accordingly, we hypothesize that internal representations for generating future video frames are better aligned with the predictive nature of human visual processing than representations of the observed video itself. To this end, we compare the alignment between human video-watching fMRI responses in the visual cortex and the internal representations from two types of video diffusion models, an autoregressive (AR) model and its non-AR base model. We first conduct a within-model analysis of the AR video diffusion model and show that the representations for future video generation align better with the visual cortex than the representations of the observed video. We then compare the internal representations of the AR model with those of its non-AR base model and again show that the representations for future video generation align better with the visual cortex than the representations for observed video reconstruction by the base model. Specifically, the alignment of observed video reconstruction is concentrated in lower-order visual cortex, whereas that of future video generation is concentrated in higher-order visual cortex. Finally, we show in a human behavioral experiment that humans prefer videos generated by amplifying the contributions of individual layers that align better with the visual cortex.
153. 【2609.38811】DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation
链接:https://arxiv.org/abs/2609.38811
作者:Md Mushfiqur Rahaman,Md Mahedi Hasan,Imtiaz Ahmed,Srinjoy Das
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:X-ray computed tomography, Metal additive manufacturing, additive manufacturing parts, Metal additive, inspected by X-ray
备注: 12 pages, 1 figure, 8 tables. Accepted at the NeurIPS 2026 Workshop on On-Device Intelligence: Foundation Models under Real-World Constraints (ODI)
点击查看摘要
Abstract:Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone carries a separate Conv-LoRA expert bank and mask decoder per defect class, each trained in its own pass, without prompts, on synthetic slices alone, updating only 4.4% of the parameters. On benchmarks that XCT-SAM reports, DCM-SAM improves on every baseline for both classes from a ViT-B backbone against their ViT-H, and reaches 64.2% pore IoU on real NIST scans having seen no real images during training. Deployment then exposes what adaptation work rarely measures: on a Qualcomm Hexagon NPU, ViT-H and ViT-L compile yet cannot allocate at 1024x1024 image resolution, since activations rather than weights exceed the device ceiling, and quantizing weights does not help. ViT-B alone runs, but the adapted encoder then fails to allocate where the stock one succeeds, until a numerically identical rewrite of the attention lets the complete DCM-SAM run in FP16 at 1024x1024, with no operator falling back to the CPU, masks within 0.01% of pixels of the FP32 reference. Code: this https URL.
154. 【2609.38810】CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models
链接:https://arxiv.org/abs/2609.38810
作者:Chunzheng Zhu,Jiaqi Zeng,Hongbo Zhao,Yihang Chen,Yijun Wang,Jianxin Lin
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:internally resolve competing, vision language models, resolve competing visual, vision language, increasingly deployed
备注:
点击查看摘要
Abstract:As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at GitHub repository.
155. 【2609.38795】Recovering Off-Policy Supervision for Speculative Decoding
链接:https://arxiv.org/abs/2609.38795
作者:Jungseob Lee,Chanjun Park,Sugyeong Eo,Hyeonseok Moon
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:off-policy token invalidates, token invalidates supervision, single off-policy token, external models, decoding are commonly
备注: 22 pages, 4 figures, 17 tables
点击查看摘要
Abstract:Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at this https URL.
156. 【2609.38777】Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models
链接:https://arxiv.org/abs/2609.38777
作者:Yuanhao Sun,Huawei Ji,Jiaxin Ding,Luoyi Fu,Xinbing Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:teacher language capabilities, visual evidence, visual, visual understanding, central goal
备注:
点击查看摘要
Abstract:A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match the teacher's answer without relying on the same visual evidence, raising the question: how can we ensure the student responds to the visual information that actually determines the answer? To this end, we propose \textbf{Cross-World On-Policy Distillation (CW-OPD)}, which explicitly supervises the student's response to changes in visual evidence. For each example, CW-OPD constructs two visual worlds that share the question and scene context but differ in answer-critical evidence, yielding different answers. We perform on-policy distillation in both worlds and distill the teacher's cross-world belief transition, encouraging the student to match not only \emph{what} the teacher predicts but also \emph{why} its prediction changes with the evidence. A gradient analysis shows that this term is invariant to errors shared by both worlds and supplies a corrective signal invisible to endpoint matching alone. In this way, CW-OPD makes reliance on the relevant visual evidence an explicit distillation target rather than an implicit consequence of output matching. To diagnose whether a model truly grounds its answers in visual evidence, we introduce CWBench, which measures cross-world consistency via Cross-World Pair Accuracy (CWPA). Experiments on Qwen3.5-4B show that CW-OPD outperforms the strongest baseline by \textbf{1.2} points on average, and the 4B student exceeds DeepSeek-V4.1 (552B) by \textbf{22.4} CWPA points on CWBench. Code is released in this https URL.
157. 【2609.38758】Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation
链接:https://arxiv.org/abs/2609.38758
作者:Abu Hanif Muhammad Syarubany,Jaehyun Jang,Siwoo Lim,Seungyeon Ryu,Chang D. Yoo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Referring Video Object, pixel-accurate mask sequence, Video Object Segmentation, Referring Video, aims to produce
备注: submitted to IEEE Access
点击查看摘要
Abstract:Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a small fixed set of initial keyframes and then relies on propagation. In long or dynamic videos, this can cause stale grounding and persistent false positives when the object composition changes (e.g., distractors enter or the target disappears/re-appears). We propose Event-Driven Refresh + Recurrence Memory (EDRRM), an enhancement that selectively re-invokes Sa2VA only at stable change points. EDRRM triggers refresh boundaries using an EMA-smoothed event score computed from tracking-derived cues (births/deaths and coarse composition/layout changes) with temporal constraints. A recurrence memory further retrieves anchor frames via CLIP similarity to re-condition the model on re-appearance events. Experiments on Ref-DAVIS17, MeViS, and ReVOS show that EDRRM achieves a competitive accuracy-efficiency trade-off relative to fixed-window and FrameDiff-SSIM baselines, maintaining comparable or superior JF scores at substantially lower average refresh-call budgets and reducing false-positive failures. End-to-end runtime analysis further confirms that the overhead introduced by tracking, CLIP-based recurrence matching, and the identifiability gate remains modest relative to the dominant Sa2VA inference cost, thereby validating the efficiency of the proposed pipeline.
158. 【2609.38755】Agentic Relative Camera Pose Estimation via Learned Ranking and Verification
链接:https://arxiv.org/abs/2609.38755
作者:Zhining Gu,Shangjie Du,Weimin Qiu,Carl Olsson,Ping Liu,Meng Tang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:including correspondence-based methods, pose, including correspondence-based, correspondence-based methods, range of approaches
备注: 22 pages, including 10 pages for the main body
点击查看摘要
Abstract:A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and recent 3D geometric foundation models. Our key observation is that no single estimator is optimal for diverse challenges, such as wide baselines, lack of texture, appearance changes, and occlusions. Further analysis reveals substantial performance variation across both benchmarks and individual image pairs, with different estimators exhibiting complementary strengths. We introduce PoseAgent, an agentic framework for relative camera pose estimation that dynamically orchestrates pose estimators through learnable ranking and verification. Given an image pair, a profiling agent first extracts appearance, semantic, and geometric features relevant to pose estimation, e.g., scene type. A learned ranking agent then predicts the relative competence of multiple pose estimators given the image-pair profile. The top-ranked estimator is executed, and its predicted pose is assessed by a learned verification agent that estimates the corresponding pose error. When verification fails, PoseAgent adaptively invokes lower-ranked estimators until a candidate is accepted or the execution budget is reached. For pose verification, our verification network predicts pose errors more accurately than prior models. For pose estimation, PoseAgent improves AUC@5 degree up to 4.2% over the strongest standalone estimator on each of ARKitScenes, MegaDepth, ScanNet++, and RealEstate10K. On ARKitScenes, PoseAgent also outperforms VLM-based agents, which include a VLM ranker with the same verifier and fallback policy. These results demonstrate the effectiveness of our learned ranking and verification.
159. 【2609.38748】Here the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation
链接:https://arxiv.org/abs/2609.38748
作者:Hanmo Chen,Chengcheng Liu,Tianxiao Chen,Zheyu Zhang,Siming Zheng,Jinwei Chen,Xu Yang,Cheng Deng,Bo Li,Peng-tao Jiang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent joint video-audio, achieved strong semantic, Recent joint, strong semantic correspondence, joint video-audio generation
备注:
点击查看摘要
Abstract:Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.
160. 【2609.38747】Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection
链接:https://arxiv.org/abs/2609.38747
作者:Junyang Xia,Luocheng Zhang,Wenwen Pan,Chifeng Zhu,Yang Yang,Xinchun Liu,Jiajun Ding
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Reference-guided camouflaged object, camouflaged object detection, object detection aims, appearance closely resembles, visual appearance closely
备注: 20 pages, including 2 pages of appendix; 9 figures and 4 tables
点击查看摘要
Abstract:Reference-guided camouflaged object detection aims to segment a target whose visual appearance closely resembles its surroundings by exploiting auxiliary reference samples. The task remains difficult because reference samples contain inconsistent target cues, while generic visual representations are not inherently aligned with the target specified by the references. To handle these problems, we present a consensus-aware multi-source fusion framework. Reference-Conditioned Dual-Backbone Fusion (RCDF) couples trainable PVTv2 query features with frozen DINOv3 representations and uses reference-conditioned correlation to select foundation-model evidence before multi-scale fusion. The framework also aggregates multiple references through cross-reference consensus aggregation and injects reference information at semantic depths matched to the query features. Extensive experiments demonstrate the effectiveness of the proposed method. The results further show that reference consensus, target-conditioned foundation features, and hierarchical decoding provide complementary improvements under the evaluation protocol. The source code will be made publicly available upon acceptance.
161. 【2609.38746】Matisse: Evidence-Space Reasoning for Active 3D Reconstruction
链接:https://arxiv.org/abs/2609.38746
作者:Xihang Yu,Kaichen Zhou,Lorenzo Shaikewitz,Clément Jambon,Xiao Zhan,Rajat Talak,Luca Carlone
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:limited computation budget, reconstruction system acquire, computation budget, system acquire, limited computation
备注: 20 pages, 7 figures, 7 tables. Project page: [this https URL](https://xihangyu630.github.io/matisse/)
点击查看摘要
Abstract:How can a 3D reconstruction system acquire and retain useful information to understand the geometry of a scene from partial views under a limited computation budget? Existing active view acquisition methods typically estimate uncertainty over observed or instantiated geometry, limiting their ability to reason about unseen structure, while long-horizon reconstruction methods often retain redundant observations. We introduce Matisse, a training-free framework that unifies active reconstruction and keyframe selection by leveraging evidence provided by a pretrained generative 3D model. Matisse estimates Evidential Uncertainty from cross-attention evidence associated with 3D latent tokens and derives an Evidential Information Gain to guide both view acquisition and keyframe selection based on the expected reduction in posterior entropy. Matisse supports multi-object scenes through occlusion-aware, object-balanced aggregation and propagates uncertainty through intermediate latents to avoid full reconstruction during planning. Matisse reduces Chamfer distance by 12.7%, 3.8%, and 9.2% on GSO30, YCB-V, and Replica, respectively, relative to the best baseline on each dataset, and achieves a $1.50\times$ end-to-end speedup over the best active reconstruction baseline on GSO30 with the same reconstruction backend. In the GSO30 keyframe selection experiment for long-horizon reconstruction, Matisse achieves comparable Chamfer distance using 14% of the input views compared with Stream3D.
162. 【2609.38721】UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
链接:https://arxiv.org/abs/2609.38721
作者:Fang Wu,Da Xing,Yanjie Huang,Junxi Wang,Ji Wang,Hejia Geng,Guancheng Wan,Bowen Zuo,Xiaomin Li,Shixiang Tang,Xinyu Xiang,Zehong Wang,Shiyi Du,Peng Xia,Shuangjia Zheng,Yining Hong,Li Erran Li,Jure Leskovec,Yejin Choi
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Modern multimodal models, single unified system, Modern multimodal, multimodal models, multimodal models bring
备注:
点击查看摘要
Abstract:Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.
163. 【2609.38717】Soft Spatial Reasoning
链接:https://arxiv.org/abs/2609.38717
作者:Rafi Ibn Sultan,Md. Sajid Alam Chowdhury,Saleh Zare Zade,Chengyin Li,Prashant Khanduri,Marco Brocanelli,Dongxiao Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Large Vision-Language Models, Large Vision-Language, Vision-Language Models, Soft Spatial Reasoning, commonly perform spatial
备注:
点击查看摘要
Abstract:Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at this https URL
164. 【2609.38716】SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
链接:https://arxiv.org/abs/2609.38716
作者:Rafi Ibn Sultan,Xiangyu Zhou,Md. Sajid Alam Chowdhury,Chengyin Li,Prashant Khanduri,Marco Brocanelli,Dongxiao Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Large Vision-Language Models, visual perception tasks, made remarkable progress, Large Vision-Language, visual space
备注:
点击查看摘要
Abstract:Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at this https URL.
165. 【2609.38714】Hard-Region Supervision: #1 on the Waymo Open Dataset 2D Video Panoptic Segmentation Leaderboard
链接:https://arxiv.org/abs/2609.38714
作者:Jinghan Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Waymo Open Dataset, Video Panoptic Segmentation, Open Dataset, Waymo Open, Panoptic Segmentation Challenge
备注:
点击查看摘要
Abstract:We describe our winning entry to the Waymo Open Dataset 2D Video Panoptic Segmentation Challenge. The task asks for a semantic class at every pixel of every frame and, for countable objects, an identity that holds across 100 frames and across five overlapping cameras. We build on DVIS++, a cascade of a segmenter, a tracker, and a refiner, as our baseline. We propose hard region supervision (HRS) to improve the baseline. In particular, we use the baseline to define the hard region as where it makes mistakes, and design a loss and an auxiliary prediction head for this region. The auxiliary head is used only in training and removed at test time, so at inference the model trained with HRS has the same architecture as the baseline. In addition, we propose three test-time steps that further improve the results: a two-model ensemble, a merge of the segmenter's output into the final panoptic map, and cross-camera identity linking. On the challenge test set, our entry reaches 0.3547 wSTQ, 0.2071 wAQ, and 0.6075 mIoU, ranking first on all three metrics. It is 3.6 wSTQ points ahead of the second entry and 2.4 points ahead of our DVIS++ baseline.
166. 【2609.38705】SCALE: Synthetic Calibration via Agreement Labeling in Embedding Space
链接:https://arxiv.org/abs/2609.38705
作者:Wenjun Liu,Saeed Hassanpour
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:AUC and accuracy, Foundation models, left untested, calibration, pathology foundation models
备注:
点击查看摘要
Abstract:Foundation models for computational pathology are usually evaluated using AUC and accuracy, while calibration is often left untested. This matters because a model can be accurate on average but still assign overly confident probabilities to cases that are difficult even for pathologists. We study calibration across eight pathology foundation models. Using pathologist agreement as a measure of diagnostic difficulty, we find that calibration error is consistently higher on low-agreement cases than on high-agreement cases. This pattern is not apparent from aggregate expected calibration error (ECE) alone. We then propose synthetic agreement calibration, a method for improving calibration without collecting multi-annotator labels. Given a trained linear probe, we select high-confidence embeddings as class anchors and interpolate between anchors from opposite classes. The interpolation weights encode a continuous notion of diagnostic ambiguity, which we use as a synthetic agreement signal to retrain the probe with agreement-aware label smoothing. On MHIST, which includes annotations from seven pathologists, synthetic agreement calibration recovers most of the calibration improvement obtained by label smoothing based on real pathologist agreement, while substantially reducing low-agreement ECE relative to the uncalibrated baseline. Discrimination metrics are preserved. On PatchCamelyon and BreakHis, public histopathology datasets without multi-annotator labels, the method improves calibration across the evaluated foundation models, whereas annotator-dependent approaches cannot be used without additional expert annotation.
167. 【2609.38691】No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation
链接:https://arxiv.org/abs/2609.38691
作者:Zejing Rao,Ketong Ren,Xiaoqiang Liu,Yiping Meng,Guoxin Zhang,Fan Tang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:dynamically modulate video, modulate video synthesis, Streaming video generators, mid-stream prompt switching, Streaming video
备注:
点击查看摘要
Abstract:Streaming video generators allow users to dynamically modulate video synthesis via mid-stream prompt switching. Existing streaming methods can respond to the updated instruction while still cutting corners, prematurely realizing goals or taking heuristic shortcuts that bypass necessary intermediate state changes needed for a plausible transition. In this study, we present SEGUE, a novel framework that makes this process explicit and trains the generator to execute these transitions faithfully. At each switch, a training-free planner parses the latest frame and prompts, writes a few segue prompts with roles and durations, and then hands control back to the user's prompt. Furthermore, to address the inherent difficulty of training causal models on short-lived temporal schedules without corrupting preparatory supervision, we introduce SPANDMD, which evaluates each active prompt using the full rollout as temporal context while retaining its DMD residual only within the prompt's assigned span. On OpenTrans-360, a benchmark of 1,800 switches that scores how the old state exits and the new one begins, SEGUE ranks first on all eight transition metrics and raises the overall score over the strongest baseline from 0.866 to 0.887. It also ranks first on four of six instruction-response metrics of StreamAV-Bench, while the planner transfers to frozen autoregressive generators without retraining.
168. 【2609.38689】EPIC: Epipolar-Consistent 360° Immersive Stereo Video Generation
链接:https://arxiv.org/abs/2609.38689
作者:Debabrata Mandal,Dongdong Fu,Jonathon Miller,William Villareal,Xi Peng,Praneeth Chakravarthula
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
关键词:circ, diverse virtual experiences, virtual experiences, Immersive displays, displays
备注:
点击查看摘要
Abstract:Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is prohibitively expensive, difficult to scale, and impractical. Generative AI models could remove this bottleneck, but today's models are built for conventional displays and cannot generate the high-resolution, stereoscopic $360^\circ$ content required for immersive viewing. Further, temporal and stereo inconsistencies that may be tolerable on conventional displays can become highly disruptive when viewed through an immersive headset. Here, we address this gap with a zero-shot generative pipeline that extends existing video diffusion models into 4K stereoscopic $360^\circ$ videos. Inspired from binocular vision and depth perception, we develop an epipolar-aware $360^\circ$ image matching metric that captures the temporal and stereo geometric inconsistencies across views. We then use this metric as a preference signal for direct preference optimization with limited training data. Our work enables $360^\circ$ stereo video generation and provides a scalable path for bringing generative content to immersive displays, allowing diverse mixed reality experiences on demand.
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
Cite as:
arXiv:2609.38689 [cs.CV]
(or
arXiv:2609.38689v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2609.38689
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
169. 【2609.38683】Unveiling the Value of Motion for Cinematic Camera Trajectories
链接:https://arxiv.org/abs/2609.38683
作者:Ziqi Zhou,Yujian Yuan,Laura Sevilla-Lara
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Cinematic camera motion, fundamental storytelling tool, Cinematic camera, storytelling tool, fundamental storytelling
备注:
点击查看摘要
Abstract:Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.
170. 【2609.38680】ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images
链接:https://arxiv.org/abs/2609.38680
作者:Shubhang Bhatnagar,Ishan Bhatnagar,Viraj Shah,Narendra Ahuja
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
关键词:synthetic images, real photos, model personalized, model, subject
备注:
点击查看摘要
Abstract:Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.
171. 【2609.38660】Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
链接:https://arxiv.org/abs/2609.38660
作者:Haibo Jin,Xinjie Li,Najmeh Sadoughi,Yang Liu,Yibo Wang,Zhu Liu,Yuzong Liu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)
关键词:Long-form subtitle translation, translation requires reasoning, cultural context spanning, subtitle translation requires, maintaining consistent terminology
备注: 49 pages
点击查看摘要
Abstract:Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On the public MuSC benchmark, SMART obtains the best model result across all four language pairs and also achieves the best human-evaluation result, with an overall score of 4.50/5.
172. 【2609.38642】ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code
链接:https://arxiv.org/abs/2609.38642
作者:Jiaxiang Tang,Yi Zhou,Chad DeLuca,Rogerio Feris,Ahmed Khalil Omran,Zhi-Li Zhang,Pengyuan Li,Ali Anwar
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:altering unrelated content, requested visual change, editing requires cross-modal, cross-modal edit grounding, requires cross-modal edit
备注:
点击查看摘要
Abstract:Chart editing requires cross-modal edit grounding, realizing a requested visual change in the code that draws it, with necessary related updates and without altering unrelated content. Existing benchmarks emphasize either code executability or chart quality, but their metrics do not clearly distinguish request completion from missed coupled updates and gratuitous changes. We introduce ChartRevise, a structured dataset and evaluation protocol for exact program-grounded chart editing. For dataset construction, we build on the grammar of graphics to systematically cover chart-editing operations, using source-program checks to verify their applicability across chart types and libraries. To improve edit exactness, our pipeline checks individual requirements and guides repair or exclusion when they are unmet. The resulting dataset contains 92,438 records covering 344 edit types across 20 chart types and three plotting libraries. For evaluation, our reference-free protocol separately measures atomic requirement completion, identifies gratuitous changes, and detects missed coupled updates. These checks are combined with successful execution and rendering to determine exact-edit success. Across five models and four external benchmarks, fine-tuning yields relative gains of 16\% in mean requirement recall and 22\% in mean exact-edit rate.
173. 【2609.38641】Vision-Language-Action Autonomous Driving Agent with Language-based Memory
链接:https://arxiv.org/abs/2609.38641
作者:Kai Yan,Xiangyu Chen,Yulong Cao,Alex Naumann,Peter Karkus,Yan Wang,Jef Packer,Alex Schwing,Yuxiong Wang,Boris Ivanovic,Wenjie Luo,Marco Pavone
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:utilize knowledge acquired, recently emerged, utilize knowledge, knowledge acquired, acquired during vision-language
备注: 39 pages, 21 figures
点击查看摘要
Abstract:Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
174. 【2609.38637】mplate-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking
链接:https://arxiv.org/abs/2609.38637
作者:Fereshteh Aghaee Meibodi,Amir Mehdi Soufi Enayati,Shadi Alijani,Homayoun Najjaran
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Visual object tracking, Visual object, subsequent search frames, search frames share, search frames
备注:
点击查看摘要
Abstract:Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi-modal tracking where paired modalities are simultaneously available, cross-modal tracking requires localization when template and search frames originate from different active modalities. Accordingly, we introduce TSDA-Track, a Template-Search Domain Adaptation framework to reduce modality discrepancy during training. We investigate two feature alignment strategies. Pre-AFA TSDA-Track applies adversarial alignment before transformer's template-search interaction to suppress modality-specific bias. Enc-CFA TSDA-Track applies contrastive alignment to encoder representations after interaction to strengthen target-level cross-modal correspondence. Both variants retain a shared inference pipeline without modality-specific branches. Experiments on LasHeR, and zero-shot evaluations on RGBT234 and GTOT under multiple cross-modal protocols demonstrate improvements over representative state-of-the-art trackers. For instance, under the modality-switch protocol on RGBT234, Pre-AFA TSDA-Track achieves an SR/PR of 43.2/56.0, compared with 36.8/50.0 for ToMP-101 baseline. In addition, a study on Anti-UAV-024 further verifies the applicability of TSDA-Track to aerial tracking. Our study highlights the effectiveness of feature alignment domain adaptation for cross-modal tracking.
175. 【2609.38636】STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding
链接:https://arxiv.org/abs/2609.38636
作者:Nicolas Thiebaut,Nameer Hirschkind,Xiao Yu,Kyle Spence
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Scene Text Editing, Text Editing, Visual Text Editing, introduce Scene Text, Scene Text
备注: 9 pages, 5 figures, 4 tables. A version of this paper appeared in PAKDD 2026 (LNCS vol. 16618)
点击查看摘要
Abstract:We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.
176. 【2609.38622】Eulerian Motion Reconstruction for Water Scenery
链接:https://arxiv.org/abs/2609.38622
作者:Chuhan Chen,Yen-Chi Cheng,Ayush Saraf,Rajvi Shah,Tuotuo Li,Johannes Kopf,Chen Gao,Hung-Yu Tseng,Deva Ramanan,Matthew O'Toole,Changil Kim
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:immersive visual experiences, nature produces compelling, Reconstructing and animating, animating water scenery, visual experiences
备注: Project at [this https URL](https://sally-chen.github.io/eulersplats)
点击查看摘要
Abstract:Reconstructing and animating water scenery from nature produces compelling and immersive visual experiences. Previous work examined this task from the perspective of 2D video textures, with the goal of creating a looping video. In our work, we tackle the problem from a 3D perspective, creating a looping 4D dynamic reconstruction which can be interactively rendered from novel viewpoints from a single non-looping 2D source video. We represent motion as a 3D static \textit{Eulerian} motion field that advects canonical Gaussian splats that are cyclically reborn at fixed time periods, supervised using rendering losses. To model non-periodic and stochastic dynamics present in real-world scenes, we add a non-periodic, time-varying residual term to capture deviations from the static Eulerian motion field. We show quantitatively and qualitatively that our framework enables photorealistic animation of water scenes better than prior art.
177. 【2609.38620】HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding
链接:https://arxiv.org/abs/2609.38620
作者:Hanwen Cao,Wenqiang Wu,Kuang-Ting Tu,Mathias Otnes,Jeffrey Delmerico,Rui Wang,Yulun Tian,Nikolay Atanasov
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:build continuous, significant impact, Neural implicit representations, features, hierarchical neural field
备注:
点击查看摘要
Abstract:Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-fidelity 3D maps. Most existing works focus on geometric reconstruction and lack semantic information for high-level spatial understanding and task planning. Also, as the scale and complexity of the environment increase, neural representations face the challenge of maintaining computational efficiency in back-end optimization. To resolve these two challenges, we introduce a hierarchical neural field that leverages multiresolution submaps to achieve an efficient and scalable implicit representation, and a unified query and decoding mechanism to support both geometric and semantic features. More specifically, the learnable map features can be converted to the output with the query and decoding process for both training and inference. For large-scale representation, we decompose a scene into overlapping submaps and do hierarchical optimization within each local submap, thus enabling scalable computation. To further improve efficiency, we design feature encoders that predict initial hierarchical grid features to substantially reduce the time needed to optimize the submap features from scratch. To correct estimation drift among submaps, we align and fuse them entirely within the implicit feature space, leading to substantial acceleration by avoiding the need to decode the final output. Building upon this efficient hierarchical representation, we embed both geometric features and vision-language latent features into the map, and demonstrate it on both Signed Distance Field (SDF) construction and open-vocabulary object grounding. Our approach significantly improves computation and memory efficiency, maintains high estimation accuracy, and endows the robot with spatial awareness on large-scale real-world benchmarks.
178. 【2609.38616】Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance
链接:https://arxiv.org/abs/2609.38616
作者:Yanyan Zhang,Disheng Liu,Xinpeng Li,Chaoda Song,Mohsen Hariri,Debargha Ganguly,Wang Yang,Kai Ye,Bryce Grant,Vipin Chaudhary,Yu Yin
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:models enable flexible, flexible action generation, enable flexible action, robotic training data, diverse environmental elements
备注:
点击查看摘要
Abstract:While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.
179. 【2609.38615】Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation
链接:https://arxiv.org/abs/2609.38615
作者:Hongjia Zhai,Xiyu Zhang,Haoran Zhang,Zhichao Ye,Haomin Liu,Guofeng Zhang,Ian Reid,Xingxing Zuo
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:human manipulation provide, manipulation provide valuable, valuable visual experience, provide valuable visual, Egocentric videos
备注: 18 pages
点击查看摘要
Abstract:Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: this https URL.
180. 【2609.38607】After a Decade: Bringing Shadow Removal into the Real World with Agentic Training Data
链接:https://arxiv.org/abs/2609.38607
作者:Shilin Hu,Jingyi Xu,Dimitris Samaras,Hieu Le
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:established benchmarks, real world, solved on established, remains brittle, Shadow removal
备注:
点击查看摘要
Abstract:Shadow removal looks nearly solved on established benchmarks, yet remains brittle in the real world. Models have advanced; the paired training data they rely on have barely changed in nearly a decade. The reason is simple: obtaining a shadow-free target requires removing the occluder while keeping the scene, camera, and illumination otherwise unchanged, making diverse paired data difficult to capture. Meanwhile, large shadow detection datasets already contain diverse real-world images and masks, but no shadow-free targets. To turn this abundant but incomplete data into paired supervision, we propose an offline agentic workflow combining physics-motivated generation, failure detection, feedback-driven retry, candidate selection, and deterministic correction. Using this workflow, we construct AgenticShadow, a dataset of 17,138 image-mask-target triplets spanning general scenes, faces, and remote sensing. Our construction workflow reduces Color Distribution Difference by 50.5% over previous shadow removal work, while training existing shadow removal models on AgenticShadow reduces cross-domain LAB RMSE by 19.7-37.5%.
181. 【2609.38603】Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing
链接:https://arxiv.org/abs/2609.38603
作者:Rishabh Mondal,Nipun Batra,Utkarsh Mall
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:earth observation models, advanced substantially, earth observation, concept, lack interpretability
备注:
点击查看摘要
Abstract:While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability and expert interaction, they are either too expensive to train for the remote sensing domain or perform poorly without annotation. We posit that in expert domains like remote sensing, such training-free models require both fine details in both image and concept space. In image space, we propose a multiscale concept bottleneck using greedy quadtree routing to locate small concepts. In concept space, we replace contrastive vision language models with pre-trained MLLMs and present a way to get reliable concept scores from them. We introduce APERTURE that blends concept scores at the global image and native concept-scale level to give state-ofthe-art training-free model performance. To test these models, introduce SiFC, a fine-grained concept-centric dataset across three countries, with human-reviewed class-level concept maps. On SiFC, APERTURE outperforms the best training-free baselines by more than 10 percentage points in macro F1-score, and notably also outperforms supervised concept bottleneck models. Targeted component-removal tests examine whether concept scores respond to changes in visual evidence, while temporal experiments show that descriptor updates improve recognition of technological changes without retraining.
182. 【2609.38597】PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
链接:https://arxiv.org/abs/2609.38597
作者:Cong Wei,Xuanchi Ren,Bryan Chu,Weiming Ren,Huan Ling,Jiahui Huang,Laura Leal-Taixé,Sanja Fidler,Wenhu Chen,Zian Wang,Jay Zhangjie Wu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:vision-language pretraining pipelines, established vision-language pretraining, increasing visual context, visual context length, separate visual representations
备注: 31 pages, 21 figures. Project page: [this https URL](https://nv-tlabs.github.io/PixelUMM/)
点击查看摘要
Abstract:Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.
183. 【2609.38592】StereoGaussians: Feed-Forward 3D Gaussian Splatting from Stereo Images
链接:https://arxiv.org/abs/2609.38592
作者:Boyuan Tian,Huangying Zhan,Zhan Li,Shin-Fang Chng,Hanwen Yang,Zirui Wang,Yi Xu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:practical stereo-camera applications, applications require nearby-view, require nearby-view extrapolation, stereo-camera applications require, Gaussian Splatting
备注:
点击查看摘要
Abstract:Feed-forward 3D Gaussian Splatting (3DGS) enables reconstruction without per- scene optimisation, but practical stereo-camera applications require nearby-view extrapolation beyond the input views. Stereo depth anchors visible surfaces, yet rendering newly exposed regions also requires learned appearance and additional scene capacity. We introduce StereoGaussians, which predicts a metric 3DGS representation from a single calibrated stereo pair. It reuses intermediate repre- sentations from frozen pretrained stereo networks to predict Gaussian attributes, while calibrated disparity anchors the geometry. A second Gaussian layer and an expanded image canvas provide capacity for disoccluded and outside-field-of- view content. For training, we construct SceneSplat-Stereo from quality-filtered 3DGS teachers, pairing stereo inputs with nearby target views across 803 training scenes. Experiments on unseen real and photorealistic stereo benchmarks demon- strate improvements over strong view-synthesis baselines, while ablation studies support our main design choices.
184. 【2609.38591】Restoring without Forgetting: Filter-Level Continual Image Restoration via Parameter-Space Integrated Gradients
链接:https://arxiv.org/abs/2609.38591
作者:Xin Feng,Jin Zhao,Yizhen Zhang,Wenjie Pei,Fanglin Chen,Guangming Lu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Adapting image restoration, remains challenging due, Adapting image, revisiting past data, past data remains
备注:
点击查看摘要
Abstract:Adapting image restoration models to a stream of new tasks without revisiting past data remains challenging due to catastrophic forgetting. In this work, we propose Restoring without Forgetting (RwF), a filter-level continual adaptation framework for image restoration built upon a critical observation: task-specific knowledge is centered in a small subset of filters and can be separated from those reconstructing general content. RwF first performs parameter-space integrated gradients attribution to localize degradation-critical filters in a coarse-to-fine manner. It then adapts to new tasks by generating task-specific filters from a filter bank using compact factorized low-rank transformations, further augmented with cross-task attention and prototypical contrastive learning, and lastly assembles them back only at localized positions. Experiments on six restoration tasks show that RwF effectively avoids forgetting and achieves competitive restoration quality against all-in-one methods that have full data access, and outperforms LoRA-style adaptation with $\sim$10$\times$ fewer additional parameters. Code is available at this https URL.
185. 【2609.38578】Retargeting Motions to Diverse Skeletons via Learnable Flattening
链接:https://arxiv.org/abs/2609.38578
作者:Kia-Jüng Yang,Fabian H. Sinz,Paweł A. Pierzchlewicz
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Cross-structural motion retargeting, Cross-structural motion, aims to transfer, recent Transformer-based attempts, Cross-structural
备注: 24 pages, 9 figures
点击查看摘要
Abstract:Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by $43-47\%$ over current benchmarks. A user study ($n = 37$), including expert animators, further ranks our approach highest in motion alignment and physical plausibility ($p 0.05$). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.
186. 【2609.38562】LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation
链接:https://arxiv.org/abs/2609.38562
作者:Byoungwoo Park,Jaemoo Choi,Juho Lee,Yongxin Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:creation require coherent, require coherent scene, coherent scene evolution, long-take video creation, video creation require
备注:
点击查看摘要
Abstract:World models, game simulators, and long-take video creation require coherent scene evolution and sustained dynamics over extended durations. Autoregressive (AR) video diffusion provides a natural framework for long-horizon generation, yet extended rollouts often become near-static or lose visual quality. We hypothesize that these failures reflect the limited guidance provided by short-video supervision on how ongoing scene dynamics develops over longer durations. This motivates us to introduce LongTake, a two-stage training pipeline built around Long-Horizon Teacher Forcing (TF) on curated real long videos. Long-Horizon TF trains the AR model to predict later frames conditioned on long ground-truth video prefixes, extending direct supervision beyond the short training horizon. This supervision is designed to help the model sustain dynamics and preserve visual quality during long-horizon generation. Our central finding is that this training stage strengthens direct initialization for distribution matching distillation (DMD) under student self-rollout, without the intermediate few-step distillation stage used in standard pipelines. Under the same five-second DMD training setup, our initialization yields substantially higher dynamic degree than short horizon TF initialization on 30-second rollouts at comparable aesthetic quality, and surpasses the evaluated baselines in both measures. Hybrid DMD further reuses this teacher to extend supervision to later frames of the self-rollout while retaining bidirectional joint supervision over the initial window. On long-horizon self-rollouts, LongTake lies on the Pareto front of dynamic degree and aesthetic quality, and Hybrid DMD attains the highest dynamic degree among evaluated methods at both 30s and 60s.
187. 【2609.38560】Detail in Context: A Dual-Scale Machine Learning Framework for Mycosis Fungoides Detection
链接:https://arxiv.org/abs/2609.38560
作者:Mohamed Hazem,Tarek Waleed,Omar Khaled,Nada Omar,Mahmoud Raslan,Marwa Mohamed Fawzy,Aya Fahim,Rania M. Mogawer,Ahmed Mourad,Kariman Mansour,Muhammad Rushdi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:cutaneous T-cell lymphoma, early stages due, Mycosis fungoides, cutaneous T-cell, T-cell lymphoma
备注:
点击查看摘要
Abstract:Mycosis fungoides (MF) is a rare form of cutaneous T-cell lymphoma that is often misdiagnosed in early stages due to its visual similarity to benign inflammatory dermatoses. Early and accurate diagnosis is critical for improving patient outcomes. In this paper, we propose a comprehensive diagnostic framework for automated MF detection that combines dual- scale histopathological image analysis with deep learning. To distinguish MF from other lymphoproliferative skin conditions, the proposed approach leverages a late-fusion ensemble of dual- magnification (10x and 20x) convolutional neural networks (CNNs), complemented by a random forest classifier trained on 16 clinical features. Experimental results on an expanded dataset of 6,267 images (4,306 MF; 1,961 Non-MF) across 463 patients demonstrate that strong detection performance is obtained by prioritizing higher-resolution cytological details (20x) within broader architectural context (10x). The image-based late-fusion model achieves an accuracy of 83.58% and a sensitivity of 89.13%, while the clinical random forest model achieves an accuracy of 96.6% and sensitivity of 93.8%, highlighting the po- tential of this multimodal framework as a robust clinical decision support system in dermatology. This framework addresses two distinct clinical objectives: an image-based dual-scale pipeline optimized for the early diagnostic screening of MF versus non- MF dermatoses, and a complementary clinical metadata model designed for the subsequent staging of confirmed MF cases (patch/plaque versus tumor)
188. 【2609.38541】hinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing
链接:https://arxiv.org/abs/2609.38541
作者:Donghao Zhou,Haoyang He,Fan Zhang,Hao Yang,Guisheng Liu,Xin Gao,Zhongwei Wan,Xingyuan Bu,Jie Wang,Qiangpeng Yang,Shilei Wen,Chi-Wing Fu,Pheng-Ann Heng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:made significant progress, multimodal large language, Instruction-guided video editing, large language models, video editing
备注: Project page: [this https URL](https://correr-zhou.github.io/ThinkV2V)
点击查看摘要
Abstract:Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.
189. 【2609.38519】GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction
链接:https://arxiv.org/abs/2609.38519
作者:Sheng Zhao,Weikai Lin,Yuhao Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Egocentric gaze prediction, gaze prediction enables, remains challenging, inherently stochastic, Egocentric gaze
备注: Accepted at NeurIPS 2026. 25 pages
点击查看摘要
Abstract:Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.
190. 【2609.38494】What to Attend, What to Keep: Skill-Conditioned Visuotactile Representation with Progress-Guided Event Memory
链接:https://arxiv.org/abs/2609.38494
作者:Amir-Hossein Shahidzadeh,Seungjae Lee,Eadom Dessalene,Shanthosh Raaj Mohanram Mageswari,Soroush Etemad,Furong Huang,Cornelia Fermüller,Yiannis Aloimonos
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Robotic manipulation integrates, vision guides reaching, Robotic manipulation, decides grasping, guides reaching
备注:
点击查看摘要
Abstract:Robotic manipulation integrates vision, touch, and language, whose importance shifts across stages: vision guides reaching, while touch, through its evolution over time, decides grasping, alignment, and contact. Yet existing multi-modal manipulation policies typically use fixed temporal contexts and fusion strategies, despite shifts in what each modality contributes across different skills. We study how vision and touch should be combined at the level of primitive skills, asking what each skill needs from each sensor, and propose a skill-conditioned representation in which the queried skill conditions fusion over modality-specific short-term observation tokens while attending to a sparse event memory that retains terminal observations from the last $K$ executed skills. Evaluated by skill progress estimation on three contact-rich tasks, it reduces slip-detection delay by 87% against fine-tuned SOTA progress models, twist-completion delay by 67.5% against a vision-only ablation, and progress error on a blind search task by 92% through sparse event memory. Gains concentrate exactly where completion is defined by contact or task history. More broadly, our results suggest that observation formation not only policy architecture is a central challenge in multi-modal representation. Project Website: this http URL
191. 【2609.38488】Gaussian Stippling: Efficient Sorting-Free 3D Gaussian Rendering through Hybrid Sampling and Spatiotemporal Reconstruction
链接:https://arxiv.org/abs/2609.38488
作者:Zijian Huang,Suiliang Mai,Chuankun Zheng,Yuan Meng,Yuchi Huo
类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
关键词:requires depth sorting, correctly render overlapping, overlapping Gaussian primitives, ordered alpha blending, render overlapping Gaussian
备注: Preprint
点击查看摘要
Abstract:Conventional 3D Gaussian Splatting (3DGS) requires depth sorting and ordered alpha blending to correctly render overlapping Gaussian primitives. Stochastic transparency enables sorting-free rendering by replacing fractional alpha contributions with discrete stochastic visibility samples, but produces substantial spatial and temporal noise at low sample counts. We refer to this conversion from continuous Gaussian splats to discrete visibility samples as \textit{Gaussian Stippling}. Based on this, we present an efficient order-independent rendering and reconstruction framework that operates directly on unmodified 3DGS assets. Our method adaptively integrates primitive-based and fragment-based stippling, leveraging their complementary strengths across different rendering regimes to significantly improve rendering throughput. To recover high-quality images from sparse stochastic samples, we further introduce a lightweight Gaussian-aware spatiotemporal reconstruction network. By exploiting the Gaussian attributes retained by each stipple, the network aggregates structured stochastic clues across both space and time, effectively suppressing stippling noise. Experiments show that our hybrid Gaussian stippling method, coupled with a spatiotemporal reconstruction network trained on diverse scenes, generalizes to unseen scenes and enables interactive, temporally stable, and visually plausible rendering on mobile devices without retraining or preprocessing. With scene-specific training and appropriately scaled sampling and network capacity, our method further outperforms the baselines in visual quality, offering a high-fidelity configuration for quality-prioritized applications.
192. 【2609.38487】Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment
链接:https://arxiv.org/abs/2609.38487
作者:Sheng Zhao,Weikai Lin,Yuhao Zhu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:machine vision tasks, Judging image quality, everyday human tasks, Judging image, image quality judgment
备注: Accepted at NeurIPS 2026. 25 pages
点击查看摘要
Abstract:Judging image quality is not only ecologically relevant to everyday human tasks, but also underpins many machine vision tasks such as image generation. This paper proposes a framework to understand the inherent perceptual space underlying image quality judgment in humans. We propose a multi-dimensional observer model that represents images as distributions in a latent perceptual space and that models human judgment as comparing noisy samples. Being constrained by neural representations in the primate ventral stream and fit to large-scale behavioral data, the model enables analysis of perceptual structure while matching the predictive power of existing metrics. Using this model, we find that the perceptual spaces needed to account for image quality judgment in humans are extremely low-dimensional compared to the image space even when considering its sparsity. The exact structure of the space (e.g., dimensionalities, information encoded) varies between low-level and high-level quality judgments, suggesting that, despite a shared retinal encoding in the beginning, humans selectively construct task-dependent perceptual spaces in visual decision making.
193. 【2609.38485】Beyond Layers: Position-Resolved Gradient Conflict and Position-Aware Modulation for Unified Multimodal Models
链接:https://arxiv.org/abs/2609.38485
作者:Shuyang Jiang,Fucheng Deng,Yuchuan Luo,Zhenyu Wu
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Unified multimodal models, train image understanding, autoregressive image generation, Unified multimodal, train image
备注: 20 pages, 10 figures, 6 tables. Code will be made publicly available upon acceptance
点击查看摘要
Abstract:Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts, measuring conflict per layer and resolving it by separating parameters. We argue that this resolution hides an orthogonal axis. Generation in a UMM is next-token prediction over a raster sequence of visual tokens whose roles vary systematically with position, so how strongly a generation gradient interferes with understanding should depend on where in the sequence it originates. We introduce a position-resolved interference map that attributes understanding-generation gradient conflict to visual-token positions within every layer, computed from a single backward pass at $1.2\times$ the cost of a standard backward pass. On Show-o and Janus-Pro, position explains a large share of conflict variance after controlling for depth (partial $\eta^2=0.31$ vs. $0.35$ for layer on Show-o; $0.15$ vs. $0.30$ on Janus-Pro): the first quarter of the sequence has a mean gradient cosine of $-0.18$ against understanding, the last quarter $-0.02$. The dependence survives per-position gradient-norm normalization, retaining $80%$ of its effect size, and conflict strength tracks semantic content (Spearman $\rho=0.64$). Building on the map, we propose position-aware modulation (PAM), which removes the anti-aligned component of generation gradients only at high-conflict positions without changing the architecture. Under a matched trainable-parameter budget, PAM improves over layer-wise separation by $+21$ MME and $+2.4$ GenEval points on Show-o while matching it on POPE and overall FID; a random-position control recovers about $31%$ of the gain. Position-based and layer-based separation are complementary degrees of freedom and can be combined.
194. 【2609.38479】Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing
链接:https://arxiv.org/abs/2609.38479
作者:Simon Parkinson,Paloma Liu,Wei Zheng,Mohammadreza Sheikhfathollahi
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:explicit natural-language intermediate, natural-language intermediate representation, paper presents, presents an explainable, perceived safety
备注:
点击查看摘要
Abstract:This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at $r=0.262$, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44\% stronger on the supervised benchmark did not produce measurable improvement in the field ($r=0.250$, $p=0.84$), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1\,km, reaching a median increase of 12.78\% in low-risk route length for a median detour of 2.73\% on journeys of 3 to 6 km.
195. 【2609.38476】Curating Synthetic Data for Task-Specific Visual Perception
链接:https://arxiv.org/abs/2609.38476
作者:Saptarshi Neil Sinha,Paul Julius Kühn,Michael Weinmann
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:valuable where general-purpose, general-purpose datasets, datasets cannot provide, provide the domain-specific, domain-specific priors
备注:
点击查看摘要
Abstract:Synthetic data are most valuable where general-purpose datasets cannot provide the domain-specific priors a task requires, and where manual annotation is expensive, imprecise, or infeasible. In this article we argue that the central question for specialized vision systems is not how to generate more data, but which data to generate. We therefore discuss curated synthetic data, whose scene content, appearance variations, sensing characteristics, and annotations are deliberately designed around a given task. We examine three complementary curation paradigms. Procedural rendering offers explicit control over scene parameters and the annotations follow by construction. Physically-based simulation encodes the mechanism behind an observed effect and yields exactly aligned supervision pairs. Generative AI learns sensor-specific appearance from small real seed sets and attains plausible realism, though it remains prone to hallucination and to inaccurate annotation. These paradigms are illustrated with examples from industrial surface defect detection, restoration of degraded digitized autochrome plates, and 6DoF pose estimation from RGB and event data. Using these examples, we analyze different data regimes and training strategies that combine synthetic and real data across these paradigms. We conclude that curated synthetic data are best understood as a complement to real observations, and that hybrid pipelines combining controllable supervision with learned appearance are the most promising direction for reliable sim-to-real transfer.
196. 【2609.38466】PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation
链接:https://arxiv.org/abs/2609.38466
作者:Chuqiao Li,Xianghui Xie,Yong Cao,Andreas Geiger,Gerard Pons-Moll
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Text-conditioned full-body human-object, remaining precisely coordinated, Text-conditioned full-body, requires synthesizing human, generation requires synthesizing
备注:
点击查看摘要
Abstract:Text-conditioned full-body human-object interaction (HOI) generation requires synthesizing human motion and object trajectories that match the input text while remaining precisely coordinated over time. Most methods represent the human and object as separate trajectories and predict the global human-object couplings. Learning this complex, dynamically changing relationship implicitly, however, often yields object drift, missed contact, and penetration. We introduce PAMI, a Part-Anchored Motion framework for Interaction generation. Inspired by the classic Hough Transform, our key idea is to localize object motion by letting body-part anchors vote for it: we express object motion relative to multiple body-part anchors and use PamiVAE to learn an interaction latent space, decoding frame-wise weights that aggregate these part-specific votes. Building on this representation, PAMI generates interactions in a coarse-to-fine hierarchy. PamiGen first generates a coarse human-object interaction from text in this structured latent space, and PamiRefiner then recursively resolves fine-grained contact geometry using a hybrid surface-sensing representation, combining long-range probes that capture overall body-part influence with short-range sensors that resolve detailed contacts near the object surface. Experiments on InterAct show that PAMI generates more faithful interactions and more accurate human-relative object motion than previous methods, achieving 14.5% higher contact recall than the previous state of the art. Extensive ablations validate the contributions of both the part-anchored voting representation and hybrid surface-sensing refinement.
197. 【2609.38465】Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models
链接:https://arxiv.org/abs/2609.38465
作者:Shuyang Jiang,Fucheng Deng,Yuchuan Luo,Zhenyu Wu
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Unified multimodal models, Unified multimodal, multimodal models, increasingly designed, designed around gradient
备注: 23 pages, 10 figures, 5 tables. Code will be made publicly available upon acceptance
点击查看摘要
Abstract:Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that reducing these metrics improves the downstream understanding-generation trade-off has never been tested directly. We audit it in a controlled testbed, GRIDUMM, which mirrors key structural ingredients of UMM training while making the ground-truth trade-off exactly computable. Across 63 configurations and 372 measured checkpoints, no directional conflict metric reaches an absolute Spearman correlation of 0.3 with a confidence interval excluding zero for conflict measured during training against the eventual trade-off. A dose-response intervention that monotonically suppresses conflict leaves the trade-off flat, separating correlation from causation. The norm ratio is a generation-failure detector and becomes null among configurations that master generation. Functional interference measures outperform directional conflict metrics, while training loss tracks the trade-off strongly. Our results do not show that conflict is useless; they show that its validity as a diagnostic target must be established, not assumed, and we release the audit protocol as a reusable standard.
198. 【2609.38463】rafficSignBench: Rule-Centric Closed-Loop Evaluation of Traffic-Sign Compliance in Autonomous Driving
链接:https://arxiv.org/abs/2609.38463
作者:Victoria Smirnova,Viktoriia Zinkovich,Gregorii Bukhtuev,Artem Belyaev,Andrey Kuznetsov,Denis Shepelev,Vlad Shakhuro
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:destination rate, collision rate, explicitly measure compliance, typically evaluated, evaluated using aggregate
备注:
点击查看摘要
Abstract:Autonomous driving planners are typically evaluated using aggregate metrics such as driving score, destination rate, and collision rate, which do not explicitly measure compliance with traffic rules. As a result, planners can achieve high benchmark scores while still exhibiting unsafe or illegal behaviors, limiting their applicability to real-world deployment. To address this gap, we introduce TrafficSignBench, a large-scale, traffic sign-centric benchmark for systematic and interpretable evaluation of traffic-rule compliance in autonomous driving. Our framework combines real-map-based simulation for realistic road layouts with rule-targeted procedural scenario generation for scalable and balanced coverage of underrepresented rules. We implement traffic rules corresponding to 34 traffic signs, each equipped with an automatic rule checker for detecting violations during closed-loop execution. This design yields 29,000 diverse road scenes and 29 distinct testing scenario types, enabling controlled evaluation of rule-specific planner behavior. We construct 5,800 testing scenes and demonstrate that current autonomous driving planners can exhibit poor traffic-rule compliance despite strong performance on standard evaluation metrics. To address this limitation, we transform existing planners into rule-compliant trajectory experts via explicit traffic-sign constraints, enabling scalable generation of high-quality oracle trajectories for fine-tuning.
199. 【2609.38444】Audible World Models: Spatially Aware Sound Generation for 3D Worlds
链接:https://arxiv.org/abs/2609.38444
作者:Duowen Chen,Jinjin He,Gouthaman KV,Sandeep Bangalore Venkatesh,Bo Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
关键词:create visually rich, image-conditioned world generators, soundtracks synthesized solely, Audible World Models, visually rich
备注: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
点击查看摘要
Abstract:Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
200. 【2609.38443】BIND: Binding 3D Robot Actions to 2D Image Features
链接:https://arxiv.org/abs/2609.38443
作者:Cameron Smith,Arsh Tangri,Vitor Guizilini,Yue Wang,Zubair Irshad,Sergey Zakharov
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:yielding strong data, visuomotor robot policies, yielding strong, representation for visuomotor, image features
备注:
点击查看摘要
Abstract:We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yielding strong data efficiency gains and robustness to out-of-distribution object positions and camera viewpoints. The action heads of current robot policies are typically formulated as an MLP regression from a single global feature vector produced by a pre-trained vision encoder. This global formulation requires the policy network to discover, from demonstrations alone, the relationship between target robot actions and the image features they project onto. The consequence is that although modern image features are semantically descriptive, spatially robust, and even multiview-consistent, the policies built on them are brittle to subtle changes in camera viewpoint and object placement--and surprisingly data-inefficient. BIND closes this gap by supplying the action-feature relationship through camera geometry rather than learning: it discretizes a volume of candidate end effector positions, attaches each candidate to the pre-trained features at its projection in each camera view, and selects actions by scoring each candidate's position and image-bound feature combination. On a real robot, we study data efficiency and out-of-distribution robustness to unseen object positions and camera viewpoints, as well as general long-horizon task execution and dexterity. We find BIND to be highly data-efficient and robust: it achieves near-perfect success on tasks with as few as 5 demonstrations, and degrades gracefully under steep camera-viewpoint shifts and held-out object positions where coordinate-regression baselines completely fail.
201. 【2609.38428】MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary
链接:https://arxiv.org/abs/2609.38428
作者:Shengyun Zhong,Xinkang Zhao,Ziyuan Chu,Linchao Zhu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Online Battle Arena, Multiplayer Online Battle, Battle Arena, Multiplayer Online, Online Battle
备注: 30 pages, 12 figures
点击查看摘要
Abstract:Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at this https URL.
202. 【2609.38426】LoopVL: Recurrent Visual Intelligence
链接:https://arxiv.org/abs/2609.38426
作者:Zhe Qian,Ziyang Gong,Zhongxing Xu,Hehan Li,Zhonghua Wang,Fei Luo,Mingxuan Wang,Xue Yang,Shiwei liu,Yanbiao Ma,Junchi Yan,Jungong Han
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Loop Transformers, extended to vision, effectively extended, Transformers, LoopVL
备注:
点击查看摘要
Abstract:We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
203. 【2609.38413】VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding
链接:https://arxiv.org/abs/2609.38413
作者:Susan Liang,Jianmin Wu,Daxiang Dong
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision-language models, prohibitively expensive, hour-long videos, frozen VLM, Vision-language
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped around a frozen VLM, address this by observing the video selectively, yet existing harnesses are hand-crafted by experts through slow build-and-test cycles. We propose VidHarness, a framework that automates harness design for cost-efficient long video understanding, in which a harness proposer iteratively evolves harnesses based on execution feedback from an evolution environment. To escape the local optima of greedy refinement, we organize the evolution as Monte Carlo tree search (MCTS), and to reduce the evaluation cost, we integrate uncertainty-aware multi-fidelity validation, which screens new harnesses on a few questions and promotes only the promising ones. Since the best harness varies with the frame budget, we further introduce a mixture-of-harness that routes each question to a harness specialized for its budget. VidHarness sets new state-of-the-art results on LongVideoBench, Video-MME, and Video-Holmes, outperforms the strongest hand-crafted video agent by up to $11.2$ points, and generalizes to the knowledge-intensive benchmarks Video-MMMU and MMVU with fewer than half of the frames of uniform sampling.
204. 【2609.38391】am MSU GenText-Forensics Challenge 2026 Technical Report
链接:https://arxiv.org/abs/2609.38391
作者:Kirill Koltsov,Aleksandr Gushchin,Dmitriy Vatolin,Anastasia Antsiferova
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:OCR LLM pipelines, OCR LLM, modern attacks alter, target the OCR, simple pixel-level manipulation
备注:
点击查看摘要
Abstract:Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decide whether a multilingual text image is forged, but also localize the point of manipulation, identify the attack type, and produce a human-readable forensic report with supporting evidence. We present our solution, a decomposed chain-of-thought (CoT) pipeline that combines a document tampering detector (DTD) with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. DTD produces tampering probability maps that are converted into numbered candidate regions; a first model (the Filterer) validates these regions and assigns a preliminary forgery type, while a second model (the Semantic Detective) merges and re-grounds the surviving regions, searches for purely semantic anomalies that are invisible to pixel-level detectors, and writes the final report. Both models are trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher that has access to ground-truth masks and reports. Our approach secured third place in the ACM MM 2026 GenText-Forensics challenge. We describe the data preparation, test-time augmentation, region rendering, distillation protocol, and training configuration in detail, and report ablations over detector thresholds, prompt designs, and pipeline decompositions.
205. 【2609.38377】PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos
链接:https://arxiv.org/abs/2609.38377
作者:Max Ku,Jiaojiao Fan,Zekun Hao,Francesco Ferroni,Heng Wang,Wenhu Chen,Ming-Yu Liu,Prithvijit Chattopadhyay
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:generated videos remains, generated videos, videos remains, remains a fundamental, physical consistency
备注: NeurIPS 2026 poster
点击查看摘要
Abstract:Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe outperforms prior methods on most pairwise benchmarks spanning real-generated and generated-generated pairs under varying correspondence, with the largest gains in no-correspondence and generated-generated settings where existing fine-tuned evaluators degrade sharply. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating that scores are both well ordered and anchored to a stable [0, 1] scale. Further, despite being trained on supervision indicative of physical consistency, without explicit general-preference labels, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.
206. 【2609.38368】Composition, Not Conversation: VLMs Lose the Scene, Not the Thread
链接:https://arxiv.org/abs/2609.38368
作者:L. D. M. S. Sai Teja,Ufaq Khan,N. Siva Gopala Krishna,Satyajit Tourani,Ashshak Sharifdeen,Fida Mohammad Thoker,Bernard Ghanem,Muhammad Haris Khan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision-language models, increasingly reason, revealed over time, Vision-language, VQA benchmarks present
备注: 33 pages, 10 figures, 11 tables. Code: [this https URL](https://github.com/lost-in-layers/composition-not-conversation) . Dataset: [this https URL](https://huggingface.co/lost-in-layer/Layered-VQA)
点击查看摘要
Abstract:Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA benchmarks present the complete image and question at once. We ask what models lose when the same information is fragmented. We introduce Layered-VQA, with 93 scenes and 300 questions. Each image is decomposed into ordered RGBA layers that exactly recompose the original scene, and each question is annotated with supporting, minimal-sufficient, and distractor layers. We evaluate eleven open-weight VLMs from 3B to 32B parameters and two proprietary models with a scale of 187,200 conversations, graded by 1.74M open-model cross-judgments. We find three consistent failures. Loss in Composition: fragmenting the question has a small effect, but fragmenting the scene substantially reduces accuracy; recomposing the same layers largely restores performance. Oracle Inversion: even oracle-selected sufficient evidence can perform worse than the complete scene. Loss in Grounding: as more evidence is required, grounding degrades much faster than answer accuracy. Together, these results show that having the right visual evidence is not enough. How that evidence is composed and presented determines whether models can use and ground it. The right evidence is not enough: VLMs need the scene it came from.
207. 【2609.38362】Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
链接:https://arxiv.org/abs/2609.38362
作者:Hung-Jen Chen,Yu-Heng Ho,Ting-Yao Huang,Po-Hsiang Hsu,Li-Yu Chen,Chun-Yi Lee,Min Sun
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generative vision-language models, strong zero-shot performance, Generative vision-language, required discriminative features, LLaVA achieve strong
备注: Accepted by ECCV 2026
点击查看摘要
Abstract:Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard adaptation methods cannot overcome this representational absence because they operate within the encoder's existing feature space. However, VLMs retain a robust descriptive capacity even when discrimination collapses: a model that cannot classify a medical scan can still articulate its visual patterns. Exploiting this asymmetry, we introduce Inductive Visual Logic (IVL), a training-free framework that constructs classification knowledge from the model's surviving descriptive ability. IVL extracts visual traits from few-shot support images through dual-mode prompting, combining semantic descriptions with primitive visual observations, and organizes them into per-class trait dictionaries. At inference, hierarchical filtering identifies spatially grounded trait evidence for classification. Across multiple distant-OOD benchmarks, IVL achieves the highest aggregate accuracy under two VLM backbones while producing interpretable, trait-traceable predictions.
208. 【2609.38347】rackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos
链接:https://arxiv.org/abs/2609.38347
作者:Patt Phurtivilai,Zhiyang Dou,Yifan Wu,Kinfung Chu,Yuan Liu,Lei Yang,Wenping Wang,Taku Komura
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Quantifying collective fish, requires accurate trajectories, behavior requires accurate, remains challenging due, Quantifying collective
备注: to be published in the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
点击查看摘要
Abstract:Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame. To make these associations identity-aware, TrackFish3D introduces a self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The resulting model is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance features, or test-time optimization. On our benchmark, TrackFish3D improves 3D Multi-Object Tracking Accuracy from 87.7% for the strongest baseline to 95.8%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. TrackFish3D also generalizes beyond fish, achieving strong results on real-world bird tracking.
209. 【2609.38343】Learning Semantic Inpainting for Animatable Gaussian Head Avatars
链接:https://arxiv.org/abs/2609.38343
作者:Pilseo Park,Fizza Rubab,Yiying Tong
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:learning Semantic Inpainting, present SInGA, Semantic Inpainting, learning Semantic, animatable Gaussian head
备注:
点击查看摘要
Abstract:We present SInGA, a novel method for learning Semantic Inpainting for animatable Gaussian head Avatars from a single image. Existing avatar approaches often rely on multi-view observations and lack effective handling of unobserved regions in single-view settings, limiting their applicability in such scenarios. To address this, we propose a semantic inpainting framework defined in UV space for completing unobserved facial regions. Our key insight lies in the structured topology of the UV representation, which provides consistent spatial correspondences and enables reliable completion of identity-specific features using the inherent symmetry cues of human faces. We extract features from observed regions and use them to complete unobserved regions. The completed representation is then used to regress Gaussian attributes, effectively performing Gaussian inpainting. In addition, instead of relying on a single Gaussian at each surface or pixel location, we stack multiple Gaussians to enhance detail. The resulting avatar generalizes across identities without requiring per-identity optimization and can be animated with driving inputs. Experimental results show that our method generates high-quality head avatars with improved completeness and identity preservation, while supporting realistic animation and consistent rendering from unobserved views.
210. 【2609.38329】ExploreNet: Learning Where to Explore in Diffusion GRPO
链接:https://arxiv.org/abs/2609.38329
作者:Shuyue Stella Li,Xiaochuang Han,Yulia Tsvetkov,Luke Zettlemoyer
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:isotropic Gaussian noise, Gaussian noise added, isotropic Gaussian, post-train image generators, Gaussian noise
备注: 23 pages, 15 tables, 8 figures
点击查看摘要
Abstract:Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate. Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.
211. 【2609.38325】Strike a Chord! Modal Kinetic Typography
链接:https://arxiv.org/abs/2609.38325
作者:Maham Tanveer,Jiyeon Han,Nanxuan Zhao,Hao Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:keeping it legible, express a semantic, glyph, introduce modal kinetic, modal kinetic typography
备注: Project page: [this https URL](https://strikeachordkt.github.io/strikeachord/)
点击查看摘要
Abstract:We introduce modal kinetic typography, which animates a vector glyph to express a semantic concept while keeping it legible. Our key idea is to build motion from the glyph's natural vibration modes. Specifically, a finite-element eigenproblem assembled from the vector outline yields the glyph's softest modes, for the whole letter and for each of its parts, allowing it to bend. The problem's zero-energy solutions, i.e., rigid translations and rotations, are applied in closed form to each part, allowing parts to also move as blocks. To animate the glyph, a frozen video diffusion model supervises only the modes' amplitudes and phases. Our modal approach addresses two weaknesses of prior work. Free-form point optimization under video score distillation (SDS) moves each point and frame independently along noisy gradients, tearing the outline and causing jitter. In contrast, our modes are smooth along the outline and driven by a few whole-cycle harmonics, which restricts these gradients to smooth, seamlessly looping motion. On the other hand, structured alternatives rely on skeletons or keypoints from category-specific priors, whereas our modes come from the glyph itself; the only prior is a list naming each letter's moving parts, generated once for the whole alphabet by a language model. In modal kinetic typography, shape and motion are disentangled by construction: a single base outline is sculpted toward the concept, and the modal drive cannot alter it, so a letter can also be animated without being reshaped. Our method produces more articulated and smoother motion than Dynamic Typography and AniClipart at comparable or better concept alignment, with less glyph tearing than Dynamic Typography, and is preferred by human raters, including in a frozen-shape setting where motion alone must carry the concept. Our results were also preferred over Astra (GPT-6) by human raters.
212. 【2609.38298】It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $ε\leq 4/255$
链接:https://arxiv.org/abs/2609.38298
作者:Binchi Zhang,Atrisha Sarkar,Apurva Narayan
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Vision Language Models, Vision Language, safety-critical scenarios, evaluating their trustworthiness, widely deployed
备注:
点击查看摘要
Abstract:Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limited success at $\varepsilon \leq 4/255$. Therefore, VLMs seems robust to perturbations in this range. We show that this robustness does not hold, as targeted semantic substitution succeeds within the same range. Specifically, we align each stream of the source image with its counterpart in the target image in the victim VLM's post-merger token space, operating under a white-box threat model. We evaluate under a strict success criterion, requiring the model to simultaneously name the target, confirm its presence, and deny the source. In images, target semantics appear at $\varepsilon = 2/255$ and complete replacement reaches 38\% at $\varepsilon = 4/255$. On video, complete replacement reaches 35.9\% at $\varepsilon = 1/255$. We also observe a phenomenon of \textit{semantic fusion}, where Large Language Model (LLM) rationalizes contradictory visual signals into a coherent narrative.
213. 【2609.38285】GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
链接:https://arxiv.org/abs/2609.38285
作者:Hongbo Wang,Zihan Lin,Wenkui Yang,Shiran Ge,Yuang Ai,Jie Cao,Huaibo Huang,Ran He
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Vision-language models, fail to respond, Vision-language, relation, failures requires supervision
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on individual answers or ordinal preferences. Therefore, we introduce GaugeVLM, which makes this structure explicit through controlled object and camera interventions in explicit 3D scenes, producing linked observations with measured differences between spatial relations and shared truths across views. To translate this structure into learning signals, its core objective, GaugeDPO, converts measured errors into preference margins, directly supervises correct canonical rankings across views, and links intervention-induced answer-odds contrasts to measured relation changes with view-specific scales. Our analysis bounds canonical prediction error and establishes that the cross-view and intervention constraints can be jointly satisfied. Empirically, GaugeVLM improves all 10 established spatial metrics over supervised fine-tuning across three VLM backbones, with the main 7B model gaining 15.0 and 18.9 percentage points on MSMU distance and QSpatial+, respectively. These gains also extend to autonomous driving and embodied reasoning, demonstrating the robust generalization across domains.
214. 【2609.38278】Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning
链接:https://arxiv.org/abs/2609.38278
作者:Anthony Fuller,Scott C. Lowe,Daniel G. Kyrollos,Graham W. Taylor,Evan Shelhamer,James R. Green
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Self-supervised learning, autoencoder SSL framework, SSL framework learns, annotations and makes, makes models
备注:
点击查看摘要
Abstract:Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MAE) are the most successful instantiation of this framework: they encode a random subset of patches, then decode the masked-out patches. In this work, we introduce key modifications to improve MAEs. Our method augments an image in two different ways, then masks and encodes each view separately. It then exchanges the global representations (CLS tokens) between views before decoding the masked patches. By design, our Masked Swingers encourages learning a view-agnostic summary of the image to facilitate efficient transfer. We perform extensive experiments, and find Masked Swingers outperforms MAE by +3-5% on ImageNet-1K kNN and provides large gains on fine-grained tasks, e.g., relative gains of +45% on instance retrieval, +22% on animal re-ID, and +76% on Omniglot character recognition. To boot, Swingers reduces error -64% relative to MAE on three new state-probing datasets, opening the door to world modeling. Welcome to our Swingers party.
215. 【2609.38271】Evaluating Multi-Task Morphological Concept Learning for Pulmonary Nodule Malignancy Assessment in 3D CT
链接:https://arxiv.org/abs/2609.38271
作者:Namitha Narayanan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Lung Image Database, Image Database Consortium, Image Database Resource, Database Resource Initiative, malignancy risk
备注: 8 pages, 3 figures, 5 tables
点击查看摘要
Abstract:Morphological characteristics such as spiculation and lobulation play an important role in assessing pulmonary nodules on computed tomography (CT), particularly in relation to malignancy risk. This study examines whether learning radiologist-annotated morphological features together with malignancy risk from lesion-centred 3D CT volumes improves classification performance. The Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) dataset was used, comprising 3,918 reader-level nodule annotations from 742 patients after excluding indeterminate malignancy ratings. Patient-level splitting was used for training, validation, and testing, with 112 patients and 628 reader annotations in the held-out test set. A single-task 3D convolutional neural network was compared with a multi-task model predicting malignancy risk, spiculation, and lobulation. The single-task model achieved a balanced accuracy of 0.548 and receiver operating characteristic area under the curve (ROC-AUC) of 0.552, while the multi-task model achieved 0.539 and 0.558, respectively. Patient-level bootstrap analysis showed an ROC-AUC difference of 0.005 (95% confidence interval (CI): -0.087 to 0.090) and a balanced-accuracy difference of -0.009 (95% CI: -0.067 to 0.043). The auxiliary tasks were strongly imbalanced and showed limited predictive performance. Overall, including morphological features did not clearly improve malignancy-risk classification, showing the importance of class balance, label formulation, and reader-level annotation structure in multi-task pulmonary CT analysis.
216. 【2609.38182】EmAvatar: Multimodal Empathetic Response Generation via Conflict Resolution and Expressive Guidance
链接:https://arxiv.org/abs/2609.38182
作者:Xiaolin Chen,Xuemeng Song,Jinlan Fu,Weili Guan,Mong-Li Lee,Wynne Hsu
类目:Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:Avatar-based multimodal empathetic, recognize user emotions, Avatar-based multimodal, empathetic response generation, response generation
备注:
点击查看摘要
Abstract:Avatar-based multimodal empathetic response generation has emerged as a pivotal capability in human-centric systems, aiming to recognize user emotions and synthesize responses with synchronized text, audio, and talking-face video. Despite recent progress, existing methods still suffer from three critical limitations: (1) overlooking conflicting emotions across modalities, (2) lacking explicit multimodal synthesis guidance, and (3) neglecting inherent error propagation of multimodal response generation. To address these limitations, we propose EmAvatar, a novel framework for precise emotion perception and expressive response generation. It first performs deliberative multimodal emotion recognition by exposing inter-modal prediction conflicts and then initiates a multi-round QA process between a Conflict Inspector and an Evidence Collector to gather evidence for conflict resolution, leading to a robust, evidence-aware prediction. Regarding response generation, EmAvatar first synthesizes a composite script that couples the textual response with an expressive instruction. Moreover, to ensure high-quality synthesis, an iterative refinement mechanism evaluates and revises the script until it aligns with predefined criteria, serving as reliable guidance for subsequent audio and video synthesis. Extensive experiments across four tasks demonstrate that EmAvatar outperforms state-of-the-art methods. Our code will be publicly released.
217. 【2606.04857】Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence
链接:https://arxiv.org/abs/2606.04857
作者:Haolu Liu,Xiyue Wang,Xuanting Xie,Liangjian Wen,Zhao Kang
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)
关键词:Incomplete multi-view clustering, Incomplete multi-view, retraining separate models, typically evaluated, separate models
备注: Accepted by NeurIPS 2026 as a poster paper
点击查看摘要
Abstract:Incomplete multi-view clustering (IMVC) is typically evaluated by retraining separate models under different missing-view configurations. Evaluations indexed only by nominal missing rate can overlook differences in observation structure across missing-view protocols. We show that missing-data protocols with identical nominal missing rates can induce substantially different learning regimes, differing by approximately 50-fold in the proportion of fully observed samples. We formalize this phenomenon as protocol divergence, which quantifies structural disparities among missing-view protocols beyond marginal missing rates. Furthermore, we analyze support-gated reconstruction mechanisms and show that their optimization contribution is inherently limited by the frequency of eligible observations under explicit normalization and optimization conditions. Based on these observations, we propose CRAFT (Co-occurrence-free Robust Attention-masked Fusion Transformer), a train-once framework that combines representation learning with an architecture designed to process missing-view inputs. CRAFT combines (i) per-sample forward computation using each sample's observed views and shared parameters, and (ii) mask-aware fusion over nonempty observed-view subsets. The deployment evaluation starts from training data with all views available and reuses one final checkpoint per dataset and seed across missing-view protocols without retraining. Experiments on CUB and MultiFashion show that CRAFT achieves the strongest performance in 12 out of 13 information-matched settings. Additional deployment experiments across seven benchmarks and sixteen missing configurations demonstrate substantial computational savings through checkpoint reuse while maintaining competitive clustering performance. Code and evaluation tools: this https URL and this https URL.
218. 【2602.19001】Life-Bench: A Benchmark and Knowledge Graph Framework for Multimodal Personalization Beyond Concept Recognition
链接:https://arxiv.org/abs/2602.19001
作者:Xia Hu,Honglei Zhuang,Brian Potetz,Alireza Fathi,Bo Hu,Babak Samari,Howard Zhou
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:large language models, language models increasingly, models increasingly power, primarily target concept-level, power personal assistants
备注:
点击查看摘要
Abstract:As large language models increasingly power personal assistants, users expect them to reason over multimodal life histories, from recognizing people to understanding events to aggregating patterns, yet existing benchmarks primarily target concept-level recognition. We introduce Life-Bench, a fully synthetic, human-verified multimodal benchmark of over 11,800 question-answer pairs across 10 tasks, organized by required evidence scope: concept identification, event understanding, and aggregated reasoning. The benchmark's photo-centric personal histories are distributionally aligned with real user accounts under embedding statistic. The interconnected structure of personal data invites graph-based solutions; we propose LifeGraph, a personal knowledge graph framework providing structured retrieval with on-demand access to source visual evidence, showing particular promise on event and aggregated tasks. Systematic evaluation of four retrieval paradigms on Life-Bench demonstrates that accuracy degrades sharply with evidence scope, falling below 0.40 on aggregated tasks, and that no single paradigm dominates across categories. Performance beyond concept recognition remains modest for all evaluated methods, establishing personalization over multimodal histories as an open challenge and Life-Bench as a testbed for future progress.
219. 【2609.40083】ssue Detection Determines False Positives in Diffusion-Based Histopathology Artifact Detection
链接:https://arxiv.org/abs/2609.40083
作者:Konstantinos Moutselos,Ilias Maglogiannis
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:whole-slide images learn, whole-slide images, flag departures, tissue, One-class artifact detectors
备注: 29 pages, 3 figures, 4 tables, including supplementary material. Submitted to Computerized Medical Imaging and Graphics. Code, data and models: [this https URL](https://doi.org/10.5281/zenodo.23016733)
点击查看摘要
Abstract:One-class artifact detectors for whole-slide images learn normal tissue from a clean training pool and flag departures from it. The pool is built by a preprocessing pipeline whose tissue-detection step is usually treated as neutral. We tested whether it is. On 16 annotated TCGA slides, we rebuilt the clean pool of a diffusion-based detector with different tissue detection methods and compared the resulting models in a four-fold cross-validation. Per-slide saturation-Otsu detection excluded normal tissue, chiefly tissue with large clear spaces such as adipose tissue and alveolar parenchyma, and on slides with thick marker ink kept the ink while excluding ordinary tissue. Replacing it with entropy-based detection reduced the false-positive fraction on held-out clean slides from 0.102 to 0.016, in every fold and with a second training seed, without loss of sensitivity; the gain came from the composition of the pool, not its size. Across three tissue detection methods, false positives followed the fraction of such clear-space tissue in the pool, a statistic that needs no labels or training (0.103, 0.016 and 0.009). The effect did not carry over at the same size to a nearest-neighbour detector on foundation-model features. On an external cohort, the curated pool lowered clean-control false positives by about 20%, far less than within TCGA, and the remaining cross-center loss was not explained by stain differences. For one-class quality control, tissue detection decides what the model learns as normal and should be chosen and reported accordingly.
220. 【2609.40043】MAGiDiff: Sampling the Photospheric Vector Field from UV/EUV Filtergrams
链接:https://arxiv.org/abs/2609.40043
作者:Ruoyu Wang(NYU),David Fouhey(NYU)
类目:olar and Stellar Astrophysics (astro-ph.SR); Instrumentation and Methods for Astrophysics (astro-ph.IM); Computer Vision and Pattern Recognition (cs.CV)
关键词:forecasting solar activity, foundational to modeling, Photospheric vector, photospheric vector magnetograms, full Stokes vector
备注:
点击查看摘要
Abstract:Photospheric vector magnetic fields are foundational to modeling, understanding, and forecasting solar activity. These data are usually produced by inverting and disambiguating the full Stokes vector at multiple passbands, which is demanding. Here, we investigate how well we can estimate photospheric vector magnetograms from UV/EUV filtergrams. This problem is challenging and intrinsically ambiguous without polarization information, as the mapping from UV/EUV intensity to the magnetic field is indirect and ill-posed. We introduce MAGiDiff, a machine-learning-based method that uses denoising diffusion models to estimate vector magnetograms from UV/EUV filtergrams. As input, MAGiDiff takes a stack of filtergrams from the Solar Dynamics Observatory (SDO) / Atmospheric Imaging Assembly (AIA); as output, it is trained to estimate the disambiguated vector magnetogram as seen by Hinode / Solar Optical Telescope-Spectro-Polarimeter (SOT-SP). We show that MAGiDiff can accurately mimic the Hinode ground-truth. Additionally, we probe MAGiDiff's understanding of the physical structure and magnetic connectivity. On full-disk, we show that it produces plausible structures for active regions. MAGiDiff generalizes across solar cycles despite hemispheric polarity reversal, and can be fine-tuned to other EUV instruments including STEREO/EUVI and GOES-R/SUVI. While clearly not a substitute for a dedicated instrument, MAGiDiff opens the door to new capabilities.
221. 【2609.38644】Joint Supervised and Self-Supervised Training with Acquisition-Robust Techniques for Accelerated 4D Flow MRI Reconstruction
链接:https://arxiv.org/abs/2609.38644
作者:Mengyuan Xue,Bochun Mei
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:flow MRI measures, MRI measures time-resolved, requires long acquisition, three-directional blood velocity, long acquisition times
备注:
点击查看摘要
Abstract:4D flow MRI measures time-resolved, three-directional blood velocity but requires long acquisition times, and its diagnostic signal is carried by the phase difference \emph{between} velocity encodings, not by image magnitude. Recent work has developed a per-encoding variational network to address image reconstruction in this field. In this work, we incorporate a joint supervised and self-supervised training regime and utilize both magnitude and velocity data during supervision. At the same time, we add multiple acquisition-robust and conditioning strategies based on the acceleration factors. On the CMRx4DFlow~2026 aortic dataset (1.5 and 3T), our model lowers RelErr by $38$--$50\%$ and AngErr by $7.0$--$8.7^\circ$ against a training-matched baseline across $R{=}10$--$50$, improving on every held-out subject at every acceleration. Our model also shows strong generalization ability to transfer on out-of-distribution data by employing the joint training scheme, with an increase of $7.2\%$ in SSIM and decrease of $38\%$ and $31\%$ in AngErr and RelErr respectively.
222. 【2609.38635】SGL: Teacher-Student Graph Learning for 3DGS Compression
链接:https://arxiv.org/abs/2609.38635
作者:Matin Bani Saedi,Matthew Kyan,Gene Cheung
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian Splatting, Gaussian primitives, view synthesis, Gaussian, Splatting
备注: 5 pages, 2 figures. Submitted to IEEE ICASSP 2027
点击查看摘要
Abstract:3D Gaussian Splatting (3DGS) is a popular representation for novel view synthesis. However, 3DGS contains millions of Gaussian primitives, each with rich attributes, resulting in large file sizes. We propose a novel 3DGS compression method based on Teacher-Student Graph Learning (TSGL) that operates directly on a trained model, without 3DGS retraining or access to training images. Specifically, for each block of Gaussian primitives, using decoded positions and DC spherical harmonic (SH) coefficients as predictors, we learn a signal-dependent geometry graph G encoding the pairwise similarities between neighbouring Gaussians via a teacher-student model. Given G, we perform Graph Fourier Transform (GFT) on the remaining attributes, so that signal energies are predominantly projected into the low-frequency coefficients for compact representation. On three standard benchmarks, the method reaches 27x to 33x compression with less than 0.6 dB of PSNR loss, improving on recent post-training compression methods in both size and rendering quality.
223. 【2609.38419】Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models
链接:https://arxiv.org/abs/2609.38419
作者:Ümit Mert Çağlar,Alptekin Temizel
类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Eighteenth International Conference, death among gastrointestinal, gastrointestinal cancers, Colorectal cancer, SPIE Eighteenth International
备注: SPIE Eighteenth International Conference on Machine Vision (ICMV 2025), Paris, France
点击查看摘要
Abstract:Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this cancer and increasing the survival rates. Although CRC is more common in developed regions, its occurrence is also increasing in developing regions as well. CRC diagnosis relies on histopathology assessment post-biopsy. Automated deep learning algorithms can significantly reduce diagnosis time, enhancing efficiency and supporting timely clinical decisions. We present an automated segmentation pipeline for whole-slide histopathology images that labels tumor grades 1-3 and normal mucosa. It utilizes dense prediction transformers with various encoder backbones, overlapping patches, and test-time augmentation. An adaptive augmentation policy, guided by large language models, further improves training. Top models were ensembled via soft voting, and mask refining post-processing steps, Gaussian blurring, morphological closing, and connected components analysis. On a colorectal cancer grade dataset, our method improved the F1 score from 62.92 to 69.84. Code is available here: this http URL
Comments:
SPIE Eighteenth International Conference on Machine Vision (ICMV 2025), Paris, France
Subjects:
Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2609.38419 [eess.IV]
(or
arXiv:2609.38419v1 [eess.IV] for this version)
https://doi.org/10.48550/arXiv.2609.38419
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Journalreference:
Proc. SPIE 14114, Eighteenth International Conference on Machine Vision (ICMV 2025), 141140I (25 Feb 2026)
Related DOI:
https://doi.org/10.1117/12.3096537
Focus to learn more
DOI(s) linking to related resources</p>
224. 【2609.38265】Raw Imagery Impacting Your AI: Should You Care?
链接:https://arxiv.org/abs/2609.38265
作者:Adrien Dorise,Marjorie Bellizzi,Stéphane May
类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:improve mission reactivity, Modulation Transfer Function, Ground Sampling Distance, gaining interest, interest for space
备注: Accepted at OBPDC 2026
点击查看摘要
Abstract:Onboard AI is gaining interest for space applications such as vessel, wildfire, and cloud detection, where real-time processing can improve mission reactivity and reduce downlink needs. However, onboard models may operate on raw or minimally processed imagery rather than on restored ground products. This study evaluates how image degradation affects object detection by varying Signal-to-Noise Ratio (SNR), Modulation Transfer Function (MTF) at Nyquist, and Ground Sampling Distance (GSD). Controlled degradations are applied to Very High Resolution Maxar imagery, and three lightweight detectors, YOLOv5s, YOLOX-S, and NanoDet, are evaluated on the resulting operating points. The results show that the impact of image quality depends on the degradation mechanism, and that increasing degradation does not necessarily lead to a proportional decrease in vessel detection performance. GSD produces the most consistent performance shift, while MTF and SNR effects depend more on the model and resolution. Severe combinations of blur and noise produce the largest losses. These results provide task-level information that can support sensor, processing, and AI trade-offs for future onboard systems.
225. 【2609.35470】Representation Risk in Pretrained Image Encoders
链接:https://arxiv.org/abs/2609.35470
作者:Ardyn Nordstrom,Morgan Nordstrom,Vamuyan Sesay,Matthew D. Webb
类目:General Economics (econ.GN); Computer Vision and Pattern Recognition (cs.CV)
关键词:Applied researchers increasingly, researchers increasingly convert, Applied researchers, increasingly convert images, downstream prediction model
备注:
点击查看摘要
Abstract:Applied researchers increasingly convert images into features with pretrained encoders, then use those features in a downstream prediction model. The encoder is often treated as an implementation detail. We show that it can instead be a consequential source of model uncertainty. We call this uncertainty representation risk: plausible pretrained encoders map the same images into different feature spaces and can yield sharply different out-of-sample conclusions from predictive performance. We compare ten modern and legacy frozen encoders across applications involving house prices, racehorse performance, breast-cancer histology, chest radiographs, continuous facial age, and rice disease. With common dimension control, heads, and group-safe splits, validation selects SigLIP 2 for houses, raising test $R^2$ from 0.396 for ResNet50 to 0.629, and DINOv2 for horses, raising $R^2$ from 0.029 to 0.105. No encoder is best in every task. Candidate procedures are constructed using training data and compared on a separate validation partition. The selected procedure reaches 0.658 for houses and 0.979 accuracy for pneumonia. Fixed-split gains are small for horses and rice, while repeated partitions reveal instability in horse feature union. Continuous age selects SigLIP 2 at 4.786 years MAE. The principal representation gaps persist with neural heads, similarly sized DINOv2 and ViT models, and limited adaptation. These results support a simple workflow: benchmark plausible representations, select on locked validation data, combine only when separate validation evidence justifies the additional cost, and report paired and split-level uncertainty. We implement this workflow in LOOKAGAIN-ML, the software package used to conduct the analyses in this paper.

