本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新647篇论文,其中:

  • 自然语言处理87
  • 信息检索8
  • 计算机视觉100

自然语言处理

1. 【2609.22081】Cross-sector generalization of accident-process role classification in occupational accident narratives

链接https://arxiv.org/abs/2609.22081

作者:Aho Yapi,Pierre Latouche,Arnaud Guillin,Yan Bailly

类目:Computation and Language (cs.CL)

关键词:Occupational accident narratives, Occupational accident, French occupational accident, valuable information, accident narratives

备注

点击查看摘要

Abstract:Occupational accident narratives contain valuable information about work situations, unfavourable conditions, accident events, and their consequences. Automatically structuring these narratives can facilitate large-scale accident analysis and support occupational risk prevention. However, the terminology and writing styles used to describe accidents vary considerably across sectors and organisations, raising questions about the ability of automated coding systems to generalize beyond their training domain. In this paper, we evaluate the cross-sector generalization of accident-process role classification in French occupational accident narratives. We construct an expert-annotated corpus in which factual units are classified into four roles: work situation (A0), explicitly reported unfavourable condition (A1), accident event or deviation (B), and reported consequence (C). The role classifiers are developed and selected exclusively on 42,244 factual units extracted from 6,040 construction-sector narratives and are then evaluated on unseen corpora from the metallurgy and chemistry--plastics sectors, as well as on an independently collected company corpus, without retraining or target-domain tuning of the role classifier. We compare frozen pretrained representations with task-specific fine-tuning and supervised representation-learning strategies. The results show that task-specific adaptation consistently improves cross-domain transfer over frozen representations. Across repeated training runs, the three leading task-adapted strategies achieved average balanced accuracies between 85.6% and 85.8% across the three target corpora. These findings support the development of transferable assisted-coding systems capable of consistently structuring heterogeneous occupational accident narratives for expert review and cross-sector prevention analysis.

2. 【2609.22056】Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

链接https://arxiv.org/abs/2609.22056

作者:Andre Bacellar

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:structurally predictable subpopulations, distributed across queries, predictable subpopulations, uniformly distributed, cluster in structurally

备注: 8 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the other. We instantiate these principles in RegimeAbstain, which computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query-ANN structural features, all available without any additional LLM call, and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLM-judge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE=0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only -0.5pp AUC loss, confirming the domain-agnostic structure of regime features.

3. 【2609.22043】An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

链接https://arxiv.org/abs/2609.22043

作者:Yiming Zhang,Jinghong Zhang,Haoran Zhao,Yiren Ma,Chunlei Zhao

类目:Computation and Language (cs.CL)

关键词:Memory Decision Layer, comparatively little attention, large language models, focused predominantly, predominantly on efficient

备注: 17 pages, 6 figures, 10 tables

点击查看摘要

Abstract:Memory systems for large language models have focused predominantly on efficient retrieval, whereas the decision of whether retrieved memories should be trusted has received comparatively little attention. When the memory store contains conflicting positions, standard retrieval-augmented generation (RAG) blindly injects memories and amplifies hallucinations: in models susceptible to memory injection, the RAG hallucination rate under conflicting memories is markedly higher than that of a memory-free baseline. Inspired by memory signaling mechanisms in the prefrontal cortex, we propose the Memory Decision Layer (MDL), a zero-parameter memory decision controller situated between the retrieval and generation stages. Its core is a three-signal complementary encoder that fuses relevance, reliability, and task risk through QR-based orthogonal subspace projection and a meta-working-memory signal into an interpretable decision representation that quantifies the trustworthiness of retrieved memories. Building on this encoder, MDL explicitly decouples confidence from consistency and introduces risk inversion and explicit abstention. Evaluations on mainstream large language models and multiple open-source datasets show that MDL reduces the hallucination rate under conflicting memories by about 56.04% in general scenarios and approaches zero hallucination in high-risk scenarios. The controller is fully white-box: it relies purely on geometric operations, requires no trained parameters, and adds only about 0.14 ms per decision -- roughly 50x faster than the embedding-retrieval step that precedes it and four to five orders of magnitude faster than an LLM self-evaluation call.

4. 【2609.22038】QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

链接https://arxiv.org/abs/2609.22038

作者:Rawan El Ghali,Umm Kulsoom,Anas Madkoor,Dima Faris Alsaudi,Roaa Abdelmagid,Roaa Ibrahim,Raghad Mousa,Hamza Aljaji,Abdullah Khanafer,Abdallah Alkanani,Salah Feras Alali,Rawan Khaled Mohamed,Ehsaneddin Asgari

类目:Computation and Language (cs.CL)

关键词:Existing Quranic benchmarks, Quranic benchmarks center, multiple dimensions, linguistic complexity, Quranic

备注

点击查看摘要

Abstract:We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom's cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting dataset comprises 980 human-reviewed questions, each issued in both open-ended and multiple-choice form. We benchmark 12 systems on these items and find that the Islamic-specialized model leads, yet every system scores higher on multiple-choice accuracy (average 84%) than open-ended answer quality (average 60%): the two rankings agree closely (Kendall's {\tau}=0.73), but multiple-choice scoring hides failures that surface only once answer choices are removed. QuranicMMLU thus offers a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.

5. 【2609.22008】DiaVLo: Diagnosing Behaviours of Vision-Language Models

链接https://arxiv.org/abs/2609.22008

作者:Lorenzo Corti,Jie Yang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Vision-language models, rely on storing, storing and transferring, transferring appropriate information, Vision-language

备注: 34 pages. To appear in EMNLP 2026 (findings)

点击查看摘要

Abstract:Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with model performance and provide context for measured performance. DiaVLo surfaced behaviours that are clearly aligned and misaligned, alongside patterns in how VLMs perceive, organise, and prioritise concepts.

6. 【2609.22005】Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

链接https://arxiv.org/abs/2609.22005

作者:Richard Zhe Wang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:reportedly improves language, prior studies disagree, attention reportedly improves, language model pretraining, improves language model

备注: 21 pages (8 pages main text plus appendices), 5 figures, 12 tables

点击查看摘要

Abstract:Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that softmax attention lacks: abstention and noise filtering. The first is abstention, which allows an attention head to output nothing, bypassing the requirement that attention weights must sum to one. The second is noise filtering, which allows the value pathway of an attention head to suppress interference from superposed features in the residual stream. In our experiments in matched models from 10M to 350M parameters, we supply abstention through a learned per-head sink logit in the softmax and noise filtering through a gate on each value. We report three empirical findings. First, the benefit of abstention, measured as the reduction in validation loss relative to a matched baseline, declines as models grow, whereas the benefit of noise filtering increases with scale. In particular, abstention accounts for nearly all of the gain from gating at 10M and filtering for most of it at 350M. Second, the best model at every scale is the one with both primitives built in. Third, injecting controlled interference into the values a head reads confirms that the gate removes such interference, and reveals that each of the two gate forms we study has a characteristic blind spot. Supplying both primitives adds negligible parameters and remains compatible with the key-value cache.

7. 【2609.22000】RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

链接https://arxiv.org/abs/2609.22000

作者:Shuai Bai,Jiayong Deng,Yikun Fu,Chang Gao,Xuhao Hu,Mianqiu Huang,Yizhen Jiang,Yuheng Jing,Dehui Kong,Keliang Li,Ning Li,Wanli Li,Dayiheng Liu,Dunjie Lu,Changwei Luo,Que Shen,Zheyuan Wang,Zijian Wang,Jie Wu,Gao Wu,Zhihui Xie,Rui Xie,Haiyang Xu,An Yang,Jiakang Yuan,Yanming Zhang,Jiajun Zhang,Xi Zhang,Zhenru Zhang,Zhuo Zhen,Mingkang Zhu,Bowen Zhou

类目:Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:separate lines, command line, development through code, software development, graphical interaction

备注

点击查看摘要

Abstract:Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software, and run and visually verify their artifacts. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web, plus a unified harness with native GUI control and coding tools. The running reference serves as an oracle for hidden behavioral tests, providing execution-grounded rewards. We scale trajectory generation with high-quality open-source applications. Models trained on these trajectories improve across five out-of-distribution coding and hybrid computer-use benchmarks and more frequently verify their rendered outputs, providing evidence of transfer beyond recreation. For held-out evaluation, we introduce RecreationBench, comprising 250 diverse tasks across domains and platforms. Reference-grounded programmatic and visual assertions cover action-conditioned outcomes at multiple interaction depths; each is validated on the reference and by human reviewers before the suite is frozen for automatic scoring. GPT-6 Astra leads at 58.1% overall, but passes all programmatic tests on just 2.8% of tasks. Agents reproduce static interface structure more reliably than interactions and computed outputs, while generated applications remain smaller and more monolithic than their references. We release the benchmark, environments, and test suites.

8. 【2609.21992】Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment

链接https://arxiv.org/abs/2609.21992

作者:Maciej Skorski

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (stat.ML)

关键词:ethics treats annotator, single annotator flags, computational ethics treats, treats annotator disagreement, treats annotator

备注: accepted to UncertaiNLP @ EMNLP 2026

点击查看摘要

Abstract:Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items -- pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) -- while the stricter majority and two-vote rules miss 63-83% of true positives.

Comments:
accepted to UncertaiNLP @ EMNLP 2026

Subjects:

Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (stat.ML)

MSC classes:
62F15, 94A17, 91E10

ACMclasses:
I.2.7; G.3; J.4

Cite as:
arXiv:2609.21992 [cs.CL]

(or
arXiv:2609.21992v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.21992

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
9. 【2609.21967】NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

链接https://arxiv.org/abs/2609.21967

作者:Jagadeesh Balam,Travis Bartley,Edresson Casanova,Sanjay Chauhan,Chen Chen,Zhehuai Chen,Zijia Chen,Francesco Ciannella,Slyne Deng,Mikyas Desta,Harishchandra Dubey,Slim Essid,Nourchene Ferchichi,Boris Ginsburg,Mariana Graterol Fuenmayor,Negar Habibi,Kevin Hu,Anand Joseph,Viraj Karandikar,Myungjong Kim,Viacheslav Klimkov,Seelan Lakshmi Narasimhan,Lily Lee,Jason Li,Eileen Long,Ameya Mahabaleshwarkar,Aditya Malte,Adi Margolin,Sasha Meister,Valentin Mendelev,Oluwatobi Olabiyi,Ankita Pasad,Yifan Peng,Elena Rastorgueva,Jayda Ritchie,Jason Roche,Nikhil Srihari,Yuanhang Su,Yoshi Suhara,Viet Anh Trinh,Jinhan Wang,Piotr Zelasko,Hui Wang,Puhui Meng,Chaosen Zhang,Yunsheng Liu,Shawn Wang,Wenjing Li,Zhonglei He

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:native tool-calling capabilities, introduce NemotronLabs VoiceChat, NemotronLabs VoiceChat, native tool-calling, streaming TTS decoder

备注

点击查看摘要

Abstract:We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93\% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5\% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.

10. 【2609.21888】Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

链接https://arxiv.org/abs/2609.21888

作者:Chenye Ke,Zirui Liu,Qi Liu,Yan Zhuang,Jintao Zhang,Zhenya Huang,Shijin Wang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, Detecting pretraining data, Detecting pretraining, strong generalization, large language

备注

点击查看摘要

Abstract:Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation. We further extend the mean--variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5\% and TPR@5\%FPR by up to 5.1\%, while remaining robust across diverse settings.

11. 【2609.21859】rialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization

链接https://arxiv.org/abs/2609.21859

作者:Jiacheng Lin,Zifeng Wang,Zheng Chen,Erick Scott,Ziwei Yang,Fanyang Yu,Sheng Zhong,Jimeng Sun

类目:Computation and Language (cs.CL)

关键词:development ultimately fail, entering clinical development, clinical development ultimately, drugs entering clinical, ultimately fail

备注

点击查看摘要

Abstract:Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to anticipate development risks, yet these decisions remain labor-intensive and subjective, requiring experts across clinical science, statistics, regulatory affairs, and competitive intelligence to jointly acquire, synthesize, and reason over heterogeneous evidence. Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk. TrialAtlas further learns from historical clinical trials and regulatory outcomes, including prior New Drug Applications (NDAs), to ground its decisions in accumulated development experience. To evaluate these capabilities in an authentic regulatory setting, we introduce TrialAtlasBench, constructed from 291 FDA Complete Response Letters and spanning three practical tasks: detecting trial design deficiencies, recommending actionable design improvements, and predicting technical and regulatory success. TrialAtlas achieves an F1 score of 50.0% for deficiency detection, outperforming the strongest baseline by 6.1 points, and reaches 85.3% balanced accuracy and 84.7% F1 for prediction of technical and regulatory success, improving over the best baselines by 6.7 points in balanced accuracy and 12.0 points in Cohen's kappa. In expert evaluation, 86.4% of TrialAtlas-generated concerns were judged valid, compared with 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.

12. 【2609.21857】Do Personality-Tuned LLMs Make Better Social Agents?

链接https://arxiv.org/abs/2609.21857

作者:Tim Krabbe,Xiaodan Shi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:socially interactive agents, agents and robots, offering more flexibility, rule-based systems, socially interactive

备注

点击查看摘要

Abstract:LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work investigates whether personality-aware fine-tuning can reduce this gap by improving the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone. We fine-tune two small open-weight LLMs, Qwen2.5-7B-Instruct and Ministral-8B-Instruct, using a corpus that combines personality-labelled social media posts and dialogues to create a personality-based dialogue engine for social simulation. The resulting models are evaluated across multiple social interaction scenarios using three independent LLM judges, which assess personality fidelity and provide evidence-based behavioral interpretations. We additionally quantify inter-rater agreement and lexical characteristics of the generated dialogue. Results indicate that fine-tuned models are not better at role-playing different personalities than their respective baseline models. However, low inter-rater agreement limits the confidence with which these results can be interpreted. Concerning the quality of generated texts, fine-tuned models are mostly comparable to the baselines, with fine-tuning improving the linguistic diversity of the Qwen models. While the results appear generally usable and the baseline models offer the best overall performance, future studies should place greater emphasis on the quality and domain alignment of training data for accurate personality role-playing.

13. 【2609.21844】Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts

链接https://arxiv.org/abs/2609.21844

作者:Steffen Freisinger,Philipp Seeberger,Thomas Ranzenberger,Tobias Bocklet,Korbinian Riedhammer

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:downstream NLP systems, downstream NLP, NLP systems, Long transcripts, irrelevant context

备注: Accepted at EMNLP 2026 Main Conference

点击查看摘要

Abstract:Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.

14. 【2609.21827】RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding

链接https://arxiv.org/abs/2609.21827

作者:Qiao Hu,Yepeng Weng,Bo Zhang,Takehisa Yairi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Speculative decoding accelerates, Speculative decoding, accelerates LLM inference, drafting multiple tokens, decoding accelerates LLM

备注

点击查看摘要

Abstract:Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through hierarchical structures. Dynamic-tree methods such as EAGLE-3 perform well under greedy decoding via deterministic top-K expansion and global pruning. However, in stochastic decoding (T0), this mechanism collapses the draft distribution into one-hot probabilities, causing a severe drop in acceptance rate. This creates a dilemma: dynamic-tree methods sacrifice stochastic sampling to preserve context-aware topology, while static-tree methods preserve stochastic sampling with context-agnostic structures. The issue arises because the same probability distribution is used for two conflicting tasks: constructing the tree and verifying tokens. This coupling makes direct injection of randomness challenging due to the resulting stochastic process. We resolve this by decoupling these roles: RheoSampling assigns a token sampled from the draft distribution a proxy probability for tree expansion and pruning alongside its true sampling probability for verification. Specifically, we inject a sampled token among the deterministic top-K slots and treat it with different probabilities during construction and verification, making RheoSampling the first dynamic-tree method with both context-aware top-K construction and stochastic sampling while maintaining losslessness. We establish the lossless guarantee through an equivalence-class analysis that compresses the stochastic tree space into tractable classes. An OT-based verification strategy and a sparse draft mechanism ensure that theoretical gains translate into practical efficiency. Experiments across LLMs and benchmarks demonstrate improvements in acceptance rate and speedup over state-of-the-art dynamic tree methods. This framework may provide a template for analyzing stochastic tree structures.

15. 【2609.21793】CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

链接https://arxiv.org/abs/2609.21793

作者:Jiale Luo,Eric Han

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, Language Models, output guard, input modification

备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions and experimental settings, have evaluated defenses largely in isolation. Here we present the first systematic study, to our knowledge, of defense combinations both within and across pipeline stages, under a consistent threat model of direct, black-box, single-turn attacks. Our decision framework standardizes evaluation through a principled attack-success-rate formulation with controlled query budgets, together with explicit fairness rules. Across 19 attacks and 15 defenses, we find that no single defense is universally best, but well-chosen combinations achieve substantial safety with minimal utility degradation, yielding practical recommendations for layered defense pipelines.

16. 【2609.21789】Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech

链接https://arxiv.org/abs/2609.21789

作者:Bernard Muller,Antonio Armando Ortiz Barrañón,LaVonne Roberts

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:multilingual dysarthria-severity systems, single aetiology-language pair, pool heterogeneous aetiologies, label space, multilingual dysarthria-severity

备注: Accepted at IEEE SLT 2026, 13-16 December 2026, Palermo, Sicily

点击查看摘要

Abstract:Most multilingual dysarthria-severity systems either train on a single aetiology-language pair or pool heterogeneous aetiologies into one label space. We test that pooling assumption with four matched HuBERT-base contrastive embedding models under a shared backbone, training recipe, corpus registry and held-out evaluation: one mixed-aetiology baseline and three aetiology-specific models for cerebral palsy (CP), Parkinson's disease (PD) and amyotrophic lateral sclerosis (ALS). Training combines clinically labelled speech with ordinal pseudo-labels from a training-free phonological profiling method [1], [2]. On speaker-disjoint, leakage-filtered held-out subsets, the per-aetiology models outperform the mixed baseline across all three target aetiologies: CP (macro F1 0.829 vs 0.676, +22.6 % relative), PD (0.715 vs 0.511, +40.0 %) and ALS (0.788 vs 0.596, +32.3 %). On CP, adding 144 SAP and 44 CDSD pseudo-labelled speakers lifts macro F1 from 0.786 to 0.829 over a clinical-only CP model (+4.3 percentage points). Training data span three to seven languages per aetiology. We position this as a controlled comparison of label-space design choices and discuss pseudo-label calibration, split hygiene, and confidence-thresholded deployment as important limitations for future work.

17. 【2609.21748】World Modeling in Transformers

链接https://arxiv.org/abs/2609.21748

作者:Pierre Beckmann,Matthieu Queloz,Andre Freitas

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:learned faithful representations, Behavioral failures, learned faithful, make a transformer, Behavioral

备注

点击查看摘要

Abstract:Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.

18. 【2609.21722】CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

链接https://arxiv.org/abs/2609.21722

作者:Yifan Wang,Junyu Lu,Qifan Wang,Shun Zhang,Chaozhuo Li,Jiahao Liu,Zhijun Cao,Lingbin Bu,Fanliang Bu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Chinese internet buzzwords, continually evolving lexicon, Chinese social media, Chinese internet, internet buzzwords

备注

点击查看摘要

Abstract:Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLMs can transfer such culturally grounded knowledge across languages and accurately convey the intended meanings in English. This cross-lingual capability is also critical for safety, as harmful expressions may obscure their offensive content through culture-specific homophony, euphemism, irony, or coded language. In this paper, we investigate the ability of advanced LLMs to understand Chinese internet buzzwords across languages. To this end, we introduce CIBuzzBench, the first benchmark for cross-lingual Chinese-to-English understanding of Chinese internet buzzwords. CIBuzzBench comprises 3,001 Chinese internet buzzwords annotated with English meaning explanations, English equivalents, category labels, and harmfulness labels. Based on these annotations, we design three evaluation tasks: Meaning Explanation, Cross-lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection. We evaluate representative state-of-the-art proprietary and Chinese LLMs under both English- and Chinese-prompting settings. Our results show that LLMs continue to struggle with the cross-lingual understanding of Chinese internet buzzwords, particularly in fine-grained non-literal interpretation, robust equivalent matching under option perturbations, and calibrated harmfulness detection. These findings highlight the persistent challenges posed by culturally grounded language phenomena for multilingual LLMs and safety-oriented evaluation. The dataset and code are available at this https URL.

19. 【2609.21683】Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

链接https://arxiv.org/abs/2609.21683

作者:Yunji Chu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)

关键词:immediately preceding behavior, listener immediately preceding, preceding behavior, depends on dialogue, dialogue context

备注: 15 pages, 2 figures, 2026 ECCV Workshop (11th ABAW) Best Student Paper Award

点击查看摘要

Abstract:Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at this https URL.

20. 【2609.21673】PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction

链接https://arxiv.org/abs/2609.21673

作者:Amartya Bhattacharya,Nikhil Singh,Neeti Pokhriyal,Soroush Vosoughi

类目:Computation and Language (cs.CL)

关键词:Probabilistic Graphical Models, Graphical Models, Bayesian Networks, expose directed structure, natural symbolic targets

备注

点击查看摘要

Abstract:Probabilistic Graphical Models (PGMs), especially Bayesian Networks (BNs), expose directed structure and probabilistic parameters, making them natural symbolic targets for neurosymbolic AI. Yet training text-to-parameterized-BN systems requires paired text-to-BN resources unavailable at scale. We introduce PRISM-BN, a controlled corpus of 5054 BN-grounded descriptions paired with discrete reference BNs containing variables, states, directed edges, root priors, and full multi-parent CPDs across five domains. The instances are derived from 50 Wikipedia-seeded backbones, and their probabilities are internally constructed benchmark targets rather than externally validated causal estimates. PRISM-BN is built with PRISM, a marginal-first pipeline that elicits marginal and local joint distributions, analytically recovers normalized CPDs, and constructs locally reparameterized subgraphs. We define a benchmark with semantic node and state alignment, conditional structural scoring, and strict full-CPD evaluation. Across six LLM extractors, Node F1 ranges from 0.56 to 0.83, conditional Edge F1 from 0.90 to 0.97, and CPD-KL from 1.11 to 3.14. Conditional state and edge recovery remain consistently strong, whereas strict full-CPD agreement remains challenging. These trends persist with independently generated GPT-5.5 references, and a human pilot corroborates structural recoverability and similar probabilistic interpretations. PRISM-BN supports separate evaluation of structural recovery and probabilistic parameter estimation.

21. 【2609.21672】Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

链接https://arxiv.org/abs/2609.21672

作者:Zhenyu Zhang,Jiudong Yang,Zhaowen Tao,Meng Chen

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language models, Large language, costly inference, suffer from slow, slow and costly

备注

点击查看摘要

Abstract:Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.

22. 【2609.21663】Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

链接https://arxiv.org/abs/2609.21663

作者:Hritika Sharma,Thibault Bañeras-Roux,Alessandra Pinto,Petr Motlicek,Hyunggu Jung,Esaú Villatoro-Tello,Somang Nam

类目:Computation and Language (cs.CL)

关键词:Word Error Rate, Automatic Speech Recognition, Word Error, Error Rate, Speech Recognition

备注

点击查看摘要

Abstract:Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations of BERTScore and SemDist, varying the language model, layer, and pooling strategy. We find that WER agrees least with human judgment among all metrics tested, that the best-performing SemDist configurations achieve the highest overall agreement, ahead of CER and BERTScore, and that no single model is best across settings. CER, despite its simplicity and low cost, remains remarkably close to these best configurations. In line with prior recommendations, our results support shifting ASR evaluation toward CER both for English and for morphosyllabic writing systems as it is a more interpretable and low-cost metric for what evaluation should actually capture, and using SemDist as a complementary evaluation.

23. 【2609.21662】When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap

链接https://arxiv.org/abs/2609.21662

作者:Gaoxiang Huang,Lei Qi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:controlling language models, Activation steering, motivating its extension, widely used approach, approach for controlling

备注

点击查看摘要

Abstract:Activation steering has become a widely used approach for controlling language models during explicit chain-of-thought (CoT) reasoning, motivating its extension to latent CoT. However, we find that steering continuous thoughts produces substantially weaker effects on subsequent language generation than steering explicit CoT, even when the hidden representations are moved by comparable amounts. We first show that task information remains identifiable in continuous thoughts. Hence, we hypothesize a \textbf{latent-to-language transition gap}, in which an intervention effect in latent space fails to transfer to language generation. Two further results support this hypothesis: the output distribution changes abruptly at the transition boundary, and task-related directions exert much weaker bidirectional control in latent CoT than in explicit CoT. These findings identify the transition interface as a central target for evaluating and designing future latent-steering methods.

24. 【2609.21655】Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces

链接https://arxiv.org/abs/2609.21655

作者:Vasudevan Nedumpozhimana,Fathima Thekkekara,John Kelleher

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:analyse how strongly, strongly different linguistic, encoded in language, model embedding spaces, language model embedding

备注: 6 pages. Accepted at the Workshop on Scientific Methods for Understanding Deep Learning (Sci4DL) at ICLR 2026

点击查看摘要

Abstract:We propose a framework to analyse how strongly different linguistic relations are linearly encoded in language model embedding spaces. We formalise linear encoding via a constrained linear approximation over related and unrelated word pairs and apply this to an extended BATS dataset covering inflectional, derivational, lexicographic, and encyclopedic relations in GloVe, RoBERTa, and ModernBERT. Our experiments show near-perfect linear encodings for inflectional and derivational relations, but substantially higher errors for lexicographic and encyclopedic relations, especially for one-to-many and many-to-many associations. We also find that RoBERTa and ModernBERT generally encode relations more linearly than GloVe. These results indicate that our framework can reveal which relational structures are most linearly accessible in embeddings, offering a compact tool for probing and comparing relational geometry across models.

25. 【2609.21651】Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

链接https://arxiv.org/abs/2609.21651

作者:Naga Ganesh,Chandrashekar M S,Lakshmi Pedapudi,Aakash Singh,Vineet Singh

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Digital Green farm, Green farm advisory, Digital Green, Green farm, farm advisory service

备注: 14 pages, 26 Tables, 12 Figures

点击查看摘要

Abstract:this http URL is Digital Green's farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no crop named, often no text at all. The service has to determine whether the picture can be used, what crop it shows, and what is wrong with it, from images taken on cheap phones in a field, in poor light and with a moving camera. The system doing this today cannot be adjusted. It has no adjustable thresholds for photograph rejection, crops and problems cannot be added, and there is no confidence cut-off to set. We study about 1.16 million photographs sent to this http URL from Ethiopia, India, Kenya and Nigeria. The production quality gate rejected 46.8% of the images it judged, over a quarter of those reaching diagnosis returned no crop name, and 35.8% of the labelled problems filed under "disease" are pests, identifiable without the crop. We therefore split the work into three stages: a quality gate (M0), a crop detector (M1), and a disease or pest detector (M2). Route A fills all three with one fine-tuned vision-language model (Qwen3-VL-4B) answering in a single call. Route B fills each with a small specialist model (DaViT, YOLO26). We replace our production GPT-4o quality gate with a small MobileNetV3 gate at 86.9% F1 in 12 ms. On one test set scored the same way for every system, a hierarchical DaViT-Base achieves 95.41% crop accuracy against 91.46% for the production baseline. It also leads on diagnosis and never declines to answer, while every language model in the comparison leaves a large share of rows with no diagnosis. The fine-tuned model retains two capabilities the specialists do not have: one call for all three stages, and a request for a better photograph when the image cannot support an answer.

Comments:
14 pages, 26 Tables, 12 Figures

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2609.21651 [cs.CV]

(or
arXiv:2609.21651v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.21651

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
26. 【2609.21637】Chinese Competitive Debating Dataset and Benchmark

链接https://arxiv.org/abs/2609.21637

作者:Zongrui Yang,Haoyuan Li,Zhongsheng Wang,Zhirui Zeng,Pengqian Han,Yi Zhou,Yuting Wang,Jiamou Liu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:rarely combine fine-grained, combine fine-grained debate, professional judgments collected, existing datasets rarely, datasets rarely combine

备注: 25 pages, 2 figures

点击查看摘要

Abstract:Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.

27. 【2609.21636】Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30

链接https://arxiv.org/abs/2609.21636

作者:Hans Andersen,David Dichas

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:Recent work applies, Recent work, target human population, Moral Foundations Questionnaire, work applies human

备注: 13 pages, 4 figures, 7 tables. Awarded best Paper Award at WNNLP 2026 (University of Oslo). Proceedings: [this https URL](https://www.uio.no/studier/emner/matnat/ifi/IN5550/v26/final-exam/wnnlp2026_proceedings.pdf)

点击查看摘要

Abstract:Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.

28. 【2609.21595】Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

链接https://arxiv.org/abs/2609.21595

作者:Abhishek Bhandari,Gaurav Harit

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Large Language, learning using Large, Devanagari script remains, Language Models

备注

点击查看摘要

Abstract:In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, comparing three in-context example retrieval strategies: domain-random selection, dense semantic retrieval, and our proposed CharBM25, which retrieves examples by character n-gram BM25 similarity over OCR inputs to target shared error patterns with the test sentence. Across a 20,000-sentence benchmark spanning five news domains, retrieval strategy is the decisive factor in correction quality: CharBM25 outperforms domain-random selection by 2.8-4.0pp absolute WER on Hindi and 2.9-3.8pp on Marathi, using character trigrams, which consistently outperform bigrams and unigrams. Scale dominates performance: Gemma-3-27B achieves WER reductions of 55.0% for Hindi and 33.3% for Marathi under CharBM25-5. Few-shot gains are capacity-gated: models below 8B do not reliably improve over the OCR baseline, and on Marathi the smallest models (3B) degrade more sentences than they improve. Marathi is persistently harder to correct than Hindi across all scales, reflecting its greater morphological complexity. These findings establish CharBM25 as an effective, GPU-free retrieval strategy that matches or exceeds dense retrieval at negligible computational cost, and show that combining it with a general-purpose LLM of 12B+ parameters delivers reliable, training-free Devanagari post-OCR correction without task-specific fine-tuning. Dataset: this https URL

29. 【2609.21562】GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

链接https://arxiv.org/abs/2609.21562

作者:Xinyu Che,Yunfei Ge,Shihao Li,Yanchen Liu,Hang Yan,Xinping Lei,Yanghai Wang,Zixuan Dong,Yifan Yao,Qianqian Xie,Letian Zhu,Jiaheng Liu

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large software projects, Coding agents, large software, Coding, Game

备注: 36 pages, 9 figures, 13 tables. Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, and Xinping Lei contributed equally. Jiaheng Liu is the corresponding author. Code and benchmark: [this https URL](https://github.com/NJU-LINK/GameLogicBench)

点击查看摘要

Abstract:Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A game can end in a valid state even after violating its rules during the run. Current game-development benchmarks replay fixed examples, score videos, or ask another model to judge the result. However, no existing benchmark checks game rules throughout execution across varied evaluator-selected scenarios while ensuring exactly reproducible verdicts. We introduce GameLogicBench, a benchmark of 72 gameplay-logic tasks in Godot projects. An automated evaluator checks each game's rules at every simulation tick. Across 403 hand-designed scenarios, seeded parameter variations produce 1,451 test cases. To ensure that the evaluator measures behavior rather than implementation choice, it must accept different correct implementations for each task while rejecting mutants, implementations with one required capability removed. The tasks span isolated mechanics, multi-system interactions, and repository-scale features. Across 20 combinations of language models and scaffolds, the best observed run solves 52.78% of tasks. Under Claude Code, all twelve models solve fewer tasks as task scope expands from isolated mechanics, through interacting systems, to repository-scale features. Agents inspect code more often and make more tool calls on repository-scale tasks than on isolated-mechanic tasks. Most unsuccessful submissions are runnable, but implement some required game behavior incorrectly. We compared versions of our benchmark evaluator built with and without validation using mutants. Without this validation, incorrect agent submissions passed. A separate analysis finds agents copying code from public repositories when network access is open. Reliable evaluation thus depends both on what the tests reject and on what external code agents can access.

30. 【2609.21554】MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

链接https://arxiv.org/abs/2609.21554

作者:Arash Lagzian,Srinivas Anumasa,Dianbo Liu

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Language Models, Large Language, revolutionized artificial intelligence, Recent advances

备注: 18 pages, 5 figures. Accepted at the ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures (MAS-2025)

点击查看摘要

Abstract:Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility - our ability to dynamically switch mental perspectives - we propose MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration), a novel inference-time creative thinking framework. MIRAGE includes a Selector that prioritizes effective conceptual perspectives (e.g., algebraic, probabilistic) and a Reasoner that sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives. Tested on GSM8K, MATH500, MMLU-Pro, and Game-of-24 benchmarks, MIRAGE consistently outperforms methods like Chain-of-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications.

31. 【2609.21490】Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations

链接https://arxiv.org/abs/2609.21490

作者:Orfeas Menis Mastromichalakis,Giorgos Filandrianos,Wafaa Mohammed,Giuseppe Attanasio,Chrysoula Zerva

类目:Computation and Language (cs.CL)

关键词:Systems Shared Task, Evaluation Systems Shared, Quality Evaluation Systems, affecting both generated, Automated Translation Quality

备注: Accepted for publication at the 11th Conference of Machine Translation (WMT26), co-located with EMNLP 2026

点击查看摘要

Abstract:Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.

32. 【2609.21401】alking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

链接https://arxiv.org/abs/2609.21401

作者:Marina Mitiaeva,Lu Xiao

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:forms remains unclear, systems produce fluent, surface forms remains, Conversational AI systems, socially appropriate responses

备注: Accepted at the 60th Hawaii International Conference on System Sciences (HICSS-60)

点击查看摘要

Abstract:Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are evaluated, trusted, and designed. This study investigates how morality, politeness, and alignment - three dimensions central to cooperative dialogue - function in human-AI interaction compared to human-human conversation. We analyze 15,881 human-ChatGPT and 10,784 human-human multi-turn dialogues, using mixed-effects models to identify which features predict turn-to-turn alignment. We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture. Moral output appears preconfigured rather than negotiated; warmth is generated without face sensitivity; linguistic convergence declines persistently. Most strikingly, the cooperative mechanisms themselves reverse direction: hedging and softening associated with greater accommodation between humans are associated with reduced alignment when produced by AI, and purity framing associated with human divergence coincides with users converging toward the AI. Agency - giving users room to shape the exchange - is the most consistent predictor of alignment across both interaction types, while lower moral assertiveness in more recent models is not accompanied by better cooperation. Together these patterns suggest that AI reproduces the surface of cooperation without the mutual adaptation that grounds it between humans - and, more surprisingly, that mechanisms sustaining human accommodation can run in reverse with AI, suggesting a turn-level view may be insufficient for interaction-level success.

33. 【2609.21392】Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

链接https://arxiv.org/abs/2609.21392

作者:Qi Chen,Yunfei Chu,Haolin He,Yifan Yang,Zihan Liu,Yuxuan Wang,Ziyang Ma,Ruiyang Xu,Meng Gao,Yinsong Yan,Ling Wang,Hui Wang,Wen Huang,Yiheng Chen,Guanrou Yang,Qiuqiang Kong,Jin Xu,Xie Chen

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)

关键词:Natural audio-visual interaction, composed text prompts, carefully composed text, Natural audio-visual, text prompts

备注

点击查看摘要

Abstract:Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs' ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.

34. 【2609.21390】Offline Multimodal Large Language Models for Decision Support in Air Operations

链接https://arxiv.org/abs/2609.21390

作者:Joao P. A. Dantas,Jelton A. Cunha,Gabriel Dietzsch

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:strict security constraints, Air operations rely, established procedures, complex rules, security constraints

备注

点击查看摘要

Abstract:Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal knowledge that remains traceable to its original sources through natural language interaction. We describe a modular retrieval-augmented architecture suitable for operation without Internet connectivity, supporting both text and image input from technical manuals. As a first step toward evaluating this architecture, we report a pilot study with four image analysts of the Brazilian Air Force, combining (i) a doctrinal knowledge assessment based on their electronic-target identification doctrine, comparing human and proposed system performance on the same test, and (ii) a measurement of the cognitive workload involved in manually producing a reconnaissance target report (Relatório de Missão de Reconhecimento - REMIR) without AI assistance. The results show a demanding manual task, especially in terms of mental demand (6.0/7) and effort (5.0/7), while the proposed system matches the human score (8/10) and completes the assessment in 7.1 minutes (compared to a human average of 26.5 minutes), establishing a baseline for future AI-assisted evaluation. Finally, we describe a future evaluation protocol to systematically compare manual and AI-assisted workflows.

35. 【2609.21387】Consistent Relexicalization of Clinical Documents using Graph-Based Approach

链接https://arxiv.org/abs/2609.21387

作者:Dipankar Das,Atri Mandal,Sandeep Singh,Tushar Shandhilya

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:facilitates robust masking, clinical NLP, retain high-fidelity, pivotal technique, facilitates robust

备注: Accepted for presentation at the Sixth International Conference on AI ML Systems (AIMLSystems 2026), Lake Como, Italy, October 6-9, 2026

点击查看摘要

Abstract:Relexicalization is a pivotal technique in clinical NLP, as it facilitates robust masking of sensitive information while synthesizing datasets that retain high-fidelity, real-world characteristics. However, preserving structural integrity, relational coherence, and temporal consistency during transformation remains a significant challenge. Existing approaches frequently rely on independent entity replacement, which results in clinical inconsistencies across longitudinal records. This reduces the value of such relexicalized datasets for downstream scientific analysis. To address these limitations, we introduce G-RELIC (Graph Based Contextual Relexicalization with Improved Consistency) which combines the power of LLMs with graphs. G-RELIC implements a graph-based mapping mechanism which optimizes for one-to-one correspondence between original and surrogate entities. It also introduces a deterministic temporal repositioning algorithm to preserve temporal consistency. Empirical evaluations on diverse, real-world clinical datasets validate that G-RELIC significantly outperforms state-of-the-art baselines. G-RELIC yields a 30.4 percentage point improvement in relational integrity (62.1% to 92.5%) and 45.9 percentage point improvement in temporal coherence (46% to 91.9%) without compromising on the recognized privacy benchmarks for clinical datasets. This maximizes the analytical utility of relexicalized datasets while minimizing re-identification risk.

36. 【2609.21383】Prediction Dynamics in Depth-Recurrent Language Models

链接https://arxiv.org/abs/2609.21383

作者:Xinyue Luo,Fei Yu

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Depth-recurrent language models, language models refine, models refine predictions, Depth-recurrent language, repeated latent updates

备注

点击查看摘要

Abstract:Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their scores continue to change? We derive a sharp margin characterization that decomposes the conservatism of a magnitude bound into common translation, direction relative to the winner, and the pairing of each competitor's update with its score gap. Across Huginn-3.5B and Ouro-1.4B, accounting for update direction and competitor pairing reduces the mean earliest qualifying depth by a further 22.5-34.4% of the total depth beyond translation removal under full answer-text scoring. This retrospective comparison uses completed trajectories. Substantial contributions also occur under label scoring. For shared predictive distributions, we separate common and contrast motion orthogonally and express the common component through candidate-set mass and within-set concentration. Common and contrast energies can attenuate at different rates, allowing a growing preference-change share to coexist with shrinking absolute updates. These findings explain finite-depth answer preservation through the geometry and composition of observed score changes.

37. 【2609.21378】ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

链接https://arxiv.org/abs/2609.21378

作者:Qiang Zhang,Ruixue Ding,Fanrui Zhang,Xi Chen,Boli Chen,Shihang Wang,Yinfeng Huang,Yi Zheng,Pengjun Xie,Kaipeng Zhang,Jiawei Liu,Zheng-Jun Zha

类目:Computation and Language (cs.CL)

关键词:large language model, substantially improved large, improved large language, reliable scalar rewards, language model

备注

点击查看摘要

Abstract:Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by replacing pointwise scoring with relative preferences. However, they still compress rich comparative feedback into a single trajectory-level reward, obscuring decisive intermediate steps and preventing successful behaviors from being consolidated into reusable skills. We propose ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning. ArenaFlow leverages tournament-based relative ranking to derive trajectory-level reward signals. Each comparison is further equipped with structured reflective evaluation, which reveals three types of supervision: pivotal success steps, reusable strategy skills, and usage attribution of retrieved skills. At the step level, ArenaFlow propagates trajectory-level advantages to high-confidence pivotal steps according to tournament survival depth, enabling more targeted optimization of local reasoning behaviors. At the skill level, ArenaFlow estimates skill utility from group-level usage attribution and maintains a global skill memory through utility-aware updating, pruning, and retrieval. The resulting high-utility skills further serve as policy priors for future exploration. Extensive experiments validate ArenaFlow's effectiveness on open-ended agent tasks.

38. 【2609.21362】Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

链接https://arxiv.org/abs/2609.21362

作者:Nghia Hieu Nguyen,Thai Bao Huynh,Binh-An Dinh-Le,Phu Gia Hoang,Dat Tien Nguyen,Kiet Van Nguyen,Ngan Luu-Thuy Nguyen

类目:Computation and Language (cs.CL)

关键词:internal phonological structure, Conventional tokenizers represent, requiring large vocabularies, Conventional tokenizers, overlooking the internal

备注: under review

点击查看摘要

Abstract:Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequence length while enabling representation sharing across phonologically related syllables. Non-phonological and unsupported units are handled through character-level fallback. This deterministic design requires no corpus-dependent vocabulary learning and yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese. Intrinsic evaluation shows that the tokenizer achieves substantially higher Rényi efficiency in both languages, represents every entry in a standard Vietnamese syllable dictionary with a Fertility of exactly one, and generally produces shorter Vietnamese sequences than existing pretrained tokenizers. We further instantiate the tokenizer in \textbf{PhonemicBERT}, which combines factorized component embeddings and reconstructs complete masked syllables using three prediction heads. Under a controlled Chinese pretraining setup, PhonemicBERT-Zh is competitive with or outperforms character, subword, and SubChar alternatives across diverse language-understanding tasks. PhonemicBERT-Vi also achieves competitive or superior results to established Vietnamese and multilingual pretrained models. These results establish phonemic factorization as a compact, efficient, and interpretable alternative to atomic and statistically segmented text representations.

39. 【2609.21349】From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers

链接https://arxiv.org/abs/2609.21349

作者:Ji-Lun Peng,Yi-Zhen Zhang,Chun-Nan Chou,Yun-Nung Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, impersonating remains challenging, shown strong potential, faithful impersonating remains, Large language

备注: Accepted by EMNLP 2026 Findings

点击查看摘要

Abstract:Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging. Existing in-context learning-based methods fail to capture how individuals react under different situations. In addition, LLM-based evaluation is difficult for obscure individuals. To address these challenges, we propose Situation--Internal state--Behavior Persona method to incorporate situation-dependent behavioral strategies. We further design an evaluation protocol that provides LLM evaluators with references about the impersonated individual. We evaluate our approach on a newly constructed dataset for the task of generating replies on social media. Experimental results show that our proposed method outperforms state-of-the-art ICL-based baselines, while our evaluation protocol achieves moderate correlation with human judgment. Besides, experiments on fictional-character benchmarks demonstrate that our proposed method is applicable beyond the social media setting. These findings suggest that incorporating behavioral information broadly improves the fidelity of role-playing for real individuals on social media or fictional characters.

40. 【2609.21340】Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

链接https://arxiv.org/abs/2609.21340

作者:Shuo Huang,Gholamreza Haffari,Xingliang Yuan,Ting Yu,Lizhen Qu

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:combine large language, Empirical identity leakage, large language models, Empirical identity, text is increasingly

备注

点击查看摘要

Abstract:Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finite-sample statistical guarantees, while training-time protections such as differential privacy are difficult to translate into release-time decisions for individual natural-language documents. We introduce Conformal Privacy Auditing(CPA), a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries. CPA outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with user-chosen confidence under exchangeability, together with an interpretable leakage proxy derived from set size. CPA supports both logit-access and sampling-only attackers, enabling audits of open-source models and proprietary API models in a unified framework. Across multiple release benchmarks and attacker configurations, CPA achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary, providing a statistically grounded basis for reporting and comparing release-time linkage risk across attacker configurations, datasets, and release mechanisms alike.

41. 【2609.21296】FairLMs: A Turnkey Library for Fairness in Language Models

链接https://arxiv.org/abs/2609.21296

作者:Jiale Zhang,Michael Larionov,Zichong Wang,Zhipeng Yin,Wenbin Zhang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:involves measuring bias, language models involves, models involves measuring, Fairness research, applying mitigation methods

备注

点击查看摘要

Abstract:Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so combining them requires reconciling model interfaces, evidence formats, access constraints, and result types before applicability can be checked or methods compared. We introduce \textbf{FairLMs}, a Python library that connects these activities through explicit declarations of model capabilities and input requirements. It provides 33 intrinsic and extrinsic metrics, 14 mitigation components spanning four intervention categories, 14 dataset and scoring-instrument diagnostics, adapters for the three Transformer architectures and supported hosted completion APIs, and benchmark loaders. Declarations are checked before execution and results carry the configuration under which they were obtained, so that compatible components can be combined, methods compared under a common protocol, and workflows extended to new models and datasets. The source code is available at: this https URL.

42. 【2609.21277】How Many Humans Is a Judge Panel Worth?

链接https://arxiv.org/abs/2609.21277

作者:Chao Li,Yingying Yu,Yunfeng Li

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:language models represent, models represent, language models, human judgments, MSE

备注: 18 pages, 10 figures, and 8 tables. Code and data: [this https URL](https://github.com/Chao1208/chaosnli-judge-votes)

点击查看摘要

Abstract:How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against empirical human label distributions, retaining disagreement that binary errors relative to one gold label collapse. We measure spectral residual diversity by matching the participation ratio of a normalized residual Gram matrix to conditionally independent human-reference draws, giving nu_H. We separately match distributional squared error, giving nu_MSE. Across three ChaosNLI tasks, the same 32-judge panels have nu_H=4.24--6.50 but nu_MSE=2.30--3.75. A spectral identity separates the eigenvalues, member energies, and averaging-direction weights that determine error. Realizable hard-label panels show that greater spectral diversity can accompany worse distribution recovery even with equal member energies and nonnegative correlations. In the observed panels, within-size ranking agreement varies sharply by task; some member additions produce conflicting changes that persist across two item halves. The consensus-direction share of centered residual variance is gamma_co=43.8% on MNLI-m and 33.7% on SNLI, quantifying shared variation retained by averaging. We provide aligned votes and analysis protocols for auditing these distinctions. Effective size is therefore a target-specific measurement: spectral diversity and distribution recovery should not be treated as interchangeable measures of panel quality or as general human-replacement rates.

43. 【2609.21247】When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces

链接https://arxiv.org/abs/2609.21247

作者:Yuxiang Liu,Jiaming Luo,Eleftheria Briakou,Colin Cherry

类目:Computation and Language (cs.CL)

关键词:Large Reasoning Models, Reasoning Models increasingly, Large Reasoning, machine translation, increasingly use intermediate

备注: Accepted to EMNLP 2026 Main

点击查看摘要

Abstract:Large Reasoning Models increasingly use intermediate traces for machine translation, but it remains unclear when such reasoning helps or hurts. We analyze reasoning traces across models, languages, domains, and datasets, focusing on reasoning language, length, and structure. We find that the best reasoning language is model-specific, reasoning length has a non-monotonic relationship with quality, and traces exhibit recurring functional patterns. To uncover these patterns, we introduce Hierarchical Meta-Summarization (HMS), a scalable framework that induces coarse- and fine-grained reasoning structures without predefined taxonomies. HMS reveals a shared organization--understanding/planning, translating/drafting, and refining/verifying--alongside domain-specific variation. Our results suggest that MT reasoning should be controlled in a model-aware, length-aware, and pattern-aware manner rather than uniformly encouraged.

44. 【2609.21231】Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

链接https://arxiv.org/abs/2609.21231

作者:Ruotian Wu,Bill E. Johnson,Gene Saunders,Osama Hamzeh,Ankit Vadehra,Pascal Poupart

类目:Computation and Language (cs.CL)

关键词:Grammatical Error Correction, Grammatical Error, Error Correction, Reference-based metrics, reference set enumerates

备注: 5 pages

点击查看摘要

Abstract:Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.

45. 【2609.21227】Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency

链接https://arxiv.org/abs/2609.21227

作者:Wenhan Yu,Wenxin Wu,Hao Wang,Lei Sha

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:incorrect factual outputs, commonly defined, Factual, factual outputs, paraphrases

备注

点击查看摘要

Abstract:Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced hallucination setting, where a model answers a factual question correctly in its original form but generates an incorrect answer under a semantically equivalent paraphrase. Such inconsistencies expose latent factual instability under semantic invariance. However, general-purpose paraphrases are often insufficient as robustness-oriented supervision: near-copy paraphrases provide weak signals, while overly diverse paraphrases may break semantic equivalence. In this paper, we propose HALLUCINATION-R1, a robustness-oriented paraphrase generation framework that learns to produce semantically faithful yet robustness-challenging paraphrases for factual consistency. Through two-stage optimization, it first stabilizes meaning-preserving and diverse paraphrasing, then rewards paraphrases that reveal factual consistency degradation in downstream QA models. Experiments on SimpleQuestions, PopQA, and TruthfulQA show that HALLUCINATION-R1 achieves a strong consistency--diversity trade-off and exposes robustness failures across multiple model families and datasets. Further analyses indicate that these failures are not reducible to surface-level artifacts or semantic drift, but reveal non-trivial factual instability under meaning-preserving variation. A lightweight fine-tuning study also shows that HALLUCINATION-R1-generated data improves robust accuracy under paraphrase variations, suggesting its utility for robustness-oriented training. Our code and models are publicly available at this https URL.

46. 【2609.21187】When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

链接https://arxiv.org/abs/2609.21187

作者:Md Tahmid Rahman Laskar,Xue-Yong Fu,Gundeep Singh,Karol Chang,Kevin Sanders,Shi Zong,Tania Habib,Julien Bouvier Tremblay,Shayna Gardiner,Harsh Saini,Matthias Lee,Elena Khasanova,Quinten McNamara,Shashi Bhushan TN

类目:Computation and Language (cs.CL)

关键词:gold interaction history, Agent models, interaction history, frequently evaluated, evaluated one decision

备注: Accepted to the REALM Workshop at EMNLP 2026

点击查看摘要

Abstract:Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is scored against a reference. We investigate whether improvement under this protocol is predictive of improved autonomous workflow execution. We study pre-SFT and supervised fine-tuned (SFT) Qwen3 models at 4B and 14B parameters and Gemma 3 models at 4B and 12B parameters on multi-turn customer-support workflows. We find that SFT consistently improves text-turn success, and that overall next-turn success increases for every model under gold-history evaluation. However, these improvements do not transfer to autonomous workflow execution. Tool-specific gains also vary across metrics and models. None of the four SFT models succeeds under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4% workflow success. Our results show that next-turn evaluation is not a reliable proxy for workflow success, motivating separate reporting of text quality, local action correctness, tool execution, and end-to-end task completion.

47. 【2609.21183】I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

链接https://arxiv.org/abs/2609.21183

作者:Amit Kumar Singh Yadav,Ritvik Shrivastava,Xuan Zhang,Seungwhan Moon,Shashank Jain,Pinar Donmez,Babak Damavandi

类目:ound (cs.SD); Computation and Language (cs.CL)

关键词:large language models, Audio large language, operate reactively, language models, large language

备注: Accepted at Interspeech 2026

点击查看摘要

Abstract:Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \texttt{interrupt} and \texttt{silent}, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication. Applied to Qwen2-Audio-7B, ISM achieves 99.6\% interrupt F1 and perfect de-duplication recall on ESC-50. On noisy Epic-Sounds kitchen audio, ISM achieves the highest interrupt F1 without domain-specific training, the only method maintaining strong onset detection without over-triggering or over-suppression. Streaming evaluation confirms real-time viability with 3.5-second average latency.

48. 【2609.21179】Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection

链接https://arxiv.org/abs/2609.21179

作者:Wen Zhang

类目:Computation and Language (cs.CL)

关键词:achieve high aggregate, rare morphological subclasses, systematic errors clustered, conceal systematic errors, benchmark datasets

备注: BabyLM 2026 Workshop @ EMNLP 2026 CR

点击查看摘要

Abstract:Neural morphological generation systems often achieve high aggregate accuracy on benchmark datasets, yet such performance can conceal systematic errors clustered in rare morphological subclasses. We present an orthography-aware diagnosis of Japanese past-tense verb inflection, treating hiragana not merely as a transcriptional medium but as a representational system that encodes morphophonological structure. Using two character-level Transformer architectures evaluated across five random seeds, we show that although both systems exceed 97% aggregate accuracy, a single structurally specific irregular subtype, verbs whose stems end in /e/ and require gemination before the past-tense suffix and make up fewer than 1% of the data, accounts for a disproportionate 30-43% share of residual errors and contributes roughly 34-48x its prevalence to total errors. We then move from diagnosis to causal isolation: controlled ablation experiments show that removing this subtype alone produces larger accuracy gains than removing all irregular verbs combined. These findings indicate that error concentration in neural morphological learning is not driven by irregularity per se, but by the interaction between extreme low-frequency morphological patterns and specific orthographic processes. We argue that morphological evaluation should incorporate fine-grained subclass analysis, and discuss implications for data-efficient, developmentally plausible language model pretraining.

49. 【2609.21154】CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

链接https://arxiv.org/abs/2609.21154

作者:Kailai He,Zhihao Wu,Linhai Zhang,Runcong Zhao,Yulan He,Jiazheng Li

类目:Computation and Language (cs.CL)

关键词:Good tutoring adapts, Good tutoring, Bayesian Knowledge Tracing, tutoring adapts, Good

备注: Accepted to EMNLP 2026

点击查看摘要

Abstract:Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer as a single bit of signal. We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner's mastery and misconceptions. This memory is updated as evidence accumulates and is used to generate the next personalised question. CoLearn has three components: (i) a persistent learner-state memory that updates per-topic mastery with a soft-evidence variant of Bayesian Knowledge Tracing, where a large language model acts as a continuous observation function; (ii) adaptive question generation that targets the learner's weakest topic and recurring misconceptions; and (iii) an evidence view that makes personalisation visible and testable through live progress visualisation and blind A/B comparison. In blind A/B evaluation, questions conditioned on this memory are preferred over non-personalised ones 68-69% of the time, and in persona simulations with hidden ground-truth mastery the agent's belief converges toward the learner's true mastery.

50. 【2609.21149】Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

链接https://arxiv.org/abs/2609.21149

作者:King Shi,Amanda Li,Jonathan Ivey,Synthia Qia Wang,Guan Gui,Hyunseo Kim,Peter Zandi,Jason Straub,Jacob Taylor,Ananya Joshi

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:AI-assisted psychiatric intake, psychiatric intake systems, health systems, AI-assisted psychiatric, routinely evaluate

备注: 7 pages, 3 figures, submitted to IAAI'27

点击查看摘要

Abstract:Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.

51. 【2609.21145】Scaling Forced Alignment to End-User Devices

链接https://arxiv.org/abs/2609.21145

作者:Lawry Sorenson,Michael Crandall,Eric K. Ringger,Stephen D. Richardson

类目:Computation and Language (cs.CL); Data Structures and Algorithms (cs.DS)

关键词:mine training data, online resources, Viterbi algorithm, mine training, training data

备注

点击查看摘要

Abstract:The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling poorly to long input sequences. We propose two optimizations to address this issue. First, we apply the Hirschberg algorithm to perform the alignment in place using linear memory. Second, we model the alignment between speech and text as a constrained random walk, allowing us to prune the search space with arbitrary confidence while accounting for transcription errors. The Hirschberg optimization reduces memory usage from 140 GB to 5 MB for three-hour inputs while producing identical alignments in one-third the time of torchaudio when both run on a CPU. We achieve an additional 2x speedup with pruning on inputs longer than 20 minutes while preserving alignment accuracy in more than 98% of tested cases.

52. 【2609.21117】From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

链接https://arxiv.org/abs/2609.21117

作者:Saki Imai,Mert İnan,Malihe Alikhani

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)

关键词:task completion time, interaction cost, interaction, outcome quality, task completion

备注: EMNLP 2026

点击查看摘要

Abstract:AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it. Motivated by economics literature, we introduce a productivity-oriented framework for evaluating human-AI collaboration as outcome quality relative to interaction cost. Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems.

53. 【2609.21096】Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing

链接https://arxiv.org/abs/2609.21096

作者:Amir Jalilifard,Anderson Rocha,Eric Wong,Marcos Medeiros Raimundo

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:effectively distinguish hallucinated, information flow patterns, examine the topology, effectively distinguish, attention graphs

备注

点击查看摘要

Abstract:In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.

54. 【2609.21094】Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

链接https://arxiv.org/abs/2609.21094

作者:Utkarsh Agarwal,Monojit Choudhury

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:weigh clashing moral, Large Language Models, exhibit hidden biases, Large Language, models exhibit hidden

备注: Accepted at the Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea. [this https URL](https://icml.cc/virtual/2026/75692)

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.

55. 【2609.21075】Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

链接https://arxiv.org/abs/2609.21075

作者:Mohit Chandra,Nabin Kim,Eli Min,Aamogh Sawant,Tanmay Sutar,Munmun De Choudhury

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:human lived experience, Reddit to seek, healthcare remains limited, professional mental healthcare, Large Language Models

备注: 25 pages, 6 figures, 17 tables

点击查看摘要

Abstract:As access to professional mental healthcare remains limited, many individuals turn to online platforms such as Reddit to seek peer support situated within human lived experience. However, a significant portion of such queries go unanswered, presenting an opportunity for using Large Language Models (LLMs) to fill this gap. While LLMs have demonstrated strong performance on clinical benchmarks, their ability to generate lived-experience informed and community-aligned peer support is underexplored. Addressing this gap, we introduce the COmmunity-centered Peer Engaged Support (COPES) dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives to mental health support seeking queries. Evaluating zero-shot and post-trained (SFT and DPO) models, we show that post-training on COPES significantly improves Strategy Alignment (50% for general-purpose models) and alignment in Emotion Tone. However, we also observe that such improvements are heterogeneous and alignment improvements vary significantly across subreddits and requested coping strategies. Furthermore, post-training induces distributional shifts, heavily favoring problem-focused recommendations while suppressing emotion-focused strategies. Together, this work shows that while curating community-driven data improves the alignment of LLM responses, model performance remains disparate across distinct sub-communities and specific mental health needs.

56. 【2609.21032】Scaling Discovery through Test-Time Communication

链接https://arxiv.org/abs/2609.21032

作者:Jongho Park,Vasilis Kontonis,Shivam Garg,Akshay Krishnamurthy,Dimitris Papailiopoulos

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:existing agentic systems, agentic systems capture, Science advances, existing agentic, agentic systems

备注: 34 pages, 12 figures

点击查看摘要

Abstract:Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.

57. 【2609.20995】Voice-Light: A Full-Duplex Cascaded Voice Agent with Causal Turn-Taking and Speculative Generation

链接https://arxiv.org/abs/2609.20995

作者:Bertil Braun

类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:Natural spoken interaction, spoken interaction requires, enter conversation history, ensure canceled audio, streaming ASR encoder

备注: 9 pages, 4 figures, 6 tables. Code, datasets, and model artifacts: [this https URL](https://github.com/BertilBraun/Voice-Light) ; live demo: [this https URL](https://voice.bertil-braun.de)

点击查看摘要

Abstract:Natural spoken interaction requires more than streaming ASR, language generation, and speech synthesis: a system must react to overlap without canceling on every acknowledgment, prepare a response before a turn is certain, and ensure canceled audio cannot enter conversation history. We present Voice-Light, a full-duplex cascaded voice agent that combines immediate acoustic onset, a causal adapter sharing a streaming ASR encoder, reversible playback control, and private speculative response generation. Structured tool calls execute concurrently with audible bridge speech, while browser acknowledgments make rendered audio authoritative for durable history. Locked evaluation on 1,673 real-conversation silence candidates found that an earlier learned completion checkpoint preserved a 2.70% false-cutoff rate but reached only 12.53% end-of-turn recall, compared with 95.60% for a Silero timing policy. The deployed system therefore retains a hybrid controller rather than claiming a learned-policy replacement. Across three unscripted operator-run microphone sessions, 36 measured response turns had a 758 ms median from final VAD endpoint to first server audio; 21 turns were below 800 ms. These sessions are an instrumented case study, not a controlled user evaluation. We release the synthetic data, model artifacts, evaluation code and summaries, source code, and deployment configuration supporting the result.

58. 【2609.20989】rustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged

链接https://arxiv.org/abs/2609.20989

作者:Aryan Ramchandra Kapadia,Eshwar Chandrasekharan,Koustuv Saha

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:personal financial guidance, people appraise, important for supporting, financial guidance, source labels

备注

点击查看摘要

Abstract:As generative AI is increasingly used as a source of personal financial guidance, understanding how people appraise such advice is important for supporting appropriate reliance. We conducted a randomized vignette experiment with 285 U.S. adults across eight financial decisions, independently varying three advice styles---AI, expert, and online community---and displayed source labels while holding the underlying recommendation consistent. Advice style most strongly shaped message and safety appraisals, Expert labels selectively increased perceived source knowledge, and decision context primarily shaped risk and safety appraisals. These appraisals were associated with downstream judgments, with models explaining 69.2% of overall quality, 75.9% of trust, and 82.9% of intended reliance. Expert-style advice also remained most preferred when shown without source labels. Our findings have implications for understanding financial advice evaluation, distinguishing the roles of advice style and source labels, and designing financial AI that supports grounded evaluation rather than simply maximizing trust.

59. 【2609.20945】$μ^2$-Bench: A Multilingual Machine Unlearning Benchmark

链接https://arxiv.org/abs/2609.20945

作者:Kyomin Hwang,Hyeonjin Kim,Hyunho Lee,Yearim Kim,Yeji Song,Nojun Kwak

类目:Computation and Language (cs.CL)

关键词:Large Language Models, indirect cross-linguistic spread, private data propagates, Multilingual Large Language, Undesired information

备注

点击查看摘要

Abstract:Undesired information such as harmful content and private data propagates through Multilingual Large Language Models (LLMs) via direct training and indirect cross-linguistic spread. Multilingual Machine Unlearning (MMU) aims to remove such information, yet its evaluation remains underexplored, leaving unclear whether unlearning truly eliminates target knowledge across all languages. To bridge this gap, we introduce $\mu^2$-Bench, an MMU benchmark that simulates the full pipeline of memorization, unlearning, and evaluation across diverse languages. It 1) spans a broad set of languages, 2) evaluates on both training and hold-out languages, and 3) assesses knowledge as dispersed across multiple languages. We show that successful MMU requires methods that reflect multilingual characteristics, and conduct analysis to provide deeper insights into MMU.

60. 【2609.20902】Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

链接https://arxiv.org/abs/2609.20902

作者:Runze Hu,Jingqi Kong,Yang Yang,Yihang Yang,Jingyao Liu,Haizhou Tang,Shanghang Zhang,Zheng Liu

类目:Computation and Language (cs.CL)

关键词:Motivational interviewing, health behavior change, elicit autonomous motivation, collaborative approach, approach to elicit

备注

点击查看摘要

Abstract:Motivational interviewing (MI) is a collaborative approach to elicit autonomous motivation for health behavior change. Generative AI (GenAI) offers new ways to deliver MI via conversational systems, but evidence on their design, assessment, and translation into interventions remains fragmented. This scoping review characterized evidence on GenAI-MI chatbots across system design, safety, MI quality, user perceptions, and intervention outcomes. We conducted a PRISMA-ScR scoping review. Nine datasets were searched for studies published or publicly available from January 1, 2015 to June 2, 2026 that used GenAI to generate MI chatbot responses or counselor utterances. Data were extracted using a predefined framework and synthesized descriptively. Forty-seven reports (48 studies) were included. Twenty (41.7%) focused on system design without direct participant use; 28 (58.3%) involved direct interaction. Most systems were text based and disembodied; 23 (47.9%) incorporated dynamic adaptation. Safety measures were unevenly reported. Among studies with direct use, 21/28 (75.0%) reported informed consent or user education. Thirty (62.5%) assessed MI quality, generally suggesting MI-consistent interactions. User perceptions were favorable, especially empathy, usability, helpfulness, and intention to use, though measures were heterogeneous. Eighteen (37.5%) reported intervention outcomes, mostly after a single session. Positive findings were more consistent for short-term motivation than sustained behavioral or functional change. GenAI-MI chatbots can deliver MI-consistent interactions perceived favorably, but evidence for sustained behavioral or functional change is limited. Future research should strengthen runtime safety monitoring, standardize MI quality assessment, and use longer-term comparative designs with behavioral and functional outcomes.

61. 【2609.20886】BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

链接https://arxiv.org/abs/2609.20886

作者:Chuxuan Hu,Yeye He,Penny Zhou,Wee Hyong Tok,Daniel Kang,Surajit Chaudhuri

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB)

关键词:enterprise decision-making, enterprise users, cornerstone of enterprise, Tableau, business questions

备注: code and data are available at \url{ [this https URL](https://github.com/Hu-Chuxuan/bi-agent) }

点击查看摘要

Abstract:Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (LLMs) in working with data, we study their ability to answer BI questions end-to-end, without requiring users to manually perform the tedious preparation steps. To do this, we harvest a large collection of real-world BI projects from public sources, and manually extract pairs of (questions, ground-truth answers) from real user dashboards. The resulting benchmark, BI-Bench, is the first benchmark to systematically study LLMs' ability on end-to-end BI. We find that even frontier LLMs perform poorly on BI-Bench, with less than 50% accuracy. To address their limitations, we design a tool-augmented BI-Agent that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages. Furthermore, we develop a post-training framework that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL). BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, and post-trained BI-Agent yields gains of up to 30 points. Our results highlight the importance of combining tool-augmented reasoning with domain-specific post-training in complex BI workflows, and point to promising directions for future research.

Comments:
code and data are available at \url{this https URL}

Subjects:

Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB)

Cite as:
arXiv:2609.20886 [cs.LG]

(or
arXiv:2609.20886v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2609.20886

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
62. 【2609.20850】MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

链接https://arxiv.org/abs/2609.20850

作者:Yueming Lyu,Yilian Shi,Haoxiang Tan,Linzhuang Zou,Qihao Wang,Guihua Yu,Jie Qin,Xin Gao,Chenyang Si,Jing Dong,Caifeng Shan

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Multimodal Large Language, Language Models, show remarkable advancements, Large Language

备注

点击查看摘要

Abstract:While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.

63. 【2609.20849】Enhancing Audio Reasoning via Semantic Summary Prediction

链接https://arxiv.org/abs/2609.20849

作者:Francesco Bonzi,Pooneh Mousavi,Cem Subakan,Mirco Ravanelli

类目:Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

关键词:Large Audio Language, reduces accuracy compared, complex question answering, Audio Language Models, Large Audio

备注: Accepted at Interspeech 2026

点击查看摘要

Abstract:Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Semantic Prediction for Audio REasoning), which introduces a register token aligned with the final conclusion using a cosine similarity loss with a Sentence-BERT embedding. This conditions the model's latent space with the target semantic goal before reasoning begins. Experiments on MMAU and MMAR with SALMONN show improved zero-shot reasoning and stronger early attention to audio without additional inference cost.

64. 【2609.20847】Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media

链接https://arxiv.org/abs/2609.20847

作者:Cris Huynh,Arlene Pham

类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:common mental health, mental health conditions, common mental, mental health, people often write

备注: 11 pages, 5 figures, 2 tables

点击查看摘要

Abstract:Anxiety is among the most common mental health conditions, and people often write about it online well before seeking clinical help. Practitioners building detection tools face a concrete choice: call a frontier commercial model, fine-tune a smaller model in-house, or deploy a conventional classifier. We compare six conditions spanning all three on a held-out Reddit test set under a single controlled protocol. We also identify a confound in how this task is evaluated. In the corpus used here, 69.3% of anxiety-labelled posts contain the word "anxiety" or a variant, roughly twice the rate of comparable conditions, so a classifier can score well by keyword matching rather than by modelling the language of the condition. We therefore evaluate every model twice, on original text and with those terms deleted, and report the difference as lexical dependence. A frontier model leads on anxiety F1 (0.846), but a 110M-parameter domain-adapted encoder reaches 0.831 with no external API dependency, and mental-health domain pretraining accounts for only 0.7 of those points. Lexical dependence spans 8.6 to 25.4 points and does not track model capability: the LoRA fine-tuned 3B model is the most keyword-dependent condition tested, above even a TF-IDF classifier, while the frontier zero-shot model is the least. Published figures on this corpus are therefore upper bounds, and the inflation is largest for the fine-tuned models such figures typically report.

65. 【2609.20846】Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models

链接https://arxiv.org/abs/2609.20846

作者:Polina Tsvilodub,Max Höth,Michael Franke,Björn Deiseroth,Carina Kauf

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:providing correct answers, provide additional evidence, large reasoning models, modern large reasoning, excel at providing

备注: 19 pages, 9 figures

点击查看摘要

Abstract:While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models' efficiency (44% shorter CoTs on average).

66. 【2609.20845】Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders

链接https://arxiv.org/abs/2609.20845

作者:Yasir Mehmood,Kashif Javed

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:text conventionally consumes, turns video, video or audio, conventionally consumes, consumes the entire

备注

点击查看摘要

Abstract:A decoder that turns video or audio into text conventionally consumes the entire input before emitting a word. Offline this is merely more than the task requires; live it is impossible, since a caption cannot wait for a match to end. Streaming systems bolt on a fixed rule such as wait-$k$, which waits for the same number of input tokens before every word, regardless of the input's length or pace. We replace the fixed offset with ZENDAYA, a schedule governed by a single continuous parameter $\gamma$. It makes the visible source prefix a closed-form function of generation progress, scaled by the input's own predicted length, so an ordinary offline decoder and a real-time streaming decoder become two endpoints of one family rather than separate models. The same scalar fixes, in closed form, the mean fraction of source consumed per emitted word, $\bar{E}(\gamma) \approx 1/(1+\gamma)$, making it at once a latency dial and an interpretable budget. We prove a structural dependency theorem: under any schedule fixed in advance and non-decreasing, no emitted token can depend on input that has not yet arrived. The guarantee holds for trained and untrained weights alike, and extends to unbounded streams under arbitrary asynchronous arrival. The empirical result is counterintuitive: seeing less can produce better text, because a flood of source dilutes attention exactly when the model has the least of its own output to anchor on. Trained from scratch across two modalities and three public corpora (Charades-STA, ActivityNet Captions, LibriHeavy), a compact 29M-parameter decoder matches or beats the fixed schedule while reading less of the source, with the sharpest gains at the lowest latencies, where a fixed offset collapses. Streaming METEOR gains are statistically significant on all three corpora.

Subjects:

Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2609.20845 [cs.CL]

(or
arXiv:2609.20845v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.20845

Focus to learn more

              arXiv-issued DOI via DataCite</p>
67. 【2609.20844】Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces

链接https://arxiv.org/abs/2609.20844

作者:Zihan Wang,Hao Wang,Boyuan Jiang,Yiqun Zhang,Shi Feng,Xiaocui Yang,Yiwen Ye,Jianghang Lin,Xiaozhong Ji,Jinghao Lin,Kai Wu

类目:Computation and Language (cs.CL)

关键词:real-world web environments, Agentic Reinforcement Learning, agents interact, rapidly over time, interact with real-world

备注

点击查看摘要

Abstract:Deepresearch (DR) agents interact with real-world web environments through multi-turn search and visit, causing their contexts to grow rapidly over time. We observe that, even after DR Agentic Reinforcement Learning (DR-RL), 61.6% of the model's remaining prediction errors can still be attributed to insufficient long-context understanding, including longcontext hallucination and failures in cross-document evidence integration. It motivates us to further break the bottleneck of DR-RL by strengthening the model's long-context ability. However, effective LongContext training requires more than simply increasing context length. To bridge the data gap, we propose `DR Rollouts to LongContext-QA (DR-to-Long)'. The method repurposes DR-RL trajectories, which naturally contain search histories, visited webpages, evidence snippets, and final-answer supervision. It then replaces the compact snippets and webpage summaries in each trajectory with the full contents of their corresponding URLs, producing substantially longer multi-document contexts while preserving the original evidence relationships. Building on DR-to-Long, we introduce DLD (DR - LongQA - DR)-RL. DLD-RL first performs a short DR-RL stage to collect rollout trajectories, which are then converted into LongQA instances at zero annotation cost. The model is subsequently optimized with LongQA-RL to strengthen LongContext ability, followed by full DR-RL to continue improving its DR capability. Experiments show that DLD-RL outperforms standard DR-RL by 7.3% on three Deepresearch benchmarks and improves performance by 13.5% on three long-context benchmarks.

68. 【2609.20843】VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

链接https://arxiv.org/abs/2609.20843

作者:Jinke Wu,Zhengpin Li,Mengzhe Jia,Yang Li,Wentao Zhang

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:achieved substantial progress, Knowledge graph question, Knowledge graph, enables models, achieved substantial

备注: Preprint

点击查看摘要

Abstract:Knowledge graph question answering (KGQA) enables models to answer natural-language questions through structured graph reasoning and has achieved substantial progress across many benchmarks and applications. Recently, multimodal KGQA (MM-KGQA) has attracted increasing attention because many questions require jointly using multimodal inputs and KG evidence. However, existing MM-KGQA methods typically use multimodal information only for starting entity grounding or evidence retrieval, after which multi-hop reasoning degenerates into text-only graph search. As a result, they cannot exploit multimodal cues that become important at intermediate hops. To address this limitation, we propose VISPATH, a visual-intent-guided path reasoning framework for MM-KGQA. VISPATH first identifies a reliable starting entity by combining multimodal grounding with graph-structural cues. It then performs intent-guided path discovery by recomputing hop-specific multimodal intent from the input, question, and current partial paths, so that each expansion is guided by the current reasoning state. The discovered paths are further refined through reasoning-chain pruning, which evaluates candidate paths as complete evidence chains based on their consistency with the question, reasoning sketch, and hop-specific intent. Finally, VISPATH checks whether the selected evidence is sufficient for answer generation. We further construct VISPATH-Bench, a benchmark for evaluating multimodal multi-hop reasoning over KGs, covering questions that require two to four hops over KG paths. Extensive experiments on VISPATH-Bench and three additional multimodal QA benchmarks show that VISPATH consistently outperforms strong baselines. Notably, with GPT-4o as the backbone, VISPATH surpasses GPT-5.4 on VISPATH-Bench, achieving a 10.6% relative improvement in average accuracy and a 13.1% improvement at 2-hop reasoning.

69. 【2609.20842】COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training

链接https://arxiv.org/abs/2609.20842

作者:Qifeng Cai,Xuanguang Pan,Hao Liang,Chang Xu,Wentao Zhang

类目:Computation and Language (cs.CL); Databases (cs.DB)

关键词:translates natural-language questions, open-source large language, real-world SQL generation, executable SQL queries, large language models

备注

点击查看摘要

Abstract:Text-to-SQL translates natural-language questions into executable SQL queries, but open-source large language models still require task-specific post-training for complex, real-world SQL generation. Effective post-training requires both training data that cover the capabilities demanded by the target task and a learning strategy that enables the model to acquire them. Existing datasets provide valuable supervision but incompletely cover SQL structures, while augmentation methods typically expand data without identifying structural gaps. Moreover, supervised fine-tuning (SFT) or reinforcement learning (RL) alone cannot dynamically address weaknesses exposed during training. We propose COAL-SQL, a unified framework combining Coverage-Guided Augmentation (CGA) and Failure-Driven Learning (FDL). CGA uses greedy selection to identify SQL structures missing from the original dataset and constructs complementary examples, improving structural coverage. FDL retains GRPO as the main optimization objective while supplying targeted supervision for unsolved examples. At the step level, it applies SFT to verified reasoning traces generated by a strong LLM for accumulated failures. At the epoch level, it retrieves structurally related examples based on accumulated failures to create targeted practice, helping the model acquire the corresponding SQL capabilities. With only 12,600 distinct post-training examples, COAL-SQL achieves 64.9% execution accuracy on the BIRD development set and outperforms baselines trained at comparable scale. The code is available at this https URL.

70. 【2609.20839】Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition

链接https://arxiv.org/abs/2609.20839

作者:Matthew Kit Khinn Teng,Haibo Zhang,Takeshi Saitoh

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)

关键词:visual speech recognition, phoneme prediction errors, speech recognition, realistic phoneme prediction, visual speech

备注: Submitted for journal publication and currently under consideration

点击查看摘要

Abstract:Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.

71. 【2609.20838】From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News

链接https://arxiv.org/abs/2609.20838

作者:Zeynep Özdemir,Murat Osmanoğlu,Sevgi Yiğit-Sert,Ömer Özgür Tanrıöver,Yılmaz Ar

类目:Computation and Language (cs.CL)

关键词:modern LLMs generate, examine how modern, generate and detect, controlled settings, detect fake

备注: 35 pages, 8 figures, 5 tables

点击查看摘要

Abstract:In this study, we examine how modern LLMs generate and detect fake news under controlled settings across four manipulation scenarios. These are open-ended generation, rewriting, manipulation prompts and attribute based prompts grounded in the journalistic discourse framework. Firstly, using seven widely adapted models, we created a synthetic fake news corpus with 14000 generated articles across these four scenarios. Then we analyzed its linguistic properties to assess how closely model-generated news resembles real news structurally and semantically. Finally, to evaluate detection performance, we conducted experiments where each model judges generated fake news, starting with a basic detection prompt and improved prompts developed through an iterative refinement process that extracts misleading patterns from real-fake pairs. Our results revealed substantial variation across models in both generating and detecting misinformation, demonstrated that the generation strategy strongly influences detectability, and show that the refined prompt does not improve and often harms detection performance. Therefore, the study provides a systematic assessment of LLMs detection capability of LLMs generated fake news across typical generation scenarios.

72. 【2609.20836】PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

链接https://arxiv.org/abs/2609.20836

作者:Mengxuan Li,Junfa Chen,Jinze Xia,Yundan Chen,Lixin Fan,Ke Liu,Keyue Shi,Haishuai Wang

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:typically require task-specific, require task-specific adaptation, physiological signal, models typically require, Physiological

备注

点击查看摘要

Abstract:Physiological signals support diverse clinical and monitoring tasks, yet existing physiological signal foundation models typically require task-specific adaptation for each task. Natural language provides a common interface for specifying different prediction objectives, but the ability of current models to follow such instructions across physiological signal modalities remains insufficiently evaluated. To address this gap, we introduce PhysioBench, a unified benchmark for physiological signal question answering. PhysioBench harmonizes annotations from 22 public datasets into 61.4 million questions across 30 tasks. Each question-answer pair is grounded in a signal segment and traceable to its source annotation. We evaluate 21 representative models, including large language models, vision-language models, time-series language models, and physiological signal foundation models under three complementary settings. The results show that none of the evaluated models achieves consistently strong performance across physiological signal modalities and tasks. The incorporation of natural language supports unified prediction across tasks, although performance remains sensitive to question formulation. Beyond these findings, PhysioBench offers an extensible platform for fine-grained analysis and future research on physiological signal understanding. Our codes are available at this https URL.

73. 【2609.20835】A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models

链接https://arxiv.org/abs/2609.20835

作者:Nicolas Turenne

类目:Computation and Language (cs.CL)

关键词:content remains undeciphered, fifteenth-century codex written, remains undeciphered, fifteenth-century codex, codex written

备注

点击查看摘要

Abstract:Background: The Voynich Manuscript is a fifteenth-century codex written in an unknown script whose content remains undeciphered. Previous studies suggest that its statistical properties resemble those of natural languages, while its illustrations - primarily plants - recall medieval herbals. Methods: We present a multidisciplinary analysis combining probabilistic modeling, phonetic decomposition, rare-event detection, and multimodal image analysis, based on a newly transliterated corpus. Word- and letter-level distributions are modeled using position-dependent probabilistic grammars, while phonetic patterns are compared across Indo-European, Semitic, and Asian languages. Image-text alignment methods based on large language models are applied to identify potential botanical correspondences. Results: The results indicate that Voynich symbols behave as letters rather than syllabic units, while word-length distributions resemble syllabic structures. Phonetic analyses show closer alignment with consonant-heavy languages such as Hebrew or Arabic than with Indo-European languages. Probabilistic modeling reproduces Zipf-like distributions and reveals extremely low probabilities for repeated initial-letter sequences, indicating a structured imitation of natural language. Image analysis suggests strong correspondences between Voynich plant illustrations and those found in Pseudo-Apuleius herbals from the Mediterranean tradition, consistent with an imitation of medieval medicinal books. Perspectives: These findings support the hypothesis that the Voynich Manuscript follows a structured generative system combining linguistic regularities and herbal knowledge, and demonstrate the value of integrating probabilistic and AI-assisted approaches in the analysis of historical manuscripts.

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.20835 [cs.CL]

(or
arXiv:2609.20835v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.20835

Focus to learn more

              arXiv-issued DOI via DataCite

Submission history From: Nicolas Turenne [view email] [v1]
Tue, 28 Jul 2026 16:51:57 UTC (3,859 KB)

74. 【2609.20834】owards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models

链接https://arxiv.org/abs/2609.20834

作者:Mostafa Anouar Ghorab,Mohamed Aymen Saied

类目:Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:rapidly evolving landscape, increasingly adopting infrastructure, adopting infrastructure models, Organizations are increasingly, emphasize scalability

备注

点击查看摘要

Abstract:In the rapidly evolving landscape of cloud-native computing, Organizations are increasingly adopting infrastructure models that emphasize scalability, flexibility, and efficiency. Kubernetes has become the de facto standard for orchestrating containerized applications in these environments. However, the inherent complexity of cloud-native ecosystems introduces significant challenges, particularly in the form of misconfigurations that can compromise both security and performance. This study explores the potential of Large Language Models (LLMs) in identifying Kubernetes misconfigurations. We introduce a comprehensive taxonomy of common misconfiguration types, offering a structured framework to better understand and categorize these issues. Additionally, we conduct an empirical evaluation of state-of-the-art detection tools to benchmark their effectiveness. Furthermore, we analyze the Kubernetes objects most prone to misconfiguration and evaluate the severity of the identified issues. By leveraging advanced machine learning techniques, including LLMs, we provide novel insights into enhancing misconfiguration detection methodologies.

75. 【2609.20833】ranssion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge

链接https://arxiv.org/abs/2609.20833

作者:Zhecheng Ren,Xuanji He,Xiaoxiao Li,Zhichen Han,Gaoyang Dong,Gaosheng Zhang,Minchuan Chen,Fengjie Zhu

类目:Computation and Language (cs.CL)

关键词:Transsion Speech Team, multilingual conversational speech, Transsion Speech, Speech Team submission, submission to Task

备注

点击查看摘要

Abstract:This paper presents the Transsion Speech Team submission to Task 1 of the MLC-SLM 2026 Challenge, which focuses on speaker-attributed transcription for multilingual conversational speech. We propose a cascaded framework consisting of three components: a speaker diarization module, a long-form multilingual ASR module, and a speaker-transcription fusion module. The diarization module is built upon DiariZen and produces speaker-homogeneous segments through local speaker activity estimation and global speaker clustering. The ASR module is based on Qwen3-Omni and generates multilingual transcriptions, while an external CTC-based alignment model provides precise word- and character-level timestamps. Finally, the fusion module combines diarization outputs with timestamped transcriptions to generate speaker-attributed STM outputs. Experimental results on the official evaluation set demonstrate the effectiveness of the proposed framework. The submitted system achieves a tcpMER of 15.41% and ranks second among all participating teams.

76. 【2609.20832】atBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar

链接https://arxiv.org/abs/2609.20832

作者:Ilshat Saetov,Dmitry Gaynullin

类目:Computation and Language (cs.CL)

关键词:Qypchaq Turkic language, ISO 639-3 tat, Turkic language written, Qypchaq Turkic, ISO 639-3

备注: 11 pages. Dataset: [this https URL](https://huggingface.co/datasets/ilchats/TatBLiMP)

点击查看摘要

Abstract:We introduce TatBLiMP, the first benchmark of linguistic minimal pairs for Tatar (tt, ISO 639-3 tat), a Qypchaq Turkic language written in Cyrillic. To our knowledge it is the first grammaticality evaluation for Tatar language models of any kind, since even the 101-language MultiBLiMP does not include Tatar. TatBLiMP covers 16 morphosyntactic phenomena in 1248 sentence pairs. Each pair differs by a single morpheme, one grammatical and one ungrammatical. A model passes a pair when it assigns higher probability to the grammatical member. Scoring compares probabilities the model already assigns, so the benchmark needs no text generation and no parser, and it runs on base models and on mid-training checkpoints. TatBLiMP adapts the phenomenon inventory and single-morpheme breaking operations of TurBLiMP to Tatar and adds one phenomenon specific to Tatar, bare-noun number after numerals and quantifiers. The grammatical member of every pair is an attested sentence from Tatar literary prose. The ungrammatical member is produced by a deterministic single-morpheme perturbation with the apertium-tat transducer. Every pair is ratified by a native speaker. A plausibility principle governs construction, so the ungrammatical member is a plausible real-world error rather than an arbitrary corruption. Across from-scratch Tatar models, cross-lingual adaptations, and frontier multilingual LLMs, the benchmark tracks focused Tatar training rather than parameter scale. A 478M from-scratch model and a 125M monolingual model lead near 0.97, a 7B adaptation trails, frontier LLMs of 30-120B parameters fall to 0.80-0.92, and a lightly tuned multilingual model is weakest. We close with the benchmark's main limitation. Its inherited taxonomy omits the morphophonology, vowel harmony and consonant assimilation, that is most salient to native speakers, and we sketch a native second layer that would add it.

77. 【2609.20831】Recursive Language Models Generalize Out of Domain

链接https://arxiv.org/abs/2609.20831

作者:Chenxiao Yang,Zhiyuan Li,David McAllester,Nathan Srebro

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:recursive language models, study when limiting, language model, language models, improves learning

备注

点击查看摘要

Abstract:We study when limiting what a language model can see improves learning. We compare standard CoT, the more general learner that reads the full trace, with recursive language models, which restricts itself by solving each subtask in an isolated context. In-distribution, this generality comes for free: CoT can efficiently simulate the recursive rule, so the IID generalization guarantee changes only by a constant factor, and recursion does not offer much. But out of domain, CoT can fit training by relying on context outside the current subtask, i.e. a shortcut that breaks once those tokens change; recursive context isolation rules out this failure mode. Even though CoT's class still covers the recursive rule, simplicity bias picks the shortcut over the truth. Thus, to go beyond distributional accuracy and truly reason, covering the right rule is not enough; this contrasts with classical learning theory.

78. 【2609.20830】Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions

链接https://arxiv.org/abs/2609.20830

作者:Sean Diab

类目:Computation and Language (cs.CL)

关键词:revise earlier content, repeated sequence-level computation, edit-based approaches obtain, earlier content, sequence-level computation

备注: 45 pages, 2 figures. Code: [this https URL](https://github.com/Sean-Diab/Reviser) Checkpoints: [this https URL](https://huggingface.co/sean-diab/reviser-checkpoints)

点击查看摘要

Abstract:Revision-capable generation is appealing because it can insert or revise earlier content, but many non-autoregressive and edit-based approaches obtain this flexibility through repeated sequence-level computation. We propose Reviser, a decoder-only Transformer that generates a response as a sequence of cursor-relative actions on a mutable canvas. At each step, Reviser predicts exactly one action token: INSERT(token), MOVE($\Delta$), or STOP, and is autoregressive over edit-history actions rather than final text order. This design enables genuinely non-monotonic generation while preserving a simple next-action interface. On a continuation benchmark, Reviser is strongly preferred to SEDD and MDLM in our arena evaluations, and trajectory statistics confirm that the model performs frequent backward moves and mid-canvas insertions rather than merely emulating end-append decoding. Against size-matched autoregressive baselines, Reviser is competitive at both the 100M and 300M scales. Under our shared FLOPs convention, Reviser also requires substantially less inference compute than representative multi-pass refinement and diffusion-style baselines.

79. 【2609.20829】SAGE: Schema-Guided LLMs for Grant Review

链接https://arxiv.org/abs/2609.20829

作者:Erik Varapaev,Andrei Chetvergov,Stepan Ukolov,Timofei Sivoraksha,Alexander Evseev,Sergey Bolovtsov

类目:Computation and Language (cs.CL)

关键词:apply detailed criteria, Aspect-Based Grant Evaluation, colleagues can inspect, reviewers must apply, supporting documents

备注: 14 pages, 2 figures, 10 tables

点击查看摘要

Abstract:Grant reviewers must apply detailed criteria to application forms, budgets, and supporting documents while producing assessments that colleagues can inspect. We present SAGE, Schema-Guided Aspect-Based Grant Evaluation, a system that translates a grant rubric into structured checks and links its judgements to evidence from the application package. We evaluate SAGE in two stages on 35 nonprofit grant applications. A post-factum comparison with 105 reviews from the original competition shows fair ordinal agreement (kappa = 0.29). The foundation then conducted a criterion-level re-review after inspecting SAGE, producing 202 assessments. In this assisted round, SAGE reached kappa = 0.58 and outperformed a one-prompt-per-criterion baseline (kappa = 0.33 on the common subset), with higher rank correlation and lower error. A claim-level audit further identifies confirmed, disputed, and unaddressed parts of the structured draft. SAGE operationalizes the review methodology by producing a detailed, evidence-linked, and auditable draft for expert correction.

80. 【2609.20828】Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

链接https://arxiv.org/abs/2609.20828

作者:Fiza Husain,Ankit Pandey,Yash Singh

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Word Error Rate, accented conversational English, miss named entities, ASR systems optimised, conversational English

备注: 5 pages, 1 figure, 1 table, accepted at Interspeech 2026

点击查看摘要

Abstract:ASR systems optimised for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both critical for language-learning feedback. We present a three-stage pipeline for speakers from India, Indonesia, and Latin America: (1) heuristic SQL filters curating entity-rich training data at 2.8x the entity density of random sampling, (2) regional LoRA adapters fine-tuned on Qwen2.5-Omni-3B producing both verbatim and corrected transcripts in a single forward pass, and (3) a six-category error taxonomy validated by an LLM-based judge (83.8% agreement, 210 human-labelled samples). The pipeline achieves 80-85% entity recall (up from 53-55%), 76-86% filler recall (up from 5%), and 6-10% WER across 6k test utterances, outperforming Whisper and a commercial ASR on entity recall while matching a zero-shot 30B model with 10x fewer parameters. Paired bootstrap tests confirm that curation alone accounts for 2.8-4.2 pp of entity recall gain (p0.0001).

81. 【2609.20827】From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators

链接https://arxiv.org/abs/2609.20827

作者:Won Seok Jang,Zonghai Yao,Hong Yu

类目:Computation and Language (cs.CL)

关键词:interactive teaching task, Hospital discharge education, Hospital discharge, interactive teaching, clinician adapts

备注

点击查看摘要

Abstract:Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.

82. 【2609.20826】ALON: A Temporally Aware Longitudinal Framework for Radiology Report Generation

链接https://arxiv.org/abs/2609.20826

作者:Nien-Tsyr Sun,Min-Chen Chen,Hui Nien Hung,Vincent S. Tseng

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:radiology report generation, produce descriptive reports, descriptive reports based, Current radiology report, detect subtle interval

备注

点击查看摘要

Abstract:Current radiology report generation (RRG) models usually produce descriptive reports based on a single examination or only the most recent prior examination, limiting their ability to perform accurate and meaningful longitudinal comparisons and detect subtle interval changes. Although recent approaches have begun to incorporate multiple prior examinations, they usually aggregate a fixed-length history without explicitly modeling the role-dependent relevance of each prior examination before fusion. To address this, we propose TALON, a Temporally Aware LONgitudinal RRG framework that adaptively integrates variable-length patient histories. The underlying Dual-Channel Temporal Fusion Module (DCTFM) compares the current examination with each prior examination through complementary similarity and change channels to capture persistent findings and interval changes, respectively. The specially designed channel-specific attention estimates the relevance of each prior examination, while a learned prior-specific gate adaptively integrates informative longitudinal evidence and suppresses redundancy. Experiments on MIMIC-CXR show that TALON outperforms the current state-of-the-art method on various clinical efficacy and graph-based metrics. When more prior examinations become available, TALON's performance on these metrics improves even further, emphasizing the strength of TALON's DCTFM in modeling longitudinal RRG across longer and more complex patient histories than existing approaches.

83. 【2609.20825】HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction

链接https://arxiv.org/abs/2609.20825

作者:Gia-Bach Nguyen,Hoang-Ha Nguyen,Tuan-Cuong Vuong,Trang Mai Xuan,Duy Quoc Ngo,Tien-Cuong Nguyen,Huan Vu,Thien Van Luong

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Electronic Health Record, Health Record data, structured Electronic Health, Electronic Health, Health Record

备注: 12 pages, 4 figures, The 15th Conference on Information Technology and its Applications

点击查看摘要

Abstract:Clinical predictive models often rely on structured Electronic Health Record data, such as time-series and procedure codes. While recent approaches have begun leveraging unstructured clinical notes, they typically encode them as flat sequences, which may lose explicit relational and temporal structure present in clinical narratives. In response, we propose HERMES, a graph-based framework that operates exclusively on clinical text while preserving clinical relationships. This approach builds on two key ideas. First, personalized Knowledge Graphs (KGs) are constructed through Large-Language-Model-guided extraction from clinical notes with Contrastive Logic Modeling that explicitly captures temporal dynamics and treatment failures and changes in outcomes. Second, a Graph Attention Network synthesizes patient representations through graph-based learning over the KGs. Experiments on MIMIC-III and MIMIC-IV for in-hospital mortality and 30-day readmission prediction show that HERMES consistently outperforms strong text-only baselines. Our findings demonstrate that explicit relational modeling with Contrastive Logic Modeling significantly advances predictive performance.

84. 【2609.20824】Do small language models know what they don't know?

链接https://arxiv.org/abs/2609.20824

作者:Prashant Mudgal

类目:Computation and Language (cs.CL)

关键词:Small Language Models, Small Language, billion parameters, Language Models, consumer hardware

备注: 9 pages, 8 figures

点击查看摘要

Abstract:We explore whether entropy-based confidence signals can be leveraged to improve the accuracy of Small Language Models (SLMs) with fewer than 3 billion parameters, running entirely on consumer hardware. We evaluate seven distinct approaches, including token-level entropy early stopping, semantic entropy estimation, and uncertainty-aware routing to larger expert models, across 7 model pairs and 5 standard NLU benchmarks. Our key finding is that token-level entropy is effectively blind in SLMs: in 91% of dataset-model combinations, mean token entropy is near zero regardless of answer correctness, rendering token-based confidence signals unusable at this scale. We demonstrate that semantic entropy, computed by generating multiple samples, clustering answers by meaning, and measuring distributional uncertainty, recovers a viable confidence signal. Using semantic entropy to selectively route uncertain queries to a larger expert model yields accuracy improvements of up to +50 percentage points. Notably, cross-family routing (e.g., SmolLM 360M to Phi-3.5-mini) averages +22.0% improvement compared to +6.8% for same-family routing, revealing that expert model quality matters more than architectural compatibility. Our results suggest that the value proposition for entropy-based methods in SLMs is not computational savings but intelligent compute allocation: spending more tokens where they matter most.

85. 【2609.21676】he Spoken Wikipedia Presentation Corpus

链接https://arxiv.org/abs/2609.21676

作者:Thomas Ranzenberger,Steffen Freisinger,Tobias Bocklet,Korbinian Riedhammer

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)

关键词:Wikipedia Presentation Corpus, Wikipedia Corpora featuring, Spoken Wikipedia Presentation, Spoken Wikipedia Corpora, Corpora featuring LLM-generated

备注: Accepted at SLT 2026

点击查看摘要

Abstract:We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content planning with rule-based design decisions. For each section, an LLM generates a slide title, bullet points, a takeaway message, and a visual description that is used to create an illustration. Rule-based matching then selects layouts, themes, and styles to produce the final slides. A vision LLM extracts slide text as Markdown. We evaluate multiple ASR and spoken language models (SLMs). The best model achieves an average micro-WER of 10.23% and an average micro-CER of 6.48% on audio-only inputs. English yields the lowest error rates, followed by German and Dutch, while performance declines across lower-resource languages. Although audio-only baselines are strong, multimodal zero-shot prompting of omni models remains challenging. The aligned slide, text, and audio data show a strong potential to improve recognition through cross-modal context.

86. 【2609.21084】he Hidden Cost of Digits: Number Normalization and WER in ASR Systems

链接https://arxiv.org/abs/2609.21084

作者:Stanisław Kacprzak,Mieszko Fraś

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:Modern automatic speech, Arabic numerals, Modern automatic, written in Arabic, automatic speech recognition

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals. This creates a need for fair comparison with models that output verbatim texts and proper processing of reference transcripts. Popular approaches often reduce text normalization to lowercase and remove punctuation, with no additional normalization applied to languages other than English. In this work, we analyze the impact of normalization of numerical expressions in the evaluation of ASR systems in various languages, using Polish as an example of a highly inflective language. We perform experiments on VoxPopuli and The Polish Parliamentary speech datasets and estimate word error rate (WER) differences for different text normalization approaches. We show that the difference due to the lack of number normalization in WER may be substantial - more than 2 percentage points, and often higher than the differences between systems in popular multilingual benchmarks.

87. 【2609.20875】Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation

链接https://arxiv.org/abs/2609.20875

作者:Simon Pals,Cristian Tejedor-Garcia

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

关键词:Parkinson disease, cost-effective severity assessment, facilitating accessible, progression tracking, severity assessment

备注: Accepted and published at IEEE SLT 2026 - IEEE Spoken Language Technology 2026. OneVoice-MSD 2026: Multilingual Speech Technologies for Motor Speech Disorders. [this https URL](https://attend.ieee.org/slt-2026/) Please cite the conference version

点击查看摘要

Abstract:Parkinson's disease (PD) often manifests through speech impairments, facilitating accessible, non-invasive, and cost-effective severity assessment for early diagnosis and progression tracking. Despite advances in speech foundation models (SFMs), their cross-lingual generalization for PD severity multi-class classification remains underexplored due to limited labeled data, a lack of explainable methods and variability across languages and datasets. In this work, we evaluate pre-trained embeddings from four state-of-the-art open-source SFMs across three datasets in zero-shot and k-shot cross-lingual settings for multi-class PD severity assessment. Our results show that pre-trained speech embeddings enable meaningful cross-lingual transfer, although performance is sensitive to dataset properties, preprocessing, and adaptation strategy. Misclassifications under these conditions related to inter-speaker variability and atypical speech patterns highlight the need for more robust feature extraction and modeling for PD severity assessment while emphasizing the importance of explainability for reliable clinical insights.

信息检索

1. 【2609.22056】Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

链接https://arxiv.org/abs/2609.22056

作者:Andre Bacellar

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:structurally predictable subpopulations, distributed across queries, predictable subpopulations, uniformly distributed, cluster in structurally

备注: 8 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the other. We instantiate these principles in RegimeAbstain, which computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query-ANN structural features, all available without any additional LLM call, and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLM-judge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE=0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only -0.5pp AUC loss, confirming the domain-agnostic structure of regime features.

2. 【2609.21863】AutoRecLab: Describe the Experiment, Get the Code!

链接https://arxiv.org/abs/2609.21863

作者:Moritz Baumgart,Philipp Meister,Justus Krell,Michael Schmidt,Bela Gipp,Joeran Beel

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:turning experimental designs, executable code remains, Empirical evaluation, central to recommender-systems, error-prone task

备注: Accepted at the 20th ACM Conference on Recommender Systems (RecSys '26), Demo Track. 4 pages, 2 figures

点击查看摘要

Abstract:Empirical evaluation is central to recommender-systems (RecSys) research, but turning experimental designs into executable code remains a manual and error-prone task. We present AutoRecLab, a Python-based autonomous RecSys lab that automates RecSys experiments from natural-language prompts. Given a research idea, AutoRecLab derives explicit experiment requirements, builds and validates a prototype, and iteratively expands it into the requested full experiment. The workflow combines retrieval-augmented generation (RAG) for documentation lookup, static type verification, and execution-steered tree search. In our demonstration, AutoRecLab autonomously implements an explicit-to-implicit feedback conversion study. In a baseline comparison across six algorithms and three datasets, 8 of 9 runs succeed at an average cost of approx- imately $1 per run with GPT-5.4-mini.

3. 【2609.21547】Do We Care About Personalization and Explainability? An Interview Study with News Recommendation Engineers

链接https://arxiv.org/abs/2609.21547

作者:Jasmin Kareem,Siddharth Mehrotra,Martijn C. Willemsen,Maarten de Rijke

类目:Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)

关键词:systems largely centers, recommender systems largely, overlooking the perspectives, model debugging, largely centers

备注: 10 pages, Accepted at ACM RecSys 2026 Main Track

点击查看摘要

Abstract:Research on explainability in recommender systems largely centers on end users, overlooking the perspectives of those who build and maintain these systems and their potential use cases such as model debugging. In this study, we examine how news engineers and related technical stakeholders perceive and implement personalization and explainability in practice. We conducted 15 semi-structured interviews across nine news organizations, spanning diverse regions in both public and private sectors, to investigate the challenges and motivations shaping their approaches. Our findings reveal that personalization is not always a straightforward or desirable choice for news organizations, as concerns around user tracking, editorial control, and resource constraints often limit its adoption. Even among organizations implementing personalized news recommender systems in production, explainability is rarely prioritized, with day-to-day operational demands frequently taking precedence over longer-term transparency goals. Definitions of explainability vary widely across organizations, though some demonstrate promising internal practices and visualization tools that facilitate communication between engineering teams and newsrooms. Based on our analysis, we provide actionable and practical guidelines for news engineers and researchers on how to adopt explainability methods within a news personalization pipeline.

4. 【2609.21475】Adaptive Preference Modeling via Explicit Indirect Relational Learning for Personalized Fashion Matching

链接https://arxiv.org/abs/2609.21475

作者:Shuiying Liao,Li Li,P. Y. Mok

类目:Information Retrieval (cs.IR)

关键词:Personalized fashion complementary, complementary recommendation requires, recommendation requires jointly, requires jointly modeling, jointly modeling user

备注

点击查看摘要

Abstract:Personalized fashion complementary recommendation requires jointly modeling user preferences and item compatibility under sparse and multimodal data conditions. Existing approaches often capture higher-order relational signals implicitly through graph propagation or rely on direct interaction data, limiting their ability to explicitly model indirect preference and compatibility relationships. To address this limitation, we propose an Adaptive Preference with Contrastive Learning framework (APCL) that explicitly models both direct and indirect relational signals within a unified recommendation architecture. Specifically, APCL constructs indirect user-item and item-item relationships through a correlation-guided adaptive aggregation mechanism and represents them as dedicated personalization and compatibility views. To improve representation learning, we further introduce a functional view contrastive learning strategy that aligns direct and indirect preference representations and direct and indirect compatibility representations, encouraging consistency across relational contexts. By integrating multimodal visual and textual information with explicit indirect relational modeling, APCL captures richer semantic characteristics while improving robustness in sparse-interaction settings. Experiments on two benchmark fashion recommendation datasets demonstrate that APCL consistently outperforms representative baseline methods.

5. 【2609.21308】Auto-Bidding with Disentangled Advertiser Profiles and Train-Free Adaptation

链接https://arxiv.org/abs/2609.21308

作者:Songyue Cai,Shan Gu,Wei Chen,Ziru Xu,Lianyu Wang,Jian Xu,Xiaofeng Zhu

类目:Information Retrieval (cs.IR)

关键词:modern advertising systems, textbf, key component, component of modern, modern advertising

备注

点击查看摘要

Abstract:Auto-bidding is a key component of modern advertising systems that provides a personalized bidding strategy for each advertiser. By characterizing each individual, profile-based methods achieve personalization and have proven effective in domains such as recommendation. However, despite the diverse bidding behavior of advertisers, their application to auto-bidding remains limited. A primary reason is that constructing and leveraging advertiser profiles face several challenges: extracting pure profiles is non-trivial, modeling common and private information simultaneously is difficult, and profile updating and cold-start adaptation remain challenging. To tackle these issues, we propose \textbf{ADAPT}, an \underline{\textbf{A}}uto-bidding framework with \underline{\textbf{D}}isentangled \underline{\textbf{A}}dvertiser \underline{\textbf{P}}rofiles and \underline{\textbf{T}}raining-free adaptation. ADAPT introduces a two-stage training paradigm and supports training-free adaptation. Specifically, (i) the stage 1 extracts pure static and dynamic profiles via contrastive learning over the advertiser memory bank; (ii) the stage 2 disentangles the dynamic profile into a common profile and a private profile, and combines them with the static profile to jointly condition the bidding strategy; (iii) once trained, ADAPT constructs profiles for new advertisers and updates profiles of existing advertisers without retraining. Our experiments on a large-scale auto-bidding benchmark demonstrate that ADAPT consistently achieves superior performance, and ablation studies further validate the effectiveness of each module. The source code will be released at this https URL.

6. 【2609.21281】Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale

链接https://arxiv.org/abs/2609.21281

作者:Hao Fu,Jichao Sun,Baiting Zhu,Qiaoling Liu,Yan Shi,Cheng Lu,Liu Liu,Yubo Wang,Xin Yao,Xiangyu Niu,Xu Dong,Wenhan Lyu,Chiyao Shen,Yinjie Huang,Minglei Chen,Shuai Ding,Li Fan,Xiao Kong

类目:Information Retrieval (cs.IR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF)

关键词:rich user intent, trillion-document scale exposes, Embedding-based retrieval, expressive personalization, user intent

备注: 10 pages, 5 figures, 9 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track

点击查看摘要

Abstract:Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization-scale paradox: hosting the full serving inventory in GPU memory is too resource intensive, while CPU compute cannot execute the same interaction-heavy model on the latency-critical path. We present a hybrid GPU-CPU co-serving system that resolves the paradox through orchestration rather than a new model class. A high-depth GPU pathway fuses retrieval and interaction pre-ranking over a curated online pool on the order of a billion documents, while a high-breadth CPU pathway searches an independently selected online inventory roughly twenty times larger with lightweight personalized scoring. Either or both pathways can run per request; candidates are deduplicated before shared downstream ranking. The system is deployed in production. A full-system A/B test against the legacy CPU-only configuration improves model-scored relevance and substantive engagement, while separate pathway experiments show positive value at their own deployment scopes. Retrieval logs show that the pathways contribute structurally distinct candidates, production serving measurements characterize their latency, and a matched capacity plan quantifies the economic rationale for assigning modeling depth to GPUs and inventory breadth to CPUs. Together, these results validate a practical, independently evolvable depth-breadth architecture for ultra-large-scale personalized search.

Comments:
10 pages, 5 figures, 9 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track

Subjects:

Information Retrieval (cs.IR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF)

Cite as:
arXiv:2609.21281 [cs.IR]

(or
arXiv:2609.21281v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2609.21281

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
7. 【2609.21257】Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

链接https://arxiv.org/abs/2609.21257

作者:Hao Fu,Baiting Zhu,Minglei Chen,Yinjie Huang,Shuai Ding

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language model, Large language, language model, evaluate model, model

备注: 9 pages, 1 figure, 8 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track

点击查看摘要

Abstract:Large language model (LLM) agents can propose, implement, and evaluate model changes. Autoresearch loops demonstrate this capability through minutes-scale iterations on a self-contained program. Online autoresearch instead spans asynchronous systems, hours-long variants, and weeks-long campaigns that can influence a product. A completed run can still support an invalid conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. We present EvoPilot, a human-gated method for long-horizon online autoresearch. Role-specific agents execute each round through a versioned domain skill and typed adapter. Durable records preserve experiments and failures; deterministic checks enforce recorded lessons. We study a 37-day campaign for the retrieval system that powers Video Deep Dive (VDD), an online experience for discovering follow-on videos after a user opens a seed video. The campaign covered seven directions and used an hourly refreshed index of hundreds of millions of videos. Earlier manual experiments had not established a benefit from an interaction head. A primitive autoresearch attempt revisited the direction but incorrectly attributed an offline hit-rate decline of 22 percentage points to the head. We then introduced EvoPilot. Its human-gated verification traced the drop to a pre-existing evaluation defect that produced output depths of 3,000 and 600. After repair, a matched comparison measured an offline improvement of 3.20 percentage points. Post-study replay and mutation tests rejected invalid comparisons while admitting valid counterparts. Durable state recovered an interrupted round, and artifact reuse avoided approximately five GPU-hours. Separately, a seven-day randomized online evaluation estimated a 0.66% relative increase in the VDD slice of Good Search Result Rate for Retention (GSRR).

Comments:
9 pages, 1 figure, 8 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track

Subjects:

Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

Cite as:
arXiv:2609.21257 [cs.IR]

(or
arXiv:2609.21257v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2609.21257

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
8. 【2609.21018】MAGIC: Marginal-Guided Compression with Optimal Transport for Efficient Visual Document Retrieval

链接https://arxiv.org/abs/2609.21018

作者:Xu Yuan,Hua Liu,Wenqi Fan,Qing Li

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:Recent visual document, MaxSim scoring overhead, patch-level vectors enable, vectors enable fine-grained, enable fine-grained evidence

备注

点击查看摘要

Abstract:Recent visual document retrieval (VDR) systems such as ColPali use multi-vector page embeddings, in which patch-level vectors enable fine-grained evidence matching but incur substantial index storage and MaxSim scoring overhead. Post-hoc merging offers a practical route to efficient VDR by reducing this cost without retraining the retriever, but its uniform reconstruction objectives are poorly aligned with the sparse, non-uniform patch usage induced by late-interaction retrieval. Under aggressive compression, this misalignment can preserve rarely used patches while concentrating retrieval activity on too few retained representatives. To address this misalignment, we propose Marginal-Guided Compression with Optimal Transport (MAGIC), a training-free post-hoc compressor for efficient retrieval with frozen multi-vector embeddings. MAGIC derives a MaxSim-induced compression surrogate and optimizes it through a two-marginal entropic optimal-transport formulation, where a retrieval-demand source marginal prioritizes high-use patches and a balanced target marginal regularizes retained-facet usage. Across ViDoRe benchmarks, keep ratios, and retrieval backbones, MAGIC consistently outperforms strong post-hoc compressors, with particularly large gains in the aggressive-compression regime; component ablations verify the complementary effects of its two marginals. We release the code at: this https URL.

计算机视觉

1. 【2609.22086】Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

链接https://arxiv.org/abs/2609.22086

作者:Hongyang Du,Lan Yan,Christian Flores,Asim Kadav

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:reliable programmatic oracle, editable artifacts emerge, long-horizon agentic task, Professional graphic design, editable artifacts

备注: 9 pages, 7 figures

点击查看摘要

Abstract:Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recurring uncovered subtasks and deepens by revising existing procedures against their own successful and failed executions, while a matched replay gate admits only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates and no human labels, grow the bank from 76 documentation-derived skills to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 points in generation quality), with 61.8% and 67.6% win rates against the no-skill agent across four specialized design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. We further show the two mechanisms are effective in combination: on 200 held-out briefs from user-traffic benchmark, widening or deepening alone reaches a 49.4% / 48.6% win rate over the no-skill agent, while their combination reaches 58.5% (p = 0.025). Procedural memory offers a practical route to continual adaptation of agents under noisy, unverifiable feedback.

2. 【2609.22083】MintAct: A Unified Visual Agent for Digital Environments

链接https://arxiv.org/abs/2609.22083

作者:Mingfei Gao,Rui Tian,Haiming Gang,Bohan Zhai,Le Zhang,Yuanzheng Gong,Di Feng,Ege Özsoy,Kaixin Ma,Vishwesh Kirthivasan,Oğuzhan Fatih Kar,Roman Bachmann,Anders Boesen Lindbo Larsen,Afshin Dehghan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unifies UI grounding, multi-step navigation, navigation across mobile, family of vision-language, visual tool

备注

点击查看摘要

Abstract:We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.

3. 【2609.22069】OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

链接https://arxiv.org/abs/2609.22069

作者:Wenxue Li,Peiyan Guan,Haoyang Jiang,Junxian Cai,Hualuo Liu,Chunjie Zhang,Chong Guan,Songlian Li,Taiyi Wu,Yongjian Yu,Xiaotong Zhao,Alan Zhao,Eric Liu,Xi Chen,Yu Liu,Lei Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:giving rise, evolving toward increasingly, increasingly general, general and versatile, reference

备注

点击查看摘要

Abstract:Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.

4. 【2609.22060】raffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features

链接https://arxiv.org/abs/2609.22060

作者:Arefeh Rezaei

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Traffic sign recognition, localize traffic signs, important perception task, advanced driver-assistance systems, localize traffic

备注

点击查看摘要

Abstract:Traffic sign recognition (TSR) is an important perception task for autonomous driving and advanced driver-assistance systems, where a system must both localize traffic signs and determine their semantic classes efficiently. This work presents a TSR system based on YOLOv2 for simultaneous detection and classification. Two complementary modifications are studied. First, YOLOv2 is extended with intermediate prediction layers, forming a branched architecture that can terminate inference early for easy cases and reduce computation time. Both whole-image and cell-wise branching strategies are investigated. Second, geometric information is introduced to reduce classification errors between visually similar signs. An unsupervised Bayesian image-segmentation method produces binary representations that are compared with class-specific geometric templates inside YOLOv2 bounding boxes. This information is used either during inference or as an additional signal during training. A dedicated dataset is constructed by combining GTSDB and GTSRB samples using seamless cloning and controlled image transformations. Experiments cover ten traffic-sign classes, with 3,000 training and 300 test samples. The selected branched architecture reports 0.647 s runtime and 0.680 mAP, compared with 0.6607 s and 0.680 mAP for baseline YOLOv2. Geometric verification during inference increases mAP to 0.713, while the geometric-feature training variant achieves 0.697 mAP with a reported runtime of 0.6608 s.

5. 【2609.22040】PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

链接https://arxiv.org/abs/2609.22040

作者:Erik Deinzer,Naya Baslan,Luca Paparusso,Narunas Vaskevicius,Peter Knott,Luigi Palmieri

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:autonomous driving operate, driving operate primarily, planning hierarchy, operate primarily, primarily through feedforward

备注

点击查看摘要

Abstract:Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions. To bridge this gap, this paper introduces PRIME, a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory. By aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, PRIME enables intent-driven perceptual attention at minimal computational cost, adding only a maximum of 29.7M parameters (0.41% of the 7.3B-parameter base model). Evaluated on the Bench2Drive closed-loop benchmark, PRIME achieves a state-of-the-art Driving Score of 82.47 (+4.73 over ORION) and a Success Rate of 60.00% (+5.38 percentage points), the highest reported Driving Score among published VLAs trained on Think2Drive demonstrations.

6. 【2609.21948】GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

链接https://arxiv.org/abs/2609.21948

作者:Yichen Liu,Puzhen Yuan,Xiang Zhu,Yanjiang Guo,Jianyu Chen

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Learning large-scale, datasets remains challenging, remains challenging due, heterogeneous action spaces, multi-embodiment datasets remains

备注

点击查看摘要

Abstract:Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at this https URL.

7. 【2609.21938】Info3R: Information-Adaptive Test-Time Training for 3D Reconstruction

链接https://arxiv.org/abs/2609.21938

作者:Sunghyun Baek,Hanna Bae,Minchan Kwon,Junmo Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recently achieved strong, recent works extend, achieved strong performance, Transformer-based models, process video streams

备注

点击查看摘要

Abstract:Transformer-based models have recently achieved strong performance on 3D reconstruction from images, and recent works extend them to process video streams in an online manner for real-world deployment. However, existing methods overlook two key signals when handling long image streams: the importance of each incoming frame and the information saturation of the model's internal state. In this paper, we propose Info3R, a novel information-adaptive test-time training method for the online 3D reconstruction. We introduce an information-aware state update that modulates the state update strength based on the redundancy and informativeness of each incoming frame. To restore the state's plasticity -- its capacity to incorporate new observations -- we propose a dynamic state reset, triggered by the cumulative magnitude of state updates and the model's prediction confidence and accompanied by an anchor-to-world alignment. Our method achieves consistent improvements on camera pose estimation, video depth estimation, and 3D reconstruction, while substantially mitigating the performance degradation in the long sequence evaluation. Notably, on KITTI Odometry, our method achieves on average 1.68x lower ATE than LongStream, demonstrating its robustness on extended outdoor sequences.

8. 【2609.21903】he Role of Radiometric Features in Cross-Site Leaf-Wood Segmentation of LiDAR Point Clouds

链接https://arxiv.org/abs/2609.21903

作者:Roman Kaharlytskyi,Derek T. Robinson,Roberto Guglielmi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:non-destructive biomass estimation, Leaf-wood segmentation, quantitative structure models, LiDAR point clouds, biomass estimation

备注: 5 pages, 6 figures. Accepted for presentation at IGARSS 2026 (IEEE International Geoscience and Remote Sensing Symposium)

点击查看摘要

Abstract:Leaf-wood segmentation of individual trees from LiDAR point clouds is essential for quantitative structure models (QSMs) used in non-destructive biomass estimation. Existing segmentation methods typically exclude radiometric features (e.g., intensity, return number) to maximize cross-sensor compatibility. We challenge this design choice by evaluating cross-site and cross-platform generalization: training on the public Heidelberg dataset (terrestrial TLS, 1550nm) and testing on a novel dataset from Ontario, Canada (RPA-LS, 905nm). Results show that geometry-only methods - including state-of-the-art deep learning models trained on high-density LiDAR datasets - fail to generalize to the sparse, top-down geometry of aerial scans, achieving F1 scores = 0.56. Incorporating radiometric features (intensity, return number, number of returns) improves F1 to 0.61, but more critically, increases wood recall by 119% from 0.16 to 0.35. Furthermore, geometry-only approaches often result in fragmented stem and branch components. We find that leveraging radiometric features preserves greater structural connectivity, resulting in more coherent architectures that are better suited for QSM reconstruction. We demonstrate that while geometric patterns are view-dependent and prone to overfitting scan patterns, radiometric features encode physical material properties that generalize across disparate sensors and environments.

9. 【2609.21887】Catena: A Comprehensive Software Suite for Large-Scale Connectomics

链接https://arxiv.org/abs/2609.21887

作者:Samia Mohinta,Pedro Gómez-Gálvez,Shi Yan Lee,Daniel Franco-Barranco,Michael Clayton,Stephan Preibisch,Jan Funke,Albert Cardona

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:gold standard datasets, densely labeled neural, labeled neural tissue, nanometer resolution, gold standard

备注

点击查看摘要

Abstract:The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet reconstructing and proofreading neuronal arbors and annotating all synapses requires pipelining multiple software tools that are often fragmented, inconsistently maintained, or proprietary, hindering reproducibility and automation. Here, we introduce Catena, an open-source, comprehensive, developer-centric software suite for connectomics that integrates modules for 3D neuron and organelle segmentation, synapse detection, microtubule tracking, and neurotransmitter inference. Catena organizes its modules in composable, chunk-wise processing pipelines in a completely documented, extensible, and adaptable design. We further reduce compute and ground-truth data requirements with pretrained machine learning models, facilitating fine-tuning. Catena ships fully containerized modules that encapsulate evolving dependencies for consistent execution across workstations and clusters. By consolidating open components, shareable models, and containerized runtimes, Catena delivers a reproducible and scalable approach to mapping cellular connectomes from electron microscopy volumes. Code and documentation: this https URL

10. 【2609.21879】Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

链接https://arxiv.org/abs/2609.21879

作者:Laurent Colbois,Sébastien Marcel

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:alongside similarity scores, produce natural language, natural language explanations, language explanations alongside, explanations alongside similarity

备注: 11 pages

点击查看摘要

Abstract:Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. However, existing evaluations of VLMs for that use case focus mostly on recognition accuracy, while the validity of generated explanations remains unquantified. In this work, we introduce a benchmarking framework for VLM-based face recognition that treats explanation quality as a core evaluation axis. We propose two criteria that explanations should satisfy: relevance, i.e., reliance on identity-stable facial features; and faithfulness, i.e., alignment with the visible image content without hallucinated features. We jointly develop a methodology enabling the quantification of relevance and faithfulness of evaluated models, based on constraining model outputs to a structured explanation format that supports automated querying and auditing. Using this framework, we benchmark several families of open-weight VLMs, jointly evaluating face verification accuracy and explanation quality. Our results highlight remaining shortcomings of produced explanations, and emphasize the need for such explanation quality metrics to get a complete picture of model performance. The proposed benchmark and open-source evaluation harness provide a foundation for proper benchmarking and future fine-tuning of explainable face recognition systems.

11. 【2609.21872】Chronosphere: Space-Time Tessellation of Local Climate Experts

链接https://arxiv.org/abs/2609.21872

作者:Daniel Cher,Eric Xing,Kexing Li,Brian Wei,Isaac Corley,Nathan Jacobs

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:spatio-temporal neural field, spatio-temporal neural, neural field, field that learns, learns representations

备注

点击查看摘要

Abstract:We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail everywhere. Global bases such as spherical harmonics spread capacity uniformly across space and time. Localized bases resolve only predefined regions. Learned tessellations adapt, but are inefficient at representing higher frequencies. Chronosphere unifies these approaches, pairing an adaptive tessellation of learnable sites on the spacetime torus $S^2\times S^1$ with a shared bank of local basis functions. Both where capacity is placed and how much detail each region carries adapt to the data, across space and time. Trained to reconstruct climatology, Chronosphere matches or leads state-of-the-art location encoders across spatial and temporal tasks, with the largest gains under spatial and temporal transfer.

12. 【2609.21866】Morphology-Aware Ambiguity Learning for Wafer Defect Decision Support

链接https://arxiv.org/abs/2609.21866

作者:Seungjun Chu,Seokhyun Chung

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:fixed-taxonomy classification problem, commonly formulated, fixed-taxonomy classification, classification problem, problem that assigns

备注

点击查看摘要

Abstract:Wafer map defect recognition is commonly formulated as a fixed-taxonomy classification problem that assigns each wafer to a single defect class. However, some wafers exhibit morphologies near class boundaries, for which forcing a single prediction may be less informative than providing plausible diagnostic alternatives. This paper proposes a morphology-aware ambiguity learning framework that supports three diagnostic actions: automatic single-class diagnosis, assisted diagnosis with two plausible defect classes, and full review. Using the radial, angular, and geometric characteristics of training wafer maps, the framework constructs a class-level ambiguity matrix representing defect-class pairs with similar morphology and plausible diagnostic alternatives. It guides the model to learn plausible alternative classes rather than treating all incorrect classes equally. During inference, the matrix determines whether an uncertain prediction can be represented by a meaningful two-class diagnostic set or should be escalated for full review. Experiments on WM-811K show that the proposed framework outperforms conventional approaches in defect recognition and diagnostic decision support, providing meaningful two-class alternatives while reserving full review for cases with unresolved ambiguity. Illustrative cost analyses further show the potential cost advantage of the proposed routing strategy. The diagnostic behavior of the framework remains consistent across different backbone architectures.

13. 【2609.21849】he Weight Is Over - Interactive Diffusion on Consumer GPUs

链接https://arxiv.org/abs/2609.21849

作者:Frieder Ganz,Maximilian Müller

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Performance (cs.PF)

关键词:LLM inference loops, language models, LLM inference, inference is booming, Abstract

备注

点击查看摘要

Abstract:On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further postprocessing that is not as standardized as LLM inference loops are. We navigate the trade-off between performance, quality, and model footprint to reach as many client devices in the wild as possible. We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.

14. 【2609.21822】Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

链接https://arxiv.org/abs/2609.21822

作者:Sarina Penquitt,Jonathan Klees,Antonia van Betteray,Parssa Jashnieh,Peter Stehr,Matthias Rottmann,Lars Schmarje

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:provide strong evidence, Pascal VOC, advanced through improved, improved architectures, architectures and open-vocabulary

备注

点击查看摘要

Abstract:While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.

15. 【2609.21811】MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention

链接https://arxiv.org/abs/2609.21811

作者:Muhammet Sami Yavuz,Sabri Mustafa Kahya,Richard R. Chen,Jana Lipkova,Benedikt Wiestler

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:remains challenging amid, Multimodal survival models, complementary prognostic information, combine complementary prognostic, fusion remains challenging

备注: Accepted at the COMPAYL 2026 Workshop on Computational Pathology and Multimodal Data at MICCAI 2026. 11 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Multimodal survival models can combine complementary prognostic information from whole-slide images and genomic profiles, but effective fusion remains challenging amid external cohort shift and computational complexity. To address these challenges, we propose MIST, multimodal survival prediction with genomic-guided histology attention. MIST represents genomic features as tokens and allows them to query compact foundation-model-derived histology context tokens before survival prediction. This design enriches molecular information with histology context rather than merging separately encoded modalities only at the final stage. Training combines discrete-time survival prediction with genomic feature masking, WSI dropout, and paired WSI-genomics contrastive alignment. Across four external evaluations in colon, renal, lung, and glioblastoma cohorts, MIST improves external C-index over standard fusion baselines in the primary comparisons. These results support genomic-guided histology attention as a compact and effective strategy for multimodal oncology outcome prediction. Our code is available at this https URL .

16. 【2609.21804】VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

链接https://arxiv.org/abs/2609.21804

作者:Qianru Li,Xuyang Chen,Xuqin Wang,Zhenghao Zhang,Hongyi Luo,Tao Wu,Daniel Cremers,Lu Liu,Yanfeng Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:semantic scene graph, compact semantic scene, long-term indoor video, long-term indoor, compact semantic

备注: 8 pages, 3 figures, 4 tables. Project page: [this https URL](https://videoreloc.github.io)

点击查看摘要

Abstract:Given a compact semantic scene graph, long-term indoor video relocalization estimates a map-frame trajectory after lighting and furniture changes. Visual methods rely on appearance and become unreliable under these changes; localizing one frame at a time from object classes and geometry instead leaves sparse, ambiguous evidence. We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene. Its run-level decision rechecks conflicting placements using evidence accumulated across connected clips, stabilizing the trajectory beyond adjacent-clip tracking. Hypothesis-first registration proposes poses from object triplets and verifies each using clip-wide object centers and box surfaces. Orientation-aware refinement uses box faces, gravity and wall directions to resolve ambiguity in camera orientation and refine the full pose. This reframes sparse-map relocalization as verification of spatially extended video queries, moving discriminative support from stored appearance to temporal context and permitting a 100 kB map of class-labelled boxes. On RIO10 and ReplicaCAD, the all-frame localization success rate at 1 m/10$^\circ$ is 73.5% and 61.1% under causal evaluation, rising to 90.6% and 74.8% with clip closure. The evaluated per-frame scene coordinate regressors reach up to 47.6% and 49.8%, respectively, with maps of 12.6-42 MB. Project page: this https URL

17. 【2609.21800】A Principled Approach to Unsupervised Anomaly Detection

链接https://arxiv.org/abs/2609.21800

作者:James Myles,Matthew Baugh,Johanna P. Müller,Bernhard Kainz,Yingzhen Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Traditional unsupervised anomaly, Traditional unsupervised, underlying generative mechanisms, unsupervised anomaly detection, normative distribution

备注: 14 pages, 2 figures, 3 tables

点击查看摘要

Abstract:Traditional unsupervised anomaly detection (UAD) methods are designed to flag or localise deviations from a normative distribution, ignoring the underlying generative mechanisms of the anomalies. Yet the nature of an anomaly is often as important as its presence. We reformulate UAD as a Bayesian inverse problem, in which the objective is to infer the most probable corruption responsible for each observation. Our framework yields a probabilistic anomaly score as the energy of the inferred corruption parameters, and serves as a principled recipe for developing new UAD algorithms. We derive several existing methods as instances of the general framework, each corresponding to the same energy score under different modelling choices. Experimentally, we study the framework's components in a controlled setting, and improve object-class AUROC on the MVTec AD dataset by 2.3% by adapting the underlying corruption model. Finally, we validate the framework on a brain MRI benchmark, achieving strong detection performance while producing estimates of pathology intensity, bias, and geometry. Code is available at this https URL.

18. 【2609.21780】PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection

链接https://arxiv.org/abs/2609.21780

作者:Xuanming Shang,Weijia Zhang,Chao Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:methods preserve fidelity, LiDAR point clouds, point clouds faces, point-based methods preserve, voxel-based methods achieve

备注: Accepted to ECCV 2026

点击查看摘要

Abstract:3D object detection from LiDAR point clouds faces a fundamental dilemma: voxel-based methods achieve efficiency at the cost of geometric quantization, while point-based methods preserve fidelity but suffer from prohibitive computational bottlenecks. Specifically, point-based architectures are crippled by slow downsampling strategies (e.g., FPS) and expensive dynamic neighbor queries (e.g., k-NN) coupled with costly continuous interactions. To tackle these systemic inefficiencies, we propose PointLAM, a highly efficient and powerful point-based architecture driven by two synergistic innovations. First, to resolve the downsampling bottleneck, we develop the Laplacian Point Sampler (LPS). LPS employs an implicit discrete Laplacian high-pass filter and Doubly Sorted Sampling to achieve fast, structure-aware foreground preservation. Second, to overcome local modeling latency, we design the Local Hadamard Aggregator (LHA). LHA decouples spatial indexing from feature representation using transient grids, and replaces complex continuous interactions with a Hadamard Gating mechanism for topology-aware, attentive modulation. By coupling this local gating with Bi-Directional Mamba (BDM) layers for global sequence modeling, we formulate the Local Attentive Mamba (LAM) block. Powered by this architecture, PointLAM achieves competitive performance on nuScenes and Waymo for point-based detectors. It rivals highly optimized voxel competitors while requiring a fraction of the computational footprint, demonstrating marked superiority in detecting small instances and handling extreme sparsity. Project page: this https URL.

19. 【2609.21770】XCalib Depth-Guided Geometric Optimization for Dense Thermal-Visible Video Registration

链接https://arxiv.org/abs/2609.21770

作者:Aurelien Godet,Gabriel Jobert,Mauro Dalla Mura

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:including image fusion, multimodal perception tasks, vital preprocessing step, including image, image fusion

备注: 10 pages, 6 figures

点击查看摘要

Abstract:Image registration is a vital preprocessing step in multimodal perception tasks, including image fusion, object detection, and semantic segmentation. In Advanced Driver- Assistance Systems (ADAS), spatial misalignment between visible (RGB) and infrared (IR) cameras -caused by non-coincident optical axes and field-of-view differences- introduces non-uniform parallax and visual ghosting. Classical keypoint-based methods are restricted to global homographies that fail under dynamic depth, while unconstrained dense flow algorithms lack structural regularization and suffer from temporal instability. In this paper, we propose XCalib, an unsupervised dense thermal-visible registration framework that bridges this gap. Rather than serving as an absolute metric calibration tool, XCalib leverages virtual pinhole camera parameterization strictly as a geometric constraint space. By optimizing effective relative pose and intrinsics alongside predicted monocular metric depth, XCalib restricts the search space of spatial displacements to physically valid projection geometries. Our key contributions are: (1) a novel registration paradigm that uses camera parameterization as an implicit regularizer for dense cross-modal warping; (2) Normalized Edges Correlation (NEC), a robust structural similarity metric tailored to cross- spectral alignment; and (3) extensive quantitative and qualitative evaluations across public ADAS datasets, demonstrating superior temporal stability and alignment accuracy over unconstrained dense flow baselines.

20. 【2609.21763】Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

链接https://arxiv.org/abs/2609.21763

作者:Mushir Akhtar,M. Tanveer,Mohd. Arshad

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:conclusion holds, medical model benchmark, model benchmark score, AUROC, benchmark score

备注: 27 pages, 7 figures, and 21 tables; includes extended methods, statistical analyses, and robustness evaluations

点击查看摘要

Abstract:A medical model's benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on 12,200 chest radiograph records from four datasets (Montgomery, Shenzhen, TBX11K, and VinDr-CXR). Five fixed prompt families yield 244,000 model--image--prompt scores. No model leads every cohort and reliability criterion. Prompt-family changes alter AUROC in 21 of 48 multiplicity-controlled comparisons. Replacing healthy controls with sick non-tuberculosis controls reduces AUROC by 0.075--0.306 across all four models. On VinDr-CXR, the three medical models distinguish tuberculosis from no-finding controls substantially better than from pneumonia or lung tumor; their AUROC point estimates for both named diseases fall below 0.5. CheXficient has documented VinDr-CXR pretraining exposure, which limits the interpretation of its results. Thresholds chosen for 95\% sensitivity on TBX11K training retain that constraint by point estimate in only four of sixteen target evaluations. A five-seed supervised source model reaches 0.999 AUROC on TBX11K validation but 0.629 on each of two external cohorts. Conservative exclusion of perceptual-overlap candidates narrows this gap without closing it. These retrospective, single-task results show that discrimination, score reliability and threshold retention support different portability claims. Evidence for chest X-ray tuberculosis screening should identify the complete evaluation specification rather than attribute clinical portability to a checkpoint alone.

21. 【2609.21754】SFVO: Decoupled Confidence-Guided Stereo-Flow Visual Odometry with Bidirectional PnP

链接https://arxiv.org/abs/2609.21754

作者:Kai Zhang,Guoyang Zhao,Jun Ma

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:achieved significant progress, existing methods focus, learning-based visual odometry, Deep learning-based visual, significant progress

备注

点击查看摘要

Abstract:Deep learning-based visual odometry (VO) has achieved significant progress, yet most existing methods focus on a monocular approach, which suffers from scale ambiguity. Stereo VO provides real metric by its nature, but remains less studied in deep learning VO due to its high computational cost and modeling complexity. Recent advances in stereo matching and optical flow estimation have made dense visual correspondence increasingly accurate and reliable, but their complementary geometric information has not been fully exploited for VO. In this paper, we present SFVO, a correspondence-driven stereo VO framework that directly builds upon pretrained stereo matching and optical flow models. SFVO exploits pretrained stereo matching and optical flow models to estimate stereo and temporal correspondences. Instead of learning pose directly from images, SFVO maps learned correspondences into geometric constraints and predicts which points are trustworthy. To improve the reliability of visual correspondence-based geometric constraints, we introduce decoupled confidence maps for rotation and translation. This design better aligns the characteristics of visual correspondence and 6-DoF transformations. Extensive experiments on outdoor and indoor datasets demonstrate that SFVO achieves robust and accurate pose estimation with strong generalization capability. The code will be released.

22. 【2609.21743】Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation

链接https://arxiv.org/abs/2609.21743

作者:Zhengshan Wang,Joshua Charles Webster-Ford,Yifei Tian,Xinxin Wang,Long Chen,Weiping Ding

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:imbalanced binary segmentation, standard objective, fail in imbalanced, Prompt Adaptation, entropy gradients vanish

备注

点击查看摘要

Abstract:Entropy minimization is a standard objective for test-time adaptation (TTA), but it can fail in imbalanced binary segmentation. Unlike image classification, dense segmentation aggregates thousands of pixel predictions, allowing the larger predicted class to dominate the update, pull minority predictions toward itself, and produce a degenerate mask as predictions saturate and their entropy gradients vanish. We theoretically establish this collapse in a shared-shift model. This analysis motivates Balanced-Anchor Prompt Adaptation (BAPA), which combines two complementary modules. The Class-Balanced Anchors (CBA) module selects high-confidence anchors separately from each predicted class and gives foreground and background equal total loss weight, preventing the larger region from dominating the update. Dynamic Prompt Adaptation (DPA) refreshes these anchors after each prediction update and optimizes only text-side prompt residuals while keeping the vision-language encoders frozen. This prompt-only update refines the foreground-background decision boundary without altering the pretrained dense visual representation. Across experiments from four domains, BAPA achieves the highest mean Dice among the evaluated methods. Factorized ablations further validate the complementary roles of CBA and DPA, supporting balanced prompt adaptation as an effective alternative to entropy minimization for test-time binary segmentation.

23. 【2609.21712】ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation

链接https://arxiv.org/abs/2609.21712

作者:Boni Hu,Xiong Wei,Haoming Huang,Yong Huang,Chenbo Wang,Yi Yang,Jiancheng Wang,Ruicheng Zhu,Zhimin Yang,Guanglai Liu,Qiaowan Jin,Dongzhuo Wang,Haiwei Kuang,Jiajun Fan,Yue Wu,Jiaxin Wei,Hao Sun,Feihong Yan,Wei Bi,Kaixuan Wang,Zichao Guo,Xiaozhi Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:repeatable closed-loop simulation, production deployment exposes, preserving scene identity, Generative world models, mixed fisheye-pinhole rig

备注: Technical Report. Videos and additional results are available at [this http URL](http://zyt-aim.github.io/ZYT-World)

点击查看摘要

Abstract:Generative world models offer controllable and repeatable closed-loop simulation for end-to-end and vision-language-action driving policies, but production deployment exposes three unresolved requirements: faithfully reproducing a mixed fisheye-pinhole rig at native resolutions; reconciling causal, per-timestep interaction with long-horizon stability and low latency; and preserving scene identity when a location is revisited. We present ZYT-World, a single architecture that natively generates four fisheye views with field of view 180° and three pinhole views. Projection-specific Plucker adapters encode camera geometry, ego-motion adaptive layer normalization provides global motion control, and a lightweight pixel-aligned layout conditions traffic participants and signals through instance-level boxes, headings and colors. Heterogeneous training combines full-rig geometric coverage with high-resolution detail. Teacher forcing, causal consistency distillation, self-rollout distribution matching distillation, and RigCritic transform a 40-step bidirectional teacher into a one-step, per-latent streaming generator, with RigCritic evaluating the seven-view rig jointly. A 19M-parameter variational autoencoder decoder (TinyVAE), W8A8 quantization, and our inference engine reduce decoding, backbone, and incremental-execution costs, respectively. Finally, cross-trajectory pairs derived from real captures train a plug-in implicit-memory module that preserves place-specific evidence. On the internal multi-view test set, the one-step model retains more than 90% of the teacher's PSNR and SSIM, while FID, FVD, and LPIPS stay within 11% of the teacher. Under the generator-only timing in Figure 2, it is 107.7 times faster than the 40-step bidirectional teacher. TinyVAE decodes 59.8 times faster than Wan. 30s rollouts and cross-trajectory revisits show the intended long-horizon and memory behavior.

24. 【2609.21709】SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

链接https://arxiv.org/abs/2609.21709

作者:Ronghui Li,Jun Dong,Zhongyuan Hu,Zunnan Xu,Jun Zhou,Liyuan Chen,Shuoling Liu,Jiangpeng Yan,Jie Guo,Xiu Li,Linchao Bao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large language models, provide limited support, Large language, sign language interaction, sign language

备注

点击查看摘要

Abstract:Large language models (LLMs) provide limited support for sign language interaction. Unifying sign language translation (SLT) and generation (SLG) to enable sign language as both input and output can reduce switching between separate models during sign-text interaction. We present SignGPT, a unified, pose-based framework for gloss-free SLT and SLG. SignGPT integrates part-aware hierarchical representations of body, hand, and facial motion into a shared language model and employs asymmetric multi-token prediction and progressive training for bidirectional modeling. We evaluate SignGPT on How2Sign (ASL) and Phoenix-2014T (DGS) through benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers assesses an LLM-mediated sign-to-sign response pipeline, highlighting the potential of unified modeling to support sign language conversation (SLC).

25. 【2609.21698】Diffusion-Based Tumor Inpainting for Renal Segmentation under Clinical Data Scarcity

链接https://arxiv.org/abs/2609.21698

作者:Ekaterina Sedykh,Salme Ussanov,Dmytro Fedorenko,Dmytro Fishman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large annotated datasets, Deep learning segmentation, requires large annotated, deployments typically offer, tumors requires large

备注

点击查看摘要

Abstract:Deep learning segmentation of renal tumors requires large annotated datasets, yet clinical deployments typically offer only a handful of tumor-positive cases from the target site. We propose a diffusion-based inpainting framework that synthesizes anatomically plausible renal tumors within healthy CT scans, requiring no additional annotation, and provide the first systematic comparison of 2D, 2.5D, and full 3D (MAISI) synthesis strategies for this task. Training the diffusion model on public data (KiTS23, KIRC) and evaluating nnU-Net segmentation on a internal cohort across three low-data regimes, we find that 2.5D and 3D augmentation substantially reduce false positives (from $\sim$18--20\% to $\sim$3--6\%) while maintaining Dice, whereas 2D provides no consistent benefit. Crucially, the proposed 2.5D method matches full 3D synthesis on every metric at substantially lower computational cost, indicating that local volumetric consistency alone is sufficient for effective augmentation in data- and resource-scarce clinical settings.

26. 【2609.21683】Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

链接https://arxiv.org/abs/2609.21683

作者:Yunji Chu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)

关键词:immediately preceding behavior, listener immediately preceding, preceding behavior, depends on dialogue, dialogue context

备注: 15 pages, 2 figures, 2026 ECCV Workshop (11th ABAW) Best Student Paper Award

点击查看摘要

Abstract:Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at this https URL.

27. 【2609.21675】DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning

链接https://arxiv.org/abs/2609.21675

作者:Wan Xu,Yuanfan Guo,Kevin Han,LaLa Chen,Wangmeng Zuo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Multimodal Large Language, Large Language, paradigms remain confined, natural-language expression space

备注

点击查看摘要

Abstract:Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural-language expression space. Consequently, they inherently incur excessive linguistic overhead, leading to information dilution and weak visual grounding. To address this challenge, we propose Dense Reasoning Trace (DRT), a paradigm that departs from natural-language-centered CoT by expressing reasoning as compact structured traces, which include concise intermediate states with symbolic connectors and disentangle visual observations from logical deductions. First, we introduce the Dense Trace Initialization to internalize the DRT reasoning mode into the model, substantially improving token efficiency while preserving visual evidence. To further enable the model to faithfully capture the logical relations within traces, we propose the Trace-Grounded Reinforcement Learning framework, which builds reference traces through a tri-perspective verification pipeline and employs Trace-Grounded GRPO with structured rewards, encouraging the model to generate concise DRT-style traces with reduced hallucination and stronger logical grounding. Extensive experiments on challenging reasoning benchmarks show that DRT achieves 5.5$\times$ token efficiency improvement while improving 1.3 accuracy points over the Qwen3-VL baseline. These findings suggest that complex multimodal reasoning may not require verbose natural-language traces, opening a more efficient path for next-generation MLLMs. Our code and data are available at: this https URL

28. 【2609.21651】Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

链接https://arxiv.org/abs/2609.21651

作者:Naga Ganesh,Chandrashekar M S,Lakshmi Pedapudi,Aakash Singh,Vineet Singh

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Digital Green farm, Green farm advisory, Digital Green, Green farm, farm advisory service

备注: 14 pages, 26 Tables, 12 Figures

点击查看摘要

Abstract:this http URL is Digital Green's farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no crop named, often no text at all. The service has to determine whether the picture can be used, what crop it shows, and what is wrong with it, from images taken on cheap phones in a field, in poor light and with a moving camera. The system doing this today cannot be adjusted. It has no adjustable thresholds for photograph rejection, crops and problems cannot be added, and there is no confidence cut-off to set. We study about 1.16 million photographs sent to this http URL from Ethiopia, India, Kenya and Nigeria. The production quality gate rejected 46.8% of the images it judged, over a quarter of those reaching diagnosis returned no crop name, and 35.8% of the labelled problems filed under "disease" are pests, identifiable without the crop. We therefore split the work into three stages: a quality gate (M0), a crop detector (M1), and a disease or pest detector (M2). Route A fills all three with one fine-tuned vision-language model (Qwen3-VL-4B) answering in a single call. Route B fills each with a small specialist model (DaViT, YOLO26). We replace our production GPT-4o quality gate with a small MobileNetV3 gate at 86.9% F1 in 12 ms. On one test set scored the same way for every system, a hierarchical DaViT-Base achieves 95.41% crop accuracy against 91.46% for the production baseline. It also leads on diagnosis and never declines to answer, while every language model in the comparison leaves a large share of rows with no diagnosis. The fine-tuned model retains two capabilities the specialists do not have: one call for all three stages, and a request for a better photograph when the image cannot support an answer.

Comments:
14 pages, 26 Tables, 12 Figures

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2609.21651 [cs.CV]

(or
arXiv:2609.21651v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.21651

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
29. 【2609.21629】Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

链接https://arxiv.org/abs/2609.21629

作者:Umar Marikkar,Sameed Husain,Muhammad Awais,Sara Atito

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:semantically distinct signal, data differs fundamentally, Multi-Channel Vision Transformers, Decoupled Vision Transformer, natural images

备注

点击查看摘要

Abstract:Multi-Channel imaging (MCI) data differs fundamentally from natural images, as each channel records a semantically distinct signal rather than a colour band. To adapt vision encoders to MCI data, Multi-Channel Vision Transformers (MC-ViTs) tokenize each channel independently and concatenate the resulting tokens into one sequence, and the channel count is no longer fixed by the architecture. Self-attention is then computed across all channel-patch tokens with no restriction on which channels attend to which, which dilutes the features of individual channels. The Decoupled Vision Transformer (DC-ViT) regulates this by separating updates computed within a channel from updates computed across channels, and by forming a representation per channel before the channels are combined. Its formulation, however, pairs tokens by spatial position, and thus requires the same visible tokens in every channel. Correspondence under independent per-channel masking is recovered by solving a linear assignment between the retained patches of each channel, which allows decoupled attention to be combined with current masked multi-channel training in its standard configuration rather than a restricted one. Across three classification and three segmentation benchmarks spanning fluorescence microscopy, imaging mass cytometry and satellite imaging, including dense prediction at high channel counts, the resulting formulation outperforms the strongest MC-ViT baseline.

30. 【2609.21628】Detection is solved, delineation is not: what governs tooth segmentation on panoramic radiographs

链接https://arxiv.org/abs/2609.21628

作者:Muhammad Rehan,Moaz Amjad,Syed Danial Ahmed,Mariam Adnan,Haider Ali

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Automatic tooth segmentation, performance remains unclear, factors govern performance, govern performance remains, computer-assisted dental diagnosis

备注: 15 pages, 6 figures, 5 tables. Code: [this https URL](https://github.com/Rehan000/opg-tooth-segmentation)

点击查看摘要

Abstract:Automatic tooth segmentation and FDI numbering on panoramic radiographs underpins computer-assisted dental diagnosis, yet which factors govern performance remains unclear. We assemble a corpus of 1,422 panoramic radiographs containing 42,142 expert-delineated tooth polygons across the 32-class FDI taxonomy, annotated by 30 dental practitioners and independently reviewed by two others, and use it to isolate input resolution, architecture and anatomical priors under a single evaluation protocol. First, resolution dominates: across a controlled 640/1024/1280 ablation, mask mAP50-95 rises 0.656 - 0.710 - 0.717 while mAP50 stays flat at ~0.982. Both gains are significant under a paired bootstrap over images (p 0.001, p = 0.024); neither mAP50 change is distinguishable from zero. Added resolution buys boundary precision, not detection. Second, architecture is nearly irrelevant in-domain: a query-based transformer with 2.1x the parameters is statistically equivalent to a one-stage detector (95% CI [-0.0064, +0.0064]), only marginally better under domain shift, 5.5x slower on CPU and not executable under standard ONNX runtimes. Third, three targeted interventions fail: a LoRA-adapted self-supervised encoder underperforms, a promptable foundation segmenter degrades masks by 39%, and globally optimal anatomical label assignment yields +0.0007 despite correcting a constraint violated in 40% of out-of-domain predictions. Zero-shot transfer to an independent multi-centre cohort, verified overlap-free, costs 62% of mask mAP50-95 but only 18% of mAP50, reproducing the dissociation. Decomposing masks along the tooth axis localises the residual error to the apical third. Boundary precision is therefore the binding constraint, and effort is better directed at resolution and acquisition diversity than at architectural novelty.

Comments:
15 pages, 6 figures, 5 tables. Code: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

Cite as:
arXiv:2609.21628 [cs.CV]

(or
arXiv:2609.21628v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.21628

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
31. 【2609.21624】Learned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video

链接https://arxiv.org/abs/2609.21624

作者:Musa Rochi,Marcel Schubert,Christoph Gebhardt

类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:Problematic internet, forced breaks, common interventions, easily circumvented, internet use affects

备注: 18 pages, 13 figures

点击查看摘要

Abstract:Problematic internet use affects a growing share of the population, yet common interventions, e.g., time limits, blocking, forced breaks, are coercive and easily circumvented. We explore a less restrictive alternative: adapting the emotional intensity of visual content. Prior work has shown that optimization can steer an image's affective content, but its per-image optimization cost makes it impractical for real-time deployment. We instead learn a model that predicts this transformation in a single forward pass: a MobileNetV4 backbone with FiLM-based emotion conditioning outputs parameters for differentiable global transformations. This replaces prior iterative optimization (80 s per image) with a single 3.7 ms forward pass. In a user study (N = 54), the model reduced viewer-reported arousal relative to unedited images, comparably to the grayscale well-being filter, while being rated higher in perceived quality. We integrate the model into an Android app that adapts Instagram video in real time, sustaining 60 fps on a Samsung Galaxy S23.

32. 【2609.21597】HAT: Hypothesis-Anchored Tracking for Video Monocular Spacecraft Pose Estimation

链接https://arxiv.org/abs/2609.21597

作者:André Lopo,Atabak Dehban,Rodrigo Ventura

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:debris removal, estimation of non-cooperative, important for on-orbit, on-orbit servicing, servicing and debris

备注: 8 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Monocular 6-DoF pose estimation of non-cooperative targets is important for on-orbit servicing and debris removal. A single-image estimator can confuse near-symmetric spacecraft orientations, and tracking can preserve an incorrect pose. We present Hypothesis-Anchored Tracking (HAT), a causal framework that uses inter-frame motion to select among competing CAD-based pose hypotheses before alignment and fusion. Rather than independently choosing the highest-scoring hypothesis in each image, HAT retains competing orientation histories and selects a pose to anchor the relative trajectory estimated by monocular SLAM. Sparse anchors and pose fusion provide per-frame estimates after initialization without revising past outputs. The method requires only a calibrated RGB sequence, a metric CAD model, and target image regions, which can be supplied by detection or segmentation. The pretrained pose and SLAM networks require no target-specific training or fine-tuning. We evaluate two versions, Mega-HAT and Pico-HAT, using MegaPose and PicoPose, on SPARK-2024, SwissCube and SHIRT, with YCB-Video assessing performance outside the space domain. Using one temporal configuration per method, the arithmetic means of the four dataset-wise comparisons show 9.4% lower mean pose error and 3.76 times the sustained input FPS for Mega-HAT relative to independent MegaPose, and 23.9% lower mean pose error and 2.42 times the FPS for Pico-HAT relative to independent PicoPose. Mega-HAT ablations on SPARK and an offline reference examine component contributions and the effect of revising past estimates.

33. 【2609.21595】Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

链接https://arxiv.org/abs/2609.21595

作者:Abhishek Bhandari,Gaurav Harit

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Large Language, learning using Large, Devanagari script remains, Language Models

备注

点击查看摘要

Abstract:In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, comparing three in-context example retrieval strategies: domain-random selection, dense semantic retrieval, and our proposed CharBM25, which retrieves examples by character n-gram BM25 similarity over OCR inputs to target shared error patterns with the test sentence. Across a 20,000-sentence benchmark spanning five news domains, retrieval strategy is the decisive factor in correction quality: CharBM25 outperforms domain-random selection by 2.8-4.0pp absolute WER on Hindi and 2.9-3.8pp on Marathi, using character trigrams, which consistently outperform bigrams and unigrams. Scale dominates performance: Gemma-3-27B achieves WER reductions of 55.0% for Hindi and 33.3% for Marathi under CharBM25-5. Few-shot gains are capacity-gated: models below 8B do not reliably improve over the OCR baseline, and on Marathi the smallest models (3B) degrade more sentences than they improve. Marathi is persistently harder to correct than Hindi across all scales, reflecting its greater morphological complexity. These findings establish CharBM25 as an effective, GPU-free retrieval strategy that matches or exceeds dense retrieval at negligible computational cost, and show that combining it with a general-purpose LLM of 12B+ parameters delivers reliable, training-free Devanagari post-OCR correction without task-specific fine-tuning. Dataset: this https URL

34. 【2609.21593】A benchmark dataset and baseline methods for four-dimensional STEM diffraction patterns

链接https://arxiv.org/abs/2609.21593

作者:Yuyan Guan,Haoran Zhang,Zian Mao,Antong Yang,Caifei Li,Jialong Wang,Chuying Ouyang,Hong Wang,Xiaoqin Zeng,Yujun Xie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Four-dimensional scanning transmission, transmission electron microscopy, yielding spatially resolved, scanning transmission electron, heterogeneous data volumes

备注: 16 pages, 5 figures. Data and trained model weights: [this https URL](https://doi.org/10.57760/sciencedb.nbsdc.00281) . Code: [this https URL](https://github.com/Gaiya69-rgb/4D-ImageNet)

点击查看摘要

Abstract:Four-dimensional scanning transmission electron microscopy (4D-STEM) records a two-dimensional diffraction pattern at each electron-probe position, yielding spatially resolved reciprocal-space information but large, heterogeneous data volumes. Here we describe 4D-ImageNet, a collection of 174,000 diffraction patterns comprising 145,000 experimental patterns selected from 29 acquisitions and 29,000 multislice simulations. The experimental data cover acquisition-level labels for Ag, Au, mixed Au-Ag, CoO, Pd and ZnO specimens across multiple fields of view, scan dimensions, camera lengths and exposure times. Each acquisition contributes 5,000 quality-ranked patterns with source scan coordinates and acquisition metadata. A set-prediction detector provides model-derived pseudo-labels for the direct-beam position and Bragg-disk centres, with a confidence score for each disk. The simulation data cover 13 crystal structures and include Euler rotations, reciprocal-space sampling and approximate low-index beam directions. A grouped mixed-domain masked-reconstruction benchmark is provided to assess leakage-resistant loading and evaluation across experimental and simulated data. The dataset is intended for representation learning, disk detection, diffraction-pattern retrieval, orientation analysis and simulation-to-experiment studies.

35. 【2609.21576】GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

链接https://arxiv.org/abs/2609.21576

作者:Pinxin Liu,Haiyang Liu,Jiahao Luo,Junhua Huang,Chunhao Zou,Luchuan Song

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Human-Computer Interaction (cs.HC)

关键词:Generating natural co-speech, embodied conversational agents, Generating natural, conversational agents, natural co-speech gestures

备注

点击查看摘要

Abstract:Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality--latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: this https URL

36. 【2609.21543】From Retrieval to Recognition:How Vision--Language Models Become OCR Specialists

链接https://arxiv.org/abs/2609.21543

作者:Yuanxiang Huangfu,Hanmeng Zhong,Linqing Chen,Jeffrey Tiong Jee Hui

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:language model acquire, model acquire specialized, acquire specialized OCR, specialized OCR ability, language model

备注

点击查看摘要

Abstract:Does a general vision--language model acquire specialized OCR ability by developing a new reading circuit or by reusing an existing mechanism? We address this question in the setting of full-sequence OCR, rather than local-answer retrieval. Using an evidence-grounded protocol with held-out causal interventions, we identify sparse and stable OCR-head sets in GLM-OCR, MinerU2.5, and PaddleOCR-VL-1.6. We then investigate the mechanistic origin of these OCR heads by comparing them with independently identified textual retrieval/copy heads in general VLMs. Across two general VLMs, visual OCR heads strongly overlap independently identified textual retrieval/copy heads, yielding untuned top-20 intersections of 73.3% and all-head Spearman correlations of 0.677-0.886. The overlap and causal interventions suggest that full-sequence OCR operates as dense sequential multimodal copy-and-paste, repeatedly retrieving visual evidence and routing it to the current output position. Finally, we examine how this shared circuit changes as a general VLM becomes an OCR specialist. Matched base-to-specialized comparisons show that OCR specialization largely preserves head identity, retaining 17-20 of the top 20 heads per task with all-head rank correlations of 0.874-0.942, while redistributing their functional and causal strengths.

37. 【2609.21541】Purification and Regulation: Comorbidity-Aware Multi-Label Few-Shot Learning for Medical Image Classification

链接https://arxiv.org/abs/2609.21541

作者:Ying-Chih Lin,Po-Chih Kuo,Yong-Sheng Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Multi-label few-shot learning, medical image analysis, Multi-label few-shot, few-shot learning, remains a significant

备注

点击查看摘要

Abstract:Multi-label few-shot learning (MLFSL) remains a significant challenge in medical image analysis (MIA). Current metric-based meta-learning methods face two critical limitations in MIA. First, conventional prototype generation often entangles irrelevant disease information, leading to contaminated prototypes and degraded performance. Second, prior studies typically enforce inter-class separability in embedding space, largely neglecting the inherent correlations among diseases. To overcome these challenges, we propose Prototype Purification and Regulation (PPR), a novel MLFSL framework for MIA. PPR first performs prototype purification by leveraging sample-level comorbidity scores to emphasize disease-specific features, producing purified prototypes that better characterize each disease. Building upon these purified prototypes, PPR further addresses the underexplored problem of inter-class prototype distance in MIA by incorporating disease-level comorbidity statistics to adaptively regulate inter-class similarity, forming a comorbidity-aware embedding space. Overall, PPR sequentially enables the model to capture pure disease features and inter-class relationships for reliable MLFSL in MIA. Extensive experiments across four chest X-ray benchmark datasets, including cross-domain evaluation, show that PPR consistently outperforms state-of-the-art methods, significantly improving disease detection while demonstrating robust generalization and clinical applicability.

38. 【2609.21522】Refine Then Fusion: Training-Free 3D Point Cloud Adaptation with Priority Refinement and Multi-Modal Knowledge Fusion

链接https://arxiv.org/abs/2609.21522

作者:Hang Cheng,Yan Chen,Mingyu Fan,Long Zeng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent pre-trained foundation, vision tasks, foundation models provide, models provide rich, provide rich multi-modal

备注

点击查看摘要

Abstract:Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representations in few-shot scenarios is limited by two fundamental challenges: High-dimensional features often contain substantial channel redundancy and task-irrelevant noise, while the reliability of different modalities varies across samples. Consequently, direct aggregation of heterogeneous representations overlooks sample-dependent modality reliability and may obscure the discriminative cues essential. To address these limitations, we propose Refine Then Fusion(RTF), a training-free framework for few-shot 3D recognition. RTF first identifies discriminative feature channels by jointly modeling inter-class similarity and intra-class stability, thereby decoupling domain-specific knowledge refinement from the cached representations of pre-trained models. It then introduces a reliability-aware fusion mechanism that estimates sample-wise modality reliability from the distribution shifts induced by feature refinement, enabling adaptive aggregation of multi-modal representations. Furthermore, RTF constructs a memory cache that integrates instance-level support features with class-level prototypes to infer query labels. Extensive experiments on five benchmarks demonstrate that RTF consistently outperforms single-modal baselines, partial-fusion variants, and existing lightweight adaptation methods, achieving state-of-the-art few-shot 3D recognition performance without gradient optimization, additional training data, auxiliary training, or parameter updates.

39. 【2609.21521】VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

链接https://arxiv.org/abs/2609.21521

作者:Changbeen Kim,Junwon Chang,Kipyo Kim,Risa Shinoda,Kuniaki Saito,Donghyun Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Video Large Language, Large Language, demonstrated strong performance, understanding remains challenging

备注

点击查看摘要

Abstract:While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.

40. 【2609.21516】2D GauSS-MI: Efficient Active Scene Reconstruction with Balanced Visual and Geometric Quality

链接https://arxiv.org/abs/2609.21516

作者:Yuhan Xie,Jia Pan

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:onboard computational resources, limited onboard computational, Gaussian Splatting Shannon, Gaussian Splatting, limited onboard

备注

点击查看摘要

Abstract:Active reconstruction requires efficient active view selection to achieve high-quality reconstruction within limited onboard computational resources. Existing methods face challenges in adequately balancing visual and geometric quality with the computational efficiency required for real-time operation. In this work, we present an active reconstruction framework based on 2D Gaussian Splatting (2DGS). We develop an efficient online 2DGS mapping pipeline for incremental RGB-D observations and introduce a probabilistic reliability model that characterizes the view-dependent reconstruction quality of individual 2D Gaussian splats. Building on this model, we formulate 2D Gaussian Splatting Shannon Mutual Information (2D GauSS-MI), a mutual-information-based metric that exploits the explicit surface orientation of 2DGS to evaluate the expected information gain of candidate views. The proposed metric enables active view selection to account for both visual and geometric reconstruction quality. We evaluate the proposed system against three state-of-the-art baselines on eight Replica scenes. Experimental results demonstrate that our method achieves a favorable balance between visual and geometric reconstruction quality with substantially lower computational cost and competitive model storage.

41. 【2609.21511】2nd Place Solution to the HANDS 2026 Workshop Challenge-Dexterous Grasp Motion Track: Single-Shot Trajectory Warping for Grasp Motion Generation

链接https://arxiv.org/abs/2609.21511

作者:Muneeb A. Khan,Woojin Kim,Shinwoo Kim,Muhammad Munsif,Binod Bhattarai,Seungryul Baek

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Dexterous Grasp Motion, conjunction with ECCV, Dexterous Grasp, address grasp motion, grasp motion generation

备注

点击查看摘要

Abstract:This report describes our 2nd place solution to the HANDS 2026 workshop challenge (Dexterous Grasp Motion track) in conjunction with ECCV 2026. In this challenge, we address grasp motion generation for the 12-DoF LinkerHand O6, aiming to produce physically plausible reach-and-lift trajectories for unseen objects from randomized initial hand poses in simulation. This task is particularly challenging because each grasp requires a per-step policy to make approximately $70$ twelve-dimensional decisions, with errors accumulating over time, while test objects and physical dynamics may differ from those encountered during training. To address these challenges, we propose editing a single successful GraspM3 demonstration instead of generating the motion step by step: a policy observes the object once and outputs a 12-D warp of the demonstration, which is then replayed open-loop. Moreover, we train the warp policy with one-step PPO over all $4{,}824$ training objects in parallel. As a result, our method achieved success rates of $94.61\%$ on the easy track, the highest of all submissions, and $57.18\%$ on the hard track of the private test set.

42. 【2609.21502】Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering

链接https://arxiv.org/abs/2609.21502

作者:Tianchen Deng,Guole Shen,Yilin Shen,Wenhua Wu,Yilin Fang,Ziqi Ma,Tianjun Zhang,Shenghai Yuan,Wolfram Burgard,Hesheng Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:enable generalizable geometric, generalizable geometric reasoning, models enable generalizable, reasoning from RGB, foundation models enable

备注

点击查看摘要

Abstract:Recent 3D foundation models enable generalizable geometric reasoning from RGB images but remain limited in persistent memory, scalability, and renderable scene modeling. We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering. Its core is an adaptive world memory mechanism that combines transformer-based gated updates with test-time temporal-spatial regulation. Learned gates control recurrent memory propagation, while temporal state evolution and spatial observation-state consistency regulate token-wise updates and forgetting over long image sequences. To support large-scale mapping, we organize memory into local submaps and integrate progressive mapping and tracking, loop closure, and SL(4)-based global refinement to maintain local accuracy and global consistency. A Gaussian reconstruction head decodes memory-enhanced features into renderable primitives, unifying camera pose estimation, dense point-cloud reconstruction, and photorealistic rendering within a single model. Experiments on public benchmarks and self-collected datasets from diverse robotic platforms demonstrate improved trajectory accuracy, reconstruction completeness, and rendering quality over existing 3D foundation reconstruction and SLAM baselines. These results support adaptive memory as a foundation for persistent robotic world modeling. The dataset and code will be made publicly available at \href{this https URL}{this https URL}.

43. 【2609.21498】VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization

链接https://arxiv.org/abs/2609.21498

作者:Yibin Zhao,Yihan Pan,Yangwen Li,Jun Nan,Jianjun Yi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:typically regress pixel-aligned, causing excessive overlap, pixel-aligned Gaussian primitives, regress pixel-aligned Gaussian, Gaussian Splatting

备注

点击查看摘要

Abstract:Recent feed-forward 3D Gaussian Splatting (3DGS) methods typically regress pixel-aligned Gaussian primitives, often causing excessive overlap and artifacts, while inaccuracies in predicted camera poses can lead to misalignment in novel-view synthesis (NVS). We present VoxelTTO, a feed-forward framework for reconstructing geometrically accurate 3DGS scenes from an arbitrary number of images and optional camera parameters. VoxelTTO aggregates dense image features into a global voxel representation and decodes Gaussians from voxel features, breaking the pixel-to-Gaussian correspondence. To exploit known camera parameters while keeping the pretrained visual foundation model (VFM) parameters frozen, we introduce test-time optimization (TTO) that adapts lightweight LoRA modules using pose supervision. We further replace vanilla 3DGS rasterization with stochastic solid volume rendering during training and inference, improving geometric fidelity. Training updates only the voxel-aligned Gaussian reconstruction modules, requiring 80 GPU hours. Experiments on Replica, Tanks and Temples, and DTU demonstrate improved RGB-D NVS and camera-pose estimation relative to prior methods.

44. 【2609.21480】OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection

链接https://arxiv.org/abs/2609.21480

作者:Alexey Bryncev,Andrey Moskalenko,Kira Shilovskaya,Ivan Kosmynin,Dmitriy Vatolin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:perceptual quality assessment, including viewport-adaptive streaming, immersive multimedia applications, saliency prediction plays, foveated rendering

备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Omnidirectional video saliency prediction plays an important role in many immersive multimedia applications, including viewport-adaptive streaming and compression, foveated rendering, mesh simplification, perceptual quality assessment. Yet progress in this area remains constrained by the cost and complexity of collecting eye-tracking data with VR headsets, which makes large-scale dataset creation difficult to extend. We present OpenSAL360, the first open-source platform for scalable, low-cost 360° video saliency collection. Unlike conventional VR-based protocols, it requires only a standard screen, mouse, and internet connection, enabling parallel saliency data collection from common crowdsourcing assessors without specialized hardware. We validate our collection protocol against seven well-established VR eye-tracking datasets and conduct ablation studies on key interface, pre-, and post-processing parameters. To demonstrate the effectiveness and scalability of the proposed methodology, we collect and publicly release a saliency dataset covering 500 omnidirectional videos annotated by 2,000+ crowdsourcing assessors, making it, to the best of our knowledge, the largest dataset in this field. We make OpenSAL360 publicly available at this https URL.

45. 【2609.21474】MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

链接https://arxiv.org/abs/2609.21474

作者:Yiguang Yang,Jiankun Peng,Xiaoming Wang,Yiran Zhang,Zhibo Fang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:diffusion Transformer forward, Transformer forward central, video diffusion Transformer, video-action co-training improves, diffusion Transformer

备注

点击查看摘要

Abstract:Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual structure needed for control. We present MT-WAM, which retains the original training objectives and adds complementary supervision for future two-dimensional point trajectories and visual features. A lightweight dual-stream branch copied from the video backbone's final blocks provides target-specific processing, while a structured attention mask prevents cross-stream attention. Motion-stream tokens supply additional dynamics conditions to the action expert. Future visual-feature prediction provides supervision in a feature space that captures object and spatial structure. This supervision trains the video backbone to provide more informative visual context for action generation under changing visual conditions, without adding visual-feature-stream tokens to action conditioning. At inference, MT-WAM uses video and motion caches computed once per replan and skips future-video prediction. Without additional embodied policy pretraining, MT-WAM achieves 98.2% success on LIBERO and 73.7% on LIBERO-Plus, exceeding Fast-WAM by 23.8 percentage points on the latter. On RoboTwin 2.0 Clean2Rand, Random success increases from 6.30% to 19.40%; across four real-world tasks, average success increases from 67.0% to 77.8%.

46. 【2609.21468】SkillIR: Evolving Scene-Aware Skills for Agentic Image Restoration

链接https://arxiv.org/abs/2609.21468

作者:Jie Shao,Shengkai Hu,Xu Zhang,Beihang Song,Yongcheng Jing,Xu Wu,Jun Wan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:paper studies agentic, studies agentic image, recover images affected, multimodal agents coordinate, agents coordinate specialized

备注

点击查看摘要

Abstract:This paper studies agentic image restoration, in which multimodal agents coordinate specialized restoration tools to recover images affected by complex degradations. Existing restoration agents often derive complete tool-use plans from the original degraded image or retrieve previously successful trajectories, providing limited support for adapting individual actions to evolving intermediate restoration states. We find that accepted tool executions can change the residual degradation state and, consequently, the applicability of subsequent tools. To address this issue, we propose SkillIR, a skill-guided framework that represents restoration experience as degradation-centered action evidence rather than complete tool-use trajectories. SkillIR consolidates context-dependent action outcomes into scene-aware restoration skills that characterize applicable conditions, expected effects, and attributable failure cases. Instead of prescribing a complete restoration plan, the retrieved skills guide one bounded action at a time within a verified residual-state loop: each tool output is treated as a candidate, committed only after transition verification, and followed by reassessment of the active residual degradations. After each rollout, the resulting evidence is used to create, refine, or patch dynamic skills, enabling accumulated restoration experience to improve decision-making for subsequent inputs. Experiments on synthetic and real-world multi-degradation datasets demonstrate that SkillIR improves restoration quality and enables more reliable and effective tool use.

47. 【2609.21462】PSEE: Progressive Sensor Event Expansion for Point-Supervised Temporal Action Localization

链接https://arxiv.org/abs/2609.21462

作者:Jiaxi Yin,Ge Wang,Han Ding,Fei Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:enabling finer-grained activity, finer-grained activity understanding, conventional action recognition, wearable sensor streams, sensor streams identifies

备注

点击查看摘要

Abstract:Temporal action localization (TAL) in wearable sensor streams identifies action classes and temporal boundaries, enabling finer-grained activity understanding than conventional action recognition. However, training typically requires costly start--end annotations for every action instance. To reduce this burden, we study point-supervised TAL, where each instance is labeled with only one timestamp and its class. We propose Progressive Sensor Event Expansion (PSEE), which combines semantic activations, sensor-specific transition evidence, and adaptive temporal ownership to recover point-supervised pseudo segments. These segments supervise standard TAL detectors without modifying their inference procedures. Cross-subject experiments on four inertial-sensing benchmarks demonstrate improved pseudo-boundary quality over adapted point-supervised baselines, compatibility with different TAL detectors, and robustness to point sampling. Code is available at this https URL.

48. 【2609.21455】CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation

链接https://arxiv.org/abs/2609.21455

作者:Haoran Qin(1),Renlong Wu(1),Tianyu Huang(1),Yukang Ding(2),Hui Li(1),Wangmeng Zuo(1) ((1) Harbin Institute of Technology, China, (2) Taobao, Alibaba Group, China)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:temporally coherent videos, demonstrated impressive capability, respect fundamental physical, coherent videos, demonstrated impressive

备注: 23 pages, 4 figures. Submitted to the 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Project page: [this https URL](https://makapic.github.io/CompAdapt/)

点击查看摘要

Abstract:While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a physics-consistent T2V framework for adaptable generation across complex real-world scenarios. It extends neural dynamics modeling beyond single-type motions to encompass composite physical behaviors, including coupled motions, multi-stage transitions, and multi-object collisions. Furthermore, CompAdapt translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial physical parameters. To generalize to novel physical environments, CompAdapt introduces dynamics-aware prior matching, achieving one-shot adaptation without retraining the core dynamics module. In addition, a physics-aware latent feature fusion module improves visual fidelity under fast and complex motion. Experiments on physics-focused T2V benchmarks demonstrate that CompAdapt improves physical consistency over both general T2V models and physics-constrained baselines, while preserving high visual quality and adaptability to unseen dynamics. The project page is available at this https URL .

49. 【2609.21449】ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling

链接https://arxiv.org/abs/2609.21449

作者:Xuancheng Zhang,Xuetao Liu,Qianying Tang,Jizhe Wang,Zhijing Cheng,Bochen Lin,Haoran Wen,Ming Li,Kun Zhan,Yu Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Tactile, Action Models bring, World Action Models, Action, providing a rich

备注

点击查看摘要

Abstract:World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.

50. 【2609.21437】hink Locally, Refine Globally for Memory-Efficient 3D Reconstruction

链接https://arxiv.org/abs/2609.21437

作者:Jingke Zhou,Chenhang Ma,Zhizhou Zhong,Mingkai Liu,Zhuang Zhou,Yicheng ji,Binghua Su,Bo Cai,Xianliang Huang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:balances local temporal, local temporal modeling, memory-efficient framework, balances local, global camera consistency

备注: 9 pages,4 figures

点击查看摘要

Abstract:We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register tokens via cross-attention to enforce scene-level constraints across the entire sequence. This design enables joint optimization of camera representations and significantly improves long-horizon pose stability without incurring the high cost of sequence-wide attention. Extensive experiments demonstrate that LoG-VGGT achieves improved depth accuracy and robust camera pose estimation across multiple long-sequence benchmarks, while delivering competitive streaming reconstruction performance.

51. 【2609.21424】P$^3$-SAM: SAM with Perceptual Parallel Prompt for Few-Shot Strip Steel Surface Defect Segmentation

链接https://arxiv.org/abs/2609.21424

作者:Qian Xu,Hang Xiong,Anpeng Wang,Sam Kwong,Cong Zhang,Runmin Cong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:strip steel surface, posed significant challenges, significant challenges distinct, steel surface defects, strip steel

备注: Accepted by ICME 2026, 6 pages, 3 figures. Corresponding authors: Anpeng Wang and Runmin Cong

点击查看摘要

Abstract:Few-shot semantic segmentation (FSS) of strip steel surface defects (S$^3$D) has posed significant challenges distinct from natural scenes. Unlike natural images, S$^3$D task exhibits unique characteristics including low local contrast, uneven illumination, and complex fine-grained texture patterns. Although recent methods based on Segment Anything Model (SAM) have shown promise in FSS on natural images by leveraging SAM's powerful pre-trained representations, these unique industrial characteristics of S$^3$D images lead to performance drop when directly applying SAM to industrial defect scenarios. In this paper, we propose a novel Perceptual Parallel Prompt (P$^3$) framework that empowers SAM, creating the P$^3$-SAM model to address these challenges through two core strategies. First, we develop a Perceptual-Optimized Encoding (POE) strategy that enhances local contrast and preserves critical texture details for S$^3$D segmentation. Second, we introduce the Parallel Prompt Generator (PPG) strategy that simultaneously generates both semantic and spatial prompts, enabling comprehensive guidance for SAM's decoder across varying images. Extensive experiments on three few-shot S$^3$D benchmarks demonstrate that P$^3$-SAM achieves state-of-the-art performance, with particularly notable improvements of 12.00% in mIoU on Surface Defects-4i dataset.

52. 【2609.21412】When Online Adaptation Hurts: Parameter-Frozen Test-Time Ensembling for Continual Medical Image Segmentation

链接https://arxiv.org/abs/2609.21412

作者:Ruijie Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:worse when sites, protocols change, scanner vendors, Medical image segmenters, cardiac MRI

备注: 7 pages, 2 figures

点击查看摘要

Abstract:Medical image segmenters often get worse when sites, scanner vendors, or protocols change. Continual test-time adaptation (CTTA) addresses this problem without target labels, but it can be impossible to update a model on a non-stationary stream and can lead to a lot of errors. We examine a more reasonable and meaningful alternative: parameter-frozen inference enhancement(PIE). We use a source-trained segmenter that learns about anatomy-preserving scale and flip views, maps their predictions back to the native location, and averages the probabilities. We do not modify the weights of the model or the normalization statistics. On a cardiac MRI stream from M\Ms, which is trained on vendor A and evaluated sequentially on vendors B, C, and D, PIE has 0.7786 mean Dice, compared to 0.7680 for source-only inference and 0.7388--0.7416 for five other online-adaptation baselines. The controlled ablations show that performance saturates at 28 views, and confidence weighting, class-prior correction, connected-component filtering, morphological refinement, and inter-slice smoothing have no effect or cause negative transfer. Qualitative results on cardiac MRI and fundus images are also consistent with the frozen ensemble keeping thinner and nested anatomical structures. These results provide a strong, stable baseline for medical CTTA and expose an important failure mode: adaptation and handcrafted refinement can be less reliable than carefully designed inference.

53. 【2609.21407】Quantization-Aware Kalman Estimation for Diffusion Sampling

链接https://arxiv.org/abs/2609.21407

作者:Qitan Shi,Cheng Jin,Jiawei Zhang,Yuantao Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Quantization offers, memory and computation, offers a practical, practical path, path to deploying

备注: 20 pages, 6 figures, 2 tables

点击查看摘要

Abstract:Quantization offers a practical path to deploying diffusion models with reduced memory and computation, but aggressive compression can cause quantized outputs to deviate substantially from their full-precision counterparts. Sampling-stage correction methods seek to compensate for such deviations during sampling, but existing approaches rely primarily on local information and underexploit trajectory history, limiting their ability to correct errors that propagate across timesteps. In this work, we formulate sampling with a quantized denoiser as an online estimation problem, using the history of quantized denoiser outputs to recover the underlying full-precision outputs required by the sampler. We propose QuAKE, a Quantization-Aware Kalman Estimator that combines a smooth trajectory prior with a conditional Gaussian observation model. At each sampling step, QuAKE recursively updates the posterior over the output window in closed form and feeds its posterior mean to the sampler. QuAKE is a lightweight plug-and-play corrector that requires no modification to the quantized network and naturally supports arbitrary high-order multistep ODE samplers. Experiments across W4A4-quantized text-to-image diffusion models show that QuAKE consistently outperforms existing methods in reducing the distributional discrepancy from full-precision sampling.

54. 【2609.21402】SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment

链接https://arxiv.org/abs/2609.21402

作者:Zhibo Zhang,Qijie Wang,Zengqiang Yan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Surgical instrument segmentation, surgical workflow analysis, Surgical instrument, plays a critical, Surgical

备注

点击查看摘要

Abstract:Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods formulate segmentation as a category-driven localization problem, limiting their ability to capture procedural context and task-dependent semantics in surgical workflows. We introduce Reasoning-Aware Surgical Instrument Segmentation (RA-SIS), a task formulation that frames segmentation as query-conditioned inference under surgical context. To benchmark this setting, we construct SurgRS, a surgical reasoning segmentation dataset consisting of 41,000 image-text pairs, which aligns instance-level masks with structured query-answer supervision to enable semantic grounding at the pixel level. Based on SurgRS, we propose Surgical Instrument Reasoning and Segmentation Assistant (SIRA), a multimodal framework that disentangles target-level and query-level semantics and integrates them with visual features through query-anchored dual alignment. By aligning query semantics with spatial features and segmentation prompts, SIRA enhances semantic-visual consistency in mask prediction. Extensive experiments on SurgRS demonstrate improvements over existing reasoning-aware baselines. Code is available at this https URL.

55. 【2609.21400】A Scene Language Model for Open-Vocabulary Scene Mapping

链接https://arxiv.org/abs/2609.21400

作者:Adam Lilja,Fabio Hübel,Siming He,Junsheng Fu,Claire Tomlin,Lars Hammarstrand,Jitendra Malik,Jonas Frey,Marco Pavone

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:scene mapping aims, scene, aims to build, scene map, mapping aims

备注

点击查看摘要

Abstract:Open-vocabulary 3D scene mapping aims to build a persistent representation of the objects in an environment. Existing systems typically rely on engineered mapping pipelines to associate observations, merge information across views, and maintain a consistent scene representation over time. Many additionally store feature-rich object representations, such as embeddings or image crops, increasing the size and complexity of the persistent memory. We introduce SceneLM, a Scene-Language Model that directly maintains a textual scene map. The full scene is represented as a structured text list of objects, which serves as the model's only persistent memory. For each input image, the model reads the current scene state and updates the map by adding, editing, and removing objects. To learn this behavior, we introduce supervision tasks for iterative scene map maintenance together with an automatic annotation pipeline that generates training data from images without human labels. We evaluate SceneLM on both a language-grounded retrieval benchmark and a localization benchmark. Across both benchmarks, the model produces a scene map that achieves competitive performance with complete mapping systems built from dedicated perception and geometric modules while producing a scene representation that is 6-12x more compact. We further show that SceneLM can be run online on an edge device through experiments on a quadruped. These results show that a persistent open-vocabulary 3D scene map can be maintained directly by a single vision-language model using only a lightweight text representation. Training and inference code is available on this https URL.

56. 【2609.21392】Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

链接https://arxiv.org/abs/2609.21392

作者:Qi Chen,Yunfei Chu,Haolin He,Yifan Yang,Zihan Liu,Yuxuan Wang,Ziyang Ma,Ruiyang Xu,Meng Gao,Yinsong Yan,Ling Wang,Hui Wang,Wen Huang,Yiheng Chen,Guanrou Yang,Qiuqiang Kong,Jin Xu,Xie Chen

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)

关键词:Natural audio-visual interaction, composed text prompts, carefully composed text, Natural audio-visual, text prompts

备注

点击查看摘要

Abstract:Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs' ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.

57. 【2609.21386】AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

链接https://arxiv.org/abs/2609.21386

作者:Seoyeon An,Hyeonseo Jang,Minsu Kim,Chanho Lee,Younghan Park,Kangwook Lee

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:advancing artificial intelligence, Comprehensive video understanding, Multimodal Large Language, Large Language Models, Comprehensive video

备注: 36 pages, 8 figures. Code: [this https URL](https://github.com/krafton-ai/agentvidbench) Dataset: [this https URL](https://huggingface.co/datasets/agentvidbench/agentvidbench)

点击查看摘要

Abstract:Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic capabilities. To bridge this gap, we introduce AgentVidBench, a multi-hop video question answering benchmark focused on evaluating the spatial, temporal, and causal reasoning capabilities of MLLM agents. Beyond standard question-answer pairs, AgentVidBench provides step-by-step solution traces to support trajectory evaluation that assesses whether agents explicitly acquire the evidence needed to justify their answers. Experiments with 12 proprietary and open-source MLLMs show that single-turn performance remains limited on AgentVidBench, while integrating these models into state-of-the-art agentic workflows generally improves performance with respect to both accuracy and trajectory scores. We further present a simple yet effective agentic strategy that serves as a competitive baseline on AgentVidBench, establishing our benchmark as a holistic testbed for future research on agentic video understanding. Code and datasets are available at this https URL and this https URL.

58. 【2609.21379】JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting

链接https://arxiv.org/abs/2609.21379

作者:Trinh Tra Giang Nguyen,Thanh Nguyen Vo,Nguyen Hoai Thuong Bui,Ha Duc Bui

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:realistic future observations, Accurate traffic forecasting, synthesizing realistic future, Accurate traffic, synthesizing realistic

备注: ECCV Workshop 2026, AI City Challenge 2026 Track 5

点击查看摘要

Abstract:Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this work, we propose a decoupled forecasting framework that separates future representation learning from video generation. A frozen V-JEPA encoder first extracts predictive latent representations from the observed traffic videos, capturing the underlying scene dynamics in a semantic latent space. A lightweight latent alignment module then projects these representations into the conditioning space of a frozen Cosmos diffusion module, enabling future video synthesis without retraining the large generative model. By freezing all foundation models and training only the lightweight alignment module, the proposed framework substantially reduces optimization complexity while preserving forecasting capability. Experimental results on the AI City Challenge 2026 Track 5 benchmark demonstrate that the proposed method achieved a score of 75.1297, ranking third in the competition. These results suggest that predictive world representations learned by V-JEPA can effectively guide downstream video generation, providing a practical and efficient alternative to end-to-end diffusion-based forecasting.

59. 【2609.21371】RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy

链接https://arxiv.org/abs/2609.21371

作者:Xinyi Che,Zheng Lian,Kuofei Fang,Xuehao Wang,Xinghai Gao,Junqing Wu,Chuyu Wu,Liyi Liu,Yanhan Huang,Keyi Xie,Haomin Ouyang,Jinyang Wu,Fan Zhang,Runhao Zeng,Xun Yang,Bin He

类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:Social Proactive Intelligence, extends proactive assistance, Proactive Intelligence, Social Proactive, extends proactive

备注

点击查看摘要

Abstract:Social Proactive Intelligence (SPI) extends proactive assistance beyond task completeness to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse scenarios. To address these gaps, we introduce RobotEQ-Video, shifting the focus from image-centric to video-centric analysis. To ensure comprehensive video coverage, we construct a hierarchical world-state taxonomy organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 16K+ labels for assessing behavior properness. Benchmark evaluation reveals that current systems remain unreliable and fall short of human performance. We further explore how world models can help tackle this task. This work advances SPI research from static images to dynamic videos and ensures more comprehensive scenario coverage during benchmarking.

60. 【2609.21369】ProTracer: Proprioception-Guided Failure Diagnosis in Robot Manipulation

链接https://arxiv.org/abs/2609.21369

作者:Chang Dong,Mehdi Hosseinzadeh,King Hang Wong,Lingqiao Liu,Francois Fraysse,Feras Dayoub,Minh Hoai Nguyen

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:robot execution deviates, valid task-completion trajectory, robot manipulation failure, binary failure detection, includes binary failure

备注: 9pages, 5 figures, 5 tables

点击查看摘要

Abstract:This paper presents a comprehensive framework for robot manipulation failure analysis that includes binary failure detection, failure categorization, explanation generation, and the additional capability of failure onset localization, which aims to identify the earliest moment at which a robot execution deviates from a valid task-completion trajectory and is ultimately followed by task failure. To address these tasks, we propose ProTracer, a training-free framework that leverages existing Vision-Language Models (VLMs) together with proprioceptive signals for failure analysis. Our method uses proprioceptive dynamics to identify temporally informative action boundaries and converts richer robot-state signals into structured natural-language descriptions that can be jointly analyzed together with visual observations by the VLM. This design combines the temporal precision of proprioceptive signals with the multimodal reasoning capabilities of modern VLMs without requiring additional model training. We further introduce FailTime, a benchmark with synchronized visual and proprioceptive observations for evaluating conventional failure diagnosis tasks as well as failure onset localization. Experiments demonstrate that ProTracer achieves strong performance across both conventional failure diagnosis tasks and the newly introduced failure onset localization task, highlighting the importance of proprioceptive reasoning for fine-grained temporal failure analysis.

61. 【2609.21363】Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models

链接https://arxiv.org/abs/2609.21363

作者:Yining Wang,Xi Li,Mi Zhang,Xiaohan Zhang,Xiaoyu You,Zhenxing Qian,Mi Wen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Multimodal large reasoning, demonstrated remarkable capabilities, complex visual understanding, Multimodal large, demonstrated remarkable

备注: NDSS 2027

点击查看摘要

Abstract:Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms.

62. 【2609.21354】Field Tracking of Insects Using a Stereoscopic Event-Based Camera Setup

链接https://arxiv.org/abs/2609.21354

作者:Pratham G. Shenwai,Martin J. Lankheet,John T. Hrynuk,Mandiyam Y. Mahadeeswara,Mandyam V. Srinivasan,Sridhar Ravi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:High-speed tracking, environments is important, High-speed, behavior and ecology, temporal resolution

备注

点击查看摘要

Abstract:High-speed tracking of small, fast-moving organisms in their natural environments is important to better understand their behavior and ecology. Traditional frame-based imaging suffers from motion blur due to low temporal resolution, and data storage limitations, propelling a search for more adaptive solutions. Event cameras, which capture changes in brightness at pixel level instead of entire frames, have emerged as a promising solution by increasing temporal resolution and data efficiency. Here, we demonstrate the use of event-based imaging with standard video-based processing methods by converting the asynchronous events into conventional video formats, allowing us to leverage the event camera's enhanced temporal detail to capture intricate insect flight movements and apply established image analysis techniques. Coupling this conversion process with a stereoscopic configuration provides continuous, low-latency, three-dimensional tracking of fast-moving subjects in field conditions. As a result, we substantially mitigate motion artifacts and achieve more accurate representations of animal movements. By making event-based imaging more readily applicable in natural field settings, our method support broader applications across animal behavior and ecological research, agricultural management, and other fields requiring high-fidelity object tracking in the wild.

63. 【2609.21351】PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR

链接https://arxiv.org/abs/2609.21351

作者:Guangyi Liu,Qianjun Huang,Boyu Hou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:suffers from frequent, frequent structural errors, Table extraction suffers, errors and semantic, frequent structural

备注: Accepted by EMNLP industry track

点击查看摘要

Abstract:Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.

64. 【2609.21347】Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization

链接https://arxiv.org/abs/2609.21347

作者:Xiangfei Guo,Hao Shi,Yufan Zhang,Zhonghua Yi,Yongqi Mao,Xiaoting Yin,Kaiwei Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent progress, enabled dense visual, Gaussian Splatting, dense visual SLAM, panoramic GS-SLAM framework

备注: Accepted to ECCV 2026. Source code : [this https URL](https://github.com/guoxf304/CubeSplat)

点击查看摘要

Abstract:Recent progress in 3D Gaussian Splatting (3DGS) has enabled dense visual SLAM with pinhole cameras, yet most pipelines are not designed for panoramic imagery. We present Cube-Splat, the first panoramic GS-SLAM framework that factorizes each 360° frame into a cubemap of four fixed-orientation virtual pinhole views sharing a single optical center. By designating the front face as the primary pose state, we accumulate gradients from all faces via an adjoint mapping, thereby enabling multi-face observations to coherently update a single state while strictly preserving cross-view geometric consistency. Concurrently, our mapping module densifies and optimizes anisotropic Gaussians using aggregated cubemap rays for high-fidelity, dense reconstruction. Furthermore, to rigorously evaluate panoramic SLAM under diverse and challenging conditions, we introduce SynPano, a highly scalable, photorealistic synthetic dataset featuring parameterized complex trajectories and multi-modal ground truth. Extensive evaluations on two public benchmarks (PALVIO and OmniBlender) and our SynPano dataset, collectively encompassing both indoor and outdoor scenes, demonstrate that Cube-Splat achieves state-of-the-art (SOTA) performance in tracking accuracy and reconstruction fidelity. Both the source code and the SynPano dataset are available at this https URL.

65. 【2609.21323】VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

链接https://arxiv.org/abs/2609.21323

作者:Hongyi Lin,Yiyao Liu,Qi Kang,Heye Huang,Yang Liu,Haris Koutsopoulos,Jinhua Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrated strong scene, strong scene understanding, Vision-language models, perception remains unclear, diverse tasks

备注: 8 pages, 4 figures

点击查看摘要

Abstract:Vision-language models (VLMs) have demonstrated strong scene understanding and semantic judgment across diverse tasks, but their appropriate role in cooperative perception remains unclear. Directly asking a VLM to regress 3D detections is unreliable and computationally expensive, whereas using it to select the output of a single source discards useful information from other agents. We introduce VeriFuse, a bounded arbitration framework for vehicle-infrastructure cooperative 3D detection. Each agent first produces detections independently. Around each vehicle and roadside proposal, VeriFuse generates source-conditioned geometric candidates and combines the original detections, their perturbations, and cross-source hypotheses into a unified candidate pool. A frozen VLM then chooses among three admissible actions: SELECT an adequate candidate; REFINE an existing anchor when an object is supported but all candidates are geometrically inadequate; or REJECT an unsupported infrastructure-only proposal. Experiments on the DAIR-V2X dataset show that VeriFuse achieves 0.494/0.357 cooperative 3D AP50/AP70 and limits the relative vehicle-side BEV AP50 drop under a 300 ms delay to 1.7%. Overall, VeriFuse assigns the VLM a clear and constrained role in cooperative perception: semantic reasoning resolves ambiguity among cross-agent hypotheses, while deterministic constraints determine the final 3D geometry.

66. 【2609.21322】S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

链接https://arxiv.org/abs/2609.21322

作者:Kui Jiang,Yiang Chen,Yan Luo,Zhaocheng Yu,Junjun Jiang,Xianming Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Heavy rainfall severely, introducing motion blur, rainfall severely degrades, severely degrades outdoor, corrupting high-frequency details

备注

点击查看摘要

Abstract:Heavy rainfall severely degrades outdoor videos by corrupting high-frequency details and introducing motion blur, critically undermining the reliability of visual tasks. Recently, State Space Models (SSMs), particularly Mamba, have emerged as efficient alternatives for vision tasks with their linear complexity and ability to model long-range dependencies. However, when confronted with the poor visual representations in rainy videos, Mamba still faces difficulties in preserving the integrity of 2D spatial semantics and modeling 3D spatio-temporal correlations. To break these limitations, we introduce S3VD, a Semantic-Guidance Spatio-Temporal Scanning framework for video deraining, featuring two key innovations: Multi-Scale Semantic Fusion (MSSF) Module and Spatio-Temporal Scanning Fusion (STSF) Module. The former integrates temporal semantic priors from DINOv2 to guide precise feature representation and counteract the loss of local semantic context inherent to Mamba's 1D flatten operation, enhancing robustness against extreme degradation. The latter introduces a spatio-temporal scanning mechanism and devises a Decoupled-Gating Mamba (DG-Mamba) layer, which employs two independent gates to adaptively control preceding and subsequent contextual information within the input clip, optimizing intra-frame and inter-frame correlation modeling. Experiments on video deraining benchmarks demonstrate the superiority of S3VD, achieving state-of-the-art performance with an average 0.84 dB PSNR improvement over Mamba-based baselines.

67. 【2609.21304】Combining Object Detection with Geometry-Aware Clustering to Distinguish Overlapping Plants in UAV Imagery

链接https://arxiv.org/abs/2609.21304

作者:Ik Jae Lee,Hieu D. Nguyen,Mahbubur Meenar,Carlos Morrison Martinez,Cameron Connelly

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unmanned aerial vehicle, Reliable plant-level information, Reliable plant-level, aerial vehicle, unmanned aerial

备注: 34 pages

点击查看摘要

Abstract:Reliable plant-level information from unmanned aerial vehicle (UAV) imagery is important for automated crop monitoring. However, in dense crop canopies, adjacent plants frequently overlap and are detected as a single object, reducing the reliability of plant-level measurements. This study presents a geometry-aware post-detection framework for resolving overlapping plant instances using standard RGB UAV imagery. The framework combines object detection with geometric clustering of plant components. Leaves or branches detected within each bush-level region are represented using two complementary geometric features: component centroids and radial intersection points (RIPs) derived from detected plant structures. K-means and Gaussian mixture models determine whether a detected region contains a single plant or two overlapping plants. Density filtering suppresses spurious radial intersections, and a post-pipeline ensemble combines spatial and directional geometric information. The framework was evaluated using UAV imagery of eggplant and tomato crops under field conditions. Centroid-based clustering achieved an F1-score of 0.89 for eggplant, while the combined centroid-RIP approach achieved the best tomato performance, with an accuracy of 0.80, precision of 1.00, and F1-score of 0.75 using K-means. Density filtering substantially improved RIP-based clustering for tomato. The proposed approach provides a lightweight, modular engineering solution that can be integrated with existing RGB UAV monitoring pipelines without additional depth sensors, pixel-level segmentation, three-dimensional reconstruction, or retraining of the primary bush detector. The results demonstrate that geometric reasoning applied to existing detector outputs can complement deep-learning-based object detection and improve plant-level interpretation in dense agricultural canopies.

Comments:
34 pages

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

ACMclasses:
I.4.9

Cite as:
arXiv:2609.21304 [cs.CV]

(or
arXiv:2609.21304v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.21304

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Hieu Nguyen [view email] [v1]
Fri, 18 Sep 2026 04:25:07 UTC (34,223 KB)

68. 【2609.21276】Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering

链接https://arxiv.org/abs/2609.21276

作者:Jia Li,Li Dai,Peng Jia,Zhenzhen Hu,Chee Seng Chan,Bingkun Bao,Richang Hong

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Circuit Board Assembly, Printed Circuit Board, automated Printed Circuit, Board Assembly, Printed Circuit

备注: 8 pages, 2 figures. Accepted to the 34th ACM International Conference on Multimedia (ACM MM 2026)

点击查看摘要

Abstract:In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and manufacturing knowledge. Although large vision-language models (VLMs) provide a promising foundation, their deployment is hindered by the domain shift between standards-derived samples and real-world production-line imagery, together with heterogeneous output spaces spanning choice-based and numerical counting tasks. To address these challenges, we propose a multimodal reasoning framework for cross-domain PCBA visual question answering. The framework converts standards-derived, real-world, and auxiliary PCB-domain data into a unified instruction format and constructs verified reasoning traces aligned with visual evidence, question semantics, candidate options, and ground-truth answers. We further introduce Task-Aware Group Relative Policy Optimization (GRPO), which moves beyond exact-match supervision by integrating multi-component semantic rewards for choice-based questions, distance-aware rewards for counting questions, and an auxiliary format reward for valid outputs. During inference, answer-option semantic consistency correction, self-consistency voting, and multi-model arbitration are combined to improve prediction robustness. The proposed system achieves an Overall Score of 83.24 on the official PCBA Standard-to-Real Grand Challenge leaderboard, demonstrating the effectiveness of task-aware reward design and robust inference for cross-domain PCBA visual question answering.

69. 【2609.21268】Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

链接https://arxiv.org/abs/2609.21268

作者:Chongbo Zhao,Jiangming Wang,Xilai Wang,Xinyu Wang,Jingyi Tang,Chunjie Hao,Pengjie Song,Yue Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:modifies target content, editing modifies target, unedited regions, Text-guided video editing, modifies target

备注: Project page: [this https URL](https://chongbozhao3-coder.github.io/Edit-VAR) . Code: [this https URL](https://github.com/chongbozhao3-coder/Edit-VAR)

点击查看摘要

Abstract:Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.

70. 【2609.21251】Geometry-Aware Diffusion Guidance via Curvature-Adaptive Tubular Correction

链接https://arxiv.org/abs/2609.21251

作者:Enze Jiang,Jinwei He,Zheng Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:provide flexible priors, samplers provide flexible, Gradient-guided diffusion samplers, conditional generation, poorly supported

备注

点击查看摘要

Abstract:Gradient-guided diffusion samplers provide flexible priors for inverse problems and conditional generation, but strong guidance can move the sampling trajectory into regions where the learned score is poorly supported. Existing tangent-projection strategies limit first-order departure from an iso-density surface, yet discard potentially useful normal motion and overlook the second-order departure induced by tangent motion on a curved surface. We introduce curvature-adaptive tubular correction (CAT), a training-free plugin that regulates both effects within a shared, noise-dependent geometric budget. CAT decomposes the guidance gradient into normal and tangent components, charges normal displacement at first order and tangent displacement according to directional curvature, and obtains their jointly optimal magnitudes from a one-dimensional dual equation. Armijo backtracking calibrates the resulting finite step against the actual guidance objective, while matrix-free directional derivatives avoid constructing the full score Jacobian. We establish local guarantees for the tubular approximation, uniqueness of the correction, and sufficient objective decrease. Across seven inverse problems on FFHQ and ImageNet, CAT improves the evaluated pixel- and latent-space host samplers, with particularly consistent gains in perceptual metrics. It also improves black hole reconstruction on InverseBench and yields the lowest FID among the compared methods at every tested classifier-free guidance scale, while maintaining stable saturation and contrast. These results support curvature-aware tubular control as a reusable mechanism for stabilizing diffusion guidance.

71. 【2609.21246】VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models

链接https://arxiv.org/abs/2609.21246

作者:Kaiwen Zhu,Dongfang Liu,Liangkai Liu

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)

关键词:map visual observations, models map visual, OOD, compromise their reliability, VLA

备注: 9 pages, 3 figures

点击查看摘要

Abstract:Vision-language-action (VLA) models map visual observations and natural-language instructions to robotic actions, but distribution shifts can compromise their reliability. Because these models may still succeed under out-of-distribution (OOD) conditions, detecting OOD inputs alone is insufficient to predict execution failure. In this paper, we introduce VLA-Scope, a two-stage framework that combines input-shift characterization with execution history to predict failure during OOD rollouts. The first stage uses pooled image and language representations to detect OOD inputs and classify their shift categories. For inputs flagged as OOD, the second stage combines the predicted category, action-prefix features, and execution progress features. A logistic regression model shared across shift categories updates failure risk as execution proceeds. We evaluate the framework with OpenVLA on ten LIBERO-Spatial tasks using leave-one-group-out cross-validation. OOD detection achieves a ROC-AUC of 0.9454, and shift classification achieves 91% accuracy. Evaluated independently of the OOD gate on all 1,400 OOD rollouts, the failure predictor achieves a ROC-AUC of 0.8497 after 60 executed actions, compared with 0.7906 without execution progress features. It also achieves a higher ROC-AUC than the evaluated ActProbe and SAFE-MLP baselines. These results suggest that combining action features with temporally aggregated execution step representations improves failure prediction under input shifts.

72. 【2609.21242】SafeStyle: Calibrated Style Residual Injection for Controllable Style-Leakage Trade-off in Diffusion Stylization

链接https://arxiv.org/abs/2609.21242

作者:Zhangping Yang,Min Li,Song Yan,Rong Gao,Xinliang Bi,Guanye Xiong,Yujie He

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reference-guided diffusion stylization, diffusion stylization aims, transfer visual style, Reference-guided diffusion, stylization aims

备注: 5pages, 6figures

点击查看摘要

Abstract:Reference-guided diffusion stylization aims to transfer visual style from a reference image while preserving the semantics specified by a text prompt. However, image conditioning often entangles transferable style cues with reference-specific content, leading to an inherent trade-off: stronger conditioning improves style fidelity but increases content leakage, whereas aggressive suppression reduces leakage at the cost of style expression. This challenge is further complicated by the distinct spatial organization of texture- and geometry-dominant styles. To address these issues, we propose SafeStyle, a training-free framework for calibrated style residual injection in frozen diffusion models. SafeStyle first estimates style-supported and content-associated subspaces from compact calibration sets, preserving their informative overlap while suppressing useless content variations. It then transports the purified style evidence over adaptive spatial granularity and constrains its effective influence through an explicit residual-norm budget. Experiments across texture- and geometry-dominant styles show that SafeStyle achieves a DINO style similarity of 0.432 while maintaining competitive text alignment. On a semantically disjoint leakage-stress benchmark, it further achieves a DINO style similarity of 0.474 with only 0.8\% semantic leakage, demonstrating an effective balance between style fidelity and reference-content suppression.

73. 【2609.21241】Multiclass Semantic Segmentation of Wildland Fire Images Using Context-Aware Centralized Copy-Paste Data Augmentation

链接https://arxiv.org/abs/2609.21241

作者:Joon Tai Kim,Nishanth Kunchala,Vishv Patel,Tianle Chen,Ziyu Dong,Daniel Ospina Acero,Roger Williams,Mrinal Kumar

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Producing accurate annotations, deep learning based, Producing accurate, learning based image, labor intensive

备注: 14 pages, 9 figures

点击查看摘要

Abstract:Producing accurate annotations for deep learning based image segmentation is both costly and labor intensive. This challenge is especially evident in wildland fire applications, where accurately labeled datasets are scarce due to the difficulty of collecting and annotating dynamic fire scenes. To address this problem, our previous work introduced the Centralized Copy-Paste Data Augmentation (CCPDA) method for semantic segmentation of wildland fire imagery, which generates artificial training samples by randomly pasting fire clusters from source images onto target images. However, random placement can produce contextually unrealistic scenes, such as fire burning on asphalt. In this paper, we present a context-aware strategy designed specifically to improve data quality and realism in small multiclass wildland fire datasets, ensuring that augmented samples remain contextually meaningful. The proposed method restricts fire placement to semantically valid target regions and selects the location whose Ash-Vegetation composition most closely matches the source context. This approach preserves existing fire regions in the target image, prevents unrealistic placements, and maintains contextual accuracy by generating images that resemble real wildland fire scenes. We evaluate the Context-Aware CCPDA strategy through numerical analysis and comparisons with other augmentation methods by a weighted sum-based multi-objective optimization (MOO) approach. The results confirm that the context-aware data augmentation strategy leads to improved segmentation performance and contextual realism, outperforming other augmentation procedures.

74. 【2609.21228】FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

链接https://arxiv.org/abs/2609.21228

作者:Zhiyuan Gao,Di Wen,Yanxiang Zhan,Mohammad Khoshnazar,Jeroen Schäfer,Kunyu Peng,Michael Beetz

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrated strong performance, diverse robotic manipulation, pretrained vision-language models, robotic manipulation tasks, VLA models

备注

点击查看摘要

Abstract:Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in precise and long-horizon manipulation. Recent methods enhance VLA models through geometric supervision and future-state prediction across the entire scene. However, these methods can suffer from redundant scene information, distracting the model from learning the geometry and dynamics relevant to the current interaction. To address this issue, we propose FOCAL-VLA, a framework that combines subtask-guided geometry distillation with implicit world modeling to learn representations of current spatial structure and future interaction dynamics. To focus geometric learning on the current subtask, we transfer geometric knowledge from VGGT to the VLA model by aligning geometry latents with features from subtask-relevant image regions. To capture the future 3D evolution of the current interaction, we incorporate implicit world modeling using Track4World features from current and future demonstration frames. The two complementary representations jointly guide action generation without running VGGT or Track4World at inference time. Experiments show that FOCAL-VLA outperforms baselines on both simulation benchmarks and real-world manipulation tasks. Project website: this https URL.

75. 【2609.21225】VGGT-CAD: Reconstructing Parametric CAD 3D Model with Geometric Grounding

链接https://arxiv.org/abs/2609.21225

作者:Chunan Yu,Tianrun Chen,Fu Shen,Cheng Chen,Lanyun Zhu,Yang Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Parametric CAD reconstruction, reconstruction requires recovering, editable modeling operations, CAD reconstruction, CAD reconstruction requires

备注

点击查看摘要

Abstract:Parametric CAD reconstruction requires recovering both precise geometry and editable modeling operations from visual observations, making it challenging under limited and ambiguous views. Existing methods mainly rely on 2D appearance cues and lack strong multi-view geometric priors. In this work, we present VGGT-CAD, a geometry-aware framework for parametric CAD reconstruction from single- and multi-view observations. We transfer pretrained 3D geometric priors into CAD reconstruction by encoding camera parameters as condition tokens and jointly modeling them with image tokens. To handle varying numbers of viewpoints, we introduce a variable-view cross-view context aggregation module that adaptively fuses multi-view features. We further develop a training-free geometry-aware view selection strategy to select complementary and reliable frames during inference. The resulting representation is decoded into CAD command sequences using a non-autoregressive decoder. We also develop VideoCAD, a large-scale multi-view video benchmark derived from existing CAD data through multi-view re-rendering. Extensive experiments demonstrate the effectiveness of VGGT-CAD for visual CAD reconstruction under different observation configurations.

76. 【2609.21219】Multi-viewpoint Geo-localization with Event Cameras

链接https://arxiv.org/abs/2609.21219

作者:Adam D. Hines,Michael Milford,Tobias Fischer

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Robot localization, ongoing challenge, challenge that demands, demands mapping, mapping and positioning

备注: 8 pages, 4 figures, 4 tables, under review

点击查看摘要

Abstract:Robot localization is an ongoing challenge that demands mapping and positioning systems that are tolerant to viewpoint change. Event cameras are attracting increasing interest and adoption in robotics; however, dealing with viewpoint variance is an under-investigated problem in existing event-based localizers. In addition, event-based datasets that emphasize viewpoint variance for challenging localization situations are scarce. Here, we introduce an event-based visual place recognition (VPR) system that performs robustly under viewpoint changes. We converted five large-scale geo-tagged datasets, conventionally used to train frame-based localization systems, into synthetic event streams using Image-to-Event (I2E) conversion, and used them to fine-tune a pre-trained event-based vision transformer backbone with a multi-loss function, yielding a system we call MegaEvent that learns viewpoint-robust features for place recognition. We achieved an average Recall@1 of 82% across three existing event-based localization datasets, leading the next best event-based method by 20 recall points, and frame-based VPR models applied directly to event frames by 8 to 26 recall points. We introduce a new, challenging dataset - Springfield-Event-VPR - which features a 3.7km walking route recorded in three camera orientations for a total of 11.1km, which MegaEvent outperforms the strongest baseline by 9 recall points. The code for MegaEvent is available at this https URL.

77. 【2609.21207】Hand-Aware Transition Modeling for Bimanual Procedural Anomaly Detection

链接https://arxiv.org/abs/2609.21207

作者:Di Wen,Jimmy Weissert,Luc Maria Scherrer,Cedric Zöllner,Kailun Yang,Ruiping Liu,Yufan Chen,Jiale Wei,Junwei Zheng,Kunyu Peng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Procedural anomaly detection, assembly requires judging, Procedural anomaly, requires judging, bimanual assembly requires

备注: 6 pages, 1 figure, 3 tables. Code: [this https URL](https://github.com/Kratos-Wen/HACT)

点击查看摘要

Abstract:Procedural anomaly detection in bimanual assembly requires judging each hand action against the execution so far. A corrective action may look unusual in isolation, while a visually plausible action can violate the order of the procedure. We present HACT, a transition model over predicted per-hand events. A role-preserving history keeps the concurrent responsibilities of both hands, and a marked temporal point process assigns each observed transition a semantic and temporal surprisal. A supervised evidence head and a two-state filter convert these surprisals into per-hand anomaly posteriors. A recovery-aware protocol on predicted events and participant-disjoint folds reports the recovery false-positive rate at an operating point selected on validation participants. On two bimanual power-tool procedures HACT has the highest AUPRC and F1 among the compared methods and the fewest recovery alarms. Applied without retraining to a different assembly order of the same product, it retains the highest AUPRC and F1. The source code is available at this https URL.

78. 【2609.21199】OnomatoBridge: Onomatopoeia Translation and Rendering Pipeline in Manga

链接https://arxiv.org/abs/2609.21199

作者:Takara Taniguchi,Wataru Shimoda,Kota Yamaguchi,Hideki Nakayama

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:white paints gaining, paints gaining popularity, comic drawn, drawn by black, black and white

备注

点击查看摘要

Abstract:Manga is a comic drawn by black and white paints gaining popularity around the world. Onomatopoeia in Manga specifically appeals to the audience with its unique visual styles, which convey sound, motion, and emotion. Visual onomatopoeia translation requires the clean replacement of Japanese onomatopoeia with onomatopoeia in the other language while preserving their visual style. Existing approaches often produce residual artifacts or style inconsistency when removing the Japanese onomatopoeia and rendering stylized English onomatopoeia. To approach these problems, we present OnomatoBridge, a filtering pipeline for visual onomatopoeia translation. We evaluate OnomatoBridge from Japanese to English on the Manga109 onomatopoeia dataset and compare it with baseline image editing models. Experimental results show that the filtered outputs by the proposed method outperform those of conventional methods. OnomatoBridge improves English text correctness by roughly 10 to 25 points and reduces residual Japanese text by about 20 to 50% in relative terms.

79. 【2609.21186】Robust Structureless Monocular Visual Inertial Initialization Exploiting Line Features and Vanishing Points

链接https://arxiv.org/abs/2609.21186

作者:Junwan Choi,Woongrae Jo,Dong-Uk Seo,Jinwoo Jeon,Hyun Myung

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:reliable visual-inertial odometry, visual-inertial odometry, Accurate initialization, essential for reliable, reliable visual-inertial

备注: 8 pages, 5 figures, Accepted to IROS 2026

点击查看摘要

Abstract:Accurate initialization is essential for reliable visual-inertial odometry (VIO), but it is often ill-conditioned under degenerate motions. Existing methods typically require restrictive excitation motions to ensure sufficient observability or rely on computationally expensive 3D structure reconstruction, limiting efficient and practical deployment. To address these limitations, we propose SLIM-init, a structureless monocular VIO initializer that directly exploits geometric constraints from tracked 2D line features without explicit 3D landmark reconstruction. Specifically, SLIM-init leverages line-derived vanishing points (VPs) as translation-invariant orientation cues to provide robust rotation-only constraints under degenerate scenarios such as low-parallax or translation-dominant motions. It further incorporates a line epipolar residual to constrain translation and a line-normal projection residual to improve the conditioning of linear alignment, enhancing the accuracy and robustness of initial state estimation. Extensive experiments on a public benchmark and challenging custom degenerate-motion sequences demonstrate improved accuracy and robustness over state-of-the-art initialization methods. The source code is available at: this https URL.

80. 【2609.21176】4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors

链接https://arxiv.org/abs/2609.21176

作者:Haitao Huang,Shenghao Zhao,Boyuan Tian,Shin-Fang Chng,Songlin Yang,Sheila Lim,Huangying Zhan,Yi Xu,Anyi Rao,Frank Guan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:paper addresses, addresses the challenges, dynamic scene synthesis, Gaussian Splatting, Gaussian

备注: Accepted to SIGGRAPH Asia TC

点击查看摘要

Abstract:This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.

81. 【2609.21095】MarsFM: Shading-Regularized Flow Matching for Martian Relief Estimation

链接https://arxiv.org/abs/2609.21095

作者:Marius F. R. Juston

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:HiRISE RED orthoimagery, single-band HiRISE RED, image-conditioned latent flow-matching, latent flow-matching model, Martian relief estimation

备注: 73 pages, 62 figures

点击查看摘要

Abstract:We present MarsFM, an image-conditioned latent flow-matching model for local Martian relief estimation from single-band HiRISE RED orthoimagery. The method combines a pretrained generative prior with stereo-derived geometric supervision and a differentiable Lunar--Lambert shading objective. Relief, normal, gradient, curvature, and ordinal terms constrain complementary aspects of terrain structure, while a positive-affine-invariant image comparison constrains rendered appearance. An evaluation comprising 2024 gathered patch records per integration-step count yields mean affine-aligned RMSE between 0.0935 and 0.0957 in normalized signed-log relief space for one to twenty Euler steps. These scores measure agreement with VAE-reconstructed references on positive-reference support. Their narrow range supports low-step inference under this protocol. Spatial, differential, and spectral diagnostics show broad terrain correspondence alongside smoothing, amplitude compression, and boundary mismatch. MarsFM provides a framework for combining learned terrain priors with image-based constraints; establishing improved physical terrain resolution requires matched baselines and independent high-resolution reference data. Data: this https URL code: this https URL.

82. 【2609.21018】MAGIC: Marginal-Guided Compression with Optimal Transport for Efficient Visual Document Retrieval

链接https://arxiv.org/abs/2609.21018

作者:Xu Yuan,Hua Liu,Wenqi Fan,Qing Li

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:Recent visual document, MaxSim scoring overhead, patch-level vectors enable, vectors enable fine-grained, enable fine-grained evidence

备注

点击查看摘要

Abstract:Recent visual document retrieval (VDR) systems such as ColPali use multi-vector page embeddings, in which patch-level vectors enable fine-grained evidence matching but incur substantial index storage and MaxSim scoring overhead. Post-hoc merging offers a practical route to efficient VDR by reducing this cost without retraining the retriever, but its uniform reconstruction objectives are poorly aligned with the sparse, non-uniform patch usage induced by late-interaction retrieval. Under aggressive compression, this misalignment can preserve rarely used patches while concentrating retrieval activity on too few retained representatives. To address this misalignment, we propose Marginal-Guided Compression with Optimal Transport (MAGIC), a training-free post-hoc compressor for efficient retrieval with frozen multi-vector embeddings. MAGIC derives a MaxSim-induced compression surrogate and optimizes it through a two-marginal entropic optimal-transport formulation, where a retrieval-demand source marginal prioritizes high-use patches and a balanced target marginal regularizes retained-facet usage. Across ViDoRe benchmarks, keep ratios, and retrieval backbones, MAGIC consistently outperforms strong post-hoc compressors, with particularly large gains in the aggressive-compression regime; component ablations verify the complementary effects of its two marginals. We release the code at: this https URL.

83. 【2609.21012】Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification

链接https://arxiv.org/abs/2609.21012

作者:Sara Miketek,Biagio Barchielli,Nadeem Iqbal Kajla,Sinem Aslan

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Artistic style classification, exploit global composition, Artistic style, spatial organisation, global composition

备注: VISART Workshop, ECCV 2026 (Oral)

点击查看摘要

Abstract:Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive only as fragmented remains, forcing recognition from incomplete, irregular, and context-limited visual evidence. We study fresco-fragment style classification using a progressive transformer-based framework. Starting from a ViT-B/16 baseline, we introduce foreground-guided masking to suppress background-only tokens, inpainting-based geometric regularisation to align irregular fragment supports with the ViT patch grid, and a supervised contrastive objective that operates on predictive distributions through a Kullback-Leibler similarity and consistently improves every branch. We combine the branches with a deliberately simple learnable logit ensemble. Experiments on CLEOPATRA and POMPAAF show that fragment-aware modelling improves over the standard ViT baseline, with the ensemble increasing accuracy from 0.604 to 0.656 and macro-F1 from 0.596 to 0.648 on CLEOPATRA, and outperforming the best single branch in four of six fragmentation settings on POMPAAF. We additionally evaluate a more complex graph-fusion variant and find that it matches the simple ensemble on POMPAAF while offering only a small, dataset-specific gain on CLEOPATRA, which does not justify its added complexity. Beyond these empirical gains, our contribution is twofold: a distribution-level contrastive objective that consistently sharpens single-branch recognition, and an interpretability analysis that verifies the models exploit genuine painted evidence, while quantifying that the inpainting-based branch draws part of its attribution from the synthesised surround.

84. 【2609.21000】Do Spinning Radar Doppler Velocity Measurements Improve Vehicle Detection and Tracking?

链接https://arxiv.org/abs/2609.21000

作者:Eric Xie,Daniil Lisus,Timothy D. Barfoot

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)

关键词:adverse weather conditions, Spinning frequency-modulated continuous-wave, autonomous vehicle perception, field of view, Doppler velocity

备注: 8 pages, 8 figures

点击查看摘要

Abstract:Spinning frequency-modulated continuous-wave (FMCW) radars have been gaining popularity in autonomous vehicle perception on account of their robustness to adverse weather conditions and 360° field of view. Recently, scanning radars have also been shown capable of generating per-azimuth Doppler velocity. In this paper, we investigate whether these Doppler velocity measurements improve spinning radar vehicle detection and tracking performance. For detection, we estimate the ego motion and use it to undo the Doppler range distortion of the radar image before passing it to a network. For tracking, we propose a new way to estimate a per-vehicle velocity and use it as a prior for the tracker's motion model. Since Doppler-enabled spinning radar data is not available in any dataset with ground-truth dynamic object labels, our first contribution is an automatic labelling pipeline that uses an ensemble of fine-tuned off-the-shelf lidar detectors to label all 643 km of the Boreas Road Trip dataset. We then transfer detections to radar, and use over 250 km of vehicle-dense sequences as ground-truth training data. By training and evaluating two state-of-the-art detectors, we show that Doppler undistortion can improve detection accuracy by up to $2.37$ points on mean average precision. Furthermore, we show that the Doppler velocity prior can improve tracking accuracy by $13.68$ points on multi-object tracking accuracy (MOTA) versus the zero-velocity initialization baseline, while achieving $99.7\%$ of the MOTA obtained using ground-truth velocities as the prior.

85. 【2609.20975】Image-Derived PM10 Estimation in Cattle Feedlot Using Machine Learning: Addressing Concentration Ranges Beyond Existing Digital Imaging Methods

链接https://arxiv.org/abs/2609.20975

作者:Sirapoom Peanusaha,Greg B. Ferguson,K. Jack Bush,Peiyang Li,Brent W. Auvermann

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Affordable dust monitoring, air quality settings, intensive livestock operations, cattle feedlot industry, urban air quality

备注

点击查看摘要

Abstract:Affordable dust monitoring remains a pressing need for the cattle feedlot industry, yet camera-based PM estimation, despite its growing body of research in urban air quality settings, has not been evaluated under the extended concentration ranges characteristic of intensive livestock operations. This study developed an image-based approach using contrast panel features and machine learning to estimate PM10 concentrations in a commercial cattle feedlot, where hourly average PM10 ranged from 250 to 1,000 ug/m^-3 and instantaneous concentrations reached 5,000 to 20,000 ug/m^-3. Grayscale images were captured during the evening dust peak period, and features including panel contrast, black and white panel pixel values, and overall image brightness were extracted. The model also incorporated recent past values from preceding images and solar zenith angle as predictors. Among the candidate models evaluated, XGBoost achieved the highest predictive performance, with an R^2 of 0.792 and a median absolute error of 103 ug/m^-3. Feature importance analysis revealed that (a) panels positioned farthest from the camera contributed most strongly to predictions and (b) that black panel pixel values were more sensitive than white panel values to changes in PM10 concentration. Prediction accuracy during the sunset transition, which coincides with the onset of the feedlot evening dust peak, remains an area for further refinement. These findings demonstrate the feasibility of image-based PM10 estimation across PM concentration ranges substantially exceeding those reported in prior urban studies and provide practical guidelines for future deployment in feedlot environments.

86. 【2609.20962】MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction

链接https://arxiv.org/abs/2609.20962

作者:Akshit Sharma,Prashant W. Patil

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:significant societal threat, harmful internet memes, internet memes poses, automated classification remains, algorithmic challenge due

备注: 10 pages, 3 figures; published in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

点击查看摘要

Abstract:The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods.

87. 【2609.20892】WM-VS: Progress-Aligned World Models for Closed-Loop Visual Servoing

链接https://arxiv.org/abs/2609.20892

作者:Guanzhong Sun,Junyi Ma,Yixuan Zhou,Yuxuan Wu,Yanzi Miao,Hesheng Wang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Closed-loop visual servoing, servoing requires predictions, visual servoing requires, Closed-loop visual, visual servoing

备注

点击查看摘要

Abstract:Closed-loop visual servoing requires predictions that indicate whether an action reduces task error, not only whether the action is plausible. We call this gap the prediction-control mismatch and introduce WM-VS, a target-centric progress-aligned world-model framework for closed-loop visual servoing. Offline target-region DINOv2 correspondences define a signed four-dimensional servo coordinate for translation, scale, and in-plane rotation. Stage 1 aligns action-conditioned latent transitions with this coordinate; Stage 2 freezes the world model and trains a reactive joint-velocity policy with action imitation, consequence supervision, and short imagined rollouts that favor error contraction. Deployment is RGB-only and reactive, without online trajectory optimization. On a real 7-DoF eye-to-hand system, WM-VS reaches a corner RMSE no larger than 10 percent of its initial value in 30/30 trials and retains this criterion at the final valid frame in 25/30 (83.33 percent). Removing future-error alignment reduces retention to 26.67 percent. The learned progress signal agrees with an external AprilTag corner error not used for training or control (mean Spearman rho = 0.8778). Without retraining, two unseen 3D targets achieve translation-error reductions of 86.48 percent and 90.27 percent and rotation-error reductions of 70.01 percent and 65.70 percent. These results link progress-aligned action consequences to repeated closed-loop correction and transfer. Code and data will be released as open source.

88. 【2609.20869】APe+ML: A Compact Structured Representation for Multi-Task Computer Vision

链接https://arxiv.org/abs/2609.20869

作者:Sergey Kurinov(1),Alexey Upatov(1) ((1) Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia)

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:Theory of Active, Active Perception, vision system based, encodes relations, relations among perceptual

备注: 39 pages, 4 figures, 11 tables. Project page: [this https URL](https://ml.comexp.net)

点击查看摘要

Abstract:We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation. TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

Comments:
39 pages, 4 figures, 11 tables. Project page: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

Cite as:
arXiv:2609.20869 [cs.CV]

(or
arXiv:2609.20869v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.20869

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
89. 【2609.20850】MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

链接https://arxiv.org/abs/2609.20850

作者:Yueming Lyu,Yilian Shi,Haoxiang Tan,Linzhuang Zou,Qihao Wang,Guihua Yu,Jie Qin,Xin Gao,Chenyang Si,Jing Dong,Caifeng Shan

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Multimodal Large Language, Language Models, show remarkable advancements, Large Language

备注

点击查看摘要

Abstract:While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.

90. 【2609.20839】Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition

链接https://arxiv.org/abs/2609.20839

作者:Matthew Kit Khinn Teng,Haibo Zhang,Takeshi Saitoh

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)

关键词:visual speech recognition, phoneme prediction errors, speech recognition, realistic phoneme prediction, visual speech

备注: Submitted for journal publication and currently under consideration

点击查看摘要

Abstract:Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.

91. 【2609.20826】ALON: A Temporally Aware Longitudinal Framework for Radiology Report Generation

链接https://arxiv.org/abs/2609.20826

作者:Nien-Tsyr Sun,Min-Chen Chen,Hui Nien Hung,Vincent S. Tseng

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:radiology report generation, produce descriptive reports, descriptive reports based, Current radiology report, detect subtle interval

备注

点击查看摘要

Abstract:Current radiology report generation (RRG) models usually produce descriptive reports based on a single examination or only the most recent prior examination, limiting their ability to perform accurate and meaningful longitudinal comparisons and detect subtle interval changes. Although recent approaches have begun to incorporate multiple prior examinations, they usually aggregate a fixed-length history without explicitly modeling the role-dependent relevance of each prior examination before fusion. To address this, we propose TALON, a Temporally Aware LONgitudinal RRG framework that adaptively integrates variable-length patient histories. The underlying Dual-Channel Temporal Fusion Module (DCTFM) compares the current examination with each prior examination through complementary similarity and change channels to capture persistent findings and interval changes, respectively. The specially designed channel-specific attention estimates the relevance of each prior examination, while a learned prior-specific gate adaptively integrates informative longitudinal evidence and suppresses redundancy. Experiments on MIMIC-CXR show that TALON outperforms the current state-of-the-art method on various clinical efficacy and graph-based metrics. When more prior examinations become available, TALON's performance on these metrics improves even further, emphasizing the strength of TALON's DCTFM in modeling longitudinal RRG across longer and more complex patient histories than existing approaches.

92. 【2609.20106】AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention

链接https://arxiv.org/abs/2609.20106

作者:Yuang Tu,Runjia Tan,Yujie Yan,Jinghan Hu,Chen Lv

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Robotic reward models, evaluate task execution, underlying task state, reward models evaluate, Robotic reward

备注: 8 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.

93. 【2609.19122】raining-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

链接https://arxiv.org/abs/2609.19122

作者:Meng'en Qin,Yinchen Liu,Mingxuan Cui,Youlu Xing

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:robust downstream prediction, Visual signals require, robust visual signal, Convolutional sparse coding, signals require compact

备注

点击查看摘要

Abstract:Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient is typically fixed and manually selected. We propose a training-adaptive convolutional sparse coding framework for robust visual signal representation. Specifically, we unfold the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and treat the sparsity coefficient as a differentiable variable jointly learned with the network parameters. From the information bottleneck perspective, this coefficient controls the trade-off between information retention and compression: the sparsity term promotes compact representations, while the reconstruction term together with task loss preserves task-relevant signal content. We further introduce a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed. Experiments on CIFAR and ImageNet demonstrate competitive clean-data recognition and greatly improved robustness under different input perturbations.

94. 【2609.21813】How Many Posterior Samples? Calibrated Stopping for Adaptive Sensing

链接https://arxiv.org/abs/2609.21813

作者:Vincent Corlay,Andriy Enttsel

类目:Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)

关键词:class labels provide, classification-oriented adaptive sensing, labels provide votes, samples characterize uncertainty, current measurement state

备注

点击查看摘要

Abstract:In classification-oriented adaptive sensing, posterior samples characterize uncertainty at the current measurement state and can serve two roles: they may guide the next sensing direction, while their class labels provide votes for the candidate classes and determine whether sensing should continue. We focus on the stopping layer that turns these votes into a declaration, without modifying the posterior sampler or sensing directions. A natural plug-in rule declares when the observed vote share exceeds a threshold. We show that this threshold is not itself a confidence guarantee: when the underlying vote mass equals the threshold, the plug-in rule declares about half the time. As alternatives, we calibrate a fixed-sample rule and a finite-horizon sequential rule to a prescribed false-declaration probability, and study exact curtailment, which stops a fixed-pool rule once its final verdict is forced. We then derive how one-round declaration probabilities determine posterior-sample cost and classification accuracy along a sensing path. On MNIST with DDRM and a fixed PCA-guided probe sequence, curtailment saves up to 62% of posterior samples. Among the evaluated rules at matched operating points, sequential stopping reduces the cost the most. At a high accuracy, that same sequential rule can trade more posterior samples for fewer measurements.

95. 【2609.21812】Classification-oriented adaptive sensing via posterior sampling

链接https://arxiv.org/abs/2609.21812

作者:Andriy Enttsel,Maxime Rousselot,Vincent Corlay

类目:ignal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV)

关键词:task-specific policy training, Recent advances, instance-adaptive compressed sensing, enabled high-performance, instance-adaptive compressed

备注

点击查看摘要

Abstract:Recent advances in diffusion models have enabled high-performance, instance-adaptive compressed sensing through posterior sampling, without task-specific policy training. Existing methods select sensing probes by maximizing total posterior signal variance and are therefore primarily reconstruction-driven. We introduce a classification-driven extension motivated by the closed-form posterior covariance of a class-conditional Gaussian mixture model, which decomposes into within-class and between-class uncertainty. Using calibrated soft classifier outputs, we estimate these uncertainty terms from diffusion posterior samples and propose a classification-oriented criterion for selecting the dominant sensing direction in the unmeasured subspace. Experiments on MNIST and CIFAR-10 compare the resulting classification accuracy, measurement cost, and reconstruction quality with those of reconstruction-oriented counterparts. The results identify regimes in which semantic posterior uncertainty yields a more favorable classification--measurement trade-off and quantify the associated reconstruction cost.

96. 【2609.21391】WS-NeRF: A Mamba-Driven World-State-Aware Adaptive Deblurring Neural Radiance Field

链接https://arxiv.org/abs/2609.21391

作者:Hang Jiang,Jinghao Wang,Yiming Zhang,Xinhong Wang,Luwei Ran,Yinfeng Yu

类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:attracted extensive attention, recent years due, Neural Radiance Fields, neural radiance field, attracted extensive

备注: Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2026 (IEEE SMC 2026)

点击查看摘要

Abstract:Neural Radiance Fields (NeRF) have attracted extensive attention in recent years due to their strong capability for high-quality 3D reconstruction and novel view synthesis from multi-view images. Existing methods usually rely on high-quality sharp inputs, while real-world image acquisition is highly susceptible to blur degradation, which severely affects the reconstruction quality of NeRF. In this paper, we propose a novel Mamba-driven world-state-aware adaptive deblurring neural radiance field, termed WS-NeRF, to address image degradation and 3D inconsistency. We formulate the alternating optimization of radiance fields as a dynamic evolution process with temporal memory, and jointly exploit comprehensive multi-dimensional world states and a mixture-of-experts mechanism to dynamically adjust the confidence of deblurring priors. Experimental results show that WS-NeRF significantly improves blurry radiance field reconstruction quality, achieving better performance on PSNR, SSIM, and LPIPS, while exhibiting more stable iterative recovery behavior.

97. 【2609.21169】Adaptive Color Grading

链接https://arxiv.org/abs/2609.21169

作者:Trevor D. Canham,Abhijith Punnappurath,Michael S. Brown

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Independent control, essential for painters, photographers and cinematographers, cinematographers to bring, Independent

备注: Accepted @ 34th Color Imaging Conference

点击查看摘要

Abstract:Independent control of tonescale regions (e.g., shadows, highlights) is essential for painters, photographers and cinematographers to bring 2D images to life. In image manipulation software this is most directly addressed by color grading modules, which use intensity thresholds to segment distinct illumination regions for local manipulation. In this work we develop an open source color grading tool and use it to annotate a large dataset of video frames with tonescale region thresholds. Using these thresholds we conduct modeling experiments with strategies based on both practitioners' conventional wisdom and machine learning. Results show that K-nearest neighbors is an effective prediction strategy, outperforming state-of-the-art end-to-end methods for image enhancement. This outcome demonstrates the benefit of focusing on a compact set of core parameters when modeling creative stylization processes. Our adaptive color grading interface and data are available at this https URL.

98. 【2609.20905】Uncertainty-driven training for three-dimensional calibrated lung nodule classification

链接https://arxiv.org/abs/2609.20905

作者:Giuseppe Tripodi,Alessandro De Rosis,Saleh Rezaeiravesh

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Monte Carlo Dropout, Evidential Deep Learning, lung nodule classification, guide loss reweighting, three-dimensional computed tomography

备注

点击查看摘要

Abstract:In this work, we present an uncertainty-driven training framework for three-dimensional computed tomography (CT) lung nodule classification, where validation-based uncertainty estimates guide loss reweighting to enhance predictive performance and probability calibration. Two Uncertainty Quantification (UQ) methods are considered: Monte Carlo Dropout (MCD) and Evidential Deep Learning (EDL). Both provide per-class uncertainty estimates that modulate the loss and encourage focus on hard or unreliable classes. The framework is evaluated with ResNet, DenseNet, EfficientNet, Vision Transformer (ViT), and Swin Transformer backbones on two datasets: the clinical LIDC-IDRI cohort and the NoduleMNIST3D benchmark. Uncertainty-driven training achieves classification performance similar to conventional training while substantially improving calibration, with an expected calibration error (ECE) reduced by up to 65% on LIDC-IDRI. EDL attains competitive performance on shallower architectures with single-pass inference, whereas MCD is more robust on deeper networks. Analysis across architectural families reveals that uncertainty-driven training benefits convolutional backbones more consistently than transformer-based architectures: EDL in particular degrades on ViT, suggesting that the Dirichlet evidence parameterisation may interact unfavourably with attention-based architectures at lower input resolutions. A posteriori temperature scaling proves highly effective across all configurations, indicating that a simple scalar calibration can be competitive even without explicit uncertainty-aware training. Our results indicate that integrating UQ into the training loop can significantly improve probabilistic calibration and support more trustworthy deployment of three-dimensional medical imaging models.

99. 【2605.15418】A Differentiable Ray-Wave Framework for Hybrid Refractive-Diffractive System Modeling and Optimization

链接https://arxiv.org/abs/2605.15418

作者:Jiazhou Cheng,Margaret Gao,Yixuan Shao,Chenkai Mao,Tom D. Milster,Jonathan A. Fan

类目:Optics (physics.optics); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Signal Processing (eess.SP); Computational Physics (physics.comp-ph)

关键词:systems combining refractive, Hybrid optical systems, disparate spatial scales, optical systems combining, wave phenomena

备注: 9 pages, 7 figures

点击查看摘要

Abstract:Hybrid optical systems combining refractive and diffractive optical responses have the potential to support new types of optical behavior, but they are difficult to model and optimize due to the disparate spatial scales and physics exhibited by ray and wave phenomena. In this work, we present a differentiable ray-wave framework for modeling hybrid refractive-diffractive optical systems that operates as a plug-and-play module within standard ray tracing pipelines. Our model uniquely applies to both planar and curvilinear diffractive surfaces and accommodates arbitrary scalar holographic profiles with high spatial frequency responses, with each simulation evaluated at a single wavelength. We analyze ray-wave modeling regimes that optimally account for the spatial frequency properties and spatial curvature of the diffractive surfaces, and we demonstrate the gradient-based end-to-end optimization of hybrid refractive-diffractive systems featuring planar and conformal diffractive surfaces. We anticipate that these modeling capabilities will enable new classes of hybrid optical systems relevant to computational imaging and display applications.

100. 【2310.03860】MultiHU-TD: Multifeature Hyperspectral Unmixing Based on Tensor Decomposition

链接https://arxiv.org/abs/2310.03860

作者:Mohamad Jouni,Mauro Dalla Mura,Lucas Drumetz,Pierre Comon

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)

关键词:representing mixed pixels, pure materials weighted, representing mixed, mixed pixels, set of pure

备注

点击查看摘要

Abstract:Hyperspectral unmixing allows representing mixed pixels as a set of pure materials weighted by their abundances. Spectral features alone are often insufficient, so it is common to rely on other features of the scene. Matrix models become insufficient when the hyperspectral image (HSI) is represented as a high-order tensor with additional features in a multimodal, multifeature framework. Tensor models such as canonical polyadic decomposition allow for this kind of unmixing but lack a general framework and interpretability of the results. In this article, we propose an interpretable methodological framework for low-rank multifeature hyperspectral unmixing based on tensor decomposition (MultiHU-TD) that incorporates the abundance sum-to-one constraint in the alternating optimization alternating direction method of multipliers (ADMM) algorithm and provide in-depth mathematical, physical, and graphical interpretation and connections with the extended linear mixing model. As additional features, we propose to incorporate mathematical morphology and reframe a previous work on neighborhood patches within MultiHU-TD. Experiments on real HSIs showcase the interpretability of the model and the analysis of the results. Python and MATLAB implementations are made available on GitHub.