本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。
统计
今日共更新742篇论文,其中:
- 自然语言处理89篇
- 信息检索19篇
- 计算机视觉114篇
自然语言处理
1. 【2609.31619】Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
链接:https://arxiv.org/abs/2609.31619
作者:Parsa Hosseini,Akasha Tigalappanavara,Sumit Nawathe,Chenrui Fan,Sourya Basu,Genta Indra Winata,Anirban Das,Soheil Feizi,Nima Chitsazan
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:long reasoning traces, making inference computationally, inference computationally expensive, Reasoning, computationally expensive
备注:
点击查看摘要
Abstract:Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.
2. 【2609.31603】User Model Extraction via Belief Self-Distillation
链接:https://arxiv.org/abs/2609.31603
作者:Ali Holmov,Yiran Huang,Kirill Bykov,Zeynep Akata
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:implicitly infer attributes, Large language models, Large language, beliefs remain difficult, implicitly infer
备注:
点击查看摘要
Abstract:Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.
3. 【2609.31587】Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
链接:https://arxiv.org/abs/2609.31587
作者:Md Shohel Arman,Igor Molybog
类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:coding agents resolve, agents resolve software, investigate whether natural-language, build the tools, tools to construct
备注: 13 pages. Code and data: [this https URL](https://github.com/haw-ai-i/roundtrip)
点击查看摘要
Abstract:We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.
4. 【2609.31571】Strategically Diverse Sampling for Self-Training
链接:https://arxiv.org/abs/2609.31571
作者:Alexander Gurung,Esmeralda S. Whitammer,Mirella Lapata
类目:Computation and Language (cs.CL)
关键词:responses meaningfully differ, LLM training, sampling IID responses, depend on repeated, meaningfully differ
备注:
点击查看摘要
Abstract:Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.
5. 【2609.31553】MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
链接:https://arxiv.org/abs/2609.31553
作者:Itzel Tlelo-Coyotecatl,Hugo Jair Escalante
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:raised Hate Speech, Hate Speech Detection, Ensuring online safety, Hate Speech, Ensuring online
备注: Preprint submitted to CIARP 2026
点击查看摘要
Abstract:Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.
6. 【2609.31511】Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge
链接:https://arxiv.org/abs/2609.31511
作者:Yahya Mohamed Elnawasany
类目:Computation and Language (cs.CL)
关键词:sourced Islamic knowledge, Standard Arabic TTS, Arabic TTS Arena, Arabic Islamic model, Arabic TTS model
备注: 6 pages, 4 tables. Deployed system: [this https URL](https://muslim.yahyaelnawasany.com) - released models: [this https URL](https://huggingface.co/NightPrince)
点击查看摘要
Abstract:We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retrieval layer routed across six Model Context Protocol servers, we report three things a research prototype typically lacks. First, a released family of fine-tuned Arabic Islamic model artifacts: an efficient tool-routing LLM (Muslim-6B-PRO, 5.94B parameters) and a Modern Standard Arabic TTS model (Fasih-TTS-V1) that ranks 5th of 17 overall and 2nd of 11 open-weight systems on the community-voted Arabic TTS Arena for MSA. Second, an account and metering layer - a free per-account turn allowance, capacity-aware refusal, and email verification deferred to the point it actually matters - that turns an open demo into an operable, abuse-resistant product. Third, a three-layer observability stack (liveness, error reporting, product analytics) built specifically around the system's characteristic failure mode: a GPU-bound agent host going silent while the web tier keeps serving normally. We report real, measured latency and accuracy figures (98.4% recitation-validation accuracy on 124 cases; end-to-end voice latency of 0.9-1.7s) and discuss the concrete engineering trade-offs and limitations of running an Islamic-knowledge voice product in production.
7. 【2609.31506】Evaluating Cultural Awareness of LLMs for Haitian Creole
链接:https://arxiv.org/abs/2609.31506
作者:Christelle Clervilsson,Yanzhu Guo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large language models, exhibit substantial performance, substantial performance disparities, Large language, exhibit substantial
备注:
点击查看摘要
Abstract:Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring portrayals of Haitian characters through hardship and resilience, showing that even positive characterizations can encode stereotypical narratives. Our code, benchmark, and evaluation framework are publicly available.
8. 【2609.31448】ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs
链接:https://arxiv.org/abs/2609.31448
作者:Junyi Gao,Yu Shi,Pingzhao Hu,Ewen M Harrison
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:models support medical, support medical text, medical text understanding, support medical, clinical time series
备注:
点击查看摘要
Abstract:Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clinical time series. Improving this ability would connect risk estimation with flexible questions about a patient's evolving condition. We introduce ViSTA, a compact adapter that incorporates irregular numerical measurements into a pretrained vision-language model's chart representations. It learns corrections to visual tokens while leaving all pretrained parameters unchanged. On MIMIC-IV, ViSTA has the highest mean scores among the compared adaptations on all four metrics for acute kidney injury and mortality prediction across models with 2-9 billion parameters. With 0.516 million trainable parameters, the 2-billion-parameter model reaches an area under the ROC curve of 0.7376 for acute kidney injury, compared with GPT-5.6 Sol's 0.7380 with text input and high reasoning effort. Training for temporal question answering yields 69.27% accuracy at 4 billion parameters with over 90% fewer trainable parameters than low-rank adaptation using charts or numerical text, at a 2.82-4.88 percentage-point accuracy gap. ViSTA extends pretrained language models to numerical prediction and temporal questions.
9. 【2609.31422】owards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis
链接:https://arxiv.org/abs/2609.31422
作者:Jakub Masłowski,Jarosław A. Chudziak
类目:Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large language model-based, language model-based multi-agent, complex decision pipelines, model-based multi-agent debate, Large language
备注: Accepted for publication at the 38th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026)
点击查看摘要
Abstract:Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are prone to fabricating smoothly written debate consensus that is not grounded in the debate's history. To address this safety gap, this paper presents empirical research and studies if the introduction of active post-debate verification can mitigate the production of such factually unsupported summaries, while still providing valuable information. Furthermore, it is examined whether explicitly signalling divergence is preferable in the absence of a reliable compromise. The Active Provenance Gate (APG) is introduced as a post-debate verification layer that treats the source as a hard constraint, analysing the debate logs, auditing each claim, and applying self-correction. In crisis simulations, the self-healing mechanism more than doubles the average data Provenance Fidelity in difficult condition scenarios, before the strict gate blocks unsupported claims and generates divergence reports. In the human study, a vast majority of the users (over 75%) preferred a report explicitly stating failure in critical scenarios, despite most of them perceiving fabricated consensus from the baseline system as more fluent. Our main contribution is the transition of data origin tracing from passive logging to active conditional blocking before publication.
10. 【2609.31403】Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
链接:https://arxiv.org/abs/2609.31403
作者:Mert İncidelen,Yamen Kashkash,Asya Berker,Murat Aydoğan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:optical character recognition, Decoy Font method, Vision-language models, character recognition, success in optical
备注: Accepted to the First Workshop on Document Intelligence and Understanding (DocInsights 2026), co-located with the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
点击查看摘要
Abstract:Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consists of 300 images, each containing text with sharp contour lines superimposed on another text with soft shading. Six recent closed-source models from three different model families were evaluated using this dataset under two different prompting conditions (naive and guided) and at two different resolutions ($512\times512$ and $64\times64$). A validation study showed that human participants could read both text layers with high accuracy. In contrast, the models, with most variants and both prompting methods, read the contour text with near-human accuracy at high resolution, but almost never fully extracted the shading text. At low resolution, the contour text could not be read by either the models or humans, while the shading text could be extracted with high accuracy. The findings indicate that the evaluated VLMs exhibit a consistent behavioral limitation when processing typographic structures containing multiple spatial frequency layers.
11. 【2609.31397】Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models
链接:https://arxiv.org/abs/2609.31397
作者:Andrea Masini,Sudipta Acharya,Paolo Bellavista,Luca Foschini,Burak Kantarci
类目:Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:enforcement requires translating, deployable traffic-management policies, requires translating high-level, translating high-level service, high-level service intents
备注: 6 pages, 6 figures, Accepted to IEEE Conference on Future Communications and Networks (FCN) 2026
点击查看摘要
Abstract:Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy specification, bridging the gap between business-level intents and executable network configurations remains complex, error-prone, and difficult to automate. This paper presents Intent2Tc, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations. The framework integrates an Active Queue Management (AQM)-based digital twin (DT) semantic model, automated metadata extraction, critique-driven refinement, and Retrieval-Augmented Generation (RAG)-based knowledge reuse to improve semantic consistency and configuration reliability. We evaluate multiple open-source large language models (LLMs) and small language models (SLMs), together with Claude Sonnet-4.6, on 100 Request for Comments (RFC) 9315-compliant traffic-shaping intents. Across both translation stages, Intent2Tc achieves high semantic fidelity, configuration accuracy, and deployment readiness, with Claude Sonnet-4.6 reaching 0.98 semantic similarity, 1.0 semantic unit coverage, and 0.045 normalized edit distance. Furthermore, RAG reduces token consumption and inference latency while enabling compact models such as Phi-4-mini to approach the performance of substantially larger models. Linux tc serves as the target configuration platform, demonstrating the practical applicability of the proposed framework.
12. 【2609.31382】Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding
链接:https://arxiv.org/abs/2609.31382
作者:Zhaoyuan Xia(1 and 2),Qinghongbing Xie(3),Yung Xiang Hue(3),Jianguang Jiang(2),Gaofeng Lu(2),Zhenyu Jiao(2),Xing Yuan(2),Dai Dai(2),Tong Mo(1),Long Zeng(3) ((1) Peking University, (2) Baidu Inc., (3) Tsinghua University)
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:understanding requires large, requires large language, scattered amid substantial, amid substantial irrelevant, large language models
备注: 23 pages, 13 figures. Zhaoyuan Xia and Qinghongbing Xie contributed equally. Corresponding authors: Dai Dai, Tong Mo, and Long Zeng. Code and data are available at [this https URL](https://github.com/X-Luffy/Highlight-Then-Summarize)
点击查看摘要
Abstract:Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first identifies source-grounded, question-relevant evidence and then integrates it into a compact, question-conditioned summary before producing the final answer. To train this behavior, we construct H2S-Dataset, comprising 6,647 examples from 11 benchmark families with an average context length of 43.9K tokens, and introduce H2S-RL, which provides process-level rewards for evidence selection and summary construction in addition to final-answer correctness. We evaluate on H2S-Bench, a seven-task long-context suite. Under a shared 128K input and 4K output budget, H2S-14B achieves an average score of 32.60, outperforming Qwen3.8-27B by 10.17 points and obtaining the strongest overall result among the evaluated open-source models. H2S-14B also achieves the highest Evidence-Summary Quality score and retains 97.1% of its 16K-budget performance with only a 4K output budget. These results show that explicitly selecting and integrating evidence improves long-context reasoning while enabling more compact generation.
13. 【2609.31342】Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers
链接:https://arxiv.org/abs/2609.31342
作者:Md Shamim Ahmed,Lukas Galke Poech,Richard Röttger
类目:Computation and Language (cs.CL)
关键词:Retrieval-augmented generation, providing external evidence, providing external, Retrieval-augmented, address outdated knowledge
备注: 17 pages, 3 figures
点击查看摘要
Abstract:Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated evidence makes a model wrong despite answering correctly without retrieval. We construct a benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy, grounded in dated official sources. Across 12 models, recent medical reversals are harder than long-established ones. More importantly, outdated retrieval flips 30% of Llama and 37% of Qwen answers even without instructions to trust the document; explicit follow instructions raise these rates to 66% and 75%. Across four open models and four domains, poisoning ranges from 17-91%, while matched up-to-date evidence is followed in 97-100% of trials. To isolate temporal applicability, we keep the historical evidence unchanged across 50 reversals and vary only the evaluation date. A clear pattern emerges: dates alone produce only modest adaptation, but when models are explicitly told when the old evidence stops applying, the larger models switch to the appropriate answer almost perfectly. Causal interventions confirm that this validity information directly shapes the final decision. The same internal components also support broader comparison tasks, suggesting that temporal applicability can recruit a general reasoning mechanism used for other comparisons. Finally, a fixed recency-aware hybrid re-ranker reduces poisoning by 4.6-10.0 points when dates are accurate, with gains that depend on reliable temporal metadata. Reliable RAG therefore requires selective trust: models must determine not only what retrieved evidence says, but whether it still applies.
14. 【2609.31341】he Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
链接:https://arxiv.org/abs/2609.31341
作者:Christoph Walser,Mauricio Fadel Argerich,Jonathan Fürst
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:process page images, parsed text depends, layout spectrum, pipeline should process, images or parsed
备注: Accepted to DocInsights at EMNLP 2026
点击查看摘要
Abstract:Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find that batching is the dominant energy lever, cutting energy per page by 38-85% at no cost in accuracy, while FP8 quantization saves 27-32% when requests are served one at a time but less than 1mWh per page (9-19%) once batching is applied. Preprocessing dominates what remains: neural OCR costs $17\times$ more energy per page than classical OCR and never reaches the Pareto frontier. Which representation wins flips with the type of document: vision--language models on layout-rich documents and small text-only models with a cheap parser on near-plain text, where they are both more accurate and cheaper than any vision--language configuration. Our work yields concrete guidelines for energy-efficient, privacy-compliant local information extraction.
15. 【2609.31264】Identifying Scientists on X
链接:https://arxiv.org/abs/2609.31264
作者:Philipp Meier,Katarina Boland,Laura Kallmeyer,Stefan Dietze
类目:Computation and Language (cs.CL)
关键词:classical knowledge order, knowledge order, growing importance, importance of science-related, science-related discourse
备注: Corrected version of Identifying Scientists on X published at Companion Publication of the 18th ACM Web Science Conference 2026
点击查看摘要
Abstract:With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forests with linguistic features and up to 0.96 using a contrastively fine-tuned DeBERTa model in an ensemble setup. Furthermore, we provide two datasets with X users labeled as scientists or non scientists and their respective tweets and user biographies.
16. 【2609.31261】MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
链接:https://arxiv.org/abs/2609.31261
作者:Michele Paolicelli,Alessandro Petruzzelli,Alessandro Franceso Maria Martina,Cataldo Musto,Giovanni Semeraro
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:long-context language modeling, language modeling, quadratic complexity, central bottleneck, bottleneck for long-context
备注:
点击查看摘要
Abstract:The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where positional relevance can decay and where broader interactions must be preserved. We introduce Mixture of Semantic Attention Regimes (MoSAR), which learns such an adaptive, controlled-decay geometry over query--key interactions. Input-conditioned query and key routers, applied after positional encoding, select mixtures over short, medium, and global regimes, inducing a continuous distance-dependent attention field rather than a fixed sparsity pattern. This geometry is learned during training and can subsequently be discretized through top-1 routing. In controlled pre-training experiments with matched 500M-parameter models, MoSAR learns a substantially lower-reach attention geometry without degrading language-modeling quality, improving perplexity over dense RoPE at the training context length. Under length extrapolation, MoSAR achieves the best perplexity among all evaluated variants, including strong baselines such as ALiBi. Moreover, the learned geometry remains stable under deterministic top-1 discretization, suggesting that it is not only adaptive, but also amenable to low-cost approximation at inference time.
17. 【2609.31255】PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding
链接:https://arxiv.org/abs/2609.31255
作者:Jeonghun Yoon,Dongchan Kim,Hongyeon Yu,Young-Bum Kim,Jaegul Choo
类目:Computation and Language (cs.CL)
关键词:extracts salient snippets, General-purpose agent memory, General-purpose agent, agent memory summarizes, salient snippets
备注: 13 pages, 6 figures, 8 tables
点击查看摘要
Abstract:General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretion, and a three-month glucose trend cannot be answered by text similarity. We present PIA, a personal intelligence agent deployed alongside a consumer health agent. PIA receives the agent's natural-language requests, decides for itself whether and how to write or read, and turns conversations into typed clinical records and records into a synthesized understanding of the user. Its memory harness consists of four controls -- extraction, memory, retrieval, and understanding -- each a domain-agnostic mechanism with a pluggable health module: schema, medical alias dictionary, knowledge graph, and temporal rules. We show how the same query receives a different answer as the memory injected into the response context deepens from one-dimensional recall, to a two-dimensional health snapshot, to a three-dimensional trajectory with causality, and report lessons from operation: self-reported health data are missing not at random, question phrasing governs the quality of synthesized understanding, and nearly a third of candidate causal links are structural noise that rules alone remove.
18. 【2609.31245】RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models
链接:https://arxiv.org/abs/2609.31245
作者:Pavithra P M Nair,Bhavik Talaviya,Shourya Bhushan,Rahul Pankajakshan,Seema Guruvadoo,Avinash Agarwal,Gilad Gressel,Krishnashree Achuthan
类目:Computation and Language (cs.CL)
关键词:large language models, comparing loan options, Individuals turn, language models, economic
备注:
点击查看摘要
Abstract:Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.
19. 【2609.31181】Where a Model Sends Its Own Repeated Token
链接:https://arxiv.org/abs/2609.31181
作者:Nicolás Vera Zúñiga
类目:Computation and Language (cs.CL)
关键词:Black-box model identification, model identification works, Black-box model, natural-language prompts, response to natural-language
备注: 8 pages, 3 tables. Companion to [arXiv:2608.10986](https://arxiv.org/abs/2608.10986) , [arXiv:2608.21315](https://arxiv.org/abs/2608.21315) and [arXiv:2609.29507](https://arxiv.org/abs/2609.29507) . Code, per-run results, pre-registrations and the findings ledger: [this https URL](https://github.com/nicoveraz/token-lattice-ca) (archived: [this https URL](https://doi.org/10.5281/zenodo.21880472) )
点击查看摘要
Abstract:Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed points -- is partially anticipated, and we report it as a failed estimand: the natural distance on it is 83% cardinality, separates a corpus manipulation by two bits in 3471 against a precision floor of zero, and attributes families at 0.5833. The second half, where the map sends tokens that are not fixed points, is unrecorded; the one paper holding those tokens logged them as a zero. Pairing on the source token removes the cardinality confound by construction (r from 0.9128 to -0.0932) and attributes families at 0.8333 -- twelve models scored against a pool of nineteen -- with chance 0.1389, across seven tokenizer groups and several corpora. Two nulls clear it: frequency-matched destinations agree at 0.1429, independent marginals at 0.0798. Family predicts agreement better than tokenizer (0.2031 against 0.1205), and recurrent architectures cluster at balanced accuracy 1.0 against a 0.7895 majority rate, or 0.90 once each model's dominant destination is excluded -- the figure we stand behind. We measure the robustness envelope: 8-bit weight rounding moves the map less than deduplicating the training corpus does (0.9004 against 0.6353, on one support), 4-bit destroys it (0.0098; 0.1812 at deployment granularity, so not a coarseness artefact), and the precision floor varies by model from 0.201 to 0.9778. All estimands and kill conditions were registered before the data, and the failed one is reported as fully as the surviving one.
20. 【2609.31169】Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting
链接:https://arxiv.org/abs/2609.31169
作者:Paweł Mąka,Piotr Andruszkiewicz,Yusuf Can Semerci,Jan Scholtes,Gerasimos Spanakis
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:incorporate additional signal, Metric-based Loss Weighting, Machine Translation aims, Multimodal Machine Translation, resolving ambiguities
备注:
点击查看摘要
Abstract:Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore this information. Therefore, increasing their visual sensitivity remains an active research area. In this work, we introduce a training method, Metric-based Loss Weighting, that improves visual grounding of translations by increasing the loss function for tokens that benefit from the accompanying image. We identify these tokens using the Point-wise Cross-mutual Information (PCXMI) metric, which compares the model's output probabilities with and without visual context. We introduce a Congruency-based PCXMI metric and experimentally show that both metrics working in combination yield the best results. We evaluate our method by fine-tuning three pretrained Multimodal Large Language Models on the task of Image-guided Machine Translation for three language directions. Metric-based Loss Weighting outperforms other tested methods on the CoMMuTE contrastive dataset, improving accuracy by up to more than 7 percentage points compared to standard fine-tuning, while maintaining strong general translation performance.
21. 【2609.31142】JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
链接:https://arxiv.org/abs/2609.31142
作者:Jianyi Hu,Hangtao Zhang,Yi Liu,Yeqi Zeng,Li Zeng,Xianlong Wang,Rui Wang,Leo Yu Zhang
类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:typed model generates, RLCD models, trained with reinforcement, reinforcement learning, learning for calibrated
备注: 33 pages, 13 figures, 19 tables. Project website: [this https URL](https://JevAdvBench.github.io/JevAdvBench/)
点击查看摘要
Abstract:Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model's own clean decision rather than against labels, and to read it against the change caused by an identical re-run. Building on this, we introduce JevAdvBench, to our knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model. On jev-1.13.0, rewording stays within 1.2 percentage points of the re-run baseline, and fields outside the schema never reach the model. In contrast, one unverified opinion appended to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that routes them to human review. Applications built on RLCD models should therefore treat the state as untrusted, argued input. Project website: this https URL
22. 【2609.31130】Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue
链接:https://arxiv.org/abs/2609.31130
作者:Amandine Decker(LORIA, UL, CNRS, SEMAGRAMME, GU),Maxime Amblard(SEMAGRAMME, LORIA),Ellen Breitholtz(GU)
类目:Computation and Language (cs.CL)
关键词:Discussion based modelling, Discussion based, empirically investigate Question, potential questions predicts, British National Corpus
备注:
点击查看摘要
Abstract:We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we construct a dataset of 7,124 questions automatically generated from utterances and preceding context from the British National Corpus, and annotated for salience and answerability. We find a robust but low positive correlation between salience and answerability in dialogue, indicating that more salient questions are more likely to be addressed. However, this effect is markedly weaker than in monologic text, suggesting that conversational structure is less predictable. We further observe that structured interactions exhibit stronger alignment between annotators than less organised dialogues.
23. 【2609.31122】LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
链接:https://arxiv.org/abs/2609.31122
作者:Irene Tallini,Lorenzo Basile,Valentino Maiorca,Francesco Locatello,Alberto Cazzaniga
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:powerful training-free paradigm, controlling large language, large language models, powerful training-free, training-free paradigm
备注:
点击查看摘要
Abstract:Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation steering to the model's own output vocabulary subspace. By identifying a property-specific linear subspace within the unembedding matrix, LocUS enforces a geometric constraint that restricts the steering transformation to a specific subspace and at the same time localizes its application to a sparse subset of attention heads. Extensive evaluations across three model families on toxicity mitigation, sentiment redirection and sycophancy suppression show that LocUS matches or outperforms state-of-the-art baselines while intervening on under 6% of parameters and better preserving general capability.
24. 【2609.31062】CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
链接:https://arxiv.org/abs/2609.31062
作者:Marko Řeháček,Vítězslav Dušek,Martin Rusinko,Vít Nováček
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:assistants promise valuable, promise valuable support, Patient-facing AI assistants, pose medical risks, Medical Urgency
备注: Accepted as a short paper at CIKM '26 (35th ACM International Conference on Information and Knowledge Management), Rome, Italy. 7 pages, 1 figure, 2 tables. Code and benchmark: [this https URL](https://github.com/mrehacek/cg-probes)
点击查看摘要
Abstract:Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential linear direction. To train the probes, we cluster 79,658 Czech oncology search queries with BERTopic and use these clusters to generate pairs of queries with contrastive risk levels via few-shot prompting. We evaluate the approach on 200 queries (90 real, 110 synthetic), each graded by two oncologists, against two open-weight LLMs and a frontier LLM. We find that urgency-based axes are recoverable as linear directions, and the probes are competitive with open-weight LLMs (no significant differences in quadratic-weighted kappa) at a fraction of the latency. Each axis yields a scalar score that clinicians can inspect and use to set escalation thresholds. The pipeline requires only search logs, axis definitions, and black-box access to the embedding model, suggesting transferability across healthcare domains. Robust validation on new queries and axes remains future work.
25. 【2609.31046】Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference
链接:https://arxiv.org/abs/2609.31046
作者:Özge Alacam,Zübeyde Demet Kirbulut Güneş,Funda Ekici,Nurcan Turan-Oluk,Dilay Dinçdemir,Hakkı Kadayıfçı,Sevinç Nihal Yeşiloğlu,Burcu Işık,Halil Tümay,Sinem Gencer
类目:Computation and Language (cs.CL)
关键词:identify knowledge gaps, science learning requires, learning requires nuanced, requires nuanced interpretation, learners identify knowledge
备注:
点击查看摘要
Abstract:Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. We investigate whether instruction-tuned large language models (LLMs) can support multidimensional analysis of collaborative sensemaking without task-specific training, and whether structured knowledge-state information improves model inference. We evaluate two mid-size LLMs on 23 richly annotated, expert-labeled episodes across prompting conditions that vary definitional scaffolding, reasoning mode, and turn structure. Without reasoning, models tend to overpredict successful sensemaking; reasoning-enabled prompting improves identification of unsuccessful cases. Knowledge-state diagnostics provide additional grounding, improving detection of unsuccessful sensemaking and increasing agreement with expert annotations. No single configuration performs best across all sensemaking dimensions, underscoring the multidimensional nature of the task.
26. 【2609.31045】KuaFu: Compressing Long User Behavior into Understanding at Billion Scale
链接:https://arxiv.org/abs/2609.31045
作者:Jiahao Hui,Lin Zhu,Yishen Hu,Jingdong Shu,Zetai Jiang,Xining Ran,Ben Tan,Yeshou Cai,Gong Chen,Haijie Gu,Jie Jiang
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Conversational agents, generative recommenders, Conversational, raw behavior, Prevailing industrial practice
备注: 12 pages, 6 figures, 3 tables
点击查看摘要
Abstract:Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.
27. 【2609.31013】Same Text, Different Numbers: The Divergence of LLM-Based Measures
链接:https://arxiv.org/abs/2609.31013
作者:Hamid Boustanifar,Sasan Mansouri
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Finance (q-fin.GN); Risk Management (q-fin.RM)
关键词:generative large language, convert corporate text, Researchers increasingly, large language models, increasingly use generative
备注: 86 pages, including an online appendix
点击查看摘要
Abstract:Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of SP 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.
28. 【2609.31009】G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
链接:https://arxiv.org/abs/2609.31009
作者:Ruikang Liu,Haoli Bai,Yuxuan Sun,Qian Zhang,Wenzheng Cai,Yanqi Hao,Feiyu Wang,Weidong Zhong,Zhuang Wang,Tong Yang,Xiangsheng Zhou
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Post-training quantization, large language models, practical approach, approach to reducing, reducing the memory
备注:
点击查看摘要
Abstract:Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian estimates before quantizing each Transformer block, G$^2$PTQ avoids the staleness of prior global methods. Furthermore, to stabilize the exact first-order compensation, we introduce a trust-region scaling mechanism that dynamically bounds the gradient step to prevent exploding weight updates. Finally, we derive efficient implementations for block-wise Hessian approximation and exact gradient compensation. Experimental results on various model families and bit-widths demonstrate that G$^2$PTQ enables better alignment with the full-precision model, outperforming state-of-the-art baselines. Code is available at: this https URL.
29. 【2609.31002】ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
链接:https://arxiv.org/abs/2609.31002
作者:Siqiao Xue,Shuxuan Liu,Ning Hu
类目:Computation and Language (cs.CL)
关键词:general web retrieval, web retrieval transfer, retrieval transfer imperfectly, ranking decisions depend, comparative product fit
备注: project page: \url{ [this https URL](https://serendipityoneinc.github.io/look-bench-page/shoprank-bench.html) }
点击查看摘要
Abstract:Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are difficult to supervise at scale: real search traffic provides authentic queries and candidates but no clean pairwise labels. We present ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, and 8B) aligned to judge-labeled shopping preference. Training pairs are labeled by a panel of reasoning large language models (LLMs) from different families acting as a preference oracle, with position-debiased judgments and agreement tiers, and the rerankers are trained on these labels. The aligned 8B flagship then serves as a distillation teacher for the efficient 4B and 0.6B models, which are fit to its scores and sharpened on judged pairs. To measure progress, we introduce ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label. ZooWork-ShopRanker-8B and -4B significantly outperform the strongest open reranker baseline, every model significantly beats its own un-aligned base, and ZooWork-ShopRanker-0.6B beats its size peer; the gains hold in both formats and extend to common MTEB benchmarks. We release the models and the dual-format ShopRank-Bench to facilitate further research.
30. 【2609.30986】Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
链接:https://arxiv.org/abs/2609.30986
作者:Geng Liu,Feng Li,Mengxiao Zhu,Francesco Pierri
类目:Computation and Language (cs.CL)
关键词:increasingly mediate information, large language models, language models increasingly, models increasingly mediate, large language
备注: 19 pages, 34 figures, 4 tables. Geng Liu and Feng Li contributed equally
点击查看摘要
Abstract:As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or causes uncertain responses to become belief-aligned incorrect answers. It also remains unclear whether anti-sycophancy interventions preserve or restore factual accuracy or merely shift responses toward uncertainty. We analyze factual sycophancy in Chinese-language information seeking using yes/no fact-checking questions. Our analysis covers 364,941 responses from three frontier Chinese-based LLMs (DeepSeek, Qwen, and Doubao) to 12,165 factual questions derived from real-world Chinese search queries. We evaluate the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses. Under incorrect user beliefs, we distinguish belief-aligned errors from losses of factual confidence, in which initially correct answers become uncertain. Patterns vary across models and reasoning settings: reasoning is not a consistent safeguard, and anti-sycophancy instructions can reduce incorrect agreement while increasing uncertainty. In Chinese-language factual question answering, avoiding agreement with false beliefs is therefore not equivalent to preserving factual accuracy, highlighting the value of transition-level evaluation. Such behavior may undermine the reliability of LLM-mediated information access by reinforcing misinformation or weakening users' confidence in factually correct answers.
31. 【2609.30984】HA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer
链接:https://arxiv.org/abs/2609.30984
作者:Seanghay Yath
类目:Computation and Language (cs.CL)
关键词:speech recognition output, spoken form, speech recognition, recognition output, Khmer text normalization
备注:
点击查看摘要
Abstract:Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur inside ordinary words. We present Tha, a Khmer text normalization and inverse text normalization toolkit built from weighted finite-state transducers. It segments and classifies a whole line in one shortest-path search, and a second transducer rejects token boundaries inside a Khmer syllable. On Google's Khmer test suite, Tha agrees with the reference on all 274 cardinals up to one spelling variant, and on 2,906 real TTS prompts, 153 of the 158 sentences it rewrites are correct. Tha is open source under the Apache 2.0 license.
32. 【2609.30977】Does Uniform Discrete Diffusion Need Time?
链接:https://arxiv.org/abs/2609.30977
作者:Chunsan Hong,Chieh-Hsin Lai,Satoshi Hayakawa,Yuhta Takida,Jong Chul Ye,Yuki Mitsufuji
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Uniform discrete diffusion, Uniform discrete, discrete diffusion models, time, explicit time conditioning
备注: Preprint
点击查看摘要
Abstract:Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence remains much closer to its original clean sequence than to competing training sequences, the empirical-optimal predictor is nearly insensitive to time over most of the diffusion trajectory, where the guarantee weakens toward the high-noise endpoint. Empirically, trained language UDMs exhibit limited time sensitivity over most of the trajectory, while time-agnostic predictors remain competitive with, and often outperform, time-conditioned models across datasets and training objectives. These results challenge the use of explicit time conditioning in UDMs: although the population optimum depends on time, explicitly conditioning on it may often be unnecessary in practice.
33. 【2609.30974】Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change
链接:https://arxiv.org/abs/2609.30974
作者:Haruka Ezoe,Ryohei Hisano
类目:Computation and Language (cs.CL); Machine Learning (stat.ML)
关键词:independently sampled period, sampled period distributions, independently sampled, sampled period, Lexical semantic change
备注:
点击查看摘要
Abstract:Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution. We introduce Coupled Usage--Sense Processes (CUSP), which derives these answers from a single marginal preserving temporal process. A hierarchical coupling relates contextual distributions through latent usage components, while Markov composition makes adjacent and longer span correspondences compatible. Displacement operators quantify change magnitude and timing, split variation exactly between movement of component centers and reorganization within components, and attribute it to transported component pairs. Word-local modes resolve distinct directions of change and their activity over time, while representative passages from attributed components ground the analysis in text. Under a Gaussian mixture specialization, we prove parametric recovery of the operators and squared distances. Synthetic experiments support the predicted rate. CUSP remains competitive on English and German DWUG and recovers controlled Janus profiles while maintaining compositionally coherent transport. A large corpus of US court opinions demonstrates transition, mode, and passage attribution in unlabeled natural text. CUSP thus makes magnitude, timing, mechanism, movement, modes, and textual evidence compatible views of one lexical history.
34. 【2609.30968】FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation
链接:https://arxiv.org/abs/2609.30968
作者:Lu Han,Jingyao Zhang,Katy Ilonka Gero,Nguyen H. Tran
类目:Computation and Language (cs.CL)
关键词:Large language models, compromise individual writing, Large language, individual writing style, personalized writing assistants
备注:
点击查看摘要
Abstract:Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tuning (PEFT) offers a data-local setting for this multi-author adaptation problem: clients keep author text local while sharing compact adapter updates. However, we show that standard aggregation can preserve continuation utility while making different authors' generations less distinguishable in style space, a failure mode we define as author-style homogenization. We evaluate author-style retention with Angular Style Classification Encoder (ASCE)-based diagnostics on our main BlogText benchmark and ASCE-independent external authorship verification. Using this protocol, we find that common federated PEFT baselines can preserve semantic utility while averaging out author-specific signals. To address this homogenization, we instantiate FAVoR (Federated Authorial Voice Retention), an author-style residual mechanism for federated PEFT. FAVoR uses a shared-private adapter design: clients upload shared-adapter updates while retaining author-specific residual corrections locally. Across BlogText and external Mythos-Reddit validation, FAVoR improves author-style retention over standard and personalized federated PEFT baselines. These gains come with small continuation-utility trade-offs and are supported by component ablations, external verification, and cold-start transfer.
35. 【2609.30935】Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models
链接:https://arxiv.org/abs/2609.30935
作者:Bing Wang,Changchun Li,Xin-Qiang Cai,Lin Yuanbo Wu,Ximing Li,Gang Niu,Masashi Sugiyama
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:large language models, inherent general-purpose knowledge, general-purpose knowledge, LLMs' general-purpose knowledge, Continual LLM fine-Tuning
备注: Accepted by NeurIPS 2026. 29 pages, 3 figures. Code: [this https URL](https://github.com/wangbing1416/EoupCT)
点击查看摘要
Abstract:Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the original data and gradients of off-the-shelf pre-training LLMs required by these methods are strictly unknown and highly diverse. To bridge this critical gap, we propose EoupCT, a novel framework designed to Estimate and Orthogonalize Unknown Pre-training gradients for Continual LLM fine-Tuning. Specifically, EoupCT estimates pre-training gradients by dynamically generating pseudo data that is most susceptible to forgetting for new tasks through a learnable soft prompt equipped with Gumbel-Softmax relaxation. Furthermore, we formulate a multi-objective optimization problem and introduce a first-order efficient Pareto optimizer that jointly optimizes LLM parameters and the soft prompt, rigorously enforcing orthogonality between new task updates and the estimated pre-training gradients. Extensive experiments across multiple LLMs demonstrate that EoupCT effectively preserves both task-specific proficiency and inherent general-purpose knowledge, successfully mitigating the catastrophic forgetting.
36. 【2609.30924】raining-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
链接:https://arxiv.org/abs/2609.30924
作者:Hikaru Asano,Yotaro Kubo,So Kuroki
类目:Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
关键词:efficient pronunciation transcription, Accurate and efficient, essential for preparing, efficient pronunciation, pronunciation transcription
备注: 5 pages, 2 figures
点击查看摘要
Abstract:Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.
37. 【2609.30914】Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces
链接:https://arxiv.org/abs/2609.30914
作者:Aman Mittal,Ferdin Sagai Don Bosco,Kasturi Venkata Srikanth,Abhishek Singh,Aditya Singh,Abhishek Chopra
类目:Distributed, Parallel, and Cluster Computing (cs.DC); Computation and Language (cs.CL); Optimization and Control (math.OC)
关键词:probabilistic amplitude evolution, representing candidate solutions, Quantum-inspired algorithms emulate, quantum mechanical principles, emulate quantum mechanical
备注:
点击查看摘要
Abstract:Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate operators. This approach offers higher optimization performance without physical qubits, and has been shown to achieve order-of-magnitude speedups (10--80$\times$) over traditional solvers on combinatorial, high-dimensional NP-hard problems. A critical barrier to adoption, however, is the lack of a unified execution framework that delivers both algorithmic performance and hardware portability. We present \textbf{Cross-Backend Quantum Inspired Evolutionary Optimizer (QIEO)}, the runtime core of BQP's BQPhy solver, which addresses this gap through a \emph{single-source-of-truth} architecture. One C++ implementation of the QIEO algorithm is compiled once per hardware target and exposed to multiple high-level languages via thin binding layers. The framework dispatches to CPU (sequential), OpenMP~5 (multi-core), CUDA (NVIDIA), and HIP (AMD) backends at runtime, adapting kernels to each device's memory hierarchy and warp/wavefront execution model. The framework's real-world utility is validated through binding demonstrations that share the identical C++ runtime. BQPhy's Python library is demonstrated on a neural network hyperparameter optimisation achieving 88.60\% test accuracy on MNIST. BQPhy's MATLAB's Toolkit is tested on wind farm layout optimisation attaining $365\,399 \pm 4\,552$~MWh/yr, which is statistically indistinguishable from particle swarm optimisation and $+7.6\%$ above genetic algorithms on a 32-variable constrained engineering problem. The Julia package tackles the Lotka--Volterra parameter estimation where BQPhy replaces native Julia solvers on the same residual, cutting mean SSE by $2.1\times$.
Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC); Computation and Language (cs.CL); Optimization and Control (math.OC)
Cite as:
arXiv:2609.30914 [cs.DC]
(or
arXiv:2609.30914v1 [cs.DC] for this version)
https://doi.org/10.48550/arXiv.2609.30914
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
38. 【2609.30906】oolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
链接:https://arxiv.org/abs/2609.30906
作者:Zhenlong Dai,Xujie Song,Zitong Wang,Tong Niu,Jian liu,Weiqiang Wang,Xiu Tang,Sai Wu,Chang Yao,Jingyuan Chen
类目:Computation and Language (cs.CL)
关键词:Large language models, natural language processing, Large language, large-scale tool selection, tool selection
备注: Accepted at NeurIPS 2026
点击查看摘要
Abstract:Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.
39. 【2609.30897】From annotation to reasoning: Culture in language models
链接:https://arxiv.org/abs/2609.30897
作者:Daniel Hershcovich,Alexander Conroy,Jens Bjerring-Hansen
类目:Computation and Language (cs.CL)
关键词:evaluate language models, Cultural, evaluate language, test factual knowledge, Abstract
备注: 8 pages, 1 table; perspective paper
点击查看摘要
Abstract:How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.
40. 【2609.30882】Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos
链接:https://arxiv.org/abs/2609.30882
作者:Yuya Wake,Sho Tsugawa,Toshiyuki Amagasa
类目:Computation and Language (cs.CL)
关键词:Large language models, Large language, language models, effectiveness may depend, assess long-form medical
备注: 15 pages. Accepted at the 18th International Conference on Advances in Social Networks Analysis and Mining (ASONAM 2026), Multidisciplinary Track, Short Paper
点击查看摘要
Abstract:Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how such transcript compression affects LLM-based veracity classification of Japanese medical YouTube videos. We compare four transcript input designs: full transcripts, LLM-generated summaries, RAPTOR-based retrievalaugmented generation (RAG), and Screening, which extracts candidate medical and health-related sentences. Using 74 long-form videos labeled as Real or Fake, we evaluate classification performance and analyze linguistic changes using J-LIWC, hedge expressions, and institutional or technical terms. The full-transcript Baseline achieved the best performance, whereas all compressed inputs increased false negatives, meaning that Fake videos were more likely to be misclassified as Real. Summary caused the largest performance drop, while Screening performed best among the compressed inputs but still omitted many medically relevant sentences. Linguistic analyses showed that these errors were not explained by a simple increase in certainty. Instead, Summary reduced affective, social, temporal, cognitive, and conversational cues, while Summary and RAG made institutional and technical terms more salient. These findings suggest that transcript compression can represent Fake videos as more coherent and authoritative inputs, thereby weakening cues needed for misinformation detection
41. 【2609.30867】Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
链接:https://arxiv.org/abs/2609.30867
作者:Yonghong Zhang,Yong Xie,Isabel M. Parra,Ricardo Correia
类目:Computation and Language (cs.CL)
关键词:assumptions remains challenging, identification assumptions remains, evaluate climate policy, studies are widely, climate policy
备注: Accepted at ClimateNLP 2026, the 3rd Workshop on Natural Language Processing meets Climate Change (EMNLP 2026). 9 pages plus appendix (21 pages total), 6 figures, 15 tables
点击查看摘要
Abstract:Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: this https URL
42. 【2609.30864】Persistent Negatives for Adversarial Black-Box On-Policy Distillation
链接:https://arxiv.org/abs/2609.30864
作者:Haixu Ma,Saad Lahrichi,Weiwei Li,Kevin Han,Weiqiang Wu,Peggy Yang,Dongzhuo Li,Ruiyi Li,Serena Li,Gedi Zhou,Mingze Gao,Abhishek Kumar,Xiangjun Fan,Lizhu Zhang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:token probabilities, discriminator, Adversarial distillation, OPD, Black-box On-Policy Distillation
备注:
点击查看摘要
Abstract:Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator compute, historical comparisons train the discriminator, while GRPO remains on-policy with fresh student responses. Our analysis identifies the Bayes-optimal reward as a teacher-to-negative log-density ratio and, under explicit assumptions, shows how persistent negatives anchor the discriminator and reduce reward-estimation MSE relative to fresh-negative training. Across two student families, three judges, and four judged-chat benchmarks, persistent-negative adversarial distillation consistently improves performance over current methods at matched discriminator compute. It also yields smoother fresh-policy discriminator trajectories, with fewer below-chance dips. These findings identify the discriminator's negative distribution as an important design axis in black-box on-policy distillation.
43. 【2609.30849】Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
链接:https://arxiv.org/abs/2609.30849
作者:Phuong Q. Le,Kemal Kurniawan,Jey Han Lau
类目:Computation and Language (cs.CL)
关键词:surface-level perturbation methods, Prior work, Prior, perturbation, perturbation methods
备注: 22 pages, 10 figures
点击查看摘要
Abstract:Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.
44. 【2609.30846】I-Parakeet: Integer-Only Conformer ASR on Mobile NPU
链接:https://arxiv.org/abs/2609.30846
作者:Taichi Nishimura
类目:Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
关键词:NVIDIA Parakeet-CTC, implementation of NVIDIA, Conformer ASR models, Modern Conformer ASR, floating-point operator
备注: Under review
点击查看摘要
Abstract:In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices because of their size, and quantized models still fall back to floating point for numerically sensitive operations. This prevents them from fully exploiting integer accelerators such as mobile NPUs. To achieve this, our contributions are threefold. First, we derive an integer formulation of the relative-positional self-attention at the core of the Conformer. We fuse its two score branches with different quantization scales and the relative shift into integer-only operations. Second, we introduce a minimax-optimized Swish approximation that minimizes the maximum error of the Swish output. Third, a layer-wise range analysis of activations yields two targeted remedies: an INT16 grid for the BatchNorm output and percentile calibration for the heavy-tailed pre-encoder activations. I-Parakeet achieves 4.97% WER on LibriSpeech test-other, running on a Qualcomm NPU at a real-time factor of 0.048, 7.5x faster than a CPU baseline.
45. 【2609.30820】Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
链接:https://arxiv.org/abs/2609.30820
作者:Nux Li
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:making low-bit quantization, transformers reuse weights, making low-bit, reuse weights, one-step GPTQ
备注: 27 pages, 5 figures
点击查看摘要
Abstract:Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Controlled experiments on linear filters and Mamba state-space models show that feedback exposure also occurs outside transformers. Grouped INT4 reveals a separate failure, calibration blindness: our one-step GPTQ baseline builds its Hessian from step-0 activations, leaving input directions used later in the recurrence nearly unweighted. Across nine checkpoints from seven looped architectures, one-step GPTQ is worse than round-to-nearest (RTN) on the primary task metric for five checkpoints. Accumulating the GPTQ Hessian across recurrence steps outperforms both one-step GPTQ and RTN on all nine checkpoints and recovers bf16-level accuracy on Huginn. These results separate two questions for PTQ on looped models: where quantization error enters the recurrence, and which states calibration sees.
46. 【2609.30802】Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
链接:https://arxiv.org/abs/2609.30802
作者:Anjila Budathoki,Manish Dhakal,Benjamin M. Ampel,Yi Ding
类目:Computation and Language (cs.CL)
关键词:Supervised Fine-Tuning, Prior research, significantly impacts, research has demonstrated, choice of prompt
备注:
点击查看摘要
Abstract:Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill this gap by analyzing how different template configurations influence the pre-existing safety alignment of the student. We observe a significant degradation of safety alignment present in the aligned base instruct-tuned model. Specifically, we find that utilizing chat templates renders the model more compliant with harmful queries compared to a non-chat template. These findings are consistent across three models: LLaMA, Gemma and Qwen model families and are evaluated across multiple safety benchmarks. We further show that using a non-chat template during distillation better preserves the base student's internal representations, while chat template distillation induces a larger representational shift. Code: this https URL
47. 【2609.30784】Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
链接:https://arxiv.org/abs/2609.30784
作者:Yotaro Kubo,Qi Sun,Yujin Tang
类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:equipping large language, large language models, audio language model, LLM, paper proposes
备注: Submitted to ICASSP
点击查看摘要
Abstract:This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.
48. 【2609.30773】Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
链接:https://arxiv.org/abs/2609.30773
作者:Manato Yaguchi,Yotaro Kubo,Hikaru Asano,So Kuroki
类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:asynchronous text backend, responsive speech frontend, architectures couple, couple a responsive, asynchronous text
备注: Submitted to ICASSP 2027. 5 pages, 1 figure, 2 tables
点击查看摘要
Abstract:Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhead when training on real conversations. We propose randomized intermediate guidance, which derives guidance directly from the conversation corpus rather than simulating backend LLM behavior. During training, target responses provide informative guidance, while randomly sampled responses provide potentially irrelevant updates during the utterance. This combination aims to teach the frontend to use backend information selectively. On synthetic dialogues, KAME trained with this recipe achieves response quality comparable to that of the LLM-generated and similarity-based baselines. Training on 3.8k hours of real conversations improves smooth turn-taking and audio-judge naturalness over synthetic-data KAME while retaining a response-quality advantage over Moshi. These results show that randomized guidance offers a practical route to combining the response-quality benefits of tandem models with natural conversational behavior learned from real speech.
49. 【2609.30739】SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages
链接:https://arxiv.org/abs/2609.30739
作者:Puja Ahmad Habibi,Faiz Assabil Firdaus,Ashvanth S,Ekapol Chuangsuwanich,Pume Tuchinda,Peerat Limkonchotiwat
类目:Computation and Language (cs.CL)
关键词:remain poorly supported, poorly supported due, region linguistic diversity, Southeast Asian languages, Multilingual text-vision embedding
备注: Accepted to ACCV 2026. Model weights and datasets are available at [this https URL](https://huggingface.co/collections/fassabilf/sea-clip-tiny-accv-2026) and code for training, evaluation, and preprocessing at [this https URL](https://github.com/fassabilf/sea-clip-tiny)
点击查看摘要
Abstract:Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10, respectively. Compared with MobileCLIP2, it improves average R@10 by 12.1 points while using 38.4% fewer parameters and lower measured CPU latency. These results highlight the importance of region-aware training for efficient multilingual text-vision models in Southeast Asia.
50. 【2609.30738】Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
链接:https://arxiv.org/abs/2609.30738
作者:Tianfang Xie,Wei Zhu
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:cache eviction methods, small observation window, PyramidKV rank tokens, rank tokens solely, cache eviction
备注: 5 pages, 1 figure, 3 tables. Submitted to IEEE ICASSP 2027
点击查看摘要
Abstract:KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $\lambda_20$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary with depth, we compare fixed global coefficients with three-segment and quadratic profiles. Only these depth profiles are searched on a development split under a $\sinh$ reparameterization. On all 16 English LongBench datasets with Mistral-7B at a budget of 64 entries per layer, a single global diversification constant improves 13 of 16 datasets (macro +1.1); the gain holds at budget 32 and narrows at 128. Per-dataset search finds no detectable layer structure on most datasets; on passage retrieval it finds a large one: a mid-layer sign flip that rewards similarity and is worth +9.6 over the baseline at budget 64 and, without re-tuning, +13.2 over the global constant at budget 128. Ablations attribute the gain to the redundancy term; replaying every accepted search state on the held-out test set separates genuine structure from tuning noise.
51. 【2609.30716】Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
链接:https://arxiv.org/abs/2609.30716
作者:Amanda Fitch
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Toggle, Toggle Hugging Face, Bibliographic Explorer Toggle, Explorer Toggle Bibliographic, Toggle Bibliographic Explorer
备注: 36 pages, 1 figure, evaluation dataset and logs released
点击查看摘要
Abstract:When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b model across a targeted behavioral suite (n = 13 items, 784 forward passes in short, single-turn contexts) using a completely counterbalanced experimental design. This setup allowed us to mathematically isolate the specific effects of source framing and reading position, while ensuring the model's natural vocabulary biases were canceled out. Across ten test conditions, we discovered the following: 1. Source framing heavily overpowers reading position. When directly competing, the semantic framing of a source (such as presenting it as an official guideline or a fresh update) had a significantly stronger impact on the model's final answer than the presentation order of the document. 2. The model favors the first document it reads, but this bias is highly variable. While the model consistently demonstrated a primacy effect (preferring the first document presented), the actual strength of this bias fluctuated by at least a factor of 5 based solely on the surface wording. 3. Overall structural repetition, not short copy-cues, drives positional bias. The model's preference for the first document is not a mechanical reaction to short, repetitive trigger phrases, such as "is [Answer]". However, the primacy effect does increase significantly when the two competing documents are structurally identical, using word-for-word verbatim templates. Introducing variation in the overall wording between the two sources reduces this positional bias.
Comments:
36 pages, 1 figure, evaluation dataset and logs released
Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:
arXiv:2609.30716 [cs.CL]
(or
arXiv:2609.30716v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2609.30716
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Amanda Fitch [view email] [v1]
Fri, 25 Sep 2026 02:43:40 UTC (94 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4, by Amanda FitchView PDFHTML (experimental)TeX Source
view license
Current browse context:
cs.CL
prev
|
next
new
|
recent
| 2026-09
Change to browse by:
cs
cs.AI
References Citations
NASA ADSGoogle Scholar
Semantic Scholar
export BibTeX citation
Loading…
BibTeX formatted citation
loading…
Data provided by:
Bookmark
checked="checked"class=“labs-tab-input”>
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.
Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)
mathjaxToggle();
We gratefully acknowledge support from
our major funders,
member institutions, ,
and all contributors.
About
Help
Contact
Subscribe
Copyright
Privacy
Accessibility
Operational Status (opens in new tab)
Major funding support from
52. 【2609.30706】LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information
链接:https://arxiv.org/abs/2609.30706
作者:Furkan Yilmaz,Habibe Aleyna Tasdemir,Muhammed Faruk Gozay
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:open counterpart Laya, counterpart Laya answer, Laya answer typed, single forward pass, TypeSafe Jev
备注: 11 pages, 3 figures, 7 tables. Code: [this https URL](https://github.com/moganai/lavoir) ; model: [this https URL](https://huggingface.co/moganai/lavoir)
点击查看摘要
Abstract:"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both the decision distribution and, for every slot, the expected gain in the probability of the correct decision if the user were asked about it. VOI targets need no human labels: gold decisions come from schema rules, an LLM only verbalizes messages and answers, a model from another family checks every text, and pairing each message with several profiles makes regression on realized gains estimate the expected gain. A Gini-impurity cap bounds the predicted value by what a calibrated model can still gain. In a controlled study, decisions on seen schemas are statistically indistinguishable from the Bayes ceiling. The final model's question policy matches a greedy oracle VOI policy on seen schemas (AUC 0.799 vs. 0.797), and with at most 0.5 questions per conversation it is 14.1 points more accurate than never asking. On real ABCD conversations, one real exchange raises accuracy by 8.3 points where LAVOIR asks and leaves it unchanged where it does not; on SGD the cap lowers the asking rate from 93% to 8.6%. On Laya's twelve benchmarks LAVOIR is above Laya's reported scores on seven, and it answers a question in 31 ms (median, GH200).
53. 【2609.30670】RACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
链接:https://arxiv.org/abs/2609.30670
作者:Yibo Ma,Qianqian Zhang,Peng Liu,Tiancheng Zhao
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:Streaming video understanding, video understanding requires, understanding requires models, report task scores, Streaming video
备注: TRACE Tech Report
点击查看摘要
Abstract:Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors explicit. TRACE combines temporally audited visual tasks with evidence timing and instruction-dependent trigger annotations, a unified causal Core--Adapter protocol that controls information availability while recording actual history processing and response events, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. On 1,240 records from 517 videos, we evaluate eight publicly available models or systems in eight configurations. We find that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms (responses emitted while no target window is currently valid and a later one remains), and missed target windows. These results show that streaming-video performance should be interpreted as execution-conditioned system behavior rather than a single score. Our benchmark and code can be accessed at \href{this https URL}{this https URL}.
54. 【2609.30657】Prompt Injection Detection for Email Agents Through Attack Chain Modeling
链接:https://arxiv.org/abs/2609.30657
作者:Ahmad Hashmi,Dhyey Patel,Yunting Yin
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:Large language model, Large language, indirect prompt injection, influence subsequent tool, untrusted email content
备注: Accepted to IEEE ICTAI 2026
点击查看摘要
Abstract:Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy. To support this framework, we derive attack chain labels from prompt injection datasets, evaluate the proposed framework under random splits, temporal phase transfer, conditional stage transfer, cross-dataset transfer, and conduct ablation studies on multiple benchmarks. Results show that random train test splits substantially overestimate robustness under distribution shift, while later tool argument stages are more predictable than earlier stages in the framework. We also show that training on harmless emails that resemble attacks helps reduce false alarms while preserving the ability to detect real attacks. Across five binary benchmarks, our framework achieves a mean F1 score of 0.406 under the strict threshold setting policy, compared with 0.216 for the strongest of five pretrained detectors evaluated without additional training. These results highlight the value of combining attack stage predictions with checks for conflicts between the user's request and instructions in retrieved emails. Our experiments also demonstrate the importance of training with challenging benign examples to balance attack detection and false alarms.
55. 【2609.30652】Recursive Self-Improvement via On-Policy Distillation for Reasoning
链接:https://arxiv.org/abs/2609.30652
作者:Shangjian Yin,Zehao Zhao,Kavosh Asadi,Rui Liu,Yuchen Lu,Shike Mei,Hang Cui,Luke Simon,Zhouxing Shi,Hamed Firooz
类目:Computation and Language (cs.CL)
关键词:teacher next-token predictions, next-token predictions, external teacher next-token, generate trajectories, student model
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvements learned by the student during training. Our primary contribution is to address this limitation with a recursive framework built around two complementary components. First, we let the privileged teacher co-evolve with the student so that revision learned in one round can guide the next, a process we refer to as Dynamic Co-Evolution (DCE). Second, because stronger revision can also make responses too verbose and self-critical, we additionally train on shorter, verified rewrites of the model's own on-policy responses. We call this complementary objective Self-Refined Concise Learning (SRCL). Overall, our comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks. Specifically, on Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, outperforming OPSD by 35.62 percentage points while reducing mean output length by 7.80% relative to DCE alone.
56. 【2609.30611】Epstein Files Engine: Agentic Search for Investigative Journalism
链接:https://arxiv.org/abs/2609.30611
作者:Duy K. Nguyen,Teresa Mondría Terol,Dylan Freedman,Zach Seward
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR)
关键词:Department of Justice, Epstein Files Engine, Jeffrey Epstein, Justice released, collection concerning Jeffrey
备注: 6 pages, 2 figures, 2 tables. Presented at the Computation + Journalism Symposium (C+J 2026)
点击查看摘要
Abstract:On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questions into Google BigQuery SQL queries across three corpora: Epstein-related releases, the Times's archive and external, Epstein-related news headlines. It used an LLM to plan queries and returned citation-rich answers a reporter could verify and trust. More than 100 journalists used the Engine, and it contributed to at least 20 published stories. We report how reporters queried it and describe Diff, our text-and-visual duplicate matching method that amplified novelty signals and allowed the Engine to surface genuinely new information. We argue that newsroom agents serve newsrooms best not as autonomous writers, but as interfaces to source material and institutional knowledge.
57. 【2609.30604】he Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
链接:https://arxiv.org/abs/2609.30604
作者:Alexander Gill,Md Farhan Ishmam,Xuyen Nguyen,Neha Bhat,Parker Henry DeYoung,Fateme Hashemi Chaleshtori,Nathan Stringham,Kenneth Marino,Ana Marasović
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Existing computer-use agent, Existing computer-use, Existing, agents, tasks
备注: 9 pages main text. Accepted to Findings of EMNLP 2026. Project page: [this https URL](https://alexgill321.github.io/KNOWS-benchmark/)
点击查看摘要
Abstract:Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirements. Each task is paired with an evaluator, a program that combines deterministic checks with LLM judgments to balance the richness, reliability, and automation tradeoff inherent to agent evaluation. We evaluate and analyze frontier computer-use agents and browser-based harnesses. They achieve moderate scores on partial-success metrics, but the best performer fully succeeds in fewer than 3% of our complex, long-horizon tasks. Failures on visual steps render the resulting artifacts unusable, even when agents complete more than 50% of other evaluation steps. Our results expose limitations of current agents acting as end-to-end assistants, and call for progress on tool use, visual understanding, and long-horizon reasoning.
58. 【2609.30563】hinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content
链接:https://arxiv.org/abs/2609.30563
作者:Ljubisa Bojic,Tijana Stanic,Joerg Matthes,Agariadne Dwinggo Samala,Bojana Dinic,Jue Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA); Social and Information Networks (cs.SI)
关键词:Platform policies, policies are increasingly, increasingly tested, tested on artificial, making agent fidelity
备注: 24 pages, 7 figures
点击查看摘要
Abstract:Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.
59. 【2609.30558】Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
链接:https://arxiv.org/abs/2609.30558
作者:Jiaqi Ding,Guorong Wu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:long-term user preferences, maintain long-term user, collapse memory behavior, user preferences, task states
备注: Accepted by EMNLP 2026 Main, code is availble at [this https URL](https://github.com/jq-ding/MemProbe)
点击查看摘要
Abstract:Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at this https URL.
60. 【2609.30547】REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
链接:https://arxiv.org/abs/2609.30547
作者:Haixu Ma,Aditya Bansal,Shubham Lohiya,Sumit Ranjan
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:digital marketing, Exact Audience sizing, Audience sizing, Exact Audience, Audience
备注: Accepted by ICDM 2026
点击查看摘要
Abstract:Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant delays, estimation errors, and poor scalability over high-dimensional profile data. We present REALMS (Real-time Exact Audience sizing via LLM-based Multi-attribute Search), a conversational system for exact audience sizing deployed in production on an enterprise customer data platform. REALMS enables marketers to query massive profile stores with millions of profiles and thousands of attributes using natural language and receive precise counts in seconds. The system introduces three key components: (1) a categorical attribute retrieval mechanism using embedding-based vector search to dynamically identify relevant schema attributes without manual configuration; (2) an LLM-powered NL2SQL pipeline with template-based in-context learning for accurate query generation over complex nested schemas; and (3) schema standardization enabling industry-agnostic deployment across diverse enterprise environments. Evaluation on real enterprise data demonstrates strong recall for attribute retrieval, high SQL execution accuracy, and low latency, which enables real-time interactive audience insights where prior methods required hours.
61. 【2609.30540】Don't CLAP: Are Music-Text Models Bag-of-Words?
链接:https://arxiv.org/abs/2609.30540
作者:Yuan-Chiao Cheng,Alexander Lerch
类目:ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:standard objective metric, text embedding capture, systems are assessed, faithfully the music, cosine similarity
备注: 5 pages, 4 figures, 1 table
点击查看摘要
Abstract:Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.
62. 【2609.30535】Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
链接:https://arxiv.org/abs/2609.30535
作者:Dries Rooryck,Alex Cai,Yonatan Belinkov,David Alvarez-Melis,Kianté Brantley
类目:Computation and Language (cs.CL)
关键词:communities often code-switch, single utterance, Children, multiple languages, Children in multilingual
备注: 17 pages, 8 figures. Accepted to the BabyLM Workshop at EMNLP 2026
点击查看摘要
Abstract:Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at this https URL.
63. 【2609.30514】Inquesto Score: A reliability Protocol For Voice Agents
链接:https://arxiv.org/abs/2609.30514
作者:Massa Baali,Bhiksha Raj
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:affect transactions, workflows where failed, failed interactions, interactions can affect, reproducible and interpretable
备注:
点击查看摘要
Abstract:Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse. Rather than combining heterogeneous metrics, IS defines explicit failure events and severity levels and evaluates the deployed voice pipeline. Timing failures, including talk-over and delayed responses, are measured directly from audio, while semantic and state-dependent failures are evaluated using scenario predicates, tool traces, and a pinned open-model judge. Diagnostic views of behavior, acoustic robustness, identity handling, and speaker groups accompany the score without being combined into it. Inquesto Score v0.1 evaluates 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations of a reference voice-agent system. Our evaluation shows that reliable measurement requires evidence beyond transcripts, explicit treatment of deployment conditions, and validation of the evaluators used to determine outcomes. We release the protocol, reference implementation, and evaluation records.
64. 【2609.30492】Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs
链接:https://arxiv.org/abs/2609.30492
作者:Sang Bin Moon,Nicole Cho,Daniel Borrajo,Sumitra Ganesh,Abolfazl Hashemi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
关键词:potentially suboptimal decision, spawn groupthink-the convergence, Language models, produce homogeneous responses, suboptimal decision
备注:
点击查看摘要
Abstract:Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform-coverage sampling, and evolutionary persona generation. Evaluations on the Alternative Uses Task (AUT), Infinity-Chat, and Divergent Association Task (DAT) show the benefits of the proposed methods across tasks and creativity objectives. On AUT, evolutionary persona generation increases response diversity by 78.8%, originality by 26.1%, flexibility by 49.5%, and holistic creativity by 13.9% over task-only prompting, while maintaining 98.5% validity; on Infinity-Chat, it nearly doubles persona-induced response separation relative to random personas. Moreover, evolutionary personas compose with creativity-optimized prompting, further increasing its response diversity by 18.6% and creativity by 6.3%. These results establish persona-set geometry as a task-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization.
65. 【2609.30483】AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth
链接:https://arxiv.org/abs/2609.30483
作者:Sheng-Tse Lin,Siyuan Zhai,Chien-Liang Kuo,Massa Baali,Bhiksha Raj
类目:ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
关键词:Audio language models, models state numbers, human opinion, language models state, numbers for acoustic
备注: 5 pages, 3 figures, 2 tables. Submitted to ICASSP 2027. Siyuan Zhai and Chien-Liang Kuo contributed equally. Code and outputs: [this https URL](https://github.com/sheng-tse/acousticlaim)
点击查看摘要
Abstract:Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets' rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.
66. 【2609.30471】CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
链接:https://arxiv.org/abs/2609.30471
作者:Mukul Chhabra,Shail Patel,Luigi Medrano
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:reference typically applies, reference-based judge penalizes, evaluation assumes, literal judge penalizes, judge penalizes
备注: 15 pages, 1 figure, 5 tables, 1 algorithm. Preprint
点击查看摘要
Abstract:Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation by retrieval confidence, casting production evaluation as selective prediction. We introduce CARGO-Bench, a perturbation-based diagnostic suite with ground truth by construction that separates leniency from discrimination. On CARGO-Bench (246 items, two judge models, 7,872 judgments), the standard reference-based judge penalizes 100% of correct entity-transplanted answers and is uninformative (discrimination index DI ~ 0); supplying the live facts without reframing changes nothing. CARGO eliminates these false penalties (0/50) while retaining near-complete contradiction recall (50/50 and 49/50), raising DI to 0.58 [0.48, 0.68]; a rubric-swap control attributes most of the effect to context-grounded dimension definitions. CARGO also exposes a limitation of its own design: the leniency that protects entity values suppresses detection of procedural corruptions (20% recall). A post-hoc fix does not close the gap, and an LLM-as-annotator study with written guidelines and adjudication shows the same blind spot. We release a preregistered protocol for extending the evaluation to expert agreement, risk-coverage, and cost on production traffic.
67. 【2609.30467】Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
链接:https://arxiv.org/abs/2609.30467
作者:Heyuan Huang,Jirui Dai,Alexandra DeLucia,Sonal Joshi,Mahsa Yarmohammadi,Jie Gao,Bernal Jiménez Gutiérrez,Mark Dredze
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Retrieval-based factuality evaluation, scalable hallucination detection, Retrieval-based factuality, high-stakes clinical settings, authoritative medical corpora
备注: Experiments' corpus knowledge cutoff date May 2026
点击查看摘要
Abstract:Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at this https URL for the full reproducibility of our results.
68. 【2609.30465】RAZOR: Pruning Replaceable Experts in LLMs
链接:https://arxiv.org/abs/2609.30465
作者:Mingyang Song,Mao Zheng
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:full expert pool, store the full, expert, models activate, expert pool
备注:
点击查看摘要
Abstract:Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25\% and 50\% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
69. 【2609.30439】Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
链接:https://arxiv.org/abs/2609.30439
作者:Bo Su,Yueru Yan,Thai Le
类目:Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
关键词:target-speaker unlearning ASR, introduce target-speaker unlearning, target-speaker unlearning, unlearning ASR, ASR
备注: 5 pages
点击查看摘要
Abstract:We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) datasets show that speech transcription accuracy for corresponding opt-out words or characters falls from 72.3% to 48.2% and from 73.6% to 27.3%, respectively, while retained speakers' transcription error rates maintain more or less the same. Our approach provides a practical solution for modern video conferencing platforms, allowing speakers to dynamically opt-out from automated AI transcriptions without forcefully leaving the meeting sessions, enabling a privacy-preserving interface for potentially millions of online meetings daily.
70. 【2609.30416】All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation
链接:https://arxiv.org/abs/2609.30416
作者:Amir Hussein,Enas Albasiri,Travis M. Bartley,Nourchene Ferchichi,Ke Hu,Harishchandra Dubey,Myungjong Kim,Zhehuai Chen,Oluwatobi Olabiyi,Sanjeev Khudanpur
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Large Language, Language Models, remains challenging due, shown strong performance
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.
71. 【2609.30414】A Unified Account of Concepts and Chunks
链接:https://arxiv.org/abs/2609.30414
作者:Karthik Singaravadivelan,Pat Langley
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Cognitive psychology, people encode, describe categories, psychology has studied, studied how people
备注: Accepted to ACS-26 (oral presentation)
点击查看摘要
Abstract:Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses a challenge for unified theories of cognition. In this paper, we review Cobweb, a computational account of categorization and concept formation, and propose an extended theory that incorporates chunks and their acquisition. The theory makes no commitments about modality, applying to any experience that decomposes into elements and relations among them. We also present \trellis/, an implementation of this theory, and illustrate its application to learning context-free grammars, which we adopt as a testbed because they involve both concept-like and chunk-like elements. In addition, we report experimental results on three synthetic grammars that demonstrate the system's ability to represent syntactic knowledge, use it to parse and generate sentences, and learn compositional structures from sample parses. We conclude by discussing related work on concepts and chunks, along with directions for future research in the area.
72. 【2609.30402】What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study
链接:https://arxiv.org/abs/2609.30402
作者:Akshit Sharma,Prashant W. Patil
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multimedia (cs.MM)
关键词:increasingly crafted, convincing by pairing, pairing a textual, textual claim, design choices
备注: Accepted at the Tenth Widening NLP Workshop (WiNLP), co-located with EMNLP 2026
点击查看摘要
Abstract:Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language backbones. Through systematic comparisons and targeted robustness analyses, we distill practical guidance on which design choices help, when do they fail silently, and what aspects of the pipeline most strongly shape model behavior, answering 4 key Research Questions (RQs). We aim to provide a reliable foundation for designing stronger and more dependable multimodal misinformation detection systems, thus contributing to the broader research community.
73. 【2609.30328】When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
链接:https://arxiv.org/abs/2609.30328
作者:Salma Roshdy Aly,Hussein Assaf,Ziad Kobti
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
关键词:report the absence, evidence, Salma Roshdy Aly, retrieved documents, Machine Learning Workshop
备注: Accepted at the 21st Women in Machine Learning Workshop (WiML), NeurIPS 2026
点击查看摘要
Abstract:When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions equally good on 78 to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly reaches 43.7%. Neither easier problems nor a larger judge changes this. Two measurements taken from the pipeline's own logs explain it without needing labels. Gating on one of them, the pipeline declines the comparisons it cannot make and raises its accuracy from 20.7 to 36.9% while still answering half of all comparisons. The contribution is not a more accurate judge, but a label-free way to tell when a judge has no basis for its answer.
Comments:
Accepted at the 21st Women in Machine Learning Workshop (WiML), NeurIPS 2026
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
Cite as:
arXiv:2609.30328 [cs.AI]
(or
arXiv:2609.30328v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2609.30328
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Salma Roshdy Aly [view email] [v1]
Wed, 23 Sep 2026 23:19:59 UTC (8 KB)
74. 【2609.30316】PALM: Point-in-Time Adaptation for Financial Language Models
链接:https://arxiv.org/abs/2609.30316
作者:Seunghan Lee,Jun Seo,Jaehoon Lee,Junhyeok Kang,Sangjun Han,Sungdong Yoo,Minjae Kim,Tae Yoon Lim,Dongwan Kang,Hwanil Choi,Soonyoung Lee,Wonbin Ahn
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computational Finance (q-fin.CP)
关键词:Language models, financial backtests suffer, look-ahead bias, asked to predict, backtests suffer
备注:
点击查看摘要
Abstract:Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrained on chronologically filtered corpora and released as one checkpoint per calendar year, each with a documented cutoff. However, each additional year costs a full pretraining run, and whether that run is necessary has never been tested. In this paper, we show that the annual pretraining run is not necessary. We instead compare each checkpoint against the newer one that replaced it, and find that the newer checkpoint scores no better on the same evaluation window. Motivated by this observation, we propose PALM (Point-in-time Adaptation for financial Language Models), a simple yet effective alternative to annual pretraining that fits a low-rank adapter on text published before the decision date without modifying any pretrained weight. We further find that a small adapter is enough to add a new period to the knowledge an old checkpoint already encodes, and that this outperforms continued pretraining. We validate PALM on a decade of financial news and on various families of PIT models, whose cutoffs span two decades and whose sizes range from 1.3 to 4.2B. Code is available at: this https URL.
75. 【2609.30298】A Benchmark Framework for Screening Automation in Systematic Reviews
链接:https://arxiv.org/abs/2609.30298
作者:Gauransh Kumar,Luciano Marchezan,Guillaume Genois,Kévin Delcourt,Eugene Syriani
类目:Computation and Language (cs.CL)
关键词:Systematic reviews, evidence-based research, time-consuming and labor-intensive, essential for evidence-based, Systematic
备注:
点击查看摘要
Abstract:Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening this http URL paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening. We also present a use case demonstrating the application of SRBench and PromptSR.
76. 【2609.30297】Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
链接:https://arxiv.org/abs/2609.30297
作者:Enrico Palumbo,Alexandre Tamborrino,Victor Ode,Ben Lacker,Adrià Casas Escoda,Jeremy Hopple,Marcus Better,James Leoni,Hugo Galvão,Hugues Bouchard,Mounia Lalmas,José Luis Redondo García,Abenezer Abebe,Ann Clifton,Anton Blomberg,Henrik Lindström,Dani Doro,Christine Doig Cardet
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:recommend Italian indie, Italian indie artists, express complex intents, recommend Italian, Italian indie
备注:
点击查看摘要
Abstract:Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a self-improvement loop to address this challenge. The synthetic data pipeline transforms single-turn prompts into realistic multi-turn conversations, enabling systematic evaluation before launch. The self-improvement loop combines variance-based contrastive optimization with iterative refinement through a coding agent, automatically identifying and fixing planning and tool-use errors. Our approach improves quality by +8% on top of a highly optimized manual prompt. The system has been productionized and significantly accelerated iteration cycles for the launch of a conversational recommendation agent at Spotify. Online A/B tests demonstrate its effectiveness, with +14% user listening, +5% increase in weekly active users, and a 5% reduction in skip rate compared to a prior experience supporting only session refinement. This work provides a practical framework for accelerating the development of conversational recommendation agents in industry.
77. 【2609.30295】SignTrace: Describe a Sign, Find the Word
链接:https://arxiv.org/abs/2609.30295
作者:Zengji Tu,Xingye Zhu,Ningjing Wang,Tingyi Huang,Yangjunfeng Zhu,Dai Wan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:formal feature codes, feature codes, learner remembers, formal feature, Chinese sign-language dictionary
备注: Includes reproducibility data and method documentation
点击查看摘要
Abstract:Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access to a Chinese sign-language dictionary. The system integrates LLM-based dictionary enrichment, action extraction, dictionary-style rewriting, seven-channel retrieval, and candidate reranking over 6,699 entries. It has been deployed for user trials and has received positive informal feedback. Evaluation on a dictionary-derived benchmark of 500 movement-description queries yields 94.0% Hit@1, 97.4% Hit@9, and a mean reciprocal rank of 0.9540. Reranking increases Hit@1 from 71.8% to 94.0%, while component analyses show the contribution of enriched entry descriptions. Median query-processing time is 13.37 seconds with six concurrent queries. By connecting everyday movement descriptions to documented signs and meanings, SignTrace provides a practical tool for identifying unfamiliar signs. Dictionary-derived wording and prior selection within the benchmark limit generalization to descriptions independently produced by users.
78. 【2609.30294】SlideLab: Audience-Centered Scientific Slide Generation and Evaluation
链接:https://arxiv.org/abs/2609.30294
作者:Vidushee Vats,Karun Sharma,Yuxia Wang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Scientific presentations, research papers, generating scientific presentations, Scientific, Abstract
备注: V1
点击查看摘要
Abstract:Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open-source baseline. We also introduce ConfArena, an audience-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide. ConfArena matches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order.
79. 【2609.30293】Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
链接:https://arxiv.org/abs/2609.30293
作者:Justice Owusu Agyemang,Michael Agyare,Kwame Opuni-Boachie Obour Agyekum,Kwame Agyeman-Prempeh Agyekum,Francisca Adoma Acheampong,Jerry John Kponyo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Model Context Protocol, Context Protocol, Model Context, connected catalogs grow, enables AI agents
备注:
点击查看摘要
Abstract:The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to $O(k)$ progressive disclosure. Cartograph combines three mechanisms: (1) operator-attested capability cards, Ed25519-signed descriptions generated under the deploying operator's control rather than ranked publisher copy; (2) Rift, a three-layer confusable-cluster analysis comprising density clustering, query-margin analysis, and token diagnosis; and (3) two-stage retrieval, which ranks servers before tools. On a 22-server, 374-tool deployment, Cartograph exposes three proxy tools instead of 374 definitions. A 49-query author-constructed benchmark yields R@5 of 0.816, compared with 0.592 for a Jaccard keyword baseline, while a measured top-5 discovery exchange uses 475 tokens rather than 42,450 under the stated full-catalog accounting. Rift identifies 49 confusable clusters, including four HIGH-risk clusters in bootstrap-generated cards. An exploratory comparison of 119 LLM-generated descriptions removes the observed zero-distance cluster but shows that mixing card-generation regimes can reduce R@5. Gateway measurements over ten trials add 5ms mean latency (0.8%) relative to direct stdio MCP calls. Cartograph is complementary to code-execution approaches: it controls which tool descriptions are surfaced and records the provenance of the descriptions used for ranking for each query.
80. 【2609.30292】A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models
链接:https://arxiv.org/abs/2609.30292
作者:Fanji Yang(1),Huiyao Chen(2),Xi Yu(1),Meishan Zhang(2),Xiaohong Xiao(3),Mingsen Deng(1) ((1) Guizhou University of Finance and Economics, (2) Harbin Institute of Technology (Shenzhen),(3) Guizhou University of Commerce)
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:http URL rise, http URL survey, http URL organize, http URL trace, http URL reviews
备注: Fanji Yang and Huiyao Chen contributed equally to this work. Accepted for publication in Information Fusion
点击查看摘要
Abstract:Online reviews shape consumer decisions, platform governance, and corporate this http URL reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust this http URL rise of large language models, or LLMs, has changed the problem in two this http URL can generate fluent and context-aware deceptive reviews, while pre-trained language models, or PLMs, and LLMs also provide stronger semantic representations for this http URL survey reviews fake review detection from an information fusion perspective, covering 211 studies published from 2018 to early this http URL organize existing work by evidence source and fusion level, covering review text, sentiment, rating behavior, temporal metadata, user-product graphs, multimodal content, external knowledge, and LLM-generated this http URL trace the development from traditional machine learning and deep learning to PLM-based and LLM-based methods, and examine how different approaches combine textual, behavioral, structural, and multimodal this http URL also analyze reported performance trends on widely used Amazon, Yelp, and OpSpam benchmark families, while noting the limitations caused by different label construction procedures, data splits, and evaluation this http URL, we identify open problems in adversarial generation, cross-domain transfer, uncertainty-aware fusion, missing-source robustness, interpretability, and trustworthy evaluation for AI-generated deceptive content.
81. 【2609.30290】Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
链接:https://arxiv.org/abs/2609.30290
作者:Haowei Liu,Hsin-Tai Wu,Yi Fang
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:pipelines often end, human annotators, Production, Cohen kappa, kappa
备注:
点击查看摘要
Abstract:Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B replacement (kappa = 0.72) lands in the same range as Claude Opus 4.7 (kappa = 0.71); the head-to-head is underpowered at n = 96, but for the deployment decision that hardly matters, since Qwen costs roughly 1/300 as much per call. Ensembling does not help for free. Pairing the weak judge with a stronger one degrades agreement, whereas three strong judges under unanimity routing reach kappa = 0.79 at 89.7% auto-coverage. Applied out-of-domain, the same audit recipe flags 25.5% of BIRD-financial's expert-authored gold SQLs as candidate gold-SQL issues under our annotation protocol. Code and pre-registration are at this https URL.
82. 【2609.30289】Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
链接:https://arxiv.org/abs/2609.30289
作者:Yufei Shi,Rujing Yao,Ang Li,Yang Wu,Zhuoren Jiang,Xiaozhong Liu
类目:Computation and Language (cs.CL)
关键词:team collaboration scenarios, memories, collaboration scenarios, heterogeneous and continually, current team consensus
备注:
点击查看摘要
Abstract:In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations, execution traces, and intermediate progress. Existing memory-augmented systems typically retrieve from all stored memories as a flat pool, ranking them by semantic relevance, importance, or recency without modeling hierarchical structure or evolving validity. As a result, they often surface semantically relevant but outdated or conflicting memories, especially individual memories that no longer align with current team consensus, instead of prioritizing currently valid memories. This is particularly problematic when collaborative LLM agents answer user questions, since their responses should be grounded in valid memories. We propose HiCoMER, a framework for hierarchical collaborative memory management and validity-aware retrieval in LLM agents. HiCoMER first maintains the validity of team and individual memories and then retrieves memories that remain valid, rather than retrieving directly from all stored memories. It consists of three components: a Hierarchical Memory Conflict Updater, a Validity-Aware Memory Retriever, and a Memory-Grounded Answer Generator. To evaluate HiCoMER, we construct two new datasets for memory-grounded question answering in collaborative settings. Experiments on both datasets show that HiCoMER consistently outperforms strong baselines by reducing outdated retrieval, preserving current team consensus, and improving downstream QA quality.
83. 【2609.30288】Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
链接:https://arxiv.org/abs/2609.30288
作者:Narges Mokhtari,Farzan Haddadi,Ebrahim Rezaii
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Transformer-based masked language, primary mechanism, mechanism for context, mix data, Transformer-based masked
备注: Submitted to IEEE Transactions on Emerging Topics in Computational Intelligence (TETCI)
点击查看摘要
Abstract:In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.
84. 【2609.30287】A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
链接:https://arxiv.org/abs/2609.30287
作者:Paweł Blicharz,Miłosz Grunwald
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:remain poorly understood, AI-generated text detectors, predictions remain poorly, achieve high accuracy, text detectors achieve
备注: Accepted to EMNLP 2026 (Main Conference)
点击查看摘要
Abstract:AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy. Bidirectional activation patching confirms this set's causal relevance: in both directions it flips predictions an order of magnitude more often than size-matched random sets. Mean-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed. Cross-generator analysis reveals a bipartite structure: instruction-tuned generators concentrate 30-36% of stable neurons in BERT's final layer while both base generators fall below 14%, consistent with a layer-12 footprint of post-training alignment. Leave-one-family-out evaluation shows the selected neurons retain 86-94% of the full-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re-identifying neurons per generator.
85. 【2609.31547】wo Conformal Constructions for Adaptive Within-Document AI-Text Screening
链接:https://arxiv.org/abs/2609.31547
作者:Marco Mandap,Jerahmeel Hipolito,Arcel Galvez,Charlie Margaret Balagtas,Michael Joshua Buluran,Jeff Roel Durmiendo,Rizzette E. Lopez
类目:Methodology (stat.ME); Computation and Language (cs.CL)
关键词:artificial intelligence, text generated, generated by artificial, study false-alert control, study false-alert
备注: 16 pages, 0 figures; theoretical manuscript; no empirical evaluation
点击查看摘要
Abstract:We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.
86. 【2609.31513】Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
链接:https://arxiv.org/abs/2609.31513
作者:Marco Mandap
类目:Methodology (stat.ME); Computation and Language (cs.CL); Applications (stat.AP)
关键词:Google Play user, Play user reviews, statistically explicit sentiment, explicit sentiment index, Google Play
备注: 16 pages, 2 tables, no figures
点击查看摘要
Abstract:We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count threshold. App-level rating histograms provide a distributional diagnostic for samples returned under different API sort orders; because star ratings are discrete, classical continuous Kolmogorov-Smirnov critical values are not used. A local-level state-space model and the Kalman filter provide a denoised temporal trend. Full proofs cover the BLUE and Gaussian maximum-likelihood result, Gaussian-conjugate shrinkage, the Glivenko-Cantelli and Donsker theorems, count transformations via the delta method, and exact Gaussian Kalman filtering. A worked three-review example shows how textual complaints can materially reduce an apparently perfect star-only score.
87. 【2609.31468】PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
链接:https://arxiv.org/abs/2609.31468
作者:Pavel Kireyev
类目:General Economics (econ.GN); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:LLMs increasingly act, preferences quietly fix, satisfy a request, increasingly act, act as purchasing
备注: Accepted to EMNLP 2026 Industry Track. 19 pages, 10 figures, 6 tables. Code and data: [this https URL](https://github.com/Pashasan/pricebench-emnlp)
点击查看摘要
Abstract:LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.
88. 【2609.31293】Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring
链接:https://arxiv.org/abs/2609.31293
作者:Zijian Lu,Sizhe Liu,Yin Zhang,Jixuan Deng,Xinrong Lin,Xinchen Yuan,Chicheng Jin,Yiping Zuo,Yuanchao Li
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
关键词:detecting Alzheimer disease, detecting Alzheimer, Alzheimer disease, non-invasive approach, Speech-based screening
备注: Accepted to NCMMSC 2026
点击查看摘要
Abstract:Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across corpora, with pause, silence, and speech rate showing high protocol sensitivity. Furthermore, while the XLM-R text baseline achieves strong average performance, its Area Under the ROC Curve (AUC) drops to 0.520 on the weakest held-out domain. A standard GroupDRO baseline reaches a 0.766 mean speaker AUC and a 0.504 worst-domain AUC under the same protocol. To address this, we propose a fusion method that integrates XLM-R text baseline scores with evidence anchors selected during training. Balanced fusion achieves a 0.785 mean speaker AUC, while anchor-heavy fusion raises the worst-case speaker AUC to 0.615. This work highlights the need to audit feature transferability and report worst-case domain robustness in cognitive speech screening.
89. 【2609.30476】Asymmetric Classifier-Free Guidance for Target-Speaker ASR
链接:https://arxiv.org/abs/2609.30476
作者:Yiwen Guan,Jacob Whitehill
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
关键词:Target-speaker automatic speech, Target-speaker automatic, automatic speech recognition, noise conditions, identify and transcribe
备注:
点击查看摘要
Abstract:Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.
信息检索
1. 【2609.31498】Retail Product Search: A Practical Approach at Target
链接:https://arxiv.org/abs/2609.31498
作者:Darshan Sonagara,Qujiaheng Zhang,Ankit Singh,Alex Li
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:directly driving customer, driving customer engagement, features in e-commerce, directly driving, business growth
备注: 10 pages, 2 figures, 6 tables
点击查看摘要
Abstract:Search is one of the most important features in e-commerce, directly driving customer engagement and business growth. A good product search system must show both relevant and desirable results. However, retail search presents unique challenges. User intent can range from exact matches to open-ended discovery. Search systems must also balance multiple goals, such as relevance, revenue, and profit, while keeping response times low. Traditional keyword-based methods often fall short in handling natural language or semantic queries. Vector search helps alleviate these issues, but it can miss key intent signals or return low-precision results. In this paper, we present the design of a hybrid search system at Target that combines lexical and vector search. We describe our approach to data processing, embedding training, precision control for the final result set, multi-channel result fusion (where we compared fusion strategies and adopted weighted interleaving), and the performance optimizations used to maintain low latency for production deployment. Our method improves offline evaluation metrics, and in online A/B testing it raised click-through rate by 0.97%, order conversion by 0.98%, and demand per visitor by 1.10% over lexical-only search, while roughly halving zero-result searches. The resulting system is deployed at scale and serves millions of guests daily.
2. 【2609.31253】Enriching Sequential Recommendation with Graph Laplacian Positional Embeddings
链接:https://arxiv.org/abs/2609.31253
作者:Ekaterina Trushkova,Artur Gimranov,Anton Lysenko
类目:Information Retrieval (cs.IR)
关键词:recommenders typically rely, Sequential recommenders typically, recommenders typically, typically rely, rely on learnable
备注: CIKM 2026
点击查看摘要
Abstract:Sequential recommenders typically rely on learnable positional embeddings to encode the order of user interactions. In this work, we ask whether this ordinal signal can be replaced by a structural one derived from the item space. We propose to use Laplacian positional embeddings in SASRec: we build an item co-occurrence graph from training interactions, compute eigenvectors of its symmetric normalized Laplacian, and use them as frozen graph-derived positional embeddings. The backbone architecture and training objective remain unchanged. Experiments on four public sequential-recommendation benchmarks show that this simple replacement improves SASRec performance on most ranking metrics and remains competitive with strong positional and temporal encoding baselines. These findings indicate that item-item graph structure can be an effective substitute for standard ordinal positional embeddings in sequential recommendation.
3. 【2609.31166】AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side
链接:https://arxiv.org/abs/2609.31166
作者:Ryoma Sato
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Databases (cs.DB); Digital Libraries (cs.DL)
关键词:Recommender systems, Recommender, traditionally been developed, user-side recommender systems, user-side recommender
备注:
点击查看摘要
Abstract:Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recommender systems have been proposed as a new paradigm for solving this problem. If users deploy their own recommender systems, they are no longer at the mercy of the platform's interests. However, building a user-side recommender system is not trivial; in particular, customizing one for oneself requires additional data. We propose AgentRecommender, a method that leverages the investigation capability and internal knowledge of LLM agents to flexibly build user-side recommender systems without additional data. AgentRecommender allows users to easily create recommender systems tailored to their own preferences.
4. 【2609.31164】SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations
链接:https://arxiv.org/abs/2609.31164
作者:Tobias Vente,Maarten Peirsman,Noah Daniëls,Hannu Toivonen,Bart Goethals
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Recommender systems engineer, predictable consumption cycles, foster active exploration, break predictable consumption, Recommender systems
备注:
点击查看摘要
Abstract:Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.
5. 【2609.31062】CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
链接:https://arxiv.org/abs/2609.31062
作者:Marko Řeháček,Vítězslav Dušek,Martin Rusinko,Vít Nováček
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:assistants promise valuable, promise valuable support, Patient-facing AI assistants, pose medical risks, Medical Urgency
备注: Accepted as a short paper at CIKM '26 (35th ACM International Conference on Information and Knowledge Management), Rome, Italy. 7 pages, 1 figure, 2 tables. Code and benchmark: [this https URL](https://github.com/mrehacek/cg-probes)
点击查看摘要
Abstract:Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential linear direction. To train the probes, we cluster 79,658 Czech oncology search queries with BERTopic and use these clusters to generate pairs of queries with contrastive risk levels via few-shot prompting. We evaluate the approach on 200 queries (90 real, 110 synthetic), each graded by two oncologists, against two open-weight LLMs and a frontier LLM. We find that urgency-based axes are recoverable as linear directions, and the probes are competitive with open-weight LLMs (no significant differences in quadratic-weighted kappa) at a fraction of the latency. Each axis yields a scalar score that clinicians can inspect and use to set escalation thresholds. The pipeline requires only search logs, axis definitions, and black-box access to the embedding model, suggesting transferability across healthcare domains. Robust validation on new queries and axes remains future work.
6. 【2609.31045】KuaFu: Compressing Long User Behavior into Understanding at Billion Scale
链接:https://arxiv.org/abs/2609.31045
作者:Jiahao Hui,Lin Zhu,Yishen Hu,Jingdong Shu,Zetai Jiang,Xining Ran,Ben Tan,Yeshou Cai,Gong Chen,Haijie Gu,Jie Jiang
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Conversational agents, generative recommenders, Conversational, raw behavior, Prevailing industrial practice
备注: 12 pages, 6 figures, 3 tables
点击查看摘要
Abstract:Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.
7. 【2609.30904】QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking
链接:https://arxiv.org/abs/2609.30904
作者:Yang Zhang,Wenhan Liu,Qiannan Zhu,Mingming Li,Yuanfei Huang
类目:Information Retrieval (cs.IR)
关键词:Passage reranking plays, plays a crucial, crucial role, role in information, information retrieval
备注: EMNLP2026 Main
点击查看摘要
Abstract:Passage reranking plays a crucial role in information retrieval by refining the ordering of candidate passages to better reflect relevance. Existing listwise LLM rerankers with Chain-of-Thought (CoT) reasoning can handle complex queries effectively, but they suffer from substantial redundancy and high latency due to sliding-window strategies, which repeatedly generate highly similar CoTs. To address this, we propose QReason, a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment. Specifically, QReason introduces a dedicated rewriter that generates a ranking-oriented reasoning query once, capturing the query's core intent while avoiding redundant reasoning, and then reuses it across all windows with a non-reasoning reranker. The rewriter is trained via a two-stage process that first uses supervised fine-tuning with relevant-passage guidance through semantic evidence to produce deeply grounded, query-focused CoTs. It then applies reinforcement learning to align CoT generation with both the inference-time setting and the reranking objective, optimizing listwise metrics and passage-level discrimination to produce reusable reasoning chains for reranking. Experiments on the BRIGHT benchmark demonstrate that QReason significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.
8. 【2609.30717】RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent
链接:https://arxiv.org/abs/2609.30717
作者:Xiao Chen,Yicheng Zhao,Yingying Wu,Zhendong Chu,Changyi Ma,Qingsong Wen,Xuan Song
类目:Information Retrieval (cs.IR)
关键词:passive filtering engines, Recent advances, resolve user intent, user intent, Model Context Protocol
备注: EMNLP 2026
点击查看摘要
Abstract:Recent advances in agentic recommender systems are shifting recommender systems from passive filtering engines to instruction-following agents that use external tools to resolve user intent. However, existing benchmarks often assume explicit user intent, simplified tool environments, or isolated function calls, leaving realistic tool orchestration for recommendation underexplored. To bridge this gap, we propose RecToolBench, a Model Context Protocol (MCP)-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. RecToolBench contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel tool calls, sequential tool chains, and hybrid tool orchestration. We construct RecToolBench with a scalable synthesize--fuzzify--judge pipeline that generates executable fuzzy recommendation tasks, and evaluates agent trajectories using rule-based execution checks and rubric-based LLM evaluation. Experiments on representative LLMs show that syntactically valid tool calls do not guarantee successful recommendations. Models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations, especially as orchestration complexity increases. Our results identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems. Our data and code are available at this https URL.
9. 【2609.30711】Recommendation World Models for Future-State Control
链接:https://arxiv.org/abs/2609.30711
作者:Jinfeng Xu,Zheyu Chen,Ziyue Peng,Jianheng Tang,Zheng Lin,Jing Yang,Puzhen Wu,Zheng Xing,Victor C. M. Leung
类目:Information Retrieval (cs.IR)
关键词:Sequential recommendation optimizes, shapes subsequent feedback, items to rank, user state, recommendation optimizes
备注:
点击查看摘要
Abstract:Sequential recommendation optimizes which items to rank, while each displayed slate also shapes subsequent feedback and user state. We study how a trained ranker can support decisions about these future consequences. We introduce UA-TWM, a utility-anchored world-model interface that constructs nearby slate actions, estimates their target-relevant consequences, and selects an alternative subject to utility constraints. The reference slate serves as a fallback when no alternative qualifies. A logged-replay instantiation combines utility and target-gain estimates with calibrated failure-risk prediction; a closed-loop instantiation uses one-step state-action prediction and updates its decisions after observed feedback. We evaluate transfer across twelve sequential backbones on MovieLens-25M and KuaiRand-Pure, and repeated target-directed interaction in KuaiSim. Attaching the interface improves Recall@20, NDCG@20, and future-state alignment for every matched logged backbone. Selection ablations reveal the utility and risk costs of aggressive target pursuit, while closed-loop diagnostics isolate the contribution of action-conditioned prediction. Local consequence modeling thus enables target-aware selection around a trained sequential ranker.
10. 【2609.30656】Component Benchmark: Hierarchical Model Profiling for Large-scale Recommendation Systems
链接:https://arxiv.org/abs/2609.30656
作者:Dharak Kharod,Yuzhen Huang,Zhou Wang,Jackie Xu,Fuzail Khan,Jacky Zhou,Hao Yan,Lidong Zhao,Xizhou Feng,Yvonne Liu,Karthik Jayaraman,Praveen Ramachandran,Vishwa Karia,Yashasvi Makin
类目:Information Retrieval (cs.IR)
关键词:models pose distinct, under-explored profiling challenges, pose distinct, Large-scale recommendation models, recommendation models pose
备注:
点击查看摘要
Abstract:Large-scale recommendation models pose distinct, under-explored profiling challenges. Most recommendation model architectures are structurally heterogeneous, intermixing memory-bandwidth-bound operations, small compute-bound dense layers, dynamic shapes from jagged categorical features, and low-arithmetic-intensity operations. Recommendation models evolve rapidly as modeling engineers experiment with compositions, often written without visibility into hardware execution characteristics. Standard profiling tools offer either end-to-end throughput or operator-level traces, but cannot attribute performance to the submodules that practitioners reason about. We present Component Benchmark (CB), a profiling system that independently characterizes each submodule performance in a hierarchical manner, providing a tree-structured, interactive visualization that brings performance clarity to ML practitioners. At its core, CB provides a simple yet extensible, submodule-based benchmarking framework with a plugin architecture that enables hierarchical performance analysis. These large-scale recommendation models are TB-scale, run on thousands of GPUs and ingest 100B examples per day. We demonstrate CB's effectiveness on common open sourced models and discuss how CB has been leveraged to accelerate modern recommendation model performance analysis and optimization.
11. 【2609.30611】Epstein Files Engine: Agentic Search for Investigative Journalism
链接:https://arxiv.org/abs/2609.30611
作者:Duy K. Nguyen,Teresa Mondría Terol,Dylan Freedman,Zach Seward
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR)
关键词:Department of Justice, Epstein Files Engine, Jeffrey Epstein, Justice released, collection concerning Jeffrey
备注: 6 pages, 2 figures, 2 tables. Presented at the Computation + Journalism Symposium (C+J 2026)
点击查看摘要
Abstract:On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questions into Google BigQuery SQL queries across three corpora: Epstein-related releases, the Times's archive and external, Epstein-related news headlines. It used an LLM to plan queries and returned citation-rich answers a reporter could verify and trust. More than 100 journalists used the Engine, and it contributed to at least 20 published stories. We report how reporters queried it and describe Diff, our text-and-visual duplicate matching method that amplified novelty signals and allowed the Engine to surface genuinely new information. We argue that newsroom agents serve newsrooms best not as autonomous writers, but as interfaces to source material and institutional knowledge.
12. 【2609.30601】Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval
链接:https://arxiv.org/abs/2609.30601
作者:Shaobo Zhang,Alice Leung,Yunxiang Ren,Ping Liu,Yuchin Juan,Qianqi Shen,Benjamin Le,Jianqiang Shen,Chengming Jiang,Ko-Cheng Wang,Vidya Krishnamurthy,Caleb Johnson,Fedor Borisyuk,Luke Simon,Jingwei Wu,Wenjing Zhang
类目:Information Retrieval (cs.IR)
关键词:Modern industrial recommender, balancing semantic relevance, industrial recommender systems, Modern industrial, balancing semantic
备注: 10 pages. To appear in the 20th ACM Conference on Recommender Systems (RecSys 2026)
点击查看摘要
Abstract:Modern industrial recommender systems must optimize across competing objectives, balancing semantic relevance with business metrics such as engagement and revenue. While bi-encoders dominate large-scale retrieval due to their efficiency, they collapse these heterogeneous signals into a single static embedding space. This design creates a fundamental limitation: once trained, the retriever cannot adapt to shifting objective priorities at serving time without retraining. Moreover, joint optimization with multi-objective losses often induces interference between objectives, leading to suboptimal trade-offs. We propose Embedding Subspace Partitioning (ESP), a retrieval framework that decomposes the embedding into task-aware subspaces and replaces the single dot product with a weighted sum of per-subspace similarities, whose weights are tunable at serving time. For Transformer bi-encoders, ESP uses the model's native end-of-sequence token as a segment delimiter, with segment-aware attention masking and position encoding resets to guarantee subspace isolation in a single forward pass. Serving is performed via GPU-accelerated exhaustive kNN over one concatenated index, eliminating the need for per-objective Approximate Nearest Neighbor (ANN) infrastructure required by multi-head approaches. We evaluate ESP on an open-source benchmark built from MS MARCO. A single ESP model traces a broad Pareto frontier, consistently outperforming strong multi-task baselines across diverse operating points. In LinkedIn's job matching platform (70M+ weekly users), ESP enabled dynamic retrieval reconfiguration and delivered significant key business metric lifts.
13. 【2609.30576】-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
链接:https://arxiv.org/abs/2609.30576
作者:Yang Liu,Noel Loo,Ali Khanafer,Shuying Sun,Akshay Soni,Zhong Wu,Linjun Yang
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:Rotary Position Embedding, including Rotary Position, Large-scale recommenders increasingly, bringing the Transformer, including Rotary
备注:
点击查看摘要
Abstract:Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and propose T-RoPE, a time-aware RoPE for sequential generative recommendation that replaces index-only rotation with timestamp-based angles, learnable temporal coefficients, multiscale frequency banks, shifted query alignment, and non-stationary key rotation. We prove that standard RoPE, even on timestamps, remains time-translation invariant and cannot distinguish seasonal contexts, and that T-RoPE breaks this invariance while preserving the RoPE interface. Across five public benchmarks, T-RoPE achieves the best result on every metric on every dataset, improving over the strongest baseline by 78--130\% in HR@10 on the sparse PixelRec data and 8--12\% across metrics on Amazon Books. On an industrial-scale e-commerce dataset with more than 6B interactions, it improves every metric over the HSTU + Time RAB backbone by 13--82\%, with ablations attributing the largest gains to multiscale frequencies ($+56\%$ NDCG@50) and non-stationary keys ($+4\%$). An online A/B test in the Shop app yields positive lifts in conversion rate ($+0.33\%$) and order count ($+0.63\%$). We also provide forward and backward algorithms whose added cost is linear in sequence length and head dimension, keeping time-aware RoPE practical for large generative recommenders.
14. 【2609.30568】Nearest but Not Dearest: Shared Curator-Feedback Infrastructure for Content-Only Search and Recommendation
链接:https://arxiv.org/abs/2609.30568
作者:Matt Sandler
类目:Information Retrieval (cs.IR); Sound (cs.SD)
关键词:LAION-CLAP joint audio-text, music-discovery platform serves, path consumes end-listener, consumes end-listener behavioral, end-listener behavioral signal
备注: 8 pages, 3 figures, 3 tables. Accepted for oral presentation at the Unified Search and Recommendation Workshop (USRW) at RecSys 2026; workshop is non-archival
点击查看摘要
Abstract:A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generation filter, and one ranking head -- and neither path consumes end-listener behavioral signal. In this content-only regime, curator judgment is the principal feedback signal available, and offline cosine similarity predicts it poorly: 38% of cosine-nearest neighbors are rejected by curators. The rejections reveal a clean partition: a majority (55%) are sound failures the encoder could address (style, tempo, mood mismatch), and a substantial minority (37%) are context failures orthogonal to the waveform (wrong language, holiday content, devotional content, rights and lyric flags). We deploy this sound-vs-context decomposition as feedback infrastructure, routing each failure mode to the layer that can absorb it: context failures to a constraint filter at candidate generation, sound failures to an embedding reweighting head at the representation layer -- both below the search/recommendation split, so a single curator loop maintains both experiences. On 1,200 curator judgments collected over two production rounds one month apart, the combined intervention reduces rejection rate from 38.17% to 28.83% (-24.5% relative, McNemar chi-squared = 22.4, p = 2.2e-6). An accounting decomposition attributes 4.08 pp of the drop to filter-eligible categories and 5.25 pp to the rest; the deployment was unblinded and compound, so this is a production accounting bound, not a causal estimate. We present this as an industrial case study rather than a validated general method, and close with lessons for content-only discovery: the failure partition is orthogonal to the paradigm partition, and corrections land in layers shared by both paradigms.
15. 【2609.30547】REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
链接:https://arxiv.org/abs/2609.30547
作者:Haixu Ma,Aditya Bansal,Shubham Lohiya,Sumit Ranjan
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:digital marketing, Exact Audience sizing, Audience sizing, Exact Audience, Audience
备注: Accepted by ICDM 2026
点击查看摘要
Abstract:Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant delays, estimation errors, and poor scalability over high-dimensional profile data. We present REALMS (Real-time Exact Audience sizing via LLM-based Multi-attribute Search), a conversational system for exact audience sizing deployed in production on an enterprise customer data platform. REALMS enables marketers to query massive profile stores with millions of profiles and thousands of attributes using natural language and receive precise counts in seconds. The system introduces three key components: (1) a categorical attribute retrieval mechanism using embedding-based vector search to dynamically identify relevant schema attributes without manual configuration; (2) an LLM-powered NL2SQL pipeline with template-based in-context learning for accurate query generation over complex nested schemas; and (3) schema standardization enabling industry-agnostic deployment across diverse enterprise environments. Evaluation on real enterprise data demonstrates strong recall for attribute retrieval, high SQL execution accuracy, and low latency, which enables real-time interactive audience insights where prior methods required hours.
16. 【2609.30541】AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
链接:https://arxiv.org/abs/2609.30541
作者:Aparajith Chandran,Juwon Kim,Saurav Jha,Pablo Castells,Florian Hottier
类目:Machine Learning (cs.LG); Information Retrieval (cs.IR)
关键词:Optimizing embedding systems, disproportionate engineering effort, demands systematic exploration, consumes disproportionate engineering, Optimizing embedding
备注: 10 pages, 3 figures. Accepted at IEEE ICDM 2026 (Applied Track)
点击查看摘要
Abstract:Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.
17. 【2609.30467】Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
链接:https://arxiv.org/abs/2609.30467
作者:Heyuan Huang,Jirui Dai,Alexandra DeLucia,Sonal Joshi,Mahsa Yarmohammadi,Jie Gao,Bernal Jiménez Gutiérrez,Mark Dredze
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Retrieval-based factuality evaluation, scalable hallucination detection, Retrieval-based factuality, high-stakes clinical settings, authoritative medical corpora
备注: Experiments' corpus knowledge cutoff date May 2026
点击查看摘要
Abstract:Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at this https URL for the full reproducibility of our results.
18. 【2609.30297】Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
链接:https://arxiv.org/abs/2609.30297
作者:Enrico Palumbo,Alexandre Tamborrino,Victor Ode,Ben Lacker,Adrià Casas Escoda,Jeremy Hopple,Marcus Better,James Leoni,Hugo Galvão,Hugues Bouchard,Mounia Lalmas,José Luis Redondo García,Abenezer Abebe,Ann Clifton,Anton Blomberg,Henrik Lindström,Dani Doro,Christine Doig Cardet
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:recommend Italian indie, Italian indie artists, express complex intents, recommend Italian, Italian indie
备注:
点击查看摘要
Abstract:Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a self-improvement loop to address this challenge. The synthetic data pipeline transforms single-turn prompts into realistic multi-turn conversations, enabling systematic evaluation before launch. The self-improvement loop combines variance-based contrastive optimization with iterative refinement through a coding agent, automatically identifying and fixing planning and tool-use errors. Our approach improves quality by +8% on top of a highly optimized manual prompt. The system has been productionized and significantly accelerated iteration cycles for the launch of a conversational recommendation agent at Spotify. Online A/B tests demonstrate its effectiveness, with +14% user listening, +5% increase in weekly active users, and a 5% reduction in skip rate compared to a prior experience supporting only session refinement. This work provides a practical framework for accelerating the development of conversational recommendation agents in industry.
19. 【2609.30295】SignTrace: Describe a Sign, Find the Word
链接:https://arxiv.org/abs/2609.30295
作者:Zengji Tu,Xingye Zhu,Ningjing Wang,Tingyi Huang,Yangjunfeng Zhu,Dai Wan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:formal feature codes, feature codes, learner remembers, formal feature, Chinese sign-language dictionary
备注: Includes reproducibility data and method documentation
点击查看摘要
Abstract:Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access to a Chinese sign-language dictionary. The system integrates LLM-based dictionary enrichment, action extraction, dictionary-style rewriting, seven-channel retrieval, and candidate reranking over 6,699 entries. It has been deployed for user trials and has received positive informal feedback. Evaluation on a dictionary-derived benchmark of 500 movement-description queries yields 94.0% Hit@1, 97.4% Hit@9, and a mean reciprocal rank of 0.9540. Reranking increases Hit@1 from 71.8% to 94.0%, while component analyses show the contribution of enriched entry descriptions. Median query-processing time is 13.37 seconds with six concurrent queries. By connecting everyday movement descriptions to documented signs and meanings, SignTrace provides a practical tool for identifying unfamiliar signs. Dictionary-derived wording and prior selection within the benchmark limit generalization to descriptions independently produced by users.
计算机视觉
1. 【2609.31620】FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
链接:https://arxiv.org/abs/2609.31620
作者:Hongyang Du,Yunfei Xie,Junjie Ye,Jiawei Yang,Xiaoyan Cong,Haodong Zhang,Yongchao Huang,Haiyu Wu,Zongxia Li,Shihang Gui,Dawei Liu,Runhao Li,Jingcheng Ni,Chen Wei,Randall Balestriero,Yue Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:strong visual representations, integrating strong visual, Representation autoencoders, visual representations, integrating strong
备注:
点击查看摘要
Abstract:Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.
2. 【2609.31595】GraphWrit3R: End-to-End 3D Scene Graph Writing
链接:https://arxiv.org/abs/2609.31595
作者:Luka Milivojevic,Nikola Popovic,Sayan Deb Sarkar,Sebastian Koch,Iro Armeni,Luc Van Gool,Danda Pani Paudel
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:environments by encoding, spatial and functional, scene graph, scene graph generation, complex environments
备注: Project page at [this https URL](https://graphwrit3r.insait.ai)
点击查看摘要
Abstract:3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.
3. 【2609.31573】How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI
链接:https://arxiv.org/abs/2609.31573
作者:Ziyao Shang,Pouya Sadeghi,Letian Jiang,Alexander Wong,Sirisha Rambhatla
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Biomedical image segmentation, faces limited annotations, cross-site distribution shifts, Biomedical image, medical image analysis
备注: 26 pages, 15 figures
点击查看摘要
Abstract:Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight alternative for semantic segmentation, achieving competitive performance with substantially fewer parameters than conventional architectures. However, the mechanisms, scaling behavior, and domain generalization abilities of INR-based segmentation remain insufficiently understood. In this work, we study these questions in the context of cross-domain brain MRI segmentation. We analyze INR-based segmentation across low-parameter regimes, comparing it with conventional pipelines in both in-domain and out-of-domain settings. Surprisingly, we find that INR-based models do not simply improve with increasing parameter budget. Their advantage is most pronounced under low-parameter and limited-augmentation settings, while U-Net-based models benefit more from larger capacity and standard augmentation. We also investigate how INRs encode semantic information in their hidden features and show that complementary segmentation-relevant structure is distributed across multiple INR layers. Building on this insight, we introduce HierINRSeg, a hierarchical INR-based architecture that aggregates multi-layer representations for improved robustness and generalization. Extensive experiments show that HierINRSeg consistently outperforms MetaSeg, a strong recent INR-based segmentation baseline, with an average improvement of 5.6 percentage points in Dice for the in-domain test set and 8.2 percentage points out-of-domain. Overall, our analysis identifies the conditions under which INR-based segmentation is most effective, providing concrete guidance for model selection and future research.
4. 【2609.31572】OC-GS: Gaussian Splatting for Irregular Turntable Capture
链接:https://arxiv.org/abs/2609.31572
作者:Jae Joong Lee,Bedrich Benes
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:dropped frames make, frames make equal-angle, make equal-angle assumptions, equal-angle assumptions unreliable, Uneven rotation
备注:
点击查看摘要
Abstract:Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores of 21.26, 19.36, and 15.83dB, respectively, exceeding all four evaluated pose-free Gaussian splatting baselines in each condition. Under a shared trainer, refining image-estimated angles improves mean foreground PSNR by 7.88dB over keeping those estimates fixed. An ablation study shows that both image-derived angle initialization and the shared motion model contribute to the improvement. On real captures, OC-GS's refinement increases mean foreground PSNR by 0.70dB. Results show that refining uncertain angles within a shared motion model improves reconstruction from sparse, irregular turntable captures.
5. 【2609.31558】Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP
链接:https://arxiv.org/abs/2609.31558
作者:Ahmed Abdelnaby,Mohamed Elmahallawy
类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
关键词:Contrastive Language, dominant vision backbone, vision backbone due, Image Pretraining, zero-shot capabilities
备注:
点击查看摘要
Abstract:Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available this https URL
6. 【2609.31553】MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
链接:https://arxiv.org/abs/2609.31553
作者:Itzel Tlelo-Coyotecatl,Hugo Jair Escalante
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:raised Hate Speech, Hate Speech Detection, Ensuring online safety, Hate Speech, Ensuring online
备注: Preprint submitted to CIARP 2026
点击查看摘要
Abstract:Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.
7. 【2609.31524】Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment
链接:https://arxiv.org/abs/2609.31524
作者:Qing Xu,Yuxiang Luo,Zhen Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:bile duct injury, laparoscopic cholecystectomy remains, cholecystectomy remains challenged, Surgical scene understanding, computer-assisted intervention
备注:
点击查看摘要
Abstract:Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety (CVS) assessment typically treat it as a holistic prediction task, mapping visual features directly to criterion-level labels. This black-box paradigm lacks explicit reasoning about anatomical relationships, limiting both interpretability and compositional generalization. To address this, we propose ReasonCVS, a structured reasoning agentic framework empowered by Vision-Language Models (VLMs) that decomposes CVS assessment into explicit, fine-grained anatomical verification. Specifically, we devise an Anatomical Scene Graph Abstraction (ASGA) that organizes anatomical entities and their spatial relationships into a structured representation. To operationalize this, we introduce a Rationale-Aware Reasoning Agent, powered by a Large Language Model (LLM) fine-tuned via rationale distillation. Functioning as a strict central decision-maker, it invokes VLM-driven Sub-criterion Verifier as a specialized perceptual tool to parse the graph and independently evaluate individual sub-criteria. Through calibrated soft reasoning, this agent synthesizes the tool-gathered distributed observations, yielding a final verdict alongside a traceable clinical rationale. Extensive experiments on the Endoscapes-CVS201 benchmark demonstrate that ReasonCVS achieves superior performance (68.1\% mAP) over state-of-the-art while providing interpretable, criterion-level explanations for reliable surgical assessment.
8. 【2609.31514】Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics
链接:https://arxiv.org/abs/2609.31514
作者:Javier Muñoz-Haro,Ruben Tolosana,Ruben Vera-Rodriguez,Aythami Morales,Julian Fierrez
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Forensic Twins, architecture emerges, Generative AI architectures, forensic, Self-Supervised Residual Learning
备注: 9 pages + Supp. Material. 3 figures
点击查看摘要
Abstract:Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task, e.g., their augmentations overwrite the micro-statistics of image formation. This paper introduces Forensic Twins, a Self-Supervised Residual Learning (SSRL) framework whose pretext task suppresses macroscopic content availability. Each image is mapped through a frozen, off-the-shelf forensic residual extractor, from which two spatially disjoint crops are drawn. Sharing no pixel, the two views retain minimal semantic structure to align, leaving a redundancy-reduction objective with a predominant common signal: the stationary fingerprint of the image acquisition pipeline. Additionally, Forensic Twins is trained exclusively on real images; no AI-generated image is observed at any stage. Experiments show that Forensic Twins attributes AI generator sources with 56.61% accuracy, i.e., 6.13% above the previous state-of-the-art zero-shot method at 375x lower latency. We also demonstrate that fitting a Gaussian Mixture Model (GMM) offline using only the real image embeddings extracted from Forensic Twins turns it into a state-of-the-art zero-shot detector, reaching 97.99% AUC across 27 unseen AI generators, including GANs, diffusion models and commercial systems. Code, weights and exact splits will be made publicly available
9. 【2609.31509】ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos
链接:https://arxiv.org/abs/2609.31509
作者:Xuanzhi Liu,Xinyi Wu,Hang Pan,Wensi Huang,Zhenyao Wu,Ruize Han,Song Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Gaussian Splatting, mixed frame quality, uneven viewpoint coverage, Reliability-aware View Allocation, uneven viewpoint
备注:
点击查看摘要
Abstract:We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since weighting cannot restore details lost to blur or distortion, ClearGS further introduces Render-Guided In-Video Restoration (RIVR). The current 3DGS render provides a pose-aligned structural candidate, a frozen no-reference restoration expert restores the corresponding raw video observation without any clean reference image, and no-reference perceptual scores select among the render, restored observation, and high-frequency fused candidate. ClearGS then applies Full-Trajectory Repair Consolidation to revisit accepted repairs and preserve details introduced early. On GS2E and GSOTM, ClearGS achieves state-of-the-art overall performance, with consistent CLIP-IQA and MUSIQ gains and LPIPS reductions in most degradation settings, without paired sharp supervision or matched clean references.
10. 【2609.31507】SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
链接:https://arxiv.org/abs/2609.31507
作者:Jiajun Jiang,Chunliang Hua,Zichun Chen,Yanxing Wu,Zeyuan Yang,Jie Song,Xiao Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:extended urban spaces, Urban uncrewed aerial, uncrewed aerial vehicle, inherently demanding long-term, urban spaces
备注: Accepted at NeurIPS 2026, Track on Evaluations and Datasets. 32 pages, 16 figures. Project page: [this https URL](https://eku127.github.io/SatNav/)
点击查看摘要
Abstract:Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav targets city-level navigation missions and uses satellite crops as approximations of UAV nadir views for visual observations. Through an automated cue-to-episode pipeline, SatNav constructs 118K episodes from 59 scenes across 18 cities, with an average trajectory length of 379 m. To stress-test long-horizon memory and geospatial reasoning, SatNav defines three task families: Boundary, Landmark, and Route, targeting loop progress tracking, landmark-based spatial grounding, and route following with counting cues. Benchmarking classical VLN agents and recent agents based on large vision-language models (LVLMs) on SatNav shows that city-scale navigation remains challenging. We further introduce SwiftVLN, a modular framework with switchable memory components, and conduct systematic memory-design ablations. Finally, satellite-to-UAV transfer experiments show that satellite-trained navigation models can operate on real-flight UAV observations, showing the practical relevance of SatNav. Our project page: this https URL
11. 【2609.31463】Uncertainty-Aware Federated Learning for Infant Movement Analysis
链接:https://arxiv.org/abs/2609.31463
作者:Edmond S. L. Ho
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Infant movement analysis, General Movement Assessment, Infant movement, neurodevelopmental disorders, General Movement
备注: Accepted at IEEE The 4th International Conference on Federated Learning Technologies and Applications (FLTA26)
点击查看摘要
Abstract:Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple institutions to be collected and stored at a single site. Such assumptions are often impractical in clinical settings due to privacy, governance, and data-sharing constraints. To address these challenges, we present, to the best of our knowledge, the first federated learning framework for automated infant movement analysis and General Movement Assessment using skeletal motion data. As a clinically relevant use case, the proposed framework is evaluated on fidgety movement classification. To quantify model confidence, Monte Carlo (MC) Dropout is employed to estimate predictive uncertainty during inference. Building upon this, we propose an Uncertainty-Aware Federated Averaging (UA-FedAvg) strategy that incorporates predictive entropy derived from MC-Dropout into the federated aggregation process, enabling client contributions to be adjusted according to their predictive uncertainty. Experiments were conducted using a cross-subject evaluation protocol under a three-client federated learning setting. Results demonstrate that federated learning substantially improves classification performance compared with independently trained local models while achieving performance approaching that of centralized training. Furthermore, UA-FedAvg and its variant incorporating validation loss generally outperform conventional FedAvg across the evaluated data-split configurations.
12. 【2609.31461】KneePreM: Towards 3D Knee MRI Foundation Models via Large-Scale Unlabeled Pretraining and Label-Efficient Fine-Tuning
链接:https://arxiv.org/abs/2609.31461
作者:Xinxin Wang,Liam Hazan,Jing Li,Simona Rabinovici-Cohen,Xiaojuan Li,Mingrui Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remain insufficiently leveraged, Large volumes, ROC AUC, knee MRI scans, unlabeled Osteoarthritis Initiative
备注:
点击查看摘要
Abstract:Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a knee-specific 3D self-supervised model, and evaluated transfer and label efficiency for classification and segmentation. Methods: A 3D U-Net masked autoencoder was pretrained on 19,011 unlabeled Osteoarthritis Initiative (OAI) MRI series from 4,791 participants. Downstream fine-tuning used full and reduced training sets for fastMRI+ two-label classification (1,172 examinations), Arthroscopic Partial Meniscectomy (APM) eight-target classification (1,716 examinations), SKM-TEA segmentation (155 examinations), and APM segmentation (25 examinations). Baselines were random initialization and SuPreM. Deployment workflow was implemented with a Model Context Protocol interface. Evaluation metrics included balanced accuracy, F1 score, ROC AUC, PR AUC, and Dice score. Statistical analysis used bootstrap confidence intervals and paired bootstrap tests for classification and Wilcoxon signed-rank tests for segmentation. Results: KneePreM achieved higher full-data macro ROC AUC than both baselines for fastMRI+ and APM (all p .001). For fastMRI+ classification, KneePreM achieved a ROC AUC of 0.722 using 50% of the training data, exceeding both full-data baselines. In APM classification, KneePreM reached a ROC AUC of 0.740 with 70% of the data, matching the full-data random baseline and outperforming SuPreM. For SKM-TEA segmentation, its 70%-data Dice of 0.838 exceeded the full-data random baseline (0.835) and both same-budget comparators. In APM segmentation, its 75%-data Dice of 0.746 exceeded the full-data random baseline (0.731) and both same-budget comparators. Conclusion: KneePreM improves transfer performance and label efficiency across knee MRI classification and segmentation tasks, particularly when labeled training data are limited.
13. 【2609.31456】Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis
链接:https://arxiv.org/abs/2609.31456
作者:Mona Gandhi,Cenk Merih Olcay,Kuan-Chieh Lo,Santiago Castro,Christopher W. Myers,Srinivasan Parthasarathy
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:underperformance remain unclear, Vision-language models, remain unclear, underperformance remain, Vision-language
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizing individual components under increasing load. We introduce COMPASS (COMPositional Analysis of SkillS), a controlled evaluation framework designed to isolate and measure the distinct factors underlying compositional failure. By comparing performance on composed captions with their decomposed counterparts , we directly quantify the cost of compositional integration across 87K image-caption pairs. Across multiple VLMs, this gap is real but partial, accounting for only part of the observed degradation. This motivates a finer-grained investigation into what additional factors govern model behavior. We analyze performance at the level of individual skills: object detection, attribute binding, and relation reasoning, using skill-targeted perturbations across 274K image-caption pairs. We find a consistent skill-specific pattern: each skill degrades primarily with the count of its own primitive type (self-load), while cross-load effects are predominantly positive, suggesting that primitives of different types provide useful grounding context. This pattern holds across standard contrastive encoders, explicitly trained compositional reasoning models, and non-contrastive architectures. These findings show that compositional degradation reflects multiple separable factors that cannot be reduced to joint reasoning alone.
14. 【2609.31454】Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality
链接:https://arxiv.org/abs/2609.31454
作者:Bradley Scott,Zeqi Luo,Edmond S. L. Ho
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Federated learning, corruption modes equally, prediction-label loss expose, prediction-label loss, modes equally
备注: Accepted at IEEE The 4th International Conference on Federated Learning Technologies and Applications (FLTA26)
点击查看摘要
Abstract:Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed against the supplied label. We test these signals against additive image noise and persistent random label flips. On ResNet-20 with CIFAR-10 and SVHN under Dirichlet partitions with data that are not independent and identically distributed (non-IID), the two corruption types behave differently. For persistent random label flips, the within-client per-sample area under the receiver operating characteristic curve (AUC) is 0.85 on CIFAR-10 and 0.95 on SVHN for prediction-label loss, while every uncertainty estimator stays at chance (0.49--0.50). This pattern is consistent with the model remaining confident in the underlying image despite the supplied label being wrong. For image noise, expected-entropy uncertainty rises above chance (0.67 on CIFAR-10 and 0.66 on SVHN), while loss responds comparably (0.64 on both). Each signal is therefore the stronger detector for a different corruption: the prediction-label loss for persistent label flips, and expected-entropy uncertainty for image noise, with its advantage becoming apparent as federation-wide corruption prevalence increases. Robust FL data-quality assessment should match the signal to the corruption rather than rely on uncertainty alone across corruption types.
15. 【2609.31452】Vision-Based 6-DoF Grasp Pose Estimation for Robot Cloth Unfolding
链接:https://arxiv.org/abs/2609.31452
作者:Domen Tabernik,Peter Nimac,Jan Jerićević,Danijel Skočaj,Andrej Gams
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:challenging task due, complex interaction dynamics, perceptual ambiguity arising, critical visual cues, challenging task
备注: Published in IEEE Transactions on Cybernetics
点击查看摘要
Abstract:Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and perceptual ambiguity arising from frequent occlusions of critical visual cues such as folds, edges, and grasp points. In this work, we tackle cloth unfolding using a regrasping-in-the-air strategy, where one manipulator holds the cloth while the other grasps it at an optimally selected point to unfold it. To this end, we propose CeDiRNet-6DoF, a deep learning framework that jointly predicts effective grasp points and the complete 6-DoF grasp pose from the observed cloth configuration. By integrating dense 3D grasp regression with segmentation and sine-cosine-encoded Euler angles, the proposed method reliably estimates the grasp configuration that maximizes the unfolded cloth area. We extensively evaluated CeDiRNet-6DoF on a bimanual robotic setup within the ICRA 2024 Cloth Competition framework, achieving state-of-the-art performance. An ablation study further validates the benefits of key design components, including joint segmentation, background randomization, and image cropping. These results establish CeDiRNet-6DoF as a robust and versatile foundation for reliable robotic cloth manipulation in unstructured environments.
16. 【2609.31451】mplateCraft: Agentic Visual Template Generation
链接:https://arxiv.org/abs/2609.31451
作者:Hongjie Yu,Zhiyuan Fan,Yuzhe Zhang,Jiangcun Du,Zhicheng Gao,Yuhong Zhang,Xiaokai Zhan,Zongshi Xie
类目:Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV)
关键词:one-click content creation, growing popularity, popularity of short, driven demand, demand for one-click
备注: 5 pages, 3 figures, 1 table. Submitted to ICASSP 2027
点击查看摘要
Abstract:The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset preparation and tool orchestration. We propose TemplateCraft, a multi-agent system that converts natural-language instructions into client-executable templates through planning, material generation, effect-workflow generation, and protocol compilation. Its Planner-Evaluator loop uses execution feedback for targeted rollback, while stage-level and long-term memory support revision without parameter updates. We evaluate TemplateCraft on TemplateBench, derived from 60 real-world templates. With the same Qwen3-VL backbone, TemplateCraft raises image/video generation success rates from 56.7%/30.0% to 66.7%/50.0% over Planner-only (best-of-three) and improves template adherence and style consistency. With additional evaluation and revision, it matches or exceeds a GPT-4o Planner-only baseline on selected metrics. Persistent assets further improve cross-input style consistency.
17. 【2609.31450】From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training
链接:https://arxiv.org/abs/2609.31450
作者:Wang Jingxin
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Medical vision-language model, Medical vision-language, commonly evaluated, Group Relative Policy, Relative Policy Optimization
备注:
点击查看摘要
Abstract:Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000 clean-test questions, language model LoRA SFT changes correct-image accuracy by +1.10 percentage points (95% paired bootstrap CI:-0.85 to +3.05), while visual-benefit events decrease by 2.40 points and image sensitivity decreases by 5.60 points. Paired records reveal 155 acquired and 203 lost visual-benefit events. Broader adaptation yields lower correct-image accuracy than language-model LoRA SFT. Standard GRPO produces mixed-reward groups and parameter updates, with an uncertain clean test accuracy change. A generation audit reveals that canonical option scores can follow a different token path from generated answers. With scores taken along the greedy generation path, the evidence target improves on the training set; its gains over standard GRPO remain inconsistent on validation data at matched training doses. Sample-level analyses trace how evidence scores, decision margins, and generated answers change during post-training. This empirical and measurement audit identifies gaps between optimization activity, target acquisition, and useful held-out visual behavior.
18. 【2609.31435】Implicit Neural Representation for Hyperspectral Video Compression
链接:https://arxiv.org/abs/2609.31435
作者:Alfredo Scalera,Paul Murray,Jaime Zabalza
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:snapshot cameras, advent of snapshot, Bjøntegaard Delta PSNR, Bjøntegaard Delta, achieving Bjøntegaard Delta
备注: Accepted at IEEE WHISPERS 2026
点击查看摘要
Abstract:With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bjøntegaard Delta PSNR gains of +4.99 dB and Bjøntegaard Delta rate of -88.88% compared to traditional hyperspectral image compression methods applied frame-by-frame. In addition to reconstruction quality, the effects on downstream task performance are measured in the form of object tracking success. Compared to video compressed with methods based on principal component analysis and JPEG2000 in low data regimes, our proposed method improves tracking area under the curve by up to 23.42% and distance precision by up to 35.56% on examples from the HOT2026 dataset.
19. 【2609.31431】AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy
链接:https://arxiv.org/abs/2609.31431
作者:Edward Gaibor,Kyriaki-Margarita Bintsi,Carmen Luz Leiva Ureta,Zayneb Bellatif,Chiara Maffei,Wenze Li,Elizabeth Hillman,Yaël Balbastre,Anastasia Yendiki
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:analyzing white-matter organization, ground truth labels, dense ground truth, Accurate segmentation, white-matter organization
备注: 11 pages, 2 figures, 2 tables. Accepted at SASHIMI 2026, held with MICCAI 2026
点击查看摘要
Abstract:Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle when tissue type, species, modality, or acquisition conditions change. We present AxonSynth, a domain-randomized synthetic-data framework for training 3D axon segmentation models without manually annotated real training volumes. AxonSynth generates dense synthetic axon labels with orientation priors that reflect realistic fiber configurations and renders them with randomized density, contrast, bias fields, blur, and noise. A three-class 3D U-Net is trained to predict background, axon sheath and intra-axonal space. We evaluate zero-shot transfer on 10 held-out light-sheet microscopy (LSM) patches from macaque and human brain samples labeled with one of three axonal markers, comparing against calibrated thresholding and Frangi filtering using overlap, corrected detection, false-positive, and topology metrics. On macaque samples, AxonSynth achieved the best corrected Dice and corrected precision (0.826 and 0.851), compared with 0.765 and 0.754 for thresholding and 0.685 and 0.762 for Frangi. On human samples, corrected Dice was comparable to thresholding (0.857 vs. 0.868), while component-count error decreased from 22,504 to 3,377. Across all held-out patches, AxonSynth reduced component-count error in 10/10 patches and Euler-characteristic error in 8/10. These results show that synthetic-label domain randomization can reduce dependence on manual axon annotation while supporting synthetic-to-real 3D segmentation.
20. 【2609.31403】Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
链接:https://arxiv.org/abs/2609.31403
作者:Mert İncidelen,Yamen Kashkash,Asya Berker,Murat Aydoğan
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:optical character recognition, Decoy Font method, Vision-language models, character recognition, success in optical
备注: Accepted to the First Workshop on Document Intelligence and Understanding (DocInsights 2026), co-located with the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
点击查看摘要
Abstract:Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consists of 300 images, each containing text with sharp contour lines superimposed on another text with soft shading. Six recent closed-source models from three different model families were evaluated using this dataset under two different prompting conditions (naive and guided) and at two different resolutions ($512\times512$ and $64\times64$). A validation study showed that human participants could read both text layers with high accuracy. In contrast, the models, with most variants and both prompting methods, read the contour text with near-human accuracy at high resolution, but almost never fully extracted the shading text. At low resolution, the contour text could not be read by either the models or humans, while the shading text could be extracted with high accuracy. The findings indicate that the evaluated VLMs exhibit a consistent behavioral limitation when processing typographic structures containing multiple spatial frequency layers.
21. 【2609.31394】InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
链接:https://arxiv.org/abs/2609.31394
作者:Xingyu Miao,Zizun Li,Baole Fang,Kaiwen Song,Tenghui Wang,Hanxue Zhang,Yating Wang,Xudong Li,Yuping He,Xueyuan Wei,Chao Gao,Xijie Yang,Yingxiang Xu,Kerui Ren,Wenqi Guo,Jianjun Zhou,Xinzhe Wang,Weiguang Zhao,Ni Yang,Zetao Cai,Yufei Xue,Hengjie Li,Zeyu He,Yuanzhen Zhou,Rong Fu,Jianyang Zhang,Siwei Cui,Fuxian Huang,Yunsong Zhou,Xing Gao,Yifei Yao,Qiaojun Yu,Kailin Li,Ming Zhou,Mu Huang,Xinyue Li,Wenze Cui,Bingqi Jiang,Xueyue Zhu,Junting Dong,Haoyu Guo,Tao Lu,Mulin Yu,Bowen Zhou,Bin Zhao,Tianfan Xue,Weinan Zhang,Chunhua Shen
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:World Action Models, generalist robot manipulation, World Action, jointly model visual, visual dynamics
备注:
点击查看摘要
Abstract:World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$\Delta$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$\Delta$ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-$\Delta$ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: this https URL
Subjects:
Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2609.31394 [cs.RO]
(or
arXiv:2609.31394v1 [cs.RO] for this version)
https://doi.org/10.48550/arXiv.2609.31394
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
22. 【2609.31383】Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization
链接:https://arxiv.org/abs/2609.31383
作者:Brayden Zhang,Mahsa Golchoubian,Igor Gilitschenski,Boris Ivanovic,Kashyap Chitta
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:open-loop behavior cloning, behavior cloning, creating a fundamental, ultimately operate, fundamental mismatch
备注:
点击查看摘要
Abstract:End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or easy for the controller to track. We observe that these inconsistencies concentrate primarily at intermediate waypoints, while the predicted endpoint remains comparatively reliable. Based on this observation, we introduce Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that anchors the trajectory to the vehicle's executed history, preserves the policy's predicted endpoint, and reshapes the intermediate waypoints to improve feasibility. ECO requires no map, privileged simulator state, or additional training, and can be inserted between a broad range of waypoint-emitting policies and their controllers. Across two closed-loop simulators, it improves the aggregate closed-loop score of all six evaluated generative and regression-based driving policies, and the gains tend to increase with how often the base plans violate motion limits. On HUGSIM, ECO improves VaVAM from 18.1 to 31.0 HD-Score (+71%), achieving 1st place on the HUGSIM Closed-Loop Driving Challenge. Similarly, on AlpaSim, ECO increases the scene scores of VaVAM and DiffusionDrive by 123% and 22%, respectively. These results show that for a broad collection of end-to-end driving models, repairing the intermediate geometry of predicted trajectories without changing the policy's predicted endpoint can substantially improve closed-loop performance.
23. 【2609.31378】ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning
链接:https://arxiv.org/abs/2609.31378
作者:Mingqian Yu,Wei-kuan Chiang,Qiurui Wang,Peilin Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:generated optical images, recent years, image translation, mainstream approaches, optical images
备注:
点击查看摘要
Abstract:In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image translation. Unlike conventional diffusion models, ContraFM-S2O learns to predict the velocity field in training and solves ODE instead of SDE during inference to improve the sampling efficiency. In addition, ContraFM-S2O replaces instantaneous velocity with average velocity along the interpolation path to realize one-step SAR-to-optical image translation and uses contrastive learning to improve the quality of the generated optical images. Experiments show our model achieves state-of-the-art on SAR2Opt and QXS datasets, outperforming baselines, and reduces inference latency via one-step generation.
24. 【2609.31374】RECAST: From Log Replay to Closed-Loop Driving Simulation with View-Complete Actors
链接:https://arxiv.org/abs/2609.31374
作者:Zijun Zhao,Liewen Liao,Kang Shen,Songan Zhang,Ming Yang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:exposing views absent, simulation requires rendered, requires rendered observations, surrounding actors move, driving simulation requires
备注: 8 pages, 5 figures
点击查看摘要
Abstract:Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded trajectories, exposing views absent from the source log. Existing data-driven simulators reconstruct dynamic actors from sparse observations, which can result in rendering artifacts under these viewpoint changes. We introduce RECAST (REconstructing Controllable Actors for Simulation and Testing), a 3D Gaussian Splatting framework that generates a view-complete actor from a single segmented vehicle observation in a driving log and registers the generated actor in the reconstructed scene. RECAST supports planner-in-the-loop rendering under controlled ego-actor interactions. To adapt an image-to-3D prior to real vehicles, we further introduce RECAR, a dataset of approximately 20K real vehicles with 600K background-free RGBA images spanning diverse vehicle colors and types. We use two-stage adaptation to improve vehicle generation from real driving-log observations. At the actor level, RECAST reduces $\mathrm{FD}_{\mathrm{incep}}$ from 9.788 to 7.992 relative to unadapted TRELLIS. At the scene level, under actor motion beyond logged trajectories, RECAST reduces $\mathrm{FD}_{\mathrm{incep}}$ from 129.35 to 112.10 and increases $\mathrm{CLIP}_{\mathrm{margin}}$ ($\times1000$) from 0.14 to 3.47 relative to Street Gaussians. We demonstrate planner-in-the-loop simulation with the image-conditioned planner GTRS-Dense. Compared with native Street Gaussians actors, RECAST increases the no-collision (NC) rate from 22.2% (12/54) to 63.0% (34/54) and the mean minimum predicted time-to-collision (TTC) from 0.798 s to 2.150 s. These experiments show that RECAST supports closed-loop planner evaluation under controlled ego-actor interactions beyond log replay. Visit our project page at this https URL
25. 【2609.31364】OpenVAM: Open-World Visual Attention Modeling with VLMs
链接:https://arxiv.org/abs/2609.31364
作者:Kiana Hooshanfar,Amirhossein Kazerouni,Alireza Hosseini,Michael Brudno,Babak Taati
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Predicting human gaze, Predicting human, visual attention modeling, Open-world Visual Attention, human-computer interaction
备注:
点击查看摘要
Abstract:Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web layouts. We, therefore, introduce OpenVAM (Open-world Visual Attention Modeling with VLMs), a unified framework that jointly addresses universality and explainability across heterogeneous domains (natural scenes, commercial imagery, and UI/web layouts) and supervision modalities. OpenVAM adopts a decoupled-but-aligned design: a dedicated dense visual pathway provides stable, spatially precise localization, while an instruction-following vision--language semantic head generates grounded what/why explanations conditioned on the same image and data-type prompts. A three-stage training strategy preserves strong localization priors while progressively introducing language grounding and improving explanation alignment via parameter-efficient adaptation without perturbing the saliency branch. We further propose a scalable pipeline to generate multi-domain saliency-reason annotations for training and systematic evaluation. Experiments across diverse datasets show that OpenVAM improves robustness under domain shift while producing image-grounded explanations that make saliency predictions more interpretable.
26. 【2609.31356】Open Vocabulary Domain Unlearning
链接:https://arxiv.org/abs/2609.31356
作者:Sumanth Udupa,Mehrtash Harandi,Yadan Luo,Mahsa Baktashmotlagh
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:idealized textbook diagrams, exhibit remarkable zero-shot, Vision-Language Models, exhibit remarkable, autonomous driving
备注: Accepted in NeurIPS 2026
点击查看摘要
Abstract:Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning solely on the specific object classes seen during the unlearning fine-tuning phase. Consequently, these methods do not unlearn the domain itself; they merely overfit to seen class-domain pairs, leaving the domain easily recognizable for unseen classes and providing a false sense of removal. We argue that true domain erasure must be class-agnostic. To address this, we formalize Open-Vocabulary Domain Unlearning (OVDU), a rigorous protocol that mandates domain forgetting must transfer to held-out classes. To solve the OVDU challenge, we propose a surgical parameter-editing framework. First, a Fisher Information mask isolates domain-sensitive weights, mathematically protecting foundational zero-shot generalization. Second, our Targeted Manifold Scattering (TMS) objective uses preference-based mining to locally scatter the forget domain's stylistic geometry. Evaluated across PACS, OfficeHome, and DomainNet, our method vastly improves open-vocabulary generalization over existing baselines. Crucially, it delivers exceptional sample efficiency, outperforming peak 8-shot baseline results with only 4 shots.
27. 【2609.31349】DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models
链接:https://arxiv.org/abs/2609.31349
作者:Haojun Xu,Jie Huang,Xin Lu,Mingchen Zhong,Zihao Fan,Linjiang Huang,Si Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Large video diffusion, diffusion models offer, models offer expressive, video diffusion models, Large video
备注:
点击查看摘要
Abstract:Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.
28. 【2609.31339】ChronoFuseGS: Multi-Temporal Gaussian Fusion with Per-Splat Persistence and Change Visualization
链接:https://arxiv.org/abs/2609.31339
作者:Tobias Batik,Diana Marin,Peter Kán,Hannes Kaufmann
类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
关键词:image sets poses, Reconstructing environments, Gaussian Splatting approach, Gaussian Splatting models, captured image sets
备注: Accepted to Pacific Graphics 2026
点击查看摘要
Abstract:Reconstructing environments where parts of the scene change between captured image sets poses a challenge for 3D scene reconstruction. We present ChronoFuseGS, a multi-temporal Gaussian Splatting approach that addresses this issue by taking multiple separately trained Gaussian Splatting models, each representing a distinct timestep and partially overlapping in geographic coverage, and merging them into a single combined model. By allowing Gaussians from one timestep to contribute to the reconstruction at other timesteps, our approach leverages data across all captured timesteps to refine persistent parts of the scene. The model supports incremental extension, allowing new timesteps to be added while preserving the existing merged reconstruction. It encodes, for each Gaussian primitive, at which timesteps it contributes to the reconstruction. To support visual exploration of the reconstructed scene, we present a change-aware visualization approach that highlights the parts of the scene that have changed across a user-defined time selection, while preserving the color of persistent parts. Since the persistence encoding operates at the Gaussian primitive level, changes are visualized at sub-object granularity rather than being limited to object-level changes. We evaluate our approach on a real-world outdoor dataset of a flood management area, captured over 7 months across eight recording days and covering seasonal vegetation changes, snow cover, and flooding events, which we make publicly available. Our results demonstrate that the combined model consistently outperforms individually trained single-timestep models in novel-view synthesis quality, recovers structural details absent in the individual reconstructions, and reliably highlights changes in fine details and sub-parts of objects and natural structures.
29. 【2609.31326】CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support
链接:https://arxiv.org/abs/2609.31326
作者:Muhammad Muhtasim Shahriar,Md. Naimur Asif Borno,Saad Aloteibi,Mohammad Ali Moni
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:requires distinguishing visually, distinguishing visually similar, visually similar neighboring, existing approaches collapse, single opaque representation
备注: Manuscript under review at Expert Systems with Applications
点击查看摘要
Abstract:Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We introduce CG-HAF, a global-local fusion framework that instead keeps this evidence explicit: averaged holistic severity probabilities from independently trained classifiers are combined with structured lesion-burden descriptors from an object detector (lesion count, detection confidence, lesion area) into a compact representation, from which a lightweight, interpretable classifier produces the final grade. On a widely used benchmark, this fusion yields a clear, statistically supported improvement over global-evidence-only baselines, with the largest gains on the most severe cases. Testing on an independent dataset with a different grading standard shows that strong within-dataset performance does not transfer automatically, and a follow-up diagnostic attributes much of this gap to mismatched grading criteria rather than detection failure alone. These findings support interpretable global-local fusion as an effective strategy for ordinal acne grading while highlighting criterion alignment as key to cross-dataset portability, with a further illustration of how the resulting severity signal can support transparent, non-diagnostic decision-making in skincare applications.
30. 【2609.31314】CytoSPM: Open-Vocabulary Cytopathology Detection with Structured Prompt Bank
链接:https://arxiv.org/abs/2609.31314
作者:Wenjie Li,Zishan Xu,Jinyang Huang,Zhengxin Nie,Shichao Kan,Yixiong Liang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Cytopathology detection requires, requires open-vocabulary recognition, open-vocabulary cytopathology detection, detection requires open-vocabulary, Cytopathology detection
备注:
点击查看摘要
Abstract:Cytopathology detection requires open-vocabulary recognition because cellular categories are fine-grained, long-tailed, and continuously evolving across different organ systems. However, existing cytology detectors are mostly single-domain and closed-set, and there is still no unified benchmark for evaluating open-vocabulary cytopathology detection. We present PentaCyto, a multi-domain benchmark covering cervical, urinary, respiratory, serous fluid, and thyroid cytology, with 24 base categories and 9 held-out novel categories. Each category is associated with structured cytomorphology prompts that describe diagnostic morphological attributes and provide clinically grounded textual knowledge. We further propose CytoSPM, an efficient detector based on a decoupled two-stage design. It first extracts reusable class-agnostic visual representations, and then performs class-aware structural prompt matching with class names and cytomorphology prompts. On PentaCyto, CytoSPM outperforms existing methods in novel-category detection and open-vocabulary detection while maintaining efficient inference.
31. 【2609.31298】UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
链接:https://arxiv.org/abs/2609.31298
作者:Lei Xin,Zeheng Wang,Jiayin Zhu,Shihong Huang,Fanhu Zeng,Changjiang Jiang,Dengbo He,Yutao Yue,Zhenglun Kong
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Autism Spectrum Disorder, complex neurodevelopmental disorder, Autism Spectrum, long-term developmental outcomes, Spectrum Disorder
备注: Accepted by ACM'MM 2026
点击查看摘要
Abstract:Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity prompt learning for robust ASD recognition under heterogeneous data variations. Specifically, UniAR leverages a large multimodal model to generate hierarchical diagnostic descriptions at the word, phrase, and sentence levels, compensating for the lack of paired clinical reports. To align the generated semantics with visual evidence, we further design a Mixture-of-Experts-based Multi-Scale Alignment Module, which dynamically matches vector-quantized visual prototypes with semantic representations at corresponding granularities. Extensive experiments on four benchmarks covering brain MRI and facial expression scenarios show that UniAR consistently outperforms existing state-of-the-art methods, achieving average accuracies of 75.9\% on MRI benchmarks and 91.6\% on facial benchmarks, while improving average Accuracy on MRI benchmarks by 1.5 percentage points and average Accuracy on facial benchmarks by 1.2 percentage points over baselines. These results demonstrate that UniAR offers a robust and interpretable framework for ASD screening under semantic scarcity.
32. 【2609.31285】MoTop: Motion-Topological Model For Micro AU Detection
链接:https://arxiv.org/abs/2609.31285
作者:Huai-Qian Khor,Mengting Wei,Yante Li,Chu Kiong Loo,Guoying Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:reveal suppressed emotions, subtle facial movements, Facial, micro-expressions are spontaneous, high-stakes environments
备注:
点击查看摘要
Abstract:Facial micro-expressions are spontaneous, brief, and subtle facial movements that reveal suppressed emotions in high-stakes environments. In contrast to classic expression analysis, detecting action unit (AU) yields a finer representation of facial movements, serving as a preliminary step before defining expression classes and other downstream tasks. Therefore, it represents a crucial upstream task in facial analysis, and improving an AU detection module increases the precision of facial analysis. Despite that, detecting AU is challenging because of the constrictive nature of the AU activation regions, leading to confusion among different AUs known as AU ambiguity. To model the fine-scale changes, we propose \textbf{MoTop}, a motion-topological model that is augmented with a learnable motion context, yielding regional soft guidance for facial activity, followed by facial landmarks that capture the fine-scale topological changes of micro AUs. To increase the micro facial landmark representations, we amplify the encoded facial landmark transitions via linear extrapolation, thereby increasing the spatial proximity of landmarks and enhancing the low-intensity landmark dynamics. In addition, we design anatomical facial clusters that enhance the hierarchical representation, facilitating multi-scale modelling of facial geometry and improving micro-topological representations. With these contributions, we have achieved state-of-the-art performance on the CD6ME protocol for the micro AU detection task.
33. 【2609.31248】Gauss What You Need: Compact Gaussian Splatting Across Scene Scales
链接:https://arxiv.org/abs/2609.31248
作者:Afif Boudaoud,Jiayi Liu,Alexandru Calotoiu,Torsten Hoefler
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian Splatting reconstructs, posed photographs called, Gaussian Splatting, Splatting reconstructs, set of posed
备注:
点击查看摘要
Abstract:3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains unresolved: configurations effective on standard benchmarks can leave larger captures with too few Gaussians to reconstruct fine details. We observe that the surface to represent, given by the capture's extent and resolution, is known before training, whereas its content complexity becomes apparent during training, through the reconstruction quality on the training views. We introduce TangoGS, which combines capture-derived model sizing with training-based adaptation: the capture determines the scale of the model, and training feedback determines its final size within that scale. Before training, TangoGS derives a learning allowance for model growth from the capture's total pixels after discounting views that re-observe the same scene points. During training, reconstruction quality guides how many Gaussians to add and remove. On 13 standard benchmark scenes, TangoGS matches the mean PSNR of the best-performing evaluated baseline, LeGS, with $48\%$ fewer Gaussians. On eight large captures, the same configuration automatically scales to larger models when necessary, achieving the highest mean PSNR among evaluated methods: $0.54$ dB above the runner-up with $2.3\times$ as many Gaussians. Together, capture-derived learning allowances and training-quality guided density control enable a state-of-the-art quality--size compromise across scene scales without retuning.
34. 【2609.31247】Geometric Inconsistency Localization in Multi-View Image Sets
链接:https://arxiv.org/abs/2609.31247
作者:Xander Staelens,Albéric Loos,Bert Ramlot,Hannes Mareen,Peter Lambert,Glenn Van Wallendael
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Multimedia (cs.MM)
关键词:produce realistic, geometric, view synthesis, NVS, geometric inconsistencies
备注: 8 pages, accepted at the Deepfake Forensics Workshop (DFF 2026) at ACM Multimedia 2026
点击查看摘要
Abstract:Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we introduce DeformView, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies. Using DeformView, we evaluate state-of-the-art MV consistency-scoring methods and show that approaches developed for NVS evaluation transfer poorly to the forensic task of geometric inconsistency localization. To address this limitation, we propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level. By learning from explicit supervision, including hard negatives from geometrically consistent yet deformed views, DEFECt3R improves localization performance and substantially reduces false positives compared to existing consistency-scoring methods. Ablation experiments further show that both feature representations and correspondence quality contribute to localization performance. Overall, our findings demonstrate that MV geometric consistency is a promising yet underexplored signal for multimedia forensics and establish a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs. Code and dataset are available at this https URL
35. 【2609.31234】WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery
链接:https://arxiv.org/abs/2609.31234
作者:Zhongyu Pang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Problem, alignment SFT, tool, SFT, Abstract
备注:
点击查看摘要
Abstract:Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A is routing-first: emission is trained, not elicited. Stage B executes conditionally: intrinsic queries enter visual answering (full-scene thumbnail; a WeaveEarth-style evidence board as an optional fixed-budget, approx. 5k-token compression interface); extrinsic queries execute tool call on original full-resolution imagery, answering from tool observations in a second, observation-masked round. Training: alignment SFT, then GRPO under reward R_WA2. Results. Alignment SFT lifts extrinsic routing from 0% to 80.75% (323/400); GRPO suppresses 9 intrinsic mis-emissions while tool selection is unchanged. The trained 2B system does not beat the zero-shot 8B baseline overall (0.263 vs. 0.250), a diagnostic contribution. Oracle attribution separates two repair ingredients: loading the observation into context lifts extrinsic answer accuracy from 0.025 to 0.425 under marker-free cross-mode returns, and the two-turn SFT stage adds a further +9.3 points to 0.518 at a small routing cost. A +/- image ablation shows emission suppression is visually grounded, and a query-register matrix shows LLM-rewritten queries cost trained checkpoints 2-11 points. Scope. All training and evaluation use the 5,000 / 3,273 / 1,000-record VagueUHR corpus (600 intrinsic + 400 tool-requiring; the base seeds synthesis and is not used for optimization). Single-pass evidence construction runs at 7.31 s per image on an RTX 4090. Code, data, and evaluation protocols will be released.
36. 【2609.31207】Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling
链接:https://arxiv.org/abs/2609.31207
作者:Guanlin Li,Shifeng Bao,Yihan Zhao,Haitao Shen,Haoyang Li,Chen Zhao,Tong Yang,Jie Tang,Jing Zhang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Achieving robust cross-embodiment, critical representation flaw, Achieving robust, hardware-specific visual geometry, learning demands overcoming
备注:
点击查看摘要
Abstract:Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.
37. 【2609.31204】FlatClip: A Geometry-Aware Surface-Level Baseline for fMRI Representation Learning
链接:https://arxiv.org/abs/2609.31204
作者:Mo Wang,Wenhao Ye,Zihan Ning,Jiayu Zuo,Junfeng Xia,Hongkai Wen,Quanying Liu
类目:Computational Engineering, Finance, and Science (cs.CE); Computer Vision and Pattern Recognition (cs.CV); Neurons and Cognition (q-bio.NC)
关键词:Recent fMRI foundation, Recent fMRI, fMRI foundation models, foundation models differ, models differ substantially
备注: NeurIPS 2026
点击查看摘要
Abstract:Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized 3D/4D architectures and costly fMRI-specific pretraining. We ask how effectively an image-pretrained encoder can reuse the spatial organization of cortical activity. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a frozen-encoder surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Across resting-state benchmarks, FlatClip serves as a competitive middle-ground representation, outperforming ROI-level baselines on HCP and ADNI tasks while remaining weaker on PPMI and below the strongest voxel-level models overall. On visual-fMRI decoding, restricting the input to visual or NSD-provided task-active cortex improves performance, highlighting the value of task-relevant cortical coverage. Spatial perturbation controls reduce the predictive performance of flatmap features under both retrained and fixed readouts, and anatomy-linked arrangements consistently outperform vertex permutations across three colormaps. Together, these results position surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models, and support the utility of anatomy-linked spatial organization for reusing image-pretrained features. Code is available at this https URL.
38. 【2609.31202】Preserve-and-Compose Training for Composed Image Retrieval
链接:https://arxiv.org/abs/2609.31202
作者:Sehyun Kwon
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:relevant visual content, aims to retrieve, motivating zero-shot CIR, satisfy a user-specified, user-specified modification
备注:
点击查看摘要
Abstract:Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the source image. PACT learns from image--text--text (ITT) triplets without target images or gallery updates, aligning composed queries with target captions while preserving source evidence through visual supervision. We further introduce Chord scoring, which combines target similarity with source-relative directional agreement in the frozen image space. Results across four ZS-CIR benchmarks show that combining target-caption supervision with source-image evidence leads to strong retrieval performance across datasets, backbone scales, and external galleries. The code is available on this https URL.
39. 【2609.31198】Light Field Primitive for Novel View Synthesis
链接:https://arxiv.org/abs/2609.31198
作者:Liang Chen,Jiahui Ning,Xun Jiang,Xing Xu,Jimmy Ren,Fenglei Fan,Heng Tao Shen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:classical two-plane parameterization, present Light Field, dense ray database, Light Field Primitives, two-plane parameterization
备注:
点击查看摘要
Abstract:We present Light Field Primitives (LFP), a formulation for novel view synthesis that replaces the dense ray database with a compact set of differentiable primitives in the classical two-plane parameterization. Each primitive condenses a group of rays into one learned record, and its response to a query is governed by how closely that query belongs to the group. Rendering a camera ray then reduces to compositing all responses it elicits, and a scene can be optimized directly from posed images and rendered in real time with rays. Beyond its competitive performance on standard benchmarks, the main advantage of LFP is structural: its primitives reside directly in the 4D ray space, so optical and appearance effects that are already operations on the light field become behaviors of a single shared renderer. With minimal changes to that renderer, LFP supports multi-scale anti-aliasing, defocus deblurring with refocusing, rendering for fisheye cameras, and even transparent object reconstruction with ray refraction, matching specialized frameworks that devote substantial machinery to these effects.
40. 【2609.31193】Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
链接:https://arxiv.org/abs/2609.31193
作者:Jihoo Jung,Youngjoon Jang,Joon Son Chung
类目:Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
关键词:Current Audio-Visual LLMs, featuring multi-speaker dialogues, videos featuring multi-speaker, struggle with reasoning, multi-speaker dialogues
备注: Accepted by NeurIPS 2026
点击查看摘要
Abstract:Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
41. 【2609.31170】askIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback
链接:https://arxiv.org/abs/2609.31170
作者:Yanjie Tu,Qingsen Yan,Axi Niu,Wenxuan Cai,Tao Hu,Wei Dong,Haokui Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:downstream task performance, task feedback refinement, downstream task, aims to improve, restoration
备注:
点击查看摘要
Abstract:Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address these challenges, we propose TaskIR, a two-stage task-driven unified image restoration framework that integrates degradation-adaptive restoration with task feedback refinement. In Stage I, a Degradation Representation Module (DRM) extracts degradation representations, enabling a Degradation-Guided Transformer Block (DGTB) to dynamically modulate feature transformations for adaptive restoration. In Stage II, a Task-to-Restoration Feedback Generation module (TRFG) transforms heterogeneous task features into restoration feedback by modeling task-representation discrepancies associated with the current restoration. Subsequently, a Selective Task Feedback Refinement module (STFR) assesses feedback relevance and selectively refines intermediate restoration features to mitigate interference with well-restored content. Extensive experiments demonstrate that TaskIR achieves competitive restoration quality and downstream task performance across diverse degradations and tasks.
42. 【2609.31160】ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation
链接:https://arxiv.org/abs/2609.31160
作者:Donghang Lyu,Zichen Zhang,Oleh Dzyubachyk,Marius Staring
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:clinical tasks, ranging from diagnosis, treatment planning, diagnosis to treatment, vascular
备注:
点击查看摘要
Abstract:Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anatomies and modalities. While the Seg- ment Anything Model (SAM) has shown promise for med- ical image segmentation, its original design does not fully exploit vascular morphology and struggles with fine-grained vascular structures, leading to suboptimal performance. In this paper, we propose ReG-SAM, a SAM-based framework tailored to 2D vessel segmentation that leverages reference graph set for enhancing vascular representations. Specifically, we introduce two modality-aware representations derived from the reference masks: graph prompt embeddings (GPEs) that encode global spatial features from graphs, and vascu- lar prototype embeddings (VPEs) that capture fine-grained modality-specific vessel characteristics from multi-scale fea- ture maps and vascular masks. Since both require vascular masks that are unavailable during inference and require robust modality-aware vascular feature representations, we construct a modality-wise vascular database and develop two reference graph-guided representation learning schemes for estimating GPEs and VPEs using samples from the database rather than ground-truth masks. Extensive experiments across 19 datasets demonstrate that ReG-SAM consistently outperforms existing baselines, even those using manual prompts, particularly on challenging thin vessels.
43. 【2609.31154】HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models
链接:https://arxiv.org/abs/2609.31154
作者:Yi Sun,Xinhao Zhong,Zhiqi Zhang,Yimin Zhou,Junhao Li,Yuxia Qiao
类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
关键词:substantially improved visual, raised increasing safety, increasing safety concerns, safety concerns due, improved visual synthesis
备注:
点击查看摘要
Abstract:Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly follow a static weight paradigm, producing a single frozen adapter that struggles to adapt to diverse prompt variations and suffers from parameter interference when scaling to multiple concepts. We propose \textbf{HyperErase}, a framework for concept erasure based on hypernetwork-driven prompt-conditioned parameter synthesis. Our approach first reframes concept erasure as prompt-conditioned parameter amortization and trains a hypernetwork to map textual descriptions to prompt-specific LoRA updates, eliminating the need for per-prompt gradient optimization or manual LoRA merging. To further improve the stability and precision of synthesized adapters, we develop a decoupled rectification strategy, which disentangles LoRA tokens into pattern and scale subspaces, applies a square-root transform to curb multiplicative over-scaling, and leverages teacher-derived canonical priors for inference-time correction. Extensive experiments across major concept categories demonstrate that HyperErase consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving performance comparable to gold-standard single-concept baselines. Furthermore, the resulting models can provide specialized LoRAs for each input prompt variation in a single forward pass without requiring gradient updates during inference. These principled and flexible framework offers a new paradigm for concept erasure in T2I models.
44. 【2609.31150】FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification
链接:https://arxiv.org/abs/2609.31150
作者:Muhammad Muhtasim Shahriar,M. M. Golam Hafiz,Saad Aloteibi,Mohammad Ali Moni
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Cross-site lung histopathology, large pathology encoders, adapting large pathology, lung histopathology classification, non-IID client data
备注: Submitted to Engineering Applications of Artificial Intelligence (Elsevier)
点击查看摘要
Abstract:Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA), Normal, and squamous cell carcinoma (SCC). FedHisto-PAST v2 combines a frozen HIBOU-B foundation model with parameter-efficient adaptation, stain-conditioned paired-view prediction and feature consistency, reliability-aware prototype learning, and adaptive federated aggregation. Experiments used a five-client, non-IID, raw-data-local simulation with fixed internal evaluation, client-level analysis, component ablations, communication accounting, and a development-influenced exploratory LungHist700 cohort. All principal methods achieved near- ceiling internal performance, which limited discrimination on the fixed split. On LungHist700, FedHisto- PAST v2 achieved a Macro-F1 of 0.728560 and a balanced accuracy of 0.730454. Higher recognition of Normal and SCC was accompanied by lower ACA recall, and calibration remained imperfect. Prediction-level consistency was the only component with a clearly supported independent contribution in the external ablation analysis. Feature consistency and prototype regularization showed no conclusive independent overall gains in Macro-F1. The framework updated 1.253841% of the model parameters. The results provide exploratory cross-dataset evidence for stain-aware, parameter-efficient federation; they do not establish formal privacy, patient-level independence, prospective deployment, or clinical validation.
45. 【2609.31148】Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos
链接:https://arxiv.org/abs/2609.31148
作者:Bowen Guo,Shiwei Gan,Yafeng Yin,Xiao Liu,Kuizhuang Liu,Zhiwei Jiang,Lei Xie
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:sign language videos, continuous sign language, sign language understanding, sign language, achieved impressive success
备注:
点击查看摘要
Abstract:Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prerequisite for downstream recognition and translation tasks. However, sentence transitions in sign language are often smooth and visually ambiguous, lacking explicit pauses or posture resets. As a result, static frame representations may fail to capture the subtle temporal changes that indicate sentence boundaries. To address this challenge, we propose \textbf{SignShift}, a difference-aware segmentation framework that explicitly models frame-to-frame feature variation as semantic cues for sentence boundary detection. First, to model the feature variation, we design a Temporal Difference Module, which incorporates full-frame, facial, and hand cues, and employs inter-frame differencing to learn multi-scale temporal variations that capture both fine-grained local kinematics and global semantic transitions. Second, to mitigate over- and under-segmentation issues, we design a Segment Count Prediction module, which predicts the number of sentences to guide boundary selection. Extensive experiments on benchmark datasets demonstrate that SignShift substantially outperforms existing methods, validating its effectiveness.
46. 【2609.31135】Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding
链接:https://arxiv.org/abs/2609.31135
作者:Alberto Presta,Michal Byra,Grzegorz Stefański,Karol Szurkowski,Eryk Kołodziejczyk,Krzysztof Arendt
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
关键词:Spatio-Temporal Video Grounding, natural language query, Video Grounding, spatio-temporal tube, aims to localize
备注: 14 pages total. 8 pages main manuscript, 3 pages references, 3 pages additional material
点击查看摘要
Abstract:Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture that addresses STVG by combining efficient pre-trained components instead of large end-to-end models. P-STVG integrates a temporal-aware video encoder based on MobileViCLIP, a spatial encoder-decoder derived from MDETR, and a shared aligned text encoder. Temporal localization is performed through either a lightweight 1D U-Net or a simple thresholding strategy, enabling the same framework to operate in both weakly supervised and zero-shot settings. Furthermore, video representations are precomputed independently of the query, yielding an indexing-friendly pipeline for efficient inference and large-scale video collections. Despite requiring fewer than 90M parameters, P-STVG performs on par with weakly supervised methods and improves on earlier zero-shot approaches at a fraction of their memory and computational cost, establishing a favorable performance-efficiency trade-off for STVG.
47. 【2609.31108】Double-stream registration with pyramid fusion for HDR video with alternating exposures
链接:https://arxiv.org/abs/2609.31108
作者:Onofre Martorell,Ivan Pereira-Sánchez,Antoni Fuentes,Antoni Buades
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:High dynamic range, ting-exposure sequences remains, sequences remains challenging, extreme luminance variation, High dynamic
备注: 5 pages, double column, IEEE format
点击查看摘要
Abstract:High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate pyramid fusion. Given three consecutive frames, our method computes optical flow directly with the central frame, while introducing a complementary midpoint displacement strategy to handle cases with severe overexposition. A pyramid fusion stage then merges the resulting radiance and LDR images into a final HDR output. Experimental results demonstrate that our approach consistently outperforms state-of-the-art methods.
48. 【2609.31103】DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models
链接:https://arxiv.org/abs/2609.31103
作者:Jiangning Wei,Yuan Yao,Miaomiao Cui,Mingsheng Li,Humen Zhong,Shuai Bai,Zhibo Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:requires linking objects, constraints requires linking, requires linking, metric constraints requires, metric
备注:
点击查看摘要
Abstract:Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. Across nine datasets, DepthEvidence achieves the highest average dense $\delta_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.
49. 【2609.31074】Band-Selection Stability and Semantic Segmentation Performance: A Study on Hyperspectral City
链接:https://arxiv.org/abs/2609.31074
作者:Jiarong Li,Imad Ali Shah,Enda Ward,Martin Glavin,Edward Jones,Brian Deegan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Resource constraints make, constraints make high-dimensional, make high-dimensional hyperspectral, high-dimensional hyperspectral imaging, hyperspectral imaging challenging
备注: Accepted for IEEE WHISPERS 2026
点击查看摘要
Abstract:Resource constraints make high-dimensional hyperspectral imaging challenging in autonomous perception, motivating the use of band selection methods. However, the sensitivity of band-selection methods to sampled data and their relationship to semantic segmentation models (SSMs) remain underexplored. This study evaluates six band selection methods on ten independently sampled, class-balanced region-of-interest (ROI) sets, yielding 60 top-25 band subsets from the Hyperspectral City V2 (128 bands: 450-950nm) dataset. Top-$K$ bands ($K\in\{3,5, ... 13\}$) from the first three ROI sets are evaluated with three SSMs against the corresponding 128-band baseline. Experiments show that intra-method stability is method-dependent: Sim-LP shows the highest stability (pairwise Jaccard similarity) and, together with JMIM+CSNR, yields the best segmentation results. Top-$K$ based SSMs remain competitive with baselines, with gains of up to 2.01 mIoU and 1.72 mF1 points, and 18-22x faster CPU inference for $K=9$. However, performance does not improve monotonically with $K$, and stability shows no consistent association with SSM performance. These findings suggest that intra-method stability is informative but an unreliable indicator of downstream segmentation performance, highlighting the need to evaluate band-selection methods across repeated samples, subset sizes, and SSMs.
50. 【2609.31050】Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion
链接:https://arxiv.org/abs/2609.31050
作者:Olga Zatsarynna,Denis Korzhenkov,Juergen Gall,Amir Habibian,Mohsen Ghafoorian
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:generation requires reducing, spatio-temporal token sequences, long spatio-temporal token, video generation requires, requires reducing
备注:
点击查看摘要
Abstract:Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty branch predicts a token-wise estimate of denoising difficulty, which is used to route uncertain tokens through dense global attention while processing more reliable tokens with efficient local attention. The resulting routing is content- and timestep-adaptive, retains global context where it matters most, and provides a single parameter for controlling the quality-efficiency trade-off. HetA-DiT is compatible with few-step distribution-matching distillation and introduces no additional Transformer evaluation at inference time by reusing uncertainty estimates from the preceding denoising step. We evaluate the method on DMD-distilled Wan2.2-5B and Wan2.1-1.3B models. Across VBench, VBench-2.0, and human preference evaluation, HetA-DiT maintains competitive generation quality while routing only approximately 20% of tokens through dense attention.
51. 【2609.31040】Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images
链接:https://arxiv.org/abs/2609.31040
作者:Tiffanie Godelaine,Manon Dausort,Karim El Khoury,Benoît Gérin,Benoît Macq,Christophe De Vleeschouwer
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:reduce pathologist workload, improving diagnosis accuracy, Automating the analysis, diagnosis accuracy, cancer diagnosis
备注: 5 pages, 2 figures
点击查看摘要
Abstract:Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for patch-level classification without requiring any annotation, yet these zero-shot (ZS) predictions remain noisy on fine-grained tasks and must be further refined. A promising direction is to refine all predictions jointly, i.e., a transductive approach. However, most existing methods are not tailored to WSIs. We thus propose SlideTIM, an adaptation to WSIs of the recent transductive approach LC-TIM, which introduces a combined spatial--latent regularizer together with a prior on the patch class distribution. The former enforces spatially and semantically close patches to receive the same predictions, while the prior calibrates the predicted class proportions. Together, they address the complex spatial organization and the strong class imbalance of WSIs. Evaluated on four histology datasets, SlideTIM consistently outperforms all TIM variants, improving the macro-F1 by +8.1pp over the best competing baseline at 1 shot. Compared to the ZS, it raises the macro-F1 by +19.4pp at 1 shot. The code will be made available after submission.
52. 【2609.31032】mpQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks
链接:https://arxiv.org/abs/2609.31032
作者:Tianmeng Fang,Jiancheng Wang,Chen Wang,Liming Wang,Wei Wang,Jiayang Liu,Xiaochun Cao
类目:Multimedia (cs.MM); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:seek more effective, effective or stealthier, stealthier attack candidates, Existing, jailbreak methods
备注: 17 pages, 4 figures, 4 tables
点击查看摘要
Abstract:Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a query-constrained candidate allocation and ranking problem and propose TempQ-Jail. The method combines heterogeneous attack mechanisms to expand candidate coverage, estimates each candidate's end-to-end attack value from security-gate passage, dangerous visual generation, preservation of the original intent, and temporal validity, and ranks candidates so that high-value attacks appear early in a limited query trajectory. We evaluate TempQ-Jail on CogVideoX-5B using 70 common viable intents derived from T2VSafetyBench and compare it with six representative T2V jailbreak methods under a unified protocol. TempQ-Jail achieves TP-ASR@5 and TP-ASR@10 of 48.9% and 65.4%, improving over the strongest baselines by 4.6 and 4.0 percentage points, respectively. It also obtains the highest AUC-TP (0.469) and the lowest AvgQ (6.3). Analyses of query trajectories, candidate allocation, failure attribution, and ablations show that TempQ-Jail more effectively identifies and prioritises candidates with complete attack potential under limited query budgets.
53. 【2609.31028】Refining Cytology Predictions with Conditional Random Fields
链接:https://arxiv.org/abs/2609.31028
作者:Manon Dausort,Tiffanie Godelaine,Karim El Khoury,Maxime Zanella,Christophe De Vleeschouwer,Benoît Macq
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:achieve strong zero-shot, cell morphology differ, morphology differ markedly, differ markedly compared, Vision-language models
备注: 5 pages, 2 figures
点击查看摘要
Abstract:Vision-language models (VLMs) achieve strong zero-shot (ZS) classification on histology images but do not perform as well on cytology, whose stains and cell morphology differ markedly compared to histology. Conditional random fields (CRFs) can refine noisy VLM predictions by propagating information across patches, but existing CRF frameworks were designed for histopathology and do not transfer to cytology datasets, released as independent patch pools spanning multiple staining protocols. We introduce CytoCRF, which adapts the pairwise terms to cytology by targeting chromatin and cytology-specific staining, and further enrich the neighborhood of each potential term by combining multiple backbones. Across ten cytology datasets, CytoCRF outperforms existing CRF frameworks at every annotation budget, reaching +13.6 percentage points over the best baseline and +33.7 over ZS with only 50 annotations. Combining information from multiple backbones brings further gains, showing that the neighborhood topology matters more than the pairwise potential computed over it.
54. 【2609.31005】RACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking
链接:https://arxiv.org/abs/2609.31005
作者:Peder Borge Hellesylt,Albert Gassol Puigjaner,Kostas Alexis,Annette Stahl
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:previously unknown environments, maps enable robots, natural language, robots to reason, reason about previously
备注:
点击查看摘要
Abstract:Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TRACKGRAPH achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7x faster and uses 3.3x less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.
55. 【2609.30997】Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
链接:https://arxiv.org/abs/2609.30997
作者:Kai Yao
类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Passive image provenance, Passive image, image provenance, Passive, source
备注: Accepted at the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026). 29 pages, including technical appendices. Code: [this https URL](https://github.com/kaikaiyao/pixels-alone-provenance)
点击查看摘要
Abstract:Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.
56. 【2609.30993】FLIP: Final Layer Inference-Time Probing for Vision-Language Models
链接:https://arxiv.org/abs/2609.30993
作者:Drandreb Earl O. Juanico,Rowel O. Atienza
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:open-weight vision-language model, final-layer inference-time probe, supports structured, vision-language model, final-layer inference-time
备注: 25 pages, 14 figures, 5 tables. Accepted at the Mechanistic Interpretability Workshop at ICML 2026, Seoul, South Korea
点击查看摘要
Abstract:We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit computation, leaving parameters, prompts, and decoding unchanged. On a controlled detection/counting probe, sweeping intervention strength reveals three regions: negligible change, a bounded interior regime in which detection recall at IoU 0.50 ($R_{50}$) improves while tolerant counting error ($\mathcal{E}_{\mathrm{count}}$) falls, and over-suppression. We formalize a four-criterion probe-and-sweep protocol for disciplining the interpretation of intervention effects: regime structure, grounding-proxy alignment, feature-coherence dependence, and failure to reproduce the same positive regime on a performance-based negative control. The post-normalization state passed to the output head is the logit-facing instantiation of this test; under a non-targeted flooring sweep it satisfies the full protocol. Raw decoder-layer interventions, including the last-block output before final normalization, and the singleton-pair left/right control fail to reproduce the Final-site signature, while same-site operators and multiple VLMs replicate it. FLIP is therefore a validation step for intervention-based mechanistic interpretability, not a steering method.
57. 【2609.30989】PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation
链接:https://arxiv.org/abs/2609.30989
作者:Lucy Fothergill,Pietro Valdastri,Dominic Jones,Duygu Sarikaya
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:critical for automa, DoF pose estimation, safe interaction, tissue operated, robotic proprioception
备注:
点击查看摘要
Abstract:Purpose: Accurate 6 DoF pose estimation of surgical tools is critical for automa- tion, robotic proprioception, and safe interaction with the tissue operated on. Kinematics-based approaches suffer from accumulated errors due to the cable- driven nature of robotic arms, while vision-based methods often rely on external markers or trackers. Although more recent vision-based advances have been pro- posed, these two-stage pose estimation methods often lack real-time robustness due to accumulated errors and computational overhead. Methods: We propose a novel end-to-end trainable model, PICO. Our model employs a multi-task learning architecture to predict segmentation and depth maps, alongside regression of translation and rotation parameters. We define two proxy tasks that enforce geometric consistency in both 2D and 3D spaces, improving accuracy and robustness. For this, we propose a projection loss, and a point-to-point loss. Results: We evaluate our method on the SurgRIPE dataset, benchmarking its performance against state-of-the-art approaches using standard 6DoF pose esti- mation metrics. Our results demonstrate consistently strong performance across all four subsets, specifically in rotation, ranking second even under occlusion. It also demonstrates comparable translational performance, remaining competitive, especially in occluded cases. Conclusion: PICO demonstrates the effectiveness of multi-task learning and geometry-aware proxy tasks for robust and reliable surgical tool pose estimation, especially in occluded scenarios, highlighting potential for future applications.
58. 【2609.30988】PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution
链接:https://arxiv.org/abs/2609.30988
作者:Xin Di,Mingyu Shi,Yuanfei Bao,Long Peng,Yue Zhao,Jiaming Guo,Renjing Pei,Xueyang Fu,Yang Cao,Zheng-Jun Zha
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Real-world image super-resolution, requires recovering perceptually, preserving faithful content, complex low-resolution observations, realistic high-resolution images
备注:
点击查看摘要
Abstract:Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing diffusion-free SR networks without introducing diffusion components at inference time? To this end, we propose PhoenixSR, a generative heterogeneous distillation framework that transfers diffusion priors to independently designed feed-forward SR networks through score-based distribution matching. Rather than aligning heterogeneous features or imitating sampled diffusion outputs, PhoenixSR uses the pretrained diffusion model as distribution-level supervision, while paired SR supervision preserves reconstruction fidelity. To make distribution matching effective for fidelity-sensitive SR, we introduce Heterogeneous Distribution Adaptation, which adapts the target score to the SR domain, improves tracking of the evolving student distribution, and anchors training with paired supervision. We further employ Directional Reliability Weighting, a lightweight residual-consistency-based reweighting strategy that reduces unstable distributional guidance. All diffusion-related components are removed after training, leaving the original student architecture and inference cost unchanged. Experiments on three SR benchmarks and six feed-forward backbones, including SwinIR, HAT, Real-ESRGAN, and SeeMoRe, show consistent perceptual improvements with largely preserved reconstruction fidelity.
59. 【2609.30987】Self-Supervised Perceptually Interpretable Monocular Depth Estimation
链接:https://arxiv.org/abs/2609.30987
作者:Zain Ul Abidin,George Dimas,Dimitris K. Iakovidis
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:requiring ground-truth supervision, ground-truth supervision, making it attractive, real-world applications, depth
备注: Published at IEEE ICIP 2026; 6 pages, 4 figures
点击查看摘要
Abstract:Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, as depth is inferred from RGB representations that obscure the impact of individual perceptual image components. This lack of transparency limits systematic analysis of failure cases and reduces confidence in safety-critical settings. This paper presents a self-supervised framework for perceptually interpretable monocular depth estimation (PIMDE), designed to associate depth predictions with distinct perceptual components of the input image. Rather than operating directly on RGB inputs, the proposed method decomposes each image into a set of perceptual feature maps (PFMs), each encoding a specific visual cue. Distinct depth estimation branches process these PFMs independently to produce depth estimates (PIDEs), which are subsequently combined through an explicit fusion strategy. This formulation allows us to examine directly the contribution of each perceptual cue to the final depth prediction. Experiments conducted on the KITTI benchmark dataset demonstrate that PIMDE achieves performance comparable to established self-supervised MDE methods while providing additional insight into how different perceptual cues influence depth estimation. These results indicate that perceptual decomposition can support interpretability without sacrificing depth estimation accuracy.
60. 【2609.30982】FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators
链接:https://arxiv.org/abs/2609.30982
作者:Kai Yao,Marc Juarez
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
关键词:deployed service, inspect model weights, opaque APIs, weights or architecture, increasingly deployed
备注: This work has been accepted for publication in the proceedings of The 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026). 22 pages, including technical appendices. Code: [this https URL](https://github.com/kaikaiyao/FARE)
点击查看摘要
Abstract:Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator---using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under the exact-model and decision-only attacks evaluated in this work.
61. 【2609.30981】STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence
链接:https://arxiv.org/abs/2609.30981
作者:Siru Zhong,Shenghan Tan,Rihong Yan,Xiaohui Lv,Yuzheng Zhuang,Shuai Tao,Wulong Liu,Haohuan Fu,Yuxuan Liang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:question answering requires, answering requires tracking, Reliable online video, Reliable online, answering requires
备注: 50 pages, 19 figures, 27 tables
点击查看摘要
Abstract:Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS. Questions are stratified by a proxy for accumulated change intensity (Low, Medium, High) and query-time answerability (Known, Uncertain). To measure reliability, we introduce STORM-BR, a harmonic metric over joint answer-status correctness that exposes abstention failures masked by aggregate accuracy, alongside STORM-BR-ATTR for uncertainty attribution. Across 14 video LLMs, online accuracy peaks at 60.3\% (mean 51.7\%), whereas STORM-BR ranges from 5.7\% to 35.6\% (mean 18.8\%), driven by pervasive overconfidence on uncertain queries. STORM-Bench shows that task accuracy masks these gaps in epistemic reliability and state tracking. Benchmark and code are available at this https URL.
62. 【2609.30980】FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models
链接:https://arxiv.org/abs/2609.30980
作者:Haoyang Li,Ruoxi Sun,Qingqing Ye,Benjamin Zi Hao Zhao,Yaxin Xiao,Jason Xue,Haibo Hu
类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:models enable data-efficient, diffusion models enable, synthesize convincing forgeries, diffusion models, models enable
备注: 19 pages, 7 figures, 14 tables; includes appendices
点击查看摘要
Abstract:Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are brittle: modest post-processing or lightweight adversarial perturbations readily suppress detection, exposing a fundamental tension between imperceptibility and robustness. We introduce FeatMark, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro-features that remain natural to humans while providing a stronger, machine-verifiable provenance signal. FeatMark builds domain-specific feature banks that encode each watermark as a compact concept program, pairing open-vocabulary semantic cues with reliable edit regions and instruction templates. It then automatically selects features that are both feasible and executable and injects them through modular, mask-guided concept editing, yielding highly localized, scene-consistent micro-edits that are difficult to perceive. We conduct extensive experiments across VGGFace2, CelebA-HQ, and WikiArt, evaluating against 10 strong watermark removal/purification attacks (including regeneration-style purification) and several bespoke adaptive attacks tailored to FeatMark, to assess perceptual fidelity, watermark detection accuracy, and robustness. We further demonstrate FeatMark's extensibility to video mimicry attacks. The results show FeatMark remains virtually impervious, withstanding all evaluated attacks with negligible bit-accuracy and fidelity degradation.
63. 【2609.30979】CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models
链接:https://arxiv.org/abs/2609.30979
作者:Linyuan Gao,Yuan Wu,Yi Chang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:lack reliable evaluation, Vision-language models, demonstrated excellent performance, visual causal reasoning, causal reasoning
备注: 21 pages, 5 figures, 12 tables
点击查看摘要
Abstract:Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 15 multimodal models show that constraint sensitivity is task- and model-dependent: intervention has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at this https URL
64. 【2609.30963】Where and When to Force: Routed Forcing for Streaming Avatars
链接:https://arxiv.org/abs/2609.30963
作者:Zihan Su,Siwen Lu,Junhao Zhuang,Zeyue Xue,Haoyang Huang,Guanghao Li,Xiaofeng Tan,Chun Yuan,Nan Duan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:avatar generation requires, streaming avatar generation, generation requires real-time, requires real-time synthesis, Distribution Matching Distillation
备注:
点击查看摘要
Abstract:Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes a reverse KL divergence, which is inherently mode-seeking: it causes the student to discard high-dynamic modes and collapse onto static outputs, compressing both dynamics and diversity of generated videos. We find that this collapse is region-heterogeneous: person regions involving pose and gesture variations suffer the largest diversity loss, the audio-driven mouth region shows a small loss, and the background remains nearly stable. Based on this observation, we propose Routed Forcing, which routes the distillation objective by semantic region and noise stage to improve dynamics and diversity while preserving visual quality. Specifically, (1) Where to Force: Semantic-Region Routing applies Data-Forcing Distillation (DFD), which supervises the student with real videos, to the person region where diversity collapse is most severe, while retaining DMD for the mouth and background to preserve lip synchronization and scene stability. (2) When to Force: Noise-Stage Routing activates DFD at high noise stages, where real video serves as effective supervision to inject diverse and dynamic motion patterns. At low noise stages, DMD is used to refine details, avoiding blur and artifacts from spatial differences between real video and student-generated video. Experiments show that Routed Forcing improves dynamics by up to 45% and diversity by 7-25% over Self Forcing, while preserving video quality and lip synchronization.
65. 【2609.30962】IDM-Net: A Lightweight Illumination-Decoupled Modulation Network for Low-Light Image Enhancement
链接:https://arxiv.org/abs/2609.30962
作者:Cheng-Yen Hsiao,Jing-Ming Guo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:RGB color space, Low-light image enhancement, remains challenging, fidelity are difficult, Low-light image
备注:
点击查看摘要
Abstract:Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and chrominance representations, they primarily optimize luminance as an enhancement target, leaving its potential as an explicit guidance prior largely unexplored during feature reconstruction. To address this limitation, we propose IDM-Net, a lightweight Illumination-Decoupled Modulation Network for low-light image enhancement. IDM-Net adopts a dual-encoder architecture consisting of a structure encoder that extracts multi-scale appearance features from the RGB image and a lightweight illumination encoder that learns illumination priors from the decoupled luminance (Y) channel. To effectively exploit these priors, we introduce an Illumination-Guided Modulation (IGM) module that injects multi-scale illumination cues into the decoder through spatially adaptive affine modulation, enabling accurate brightness restoration while preserving natural color consistency. Furthermore, we design a lightweight Feature Refinement Block (FRB) to progressively suppress degradation artifacts and recover fine-grained image details during reconstruction. Extensive experiments on multiple standard low-light image enhancement benchmarks demonstrate that IDM-Net achieves competitive performance among lightweight LLIE methods while maintaining an excellent balance between restoration quality and computational efficiency.
66. 【2609.30952】MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
链接:https://arxiv.org/abs/2609.30952
作者:Hyungjin Chung,Byeongjun Park,Joonseok Lee,Hojun Kim,Jaeho Choi,Byung-Hoon Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:understanding requires integrating, requires integrating spatial, single visible frame, video understanding requires, non-overlapping camera streams
备注: NeurIPS 2026, 23 pages, 8 figures
点击查看摘要
Abstract:Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.
67. 【2609.30947】DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry
链接:https://arxiv.org/abs/2609.30947
作者:Luca Gandolfi,Simone Nascivera,Roberto Pellerito,Rong Zou,Chiara Plizzari,Davide Scaramuzza
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:RGB-based methods remain, methods remain vulnerable, challenging illumination, DPVO and RAMP-VO, GPS-denied environments
备注:
点击查看摘要
Abstract:Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch locations and fuses their correlation evidence before motion refinement. Each tracked patch maintains image and event descriptors, and a learned scalar gate combines modality-specific correlation embeddings for each patch--frame edge before a shared recurrent refinement and bundle-adjustment update. DAPEVO also supports event-only observations, enabling continued tracking when RGB frames are sparse or unavailable, while modality-aware keyframe culling preserves scarce frame constraints. On UZH-FPV, when retaining only one in six RGB frames, DAPEVO's mean absolute trajectory error (ATE) increases by only 36%, from 1.00 to 1.36m, whereas the ATE of DPVO and RAMP-VO rises by factors of $3.7\times$ and $3.1\times$, respectively. On TartanEvent, DAPEVO similarly remains below 1m ATE at 3Hz RGB input, while DPVO and RAMP-VO exceed 9m. Under degraded RGB input on TartanEvent, DAPEVO achieves an ATE of 0.60m, compared with more than 4m for both DPVO and RAMP-VO, while also outperforming event-only DEVO at 0.87m.
68. 【2609.30946】OneWorld: Learning Consistent Physics Across Actions in World Models
链接:https://arxiv.org/abs/2609.30946
作者:Ke He,Yichen Ding,Bin Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:predict scene evolution, Action-conditioned video world, aim to predict, essential for reliable, video world models
备注: 27 pages, 4 figures
点击查看摘要
Abstract:Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain a coherent understanding of the underlying world and limiting its reliability for planning and decision-making. To address these issues, we propose OneWorld, a shared-mechanism counterfactual generation framework that jointly models multiple action-conditioned futures under a common latent physical mechanism. A physical mechanism interpreter first infers a distribution over latent mechanisms from each action-outcome branch. These distributions are then aggregated into shared-world evidence, which captures whether the branches admit a common physical explanation while accounting for uncertainty in less informative branches. This evidence constrains flow training and guides sampling, encouraging consistency in the underlying physical mechanism while preserving the distinct outcomes induced by different actions. We further introduce a multi-intervention evaluation protocol in controlled environments, following the interaction settings of ACWM-Phys, to assess whether generated futures can be jointly explained by the same physical parameters, alongside standard measures of single-rollout prediction quality. Experiments in these environments show that OneWorld improves cross-intervention physical consistency while maintaining competitive single-rollout prediction quality.
69. 【2609.30941】Spackle: Completing Large View Single Image NVS with Adaptive Gaussians
链接:https://arxiv.org/abs/2609.30941
作者:Xuanzhi Liu,Yuhe Zhou,Xinyi Wu,Zhenyao Wu,Jinghao Chen,Ruize Han,Song Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:enables photorealistic rendering, enables photorealistic, observed viewpoints, photorealistic rendering, Single-image
备注:
点击查看摘要
Abstract:Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed number of Gaussians forces resource shifts from visible to newly disoccluded areas, degrading original scene fidelity when the target view deviates significantly from the input. To address this, we propose Spackle, a lightweight residual learning framework that mit- igates capacity competition without sacrificing efficiency. Spackle operates in three stages: predicting base 3DGS attributes from given views, automatically identifying poorly reconstructed regions, and learning a residual 3DGS optimized exclusively for these areas. At inference, we combine the baseline and aug- mented Gaussians for NVS. We conduct comprehensive experiments and show that Spackle achieves state-of-the-art performance on large-view-deviation cases.
70. 【2609.30934】ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos
链接:https://arxiv.org/abs/2609.30934
作者:Hengrui Kang,Zhonghao Yan,Yuxuan Yang,Ruoyan Jing,Yuncheng Guo,Hao Chen,Kongming Liang,Zhanyu Ma,Conghui He,Weijia Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Rapid advances, advances in AI-generated, increased the risks, risks posed, posed by deceptive
备注:
点击查看摘要
Abstract:Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forgery analysis beyond binary classification but struggle to use low-level forensic cues and provide precise pixel-level grounding. Specifically, we introduce ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos. We construct ManiVid-38K, the first dataset to combine paired, open-vocabulary localized manipulations of general videos with authenticity labels, forgery masks, and anomaly explanations. It comprises about 19K manually verified real-fake video pairs, mostly at 1080P resolution, generated under 2 paradigms with 15 powerful generation models. We sample 1K pairs for ManiVidBench, balanced across six manipulation types and generation models for fair evaluation. We further propose ManiVidLens, a unified framework for explainable video forgery analysis. Its Forensic Evidence Router supplies shared low-level forensic evidence for multimodal reasoning and video segmentation. Its Prompt Distill Module converts grounding states into semantic and geometric prompts and distills spatial priors for mask decoding and full-video propagation. ManiVidLens achieves relative gains over the strongest comparison methods in artifact grounding (+21.1% mIoU; +21.3% JF) and anomaly explanation (+131.3% ROUGE-L; +9.9% CSS). Its forgery detection remains comparable to dedicated classifiers (0.914 Acc; 0.913 F1).
71. 【2609.30928】UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
链接:https://arxiv.org/abs/2609.30928
作者:Quanhao Zhu,Bo Xu,Rui Lin,Chenyuan Wang,Yu Shao,Boling Zhu,Jiuyan Sun,Liang Zhao,Hongfei Lin,Feng Xia
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:medical imaging modalities, recent large vision-language, large vision-language models, shown increasing capabilities, imaging modalities
备注:
点击查看摘要
Abstract:Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at this https URL.
72. 【2609.30865】Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting
链接:https://arxiv.org/abs/2609.30865
作者:Zijian Wu,Jinliang Wang,Zidian Lin,Ying Song,Ziqian Lu,Hanjie Ma,Zhen Ye,Mingfeng Jiang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian Splatting, bypasses computationally expensive, remains fundamentally vulnerable, remain permanently frozen, corrupt subsequent frame
备注:
点击查看摘要
Abstract:COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking remains fundamentally vulnerable to error compounding---early pairwise tracking inaccuracies both corrupt subsequent frame initializations and remain permanently frozen in the scene representation. Rather than relying on heavyweight external neural priors or treating progressive tracking through isolated heuristic fixes, we propose a unified reliability-regulated trajectory optimization framework for progressive COLMAP-free 3DGS. At its core, our framework establishes an intrinsic, self-supervised bidirectional cycle-consistency mechanism that systematically regulates progressive camera trajectory estimation across two complementary temporal horizons: (1) Forward Motion Propagation, where the online reliability signal adaptively gates first-order kinematic warm-starts of rigid motion into upcoming pairwise registrations, supplying informed directional search priors while safely intercepting untrusted transitions; and (2) Retrospective Trajectory Correction, where the same reliability signal dynamically weights relative-pose consistency constraints within a sliding window of neighboring camera poses. By governing both prospective state initialization and retrospective trajectory consolidation through a unified reliability regulator, our self-contained framework resolves progressive drift without external priors or offline preprocessing. Extensive evaluations on Tanks and Temples and CO3D-V2 benchmarks show that our method substantially improves camera trajectory accuracy and novel-view rendering quality, outperforming existing unposed baselines. Code is available at this https URL.
73. 【2609.30855】MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization
链接:https://arxiv.org/abs/2609.30855
作者:Yijian Li,Saad Bedros,Paul Bigliardi,Mei Bigliardi Qi,Vassilios Morellas,Nikolaos Papanikolopoulos
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Reliable skin lesion, Reliable skin, skin lesion segmentation, dermoscopic computer-aided diagnosis, Pattern Analysis
备注: 13 pages 4 figures
点击查看摘要
Abstract:Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscopic ABCD rule, which was not designed for dermoscopy. Dermoscopic diagnosis is grounded in Pattern Analysis, a microscopic framework structured around dermoscopic features. We propose MDSkin-Net, which incorporates cue-level Pattern Analysis priors into a hybrid CNN-Transformer architecture. At its core is a Pattern Analysis-Guided Attention Module (PAGAM) comprising three priors motivated by distinct dermoscopic cues: an improved Efficient Channel Attention (iECA), a Multi-Scale Spatial Attention (MSSA), and a Biased Asymmetry Attention (BAA). We further introduce a multi-scale spatial alignment regularization (MSAR) that uses the segmentation ground-truth mask as hierarchical soft supervision, confining the classification head to lesion-localized evidence and coupling both task pathways through a shared spatial prior. Trained exclusively on the ISIC 2017 training split without external dermoscopy data, the MDSkin-Net ensemble transfers robustly under zero-shot evaluation, reaching a Dice Similarity Coefficient (DSC) of 92.38% and a melanoma AUC of 97.84%on PH2, and a DSC of 88.92% on the ISIC 2018 Task 1 test set. On the in-domain ISIC 2017 benchmark, the ensemble attains a mean Area Under the Curve (AUC) of 91.60% across the two classification tasks (melanoma and seborrheic keratosis vs. rest), and a DSC of 84.72% for segmentation. Classification remains competitive with baselines; in-domain segmentation trails single-task specialists, yet the proposed priors and alignment regularization yield representations that generalize consistently across cohorts of different scales.
74. 【2609.30840】Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
链接:https://arxiv.org/abs/2609.30840
作者:Austin Wang,Ziheng Cheng,Lexing Ying
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:enable high-quality visual, high-quality visual generation, single network evaluation, general implicit generators, generators enable high-quality
备注:
点击查看摘要
Abstract:One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution, RWTD constructs an adaptive target that mixes separately tilted current and reference distributions. The current component incorporates improvements discovered during training, while the reference component anchors the target to the pretrained generator. RWTD realizes this target through feature-space optimal transport and fixed-point regression. Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge. Empirically, RWTD substantially improves the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80, while separate preference alignment experiments demonstrate strong cross-reward generalization that yields balanced improvements and preservation of compositional capabilities.
75. 【2609.30795】Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
链接:https://arxiv.org/abs/2609.30795
作者:Chen-Chieh Liao,Yichen Peng,Yiyi Cai,Yûi Ono,Hiroki Hanaoka,Erwin Wu,Hideki Koike,Shuichi Kurabayashi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Existing human motion, methods provide strong, Existing human, intensity remains underexplored, provide strong motion
备注:
点击查看摘要
Abstract:Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit, but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised continuous control. Given a content motion and a style motion, we construct a style direction in a learned motion-style embedding space and condition diffusion generation with a scalar intensity. The training objective combines diffusion denoising with latent intensity regularization to encourage smooth and monotonic style scaling without requiring intermediate-intensity ground-truth motions. Our framework is compatible with pretrained motion diffusion backbones and supports heterogeneous style datasets, including the multi-actor style motion dataset. To test out-of-range usability, we additionally introduce a small real-capture over-reaction extension and evaluate large-intensity behavior against these unseen targets. Experiments measure controllability, interpolation/extrapolation behavior, content preservation, and motion realism, with ablations on direction construction and loss design.
76. 【2609.30783】Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models
链接:https://arxiv.org/abs/2609.30783
作者:Tianhang Guo,Yulin He,Wei Chen,Wenjuan Zhou,Yuhang Li,Xinbiao Gan
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:fine-grained visual perception, interpret implicit textual, implicit textual queries, enable fine-grained visual, embodied agents
备注:
点击查看摘要
Abstract:Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception-token generation and also increase the effective distance between visual tokens. To address this issue, we propose LIRSeg, which fully replaces explicit CoT with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages: spatial alignment grounds the latent tokens in object-relevant visual evidence, and GRPO further optimizes them with segmentation rewards. To make these compact latent tokens more informative, we introduce three complementary mechanisms from an information perspective: extreme-advantage sampling for selecting informative training signals, decoupled exploration-stability updates for learning complementary representations, and latent diversity amplification for preventing representational collapse. Extensive experiments on benchmarks demonstrate that LIRSeg consistently improves both segmentation accuracy and reasoning efficiency. Compared with the VisionReasoner baseline, LIRSeg achieves absolute gIoU improvements of 4.9% on ReasonSeg, 7.1% on MUSE, and 4.7% on MMR, while achieving a approximately 16x reduction in reasoning tokens. Code is available in supplementary materials.
77. 【2609.30769】Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes
链接:https://arxiv.org/abs/2609.30769
作者:Rushab Rasik Karania,Tomas Maul
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:few-shot learning requires, learning requires adapting, Cross-domain few-shot learning, target-time parameter updates, parameter updates
备注:
点击查看摘要
Abstract:Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming query-specific class means. Using a shared frozen ViT-S/16 encoder, miniImageNet source training, and CUB, EuroSAT and ISIC targets, we replicate the key comparisons across five independent training seeds. In 1-shot evaluation, WIPT improves frozen ProtoNet in every run on CUB (+0.21 percentage points) and EuroSAT (+2.07), but decreases ISIC (-0.22). In 5-shot evaluation, ProtoNet remains strongest overall, while WIPT consistently improves a capacity-matched support-only Transformer on ISIC (+0.99). Joint processing of up to five queries yields no reliable accuracy gain; in a head-only 5-shot benchmark, g = 5 reduces analytical attention-token pairs by 73% and peak allocated memory by 29% relative to g = 1, although latency is non-monotonic. Across all target/shot conditions, WIPT changes uncertain ProtoNet decisions far more than confident ones, and rescue/break decomposition accounts for the observed gains and losses. Source-shift and scorer controls further show that the benefit is not universal. Overall, WIPT provides a streaming-compatible form of test-time prototype adaptation that can improve difficult low-shot cross-domain decisions without target-time optimization.
78. 【2609.30761】mo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation
链接:https://arxiv.org/abs/2609.30761
作者:Zhao Wang,Jiangtao Hu,Jack Yu,Tao Yu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:inject text semantics, limits text comprehension, existing human motion, human motion generation, existing human
备注:
点击查看摘要
Abstract:Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing multimodal diffusion transformers (MMDiT), which have shown effective joint text--visual modeling in vision generation, into HMG. However, we find that articulated motion is temporally coherent but weakly correlated across joints, in which directly applying an MMDiT with flow matching produces poorly coordinated and jerky motion. In this work, we propose Timo, a novel kinematics-aware MMDiT framework tailored for HMG. Timo combines fully shared multimodal attention for bidirectional text--motion modeling with flow matching, geometric and rotational-kinematics supervision that compares actual rotations and their changes over time, and a two-stage curriculum progressing from broad motion learning to detailed caption alignment. Further, we construct a benchmark of $40{,}025$ held-out clips from six public datasets spanning diverse actions, assessing six complementary dimensions under a common evaluator and scoring protocol. Our model substantially outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Remarkably, Timo surpasses Kimodo on five of six dimensions, achieving a $40.8$% relative improvement in the average benchmark score. Project page: this https URL. Demo page: this https URL.
79. 【2609.30758】LLPR: Location-aware learning and physics-based reconstruction for raindrop removal from a single image
链接:https://arxiv.org/abs/2609.30758
作者:Zewei He,Xingyu Liu,Xing Luo,Guizhong Fu,Zixuan Chen,Yu Chen,Jinlei Li,Zhe-Ming Lu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:CNN or Transformer, camera lenses, Transformer architectures, occlusion and distortion, adherence to windows
备注:
点击查看摘要
Abstract:Raindrops can cause occlusion and distortion in the background scenes due to their adherence to windows or camera lenses. Existing raindrop removal methods concentrate on designing sophisticated CNN or Transformer architectures to recover distorted and missing texture. In this paper, we try to integrate location information and physical model into off-the-shelf CNN or Transformer architectures to help improve their performance. Specifically, we notice that existing methods deploy a preprocessing sub-network to generate a binary or soft mask to indicate the raindrop location, which will increase the network parameters and computational complexity. In contrast, a location-aware learning branch is embedded to teach the encoder in the training phase with the capability of perceiving the position of the raindrops. Note that this location-aware learning branch can be removed during the inference process (achieving performance improvements at no cost). Furthermore, instead of directly reconstructing the raindrop-free image (i.e., background scene), we devise a physics-based reconstruction scheme to first learn the transparency matrix and the raindrop layer. The latent background layer is then reversely derived based on the physical model. By combining the above-mentioned components, we propose our location-aware learning and physics-based reconstruction (LLPR) framework for this challenging ill-posed problem. We also collect a real-world raindrop-degraded image dataset, which is challenging for single-image raindrop removal (SIRR) methods. Extensive experimental results demonstrate the effectiveness and generality of our LLPR framework, achieving superior performance against state-of-the-art SIRR methods. The code will be made available upon acceptance.
80. 【2609.30755】raining-Free Bottleneck Width Planning for Convolutional Autoencoders
链接:https://arxiv.org/abs/2609.30755
作者:Guannan Guo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multiscale Spectral Rate-Distortion, Multiscale Spectral, user-supplied spatial cuts, Spectral Rate-Distortion, normalized mean-squared error
备注:
点击查看摘要
Abstract:Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for shared linear block-convolutional autoencoders under squared error. A nested-scale dominance result motivates reporting the activation-parameter Pareto frontier alongside the minimum-latent candidate. At NMSE = 0.01 on thirteen grayscale datasets, its latent-size prediction has 0.84% mean absolute percentage error against nonlinear patch-autoencoder boundaries; ten predictions are exact and the remaining three differ by one channel. In a four-dataset deployable comparison, MS-SRD matches all retrospective external widths and all four selected models pass, without training a selector; a 46-fit validation grid and four Least-Volume fits each pass on two datasets. In a skip-closed U-shaped autoencoder at the same bound, five predictions are exact, nine are within one channel, and every failing prediction is one channel short. Experiments at looser bounds show progressively larger nonlinear savings.
81. 【2609.30741】From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching
链接:https://arxiv.org/abs/2609.30741
作者:Hongfei Zhu,Ling Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:repeats substantial visibility, shading work, nearby views, repeats substantial, substantial visibility
备注:
点击查看摘要
Abstract:Binocular rendering requires two nearby views of the same scene and therefore repeats substantial visibility and shading work. We present a 2D Gaussian Splatting (2DGS) pipeline that fully renders a dominant-eye RGB image and an alpha-weighted depth proxy, reprojects that image to the affiliated eye, and repairs uncovered pixels. Small interior gaps are interpolated, whereas larger disoccluded regions are identified as regions of interest (ROIs) and selectively re-rendered. The depth proxy reuses the alpha-blending weights computed during dominant-eye rasterization, avoiding a separate depth-rendering pass. An adaptive ROI generator localizes the required updates using reprojected image boundaries and optional connected center-hole detection. On DTU, Tanks and Temples, and MipNeRF-360, the method reduces the measured time of a sequential two-pass binocular reference by 15.5\% to 28.8\% and peak GPU memory by 6\% to 11\%. The corresponding affiliated-eye quality degradation is at most 1.3 dB PSNR, 0.02 SSIM, and 0.02 LPIPS, representing a measurable trade-off that requires application-specific perceptual validation. These results establish a practical efficiency-quality trade-off for controlled static-scene stereo rendering and motivate future evaluation under continuous motion and on physical VR hardware.
82. 【2609.30733】Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation
链接:https://arxiv.org/abs/2609.30733
作者:Shengqi Dang,Zhengxi Yu,Feilin Han,Xingyu Lan,Nan Cao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remains largely unexplored, objects remains largely, Saliency, largely unexplored, advanced in controlling
备注:
点击查看摘要
Abstract:Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the visual saliency of a specific object during text-to-image generation without requiring any visual priors. Our key insight is that visual saliency is inherently relative: boosting the saliency of a target object also depends on the global saliency distribution across all objects in the scene. Based on this insight, we propose GazeME, a lightweight framework that uses saliency-marked prompts, inserting learnable marker tokens around object descriptions to indicate which objects to visually emphasize or suppress. To learn these markers, we construct a saliency-semantics dataset that associates objects in image--prompt pairs with object-level saliency scores, and propose Saliency Prior Marker Activation (SPMA), a saliency-aware stochastic marker activation strategy that exploits relative saliency relationships for robust training. During inference, GazeME automatically inserts appropriate markers into the prompt, thereby directly enhancing the visual saliency of the target object. Extensive experiments demonstrate that GazeME effectively boosts target saliency while preserving both semantic alignment and image quality.
83. 【2609.30728】Learning Polarization Image Restoration with General Restoration Priors
链接:https://arxiv.org/abs/2609.30728
作者:Chenggong Li,Jinhao Liu,Caiyun Wu,Yidong Luo,Junchao Zhang,Degui Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:imaging captures distinctive, captures distinctive surface, Polarization imaging captures, imaging captures, captures distinctive
备注:
点击查看摘要
Abstract:Polarization imaging captures distinctive surface and geometric cues that benefit a wide range of vision tasks. However, real-world polarization acquisition is often affected by multiple coupled degradations, making image restoration essential for practical polarization vision. Existing methods are largely tailored to specific degradations and remain constrained by the limited scale and quality of polarization data. To address these limitations, we develop an all-in-one polarization restoration framework for diverse and composite degradations. We first study the impact of different polarization representations on restoration performance and identify the normalized Stokes representation as an effective choice for separating intensity and polarization information. Accordingly, we devise a dual-branch architecture that separates intensity and polarization modeling. To overcome the limitations of polarization-specific training, the intensity branch leverages pretrained general restoration priors and a mixture-of-experts extension for composite degradations, while its restoration knowledge is adaptively distilled into the symmetric polarization branch via a cross-domain feature transform. In addition, we establish a composite-degradation polarization benchmark to support all-in-one restoration research. Extensive experiments on public datasets and our proposed benchmark demonstrate the effectiveness of the proposed method.
84. 【2609.30724】EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection
链接:https://arxiv.org/abs/2609.30724
作者:Haoran Sun,Yufan Li,Qichen Zhang,Haoran Zhao,Shuqi Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Joint video moment, detection requires identifying, Joint video, requires identifying query-relevant, explicitly preserve query-relevant
备注: 5 pages, 3 tables, 1 figure. Submitted to ICASSP 2027
点击查看摘要
Abstract:Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.
85. 【2609.30722】rafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation
链接:https://arxiv.org/abs/2609.30722
作者:Xiangyu Li,Tianyi Wang,Zhihao Dou,Christian Claudel,Zhaomiao Guo
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Existing roadside traffic, visual question answering, traffic video generation, Existing roadside, video generation
备注:
点击查看摘要
Abstract:Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and 31,145 actor-centered history-future samples) with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is represented as an actor-level program describing the target actor, intended behavior, legal route, interaction order, and temporal constraints, enabling a unified evaluation interface across heterogeneous foundation models. TrafficImag evaluates four complementary validity dimensions: initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation, and considers an end-to-end counterfactual successful only when all four are satisfied. Across state-of-the-art foundation models, the strongest reasoner reaches 80.4% macro F1, the complete condition interface raises end-to-end success from 23.3% to 55.0% for the best generator. Oracle studies further show that conditional video execution is the primary remaining bottleneck. TrafficImag provides a reproducible benchmark for evaluating and diagnosing counterfactual traffic video generation beyond perceptual video quality.
86. 【2609.30709】VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control
链接:https://arxiv.org/abs/2609.30709
作者:Kemou Jiang,Maonan Wang,Xingchen Zou,Jiayue Zhu,Yuhang Fu,Sicheng Wang,Xi Chen,Yirong Chen,Zhiyong Cui
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:mitigating urban congestion, urban congestion, Traffic signal control, essential for mitigating, mitigating urban
备注: 9 pages, 7 figures
点击查看摘要
Abstract:Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.
87. 【2609.30708】Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation
链接:https://arxiv.org/abs/2609.30708
作者:Tasneem Nasser,Susanne Schmid,Roberto Souza,Naser El-Sheimy
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:large annotated datasets, populations and diseases, key challenge, scarcity of large, large annotated
备注: 7 figures, 5 tables
点击查看摘要
Abstract:A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on manual annotations. Self-supervised learning has emerged as a promising approach for developing foundation models by enabling the learning of transferable feature representations from large-scale unlabeled medical imaging datasets. In this study, we investigate voxel-level brain age prediction as a domain-specific self-supervised pretext task and compare it with image inpainting, a widely used non-domain-specific alternative. We further propose a multitask self-supervised pretraining framework that jointly optimizes both objectives to learn complementary neuroimaging representations. The pretrained models are evaluated on three downstream magnetic resonance image segmentation tasks: multiple sclerosis lesion segmentation, ischemic stroke lesion segmentation, and cortical brain structure segmentation. Overall, the proposed multitask pretraining framework consistently outperformed the single-task pretrained models and training from scratch across most experimental settings, demonstrating the benefit of combining domain-specific and general self-supervised learning pretext tasks for the development of generalizable neuroimaging foundation models.\ Code Availability: The source code used in this study is publicly available at this https URL
88. 【2609.30703】SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion
链接:https://arxiv.org/abs/2609.30703
作者:Timing Li,Yiming Sun,Boan Tao,Xiyuan Gao,Haifang Cao,Pengfei Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Spatial misregistration, Hierarchical RGB-T Alignment, misregistration and cross-modal, cross-modal discrepancies, content imbalance
备注:
点击查看摘要
Abstract:Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency propagation across stages. We propose Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion (SAGE), a unified framework integrating frequency equalization, hierarchical alignment, and subband fusion. SAGE employs invertible joint encoding and source-specific low-frequency modulation to derive structural and gain guidance while preserving source information. Hierarchical frequency collaborative alignment estimates global affine geometry from low-frequency approximations and transfers geometric and contextual cues to high-frequency correlation reasoning for reliability-aware residual refinement. Guided subband fusion jointly aggregates the aligned frequency coefficients under propagated source and alignment guidance, coordinates complementary low- and high-frequency information, and reconstructs the fused image through the inverse wavelet transform. Extensive experiments on RGB-T datasets with real-world and synthetic misalignments demonstrate consistently competitive performance in alignment and fusion, validating the effectiveness of source-anchored guidance for weakly registered RGB-T images.
89. 【2609.30698】MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning
链接:https://arxiv.org/abs/2609.30698
作者:Peipei Li,Shuhan Xia,Shengyang Liu,Zekun Li,Ran He
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Real-world multimodal misinformation, mixed forgery sources, sample-specific detection strategies, requiring sample-specific detection, Real-world multimodal
备注:
点击查看摘要
Abstract:Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By benchmarking various candidate models and methods on the sub-tasks required by mixed-source detection, we select the strongest for textual, visual, and cross-modal forgery analysis and encapsulate them as callable tools with a unified interface. On top of this toolkit, we train the LVLM agent with reinforcement learning to teach it how to use these tools to better solve mixed-source detection. Since many of the tools are specialized models whose online execution at every rollout severely limits RL efficiency, we further introduce \textbf{Tool-Execution Cache}, which pre-executes candidate tool calls and reuses their cached outputs during training. This preserves multi-step rollouts while reducing online tool execution, largely improving the training this http URL on MMFakeBench demonstrate substantial accuracy gains over the base model without explicit tool search at inference time. Ablation and efficiency analyses further validate the learned tool-use policy and show that Tool-Execution Cache reduces online tool executions during training.
90. 【2609.30682】Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
链接:https://arxiv.org/abs/2609.30682
作者:Enzhi Zhang,Du Wu,Rui Zhong,Cong Ma,Isaac Lyngaas,Amir Koushyar Ziabari,Xiao Wang,Peng Chen,Tao Luo,Toshio Endo,Fumiyoshi Shoji,Kento Sato,Kentaro Uesugi,Takayuki Nonoyama,Ryuji Kiyama,Masahiro Yoshida,Masaru Tezuka,Tetsuya Ishikawa,Satoshi Matsuoka,Masaharu Munetomo,Mohamed Wahib
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:including Masked Autoencoders, Vision Transformers, Masked Autoencoders, Self-supervised pre-training, difficult to apply
备注: Accepted to NeurIPS 2026. 22 pages, 10 figures, 6 tables
点击查看摘要
Abstract:Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.
91. 【2609.30670】RACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
链接:https://arxiv.org/abs/2609.30670
作者:Yibo Ma,Qianqian Zhang,Peng Liu,Tiancheng Zhao
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:Streaming video understanding, video understanding requires, understanding requires models, report task scores, Streaming video
备注: TRACE Tech Report
点击查看摘要
Abstract:Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors explicit. TRACE combines temporally audited visual tasks with evidence timing and instruction-dependent trigger annotations, a unified causal Core--Adapter protocol that controls information availability while recording actual history processing and response events, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. On 1,240 records from 517 videos, we evaluate eight publicly available models or systems in eight configurations. We find that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms (responses emitted while no target window is currently valid and a later one remains), and missed target windows. These results show that streaming-video performance should be interpreted as execution-conditioned system behavior rather than a single score. Our benchmark and code can be accessed at \href{this https URL}{this https URL}.
92. 【2609.30667】StarWM: Self-Supervised Trained Attention Routing for Robust World Models
链接:https://arxiv.org/abs/2609.30667
作者:Zeqiang Zhang,Fabian Wurzberger,Maximilian Otte,Daniel Schmid,Sebastian Gottwald,Arne Peter Raulf,Daniel Alexander Braun
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:faithfully capturing environmental, robust world model, capturing environmental dynamics, world models ensure, strike the balance
备注: Accepted by NeurIPS 2026
点击查看摘要
Abstract:A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant information. We propose StarWM, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies. A dual-stream decoder then restricts reconstruction to the attended regions, with stop-gradient barriers preventing interference between the two objectives. These components allows reconstruction to supervise the visual content of attended regions without contaminating the latent with non-predictive information. On DeepMind Control with dynamic video backgrounds, default (reward-free) StarWM achieves the strongest performance under random-frame distractors and substantially outperforms reconstruction-based baselines under sequential video. In addition, its reward-augmented variant matches or exceeds reconstruction-free methods on sequential video, achieving the highest overall return across all distractor regimes. Mechanistic probing confirms StarWM preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.
93. 【2609.30647】Conditional Predictive Sufficient Statistics for Visual Representation Learning
链接:https://arxiv.org/abs/2609.30647
作者:Yuzhou Hong
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:discards patch-private noise, latent factors shared, patch-private noise, visual representation, retains the latent
备注: 14 pages, 2 figures
点击查看摘要
Abstract:A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.
94. 【2609.30613】MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification
链接:https://arxiv.org/abs/2609.30613
作者:Zhexiang Li
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Vision Transformers process, Dermoscopy classifiers built, Vision Transformers, Dermoscopy classifiers, built on Vision
备注: 12 pages, 3 figures. Includes supplementary material
点击查看摘要
Abstract:Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely ask whether the retained subset still contains the lesion. This paper introduces MedTokenBudget, a supervised post-backbone token routing framework that learns to construct compact lesion-enriched representations when auxiliary lesion masks are available. Its Lesion-Aware Token Scoring (LATS) module fuses attention entropy, feature norm, and local feature contrast through a learned scorer, then routes the top-$K$ patches under a target budget. LATS is trained with budget curriculum learning, diversity regularization, attention distillation, and lesion-mask supervision. The trained router is evaluated with a lesion retention rate that directly measures how much ground-truth lesion evidence survives the token budget. On ISIC 2019, mask-supervised LATS consistently outperforms Random and ToMe at headline budgets while retaining substantially more lesion patches. Code is provided for reproducibility, and complete tabulated results are included in the supplementary material.
95. 【2609.30609】MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization
链接:https://arxiv.org/abs/2609.30609
作者:Xiangyu Kong,Wenjie Zhou,Fengping Tian,Lihua Fang,Haoqin Sun,Chenyang Lyu,Longyue Wang,Weihua Luo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:generation requires consistent, requires consistent character, stable spatial layout, video generation requires, consistent character appearance
备注: 5 pages, 2 figures
点击查看摘要
Abstract:Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoint, a Spatial Grounding agent samples views from generated camera-traversal clips and anchors each shot to the view matching its framing. As generated shots drift from the plan, an Observer records how each shot ends in a continuity memory, from which a Transition agent builds character action and spatial references for the next shot. An Orchestrator composes these inputs into each request. Since a request reveals its effect only after rendering, we train it by agentic reinforcement learning with Trunk-GDPO, which compares rendered candidates at every shot rather than once per video and continues the best as the trunk. With generator and judges frozen, MVAgent attains the highest cross-shot consistency and narrative-planning quality among the compared methods on ViMax-Bench and is preferred over the strongest agentic baseline in human evaluation.
96. 【2609.30595】Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases
链接:https://arxiv.org/abs/2609.30595
作者:Ashish Sundar,Tiankuo Hou,Zhong Fan,Chunbo Luo,Xiaoyang Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:train controllable world, Synchronised action annotations, datasets remain elusive, controllable world models, Synchronised action
备注:
点击查看摘要
Abstract:Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure induced by egomotion to obtain grounded control signals directly. Using a method as simple as principal components analysis perform this, we find that the leading components provide signed, scalable, and composable throttle--yaw controls, although the method can recover only motion axes represented in the data. To prevent a high-capacity video DiT from exploiting pixel-level supervision, an online latent critic distils a frozen decoder--tracker--PCA (Principal Components Analysis) teacher without backpropagating through the decoder or tracker. Finally we critique the use of video generation metrics to evaluate WMs and introduce an example of an alternative, reference-free evaluation method. We measure \textit{controllability}, \textit{plausibility}, \textit{conjuring} (creating objects out of thin air) and \textit{geometric integrity}, revealing failures that conventional video metrics miss. We show that most baselines follow familiar action directions but struggle to reverse or remain stationary. Our model handles both while retaining compositional control and generation quality. Despite backwards actions being less than $1\%$ of our training data, we find that the model learns to reverse, scale its response linearly, and compose throttle with steering, all simply by learning through a grounded action space.
97. 【2609.30592】QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
链接:https://arxiv.org/abs/2609.30592
作者:Nicholas Foley,Devin Marinelli,Donny Moore,Diego Enriquez,Amanda Fernandez
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:separately learned projections, learned projections decide, transformed before aggregation, decide how strongly, concentric Fibonacci spheres
备注: 9 pages, 3 figures, 2 tables
点击查看摘要
Abstract:In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion $r_{ij} = q_i^{*} \otimes q_j$ supplies both the attention logit $\operatorname{Re}(r_{ij})$ and a sandwich-product feature transport $x \mapsto r_{ij} \otimes x \otimes r_{ij}^{*}$, with messages passed over sparse kNN graphs on concentric Fibonacci spheres. Ablations that change only the targeted component show the two roles to be asymmetric. Removing the transport reduces test accuracy by about four percentage points on CIFAR-10 and CIFAR-100 (single runs per CIFAR-100 variant), while replacing the learned attention weights with uniform averaging leaves it essentially unchanged. Parameter-matched controls then remove the geometry itself: standard attention on the same graph exceeds QSV (mean $87.3\%$ vs. $85.9\%$), and the same model on a flat 2D lattice reaches $91.1\%$, within $2.1$ points of a ResNet-20 trained under the same pipeline (single run). In the coupled kernel, nearly all of the learned pairwise computation resides in the transport channel.
98. 【2609.30566】Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models
链接:https://arxiv.org/abs/2609.30566
作者:Jian Shi,John Femiani,Peter Wonka
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:inference-time sampler, atlas, pretrained model, population, diffusion model
备注:
点击查看摘要
Abstract:We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central anatomy, which we call the \emph{intrinsic atlas}. The advantage is threefold. (1) It requires no retraining. A diffusion model that has already learned a coherent population, including the released ones, yields its atlas in a single inference pass without involving deformable registration. (2) It applies to multiple domains, such as brain MRI, chest X-ray, faces, and 3D shapes. (3) It extends to subpopulations. One age-conditioned model gives an atlas at any age in its training range, and the resulting family reproduces the CSF expansion of healthy aging. Evaluated as a registration target, the intrinsic atlas is best or second-best on every dataset against classical and learned templates, and the most central template on held-out brain MRI cohorts. Atlas construction can be reframed as a byproduct of generative modeling: a diffusion model is a learned representation of population structure, and the atlas is what it already contains.
99. 【2609.30478】he Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation
链接:https://arxiv.org/abs/2609.30478
作者:Soshun Kihara,Shunsuke Yasuki,Masato Taki
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Convolutional neural networks, neural networks trained, Convolutional neural, inductive bias, RGB domain
备注: Accepted at NeurIPS 2026. All authors contributed equally. Code: [this https URL](https://github.com/snskysk/event2rgb-distillation)
点击查看摘要
Abstract:Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained underexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while acquiring a stronger dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a spectral trade-off in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs from existing robustification methods and that it functions as a useful prior for diverse downstream tasks in which shape and contour information contribute alongside other cues. The code is available at this https URL .
100. 【2609.30459】VkVIO: Cross-platform GPU Acceleration for Visual-Inertial Odometry with Vulkan
链接:https://arxiv.org/abs/2609.30459
作者:Ole Hoffmann,Mateo de Mayo,Daniel Cremers
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:good state estimation, state estimation, Perception in robotics, fundamentally relies, relies on good
备注:
点击查看摘要
Abstract:Perception in robotics and XR fundamentally relies on good state estimation. Visual-inertial odometry (VIO) and Simultaneous Localization and Mapping (VI-SLAM) are proven ways of achieving this goal in a cost-effective and accurate manner. Efficiency in these systems allows for smaller, cooler, and lighter devices. GPU acceleration is a natural approach for reducing latency, thanks to their wide availability in platforms like embedded computers, mobile phones, and XR headsets. However, previous works in the literature have limited themselves to the use of CUDA for this task, significantly reducing deployment options to a single vendor. We instead leverage the vendor-agnostic Vulkan API, originally designed for the strict performance requirements of 3D graphics applications. In this work, we present VkVIO, the first, to the best of our knowledge, cross-platform GPU-accelerated VIO method. We provide state-of-the-art accuracy with causal estimates required for real-time operation. We deploy VkVIO on a diverse range of devices spanning a workstation, a laptop, and an extremely inexpensive single-board computer, while outperforming CUDA-based systems on the same hardware. VkVIO enables possibilities for low-latency, low-power, and low-cost VIO in robotics and XR.
101. 【2609.30450】LensDesigner: A Self-Improving Agent for Optical Lens Design
链接:https://arxiv.org/abs/2609.30450
作者:Lei Sun,Haoran Liang,Dannong Xu,Yao Gao,Yuyu Geng,Jinjin Gu,Kaiwei Wang,Danda Pani Paudel,Luc Van Gool
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:non-convex optimization challenge, experience and intuition, challenge that relies, relies heavily, heavily on human
备注:
点击查看摘要
Abstract:Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Furthermore, we introduce a continuous self-evolving mechanism guided by a curriculum agent. By iteratively solving design tasks with progressively increasing difficulty, the agent autonomously extracts, accumulates, and reuses design heuristics, effectively evolving its optical lens design expertise over time. At the evaluation level, we introduce LensArena, a standardized evaluation benchmark comprising $120$ diverse optical design tasks, covering extreme configurations. Extensive experiments on this benchmark demonstrate that LensDesigner significantly outperforms publicly available baseline algorithms, achieving superior success rates and optimization efficiency. We hope this work sheds light on the emerging field of intelligent optics. The code will be publicly available.
102. 【2609.30436】WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving
链接:https://arxiv.org/abs/2609.30436
作者:Mingkai Jia,Jiaxin Guo,Zhijian Shu,Jiawei Xu,Mingxiao Li,Jintao Cheng,Ping Tan,Wei Yin
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:accurate visual prediction, world model, Driving world models, driving world model, models learn rich
备注:
点击查看摘要
Abstract:Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Trajectories (WALT), which learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model without modifying the world model itself. Rather than directly generating raw waypoints, WALT maps them into compact representations through a dual-branch trajectory autoencoder and transfers semantic knowledge from the frozen visual world model into this trajectory space, encouraging the learned action representation to capture scene-level cues relevant to future motion and planning. Beyond our proposed formulation, we systematically study latent learning based on Joint-Embedding Predictive Architectures (JEPA) and feature alignment following Representation Alignment (REPA) to investigate how trajectory-only representation learning affects downstream planning. We evaluate WALT on the NAVSIM benchmarks. Relative to the raw-waypoint baseline, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%. These results suggest that preserving world representations while extracting action-relevant information provides an effective interface for world-model-based trajectory planning.
103. 【2609.30434】ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models
链接:https://arxiv.org/abs/2609.30434
作者:Hiwa Azeez Abbas,Fatemeh Daneshfar,Moloud Abdar
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Pre-trained vision-language models, Pre-trained vision-language, test distribution shifts, categories via prompting, vision-language models
备注: 20 pages, 8 figures, 9 tables
点击查看摘要
Abstract:Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone frozen, yet many existing multimodal prompt learners couple the visual and textual branches weakly and can be brittle in low-shot regimes. We propose ProCAP, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine each other across prompt depth. To reduce overfitting under limited supervision, we parameterize prompt tokens with Gaussian means and variances and regularize them with lightweight KL and L2 penalties, and we further add a compact symmetric InfoNCE head that aligns cross-attended image features with class-level text representations in a shared low-dimensional space. Across few-shot base-to-novel generalization on 11 datasets, cross-dataset transfer, and domain generalization on ImageNet shift benchmarks, ProCAP achieves strong aggregate base-to-novel performance and competitive transfer performance while keeping the CLIP backbone unchanged.
104. 【2609.30402】What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study
链接:https://arxiv.org/abs/2609.30402
作者:Akshit Sharma,Prashant W. Patil
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multimedia (cs.MM)
关键词:increasingly crafted, convincing by pairing, pairing a textual, textual claim, design choices
备注: Accepted at the Tenth Widening NLP Workshop (WiNLP), co-located with EMNLP 2026
点击查看摘要
Abstract:Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language backbones. Through systematic comparisons and targeted robustness analyses, we distill practical guidance on which design choices help, when do they fail silently, and what aspects of the pipeline most strongly shape model behavior, answering 4 key Research Questions (RQs). We aim to provide a reliable foundation for designing stronger and more dependable multimodal misinformation detection systems, thus contributing to the broader research community.
105. 【2609.30395】CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices
链接:https://arxiv.org/abs/2609.30395
作者:Amir Zamani,Zeinab Ghasemi-Naraghi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Real-time tiny object, weak spatial evidence, Real-time tiny, tiny object detection, tiny object
备注:
点击查看摘要
Abstract:Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high-resolution spatial representations from a YOLO11m-P2 teacher to a compact YOLO11n student without altering the student's inference architecture. Unlike conventional same-scale feature distillation, CSCWD transfers supervision from teacher P2 to student P3 after feature alignment while retaining same-scale distillation at deeper pyramid levels. Under the unified seven-sequence Drone-vs-Bird validation protocol, YOLO11n-CSCWD achieves 50.17% mean average precision at an intersection-over-union threshold of 0.5 (mAP@0.5) and 59.73% recall, improving the matched CA-YOLO11n baseline by 2.92 percentage points in mAP@0.5 and 3.55 points in recall. Cross-scale alignment further increases mAP@0.5 by 2.09 points over the corresponding same-scale channel-wise distillation configuration. In zero-shot evaluation on DUT-Anti-UAV, mAP@0.5 increases from 48.29% to 50.06% without target-domain fine-tuning. This domain was included because its challenging small targets make low-latency, computationally efficient detection particularly relevant. On Raspberry Pi 5 using NCNN-FP16 at 640x640 resolution, the 2.58-million-parameter student achieves 50.32% mAP@0.5 at 82.32 ms mean wall-clock latency, or 12.15 frames per second, while retaining essentially the same runtime and memory requirements as the matched baseline. The results support cross-scale distillation for improving tiny-target detection without increasing inference-time model complexity.
106. 【2609.30393】LiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting
链接:https://arxiv.org/abs/2609.30393
作者:Vivek Pandey,Amirhossein Mollaei Khass,Nader Motee
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:Selecting informative camera, influences model parameters, observation significantly influences, significantly influences model, Gaussian Splatting
备注:
点击查看摘要
Abstract:Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated evaluations of expensive information-gain oracles as the number of candidate views increases. We propose LiTe-GS, an oracle-efficient method for next best view selection in 3D Gaussian Splatting. LiTe-GS reduces the number of information-oracle evaluations by performing randomized subset evaluation of candidate views rather than exhaustively scoring the full candidate pool. The resulting approach achieves expected $O(M\log(1/\epsilon))$ oracle complexity with respect to the number of candidate views $M$, independent of the selection cardinality $K$, while providing an explicit trade-off between oracle efficiency and approximation quality through $\epsilon$. We provide theoretical guarantees on oracle complexity and approximation performance under the proposed selection scheme. Experiments on Blender and Mip-NeRF 360 demonstrate that LiTe-GS maintains reconstruction quality comparable to Fisher-information-based baselines while substantially reducing the number of Fisher-oracle evaluations across different acquisition settings.
107. 【2609.30356】AlphaEarth distinguishes cities but compresses urban variation
链接:https://arxiv.org/abs/2609.30356
作者:Andrew Renninger
类目:Computer Vision and Pattern Recognition (cs.CV); Physics and Society (physics.soc-ph)
关键词:land cover, built form, models map Earth, map Earth surface, Cities differ
备注:
点击查看摘要
Abstract:Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do not focus on cities: globally consistent labels for urban function do not exist, and many datasets - especially land cover and land use classifications - collapse the built environment into few classes. Here we audit the representation, focusing on AlphaEarth but with broader applicability to other Earth embeddings, by probing the geometry and geography of embeddings for 1,000 urban areas in 162 countries. We find that cities occupy a shifted but overlapping region on the hypersphere, 62.7° from the global mean direction, and continent and climate predict 24.3% of variation among the mean directions of urban centres in excluded countries. Inside cities, degrees of urbanisation carry 8.9% of the variation, and what they leave holds shared directions whose local orientation varies, not one universal axis of urbanisation. Retained variation is itself unequal: dispersion within urban centres is 14.1% greater per standard deviation of national development, even after adjusting for population, land area and continent. Further controls suggest cities in developing countries present less contrast in vegetation and texture, and dispersion follows that contrast: full adjustment for it leaves at most 6.4% of the gradient. Annually, a city's representation moves nearly eight times more than redrawing its own pixels explains, and contracts where the 2022 loss of Sentinel-1B removed a pass direction. AlphaEarth's representations therefore support comparison across regions, while the differences between its annual layers are not yet validated for comparison over time.
108. 【2609.31376】owards Whole-Study Screening for Congenital Heart Disease in Fetal Ultrasound Using Multiple Instance Learning
链接:https://arxiv.org/abs/2609.31376
作者:Mohamed Azzam,Ruobing Liu,Esther C. Ugwueke,Ziyang Xu,Shibiao Wan,Alex Foy,Abraham Zabih,Jason Christensen,Neil Hamill,Ling Li,Jieqiong Wang
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:Congenital heart disease, common birth defect, cases remain undetected, current artificial-intelligence methods, artificial-intelligence methods assume
备注:
点击查看摘要
Abstract:Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated from a study, by a clinician or by a view classifier. We remove that assumption and address CHD screening directly at the level of the whole ultrasound study. We propose a two-stage framework that first learns transferable frame representations by self-supervised masked-autoencoder pre-training on unlabeled fetal ultrasound, then identifies cardiac frames with a disease-robust module and aggregates them with a transformer-based multiple instance learning (MIL) model that produces a case-level diagnosis from study-level labels alone. The model further returns its highest-scoring frames for clinician review, and a hierarchical head separates critical from non-critical CHD. On the internal test set of our multi-source development cohort (FUSE), the proposed cardiac-gated MIL model reaches an area under the curve (AUC) of 0.985 with a specificity of 0.990, outperforming the reproduced NATMED ensemble (AUC 0.861, specificity 0.600) and the FetalCLIP foundation model (AUC 0.867, specificity 0.710). On an independent external cohort, all models initially perform near chance, but label-free CORAL adaptation raises the proposed model from an AUC of 0.513 to 0.944, whereas whole-study and view-dependent baselines do not recover. These results indicate that whole-study MIL with disease-robust cardiac-frame identification is an accurate and deployable route to prenatal CHD screening.
109. 【2609.31199】Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
链接:https://arxiv.org/abs/2609.31199
作者:Antoine Lorentz,Stéphane May,Valentine Bellet,Dawa Derksen,Bastien Nespoulous
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:Large-scale Digital Surface, Digital Surface Models, Large-scale Digital, Digital Surface, Surface Models
备注:
点击查看摘要
Abstract:Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accuracy elevation measurements at a substantially higher cost. In this work, we study diffusion models conditioned both on photogrammetric DSMs and Pléiades imagery to refine vertically co-registered DSMs. We introduce a modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy, enabling stable training on LiDAR data and transfer from natural images to elevation maps. Experiments in French cities demonstrate that multimodal conditioning improves elevation accuracy, reducing Dense Urban RMSE from 6.00 to 3.45 m in the in-context cities and from 4.16 to 2.77 m in the held-out city of Bordeaux.
110. 【2609.31070】Quantum Diffusion Models for Medical Image Analysis
链接:https://arxiv.org/abs/2609.31070
作者:Francesco Aldo Venturelli,Stefano Martina,Marco Parigi,Filippo Caruso,Alba Cervera-Lierta,Miguel A. González Ballester
类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Machine Learning (cs.LG); Quantum Physics (quant-ph)
关键词:approaches exploiting principles, Quantum Machine Learning, learning approaches exploiting, devising machine learning, machine learning approaches
备注: 12 pages, 12 supplementary pages, 7 figures, 1 table, 12 supplementary figures
点击查看摘要
Abstract:Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusion Model, and evaluate its use for medical image analysis. Specifically, our method is based on a Discrete-Time Quantum Walk algorithm, executed on a real quantum device, to model the forward dynamics of the diffusion model. For the backward step of the diffusion model, we devise and evaluate a classical learning model, which is used to reversely denoise the data. In contrast with other existing attempts at applying quantum machine learning for image analysis tasks, severely limited by the size of existing quantum devices, our method allows to process real-world large size medical data. In particular, we present results on grayscale and RGB images, as well as 3D volumes of moderate sizes. We benchmark our results by reproducing an alternative classical counterpart model, based on diffusion models on discrete state spaces. By doing so, we compare the generation capabilities of both models in terms of three distinct state-of-the-art metrics in the field of image generation, showing the competitive, promising results of our approach.
111. 【2609.30866】Universal Drift Correction for Multidimensional Scanning Microscopy
链接:https://arxiv.org/abs/2609.30866
作者:Sangjoon Lee,William Millsaps,Dasol Yoon,Caitlyn Obrero,Guoliang Hu,Corrie Barnes,Cedric Lim,Andrew Barnum,Arthur R. C. McCray,Colin Ophus
类目:Materials Science (cond-mat.mtrl-sci); Computer Vision and Pattern Recognition (cs.CV)
关键词:nominal probe positions, nominal probe, scanning microscopy, positions displaced, channel-resolved spectroscopic mapping
备注:
点击查看摘要
Abstract:In scanning microscopy, drift causes the specimen to be sampled at positions displaced from the nominal probe positions. This displacement alters the spatial assignment of the recorded signals and biases quantitative measurements across two-dimensional imaging, channel-resolved spectroscopic mapping, and scan-position-resolved diffraction analysis. Here, we extend orthogonal-scan drift correction from 2D images to spectrum images and diffraction datasets. We demonstrate how to recover probe positions using either differently oriented multidimensional scans or structural reference images. The recovered positions are used either to resample the multidimensional data onto a regular grid or to assign each recorded signal to its corrected coordinate. Our method combines affine and non-rigid correction, requires no prior structural model, and is implemented as open-source, GPU-accelerated software that reduces processing times by two to three orders of magnitude, enabling routine and automated drift correction for quantitative multidimensional microscopy.
112. 【2609.30659】Image Reconstruction from Phase with Untrained Neural Priors
链接:https://arxiv.org/abs/2609.30659
作者:Ene Meco,Ahmet Enis Cetin
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:absolute intensity ambiguous, Fourier phase encodes, encodes important spatial, measured spectral magnitude, spectral magnitude requires
备注:
点击查看摘要
Abstract:Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourier-phase and spatial-support constraints with an image-specific neural prior. The first stage alternates constraint enforcement with regularized neural-prior updates, while the second performs phase/support refinement alone with guaranteed convergence. We evaluate two neural-prior implementations on the same 77 microscopy images and compare them with a constraint-only baseline. After 500 final refinement passes, the best-performing variant achieves 31.41 dB pooled PSNR, 35.75 dB mean PSNR, and 0.9531 mean SSIM, improving pooled PSNR by 1.51~dB and reducing pooled MSE by 29.3% relative to the baseline. The results demonstrate the benefit of combining neural guidance with explicit constraint refinement at the evaluated iteration budget, while showing that lower phase residual alone does not guarantee greater reconstruction accuracy.
113. 【2609.30631】Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
链接:https://arxiv.org/abs/2609.30631
作者:Rayhan Rashed,Senja Filipi,Ross Cutler
类目:Audio and Speech Processing (eess.AS); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
关键词:target-speaker extraction aims, audio-visual target-speaker extraction, bounding lookahead, Online audio-visual target-speaker, preserving speech quality
备注:
点击查看摘要
Abstract:Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
114. 【2609.30629】FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception
链接:https://arxiv.org/abs/2609.30629
作者:Rajat Bhattacharjya,Minwoo Kim,Arnab Sarkar,Tamoghno Das,Sing-Yao Wu,Eli Bozorgzadeh,Marco Levorato,Nikil Dutt
类目:ignal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Robotics (cs.RO)
关键词:Mission-critical UAVs increasingly, UAVs increasingly rely, Mission-critical UAVs, split vision-language model, vision-language model
备注: Paper is currently under review. Authors' version posted for personal use and not for redistribution
点击查看摘要
Abstract:Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard cost. We present FreshLatent, a lightweight channel-aware latent adapter that trains a power-normalized encoder-decoder through wireless corruption while keeping the surrounding VLM frozen. We formulate deployment around a mission-conditioned perception requirement and embedded interface cost, linking channel quality and communication budget to the operating conditions under which perception remains usable. At 0 dB and the tightest communication budget, FreshLatent improves gIoU and cIoU over clean split compression by 20.79 and 20.87 points, respectively. At the most adverse evaluated SNR (0 dB), across all three communication budgets, FreshLatent recovers 63.5-69.1% of the gIoU improvement achieved by a much heavier, range-trained feature-JSCC codec. On an NVIDIA Jetson AGX Xavier in 10-W mode, FreshLatent uses 37-40x fewer encoder parameters, 7.7-9.9x lower edge-interface latency, and 8.8-10.0x lower edge-interface energy than the heavier codec. Together, these results show that lightweight channel-aware adaptation can recover a substantial fraction of the robustness of a much larger communication interface while broadening quality-valid operation under constrained wireless conditions.

