本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新1109篇论文,其中:

  • 自然语言处理149篇
  • 信息检索27篇
  • 计算机视觉164篇

自然语言处理

1. 【2610.08781】IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas

链接:https://arxiv.org/abs/2610.08781

作者:Ziyu Chen,Yilun Zhao,Jiashuo Sun,Yiling Ma,Manasi Patwardhan,Arman Cohan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Scientific research, formulate new directions, begins by synthesizing, set of related, identify gaps

备注:

点击查看摘要

Abstract:Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation remains challenging, as existing approaches based on prompting or feedback lack structured supervision for how papers should be synthesized. We introduce IdeaAnchor, a paradigm for training LLMs to perform research ideation using structured specifications as privileged signals. Each IdeaAnchor instance encodes how each input paper should be synthesized into a successful idea, including their functional roles, relationships, and target synthesis criteria. We build this paradigm by mining instances from published papers, capturing how real ideas emerge from prior literature. We then train models via demonstration, self-distillation, and reinforcement learning, and further enhance generation with retrieval at inference time. Experiments show consistent improvements in ideation quality. Our analysis reveals a functional decomposition: anchor-based training strengthens creative synthesis, retrieval enhances detail elaboration, and combining both yields the best performance.

2. 【2610.08778】Sherpa: Teaching LLMs to Teach Adaptively

链接:https://arxiv.org/abs/2610.08778

作者:Weixian Xu,Yanzhe Zhang,Zora Zhiruo Wang,Changyu Chen,Diyi Yang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:capable problem solvers, Large language models, increasingly capable problem, Large language, problem solvers

备注: 32 pages, 6 figures. Code and model are available at [this https URL](https://github.com/SALT-NLP/Sherpa)

点击查看摘要

Abstract:Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.

3. 【2610.08773】AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

链接:https://arxiv.org/abs/2610.08773

作者:Sarim Hashmi,Mukul Ranjan,Kshitij Mishra,Mikhail Kuznetsov,Praneeth Vepakomma,Nils Lukas

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:complete user requests, agents complete user, Web agents complete, user goal, complete user

备注: Code at [this https URL](https://github.com/Sarim-MBZUAI/advsim2real)

点击查看摘要

Abstract:Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.

4. 【2610.08747】he Missing Minimal Pair: Stereotype Evaluation in LLMs

链接:https://arxiv.org/abs/2610.08747

作者:Nataliya Stepanova,Ivan Titov,Emily Allaway,Björn Ross

类目:Computation and Language (cs.CL)

关键词:Large Language Models, contrastive stereotype sentences, Large Language, common approach, approach to measuring

备注:

点击查看摘要

Abstract:A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at this https URL.

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2610.08747 [cs.CL]

(or
arXiv:2610.08747v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.08747

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
5. 【2610.08738】Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

链接:https://arxiv.org/abs/2610.08738

作者:Mathias Ollu,Nikos Komodakis

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Diffusion Language Models, Continuous Diffusion Language, parallel text generation, Diffusion Language, Language Models

备注: 27 pages, 10 figures

点击查看摘要

Abstract:Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations and diffusion/flow spaces. In this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead. Drawing on the discrete DLM and continuous image diffusion literature on joint diffusion, we diffuse multiple modalities in parallel. These modalities represent tokens at different semantic granularities: in our instantiation, the tokens themselves and coarser clusters obtained by clustering pretrained token embeddings. We propose a general setup that allows per-modality samplers and schedules to enhance the interplay between modalities. Applied to CoBit, this yields H-CoBit, which delivers large empirical gains across benchmarks. At dataset entropy, H-CoBit improves MAUVE and reaches a generative perplexity (GenPPL) of 49.4 on LM1B and 50.4 on OWT, improving on the baseline by 24.2 and 20.7 points and surpassing even discrete DLMs of comparable size. On GSM8K, it reaches 27.4% accuracy, outperforming prior continuous diffusion and flow-based models. We further apply H-CDLM to the flow matching model FLM, obtaining consistent gains with H-FLM and demonstrating that the framework generalizes across continuous generative paradigms. Our code will be made publicly available at this https URL .

6. 【2610.08732】A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

链接:https://arxiv.org/abs/2610.08732

作者:Alexia Allal,Hicham Randrianarivo,Sylvain Lamprier

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Generative Information Retrieval, predicts document identifiers, Generative Information, directly predicts document, shifting document retrieval

备注: 8 pages, 3 figures, 1 table

点击查看摘要

Abstract:Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs). While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID? Current approaches rely heavily on computationally expensive downstream evaluations, hindering systematic analysis and rapid iteration. In this work, we address this challenge by presenting a comprehensive study on the properties, metrics, and trade-offs that define effective numerical DocIDs. Specifically, our contributions are threefold: First, we propose a unified framework that unifies Product Quantization (PQ) and Residual Quantization (RQ), and their hybrid variants within a single design space. This enables us to systematically study key DocID properties, such as hierarchy versus parallelism, as well as the impact of hyperparameters like DocID length and codebook size. Second, we define a suite of training-free, intrinsic metrics, to quantify DocID quality and evaluate structural fidelity without the overhead of full model training. Through extensive experiments on MS MARCO 300K and NQ320K, we analyze how these structural properties influence retrieval effectiveness.

7. 【2610.08719】Holdout Best-of-N: Unbiased Evaluation and Its Cost

链接:https://arxiv.org/abs/2610.08719

作者:Shrey Shah,Yinheng Li

类目:Computation and Language (cs.CL)

关键词:select a Best-of, Reusing, Best-of, scores, unbiased

备注: 25 pages, 2 figures, 3 tables

点击查看摘要

Abstract:Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $JK$, for every pool size $M\ge N\ge2$. At $J=K-1$, the selector deepens as $K$ grows. For independent Gaussian scores with common variance and fixed $M\ge N\ge2$, the unbiased minimax risk in this regime is of order $\sigma^2/\sqrt K$, attained by Holdout; allowing bias improves the rate to $\sigma^2/K$. For two candidates, we derive the minimum-variance unbiased estimator at known variance and the sharp asymptotic unbiased minimax constant $1/(\pi\sqrt2)$, which Holdout attains without knowing the variance. The cyclic average over subsets and ties can be computed in $O(MK\log M)$ operations. At fixed selector depth, cyclic evaluation of bounded scores has $O(K^{-1})$ risk uniformly in pool size. The impossibility result concerns the fixed matrix: one additional fresh winner score permits unbiased evaluation of the all-$K$ policy.

8. 【2610.08718】When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting

链接:https://arxiv.org/abs/2610.08718

作者:Vedant Palit,Florent Draye,Nicolas Zucchet,Zhijing Jin,Bernhard Schölkopf

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:phenomenon called spurious, called spurious forgetting, remains stored, phenomenon called, called spurious

备注:

点击查看摘要

Abstract:Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.

9. 【2610.08716】Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval

链接:https://arxiv.org/abs/2610.08716

作者:Hicham Randrianarivo,Logan Renaud,Alexia Allal

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Generative retrieval trains, Generative retrieval, Generative, diffusion, identifier

备注: 13 pages, 7 figures, 11 tables

点击查看摘要

Abstract:Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers. With identifier length and training budget fixed, we decode each model in several ways. Decoding alone moves a diffusion model's Hit@1 by 6.6 to 13.7 points. Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers. The generated identifier is right for 14-21% of NQ320K queries. We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes' probabilities. It matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead in Hit@1; on NQ320K, the lead comes from the model, not beam search. Starting from one sampled identifier, one-pass scoring removes 46-83% of masked diffusion's deficit to beam search; from generate-and-match, at most a quarter. On NQ320K, every paradigm largely memorises which identifier answers which query: random identifiers keep 83-90% of the Hit@1 of residual-quantised ones. There, product-quantised identifiers lead residual-quantised ones by 3.4 points in the autoregressive model and by -0.7 to +3.6 in diffusion models; across decodings, AR's gap exceeds diffusion's by 1.5-2.3 points, around our 2-point threshold. Paradigm comparisons must report each paradigm at its own recipe and best decoding.

10. 【2610.08703】Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue

链接:https://arxiv.org/abs/2610.08703

作者:Clayton Cohn,Joyce Fonteles,Kirk Vanacore,Gianni Mazza,Candida Crawford,Tom Hooper,Gautam Biswas,Rene Kizilcec

类目:Computation and Language (cs.CL)

关键词:learners' problem-solving processes, sources of difficulty, learners' problem-solving, problem-solving processes, processes and sources

备注: Submitted to LAK27 as a short paper. Currently under review

点击查看摘要

Abstract:In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.

11. 【2610.08680】A Systematic Study of Small Language Models on Abstract Reasoning Tasks

链接:https://arxiv.org/abs/2610.08680

作者:Nur A Zarin Nishat,Jens Lehmann,Andrei Aioanei,Sahar Vahdati

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:fit distribution-specific regularities, Endpoint accuracy, distribution-specific regularities, acquired a transferable, fit distribution-specific

备注:

点击查看摘要

Abstract:Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.

12. 【2610.08675】Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

链接:https://arxiv.org/abs/2610.08675

作者:Chuhong Xu(Sofia University),Bo Su(Indiana University),Ziyao Chen(University of California, San Diego),Ruiyang Xu(Northeastern University),Shimeng Dai(Michigan State University),Xinyu Qiu(Northeastern University)

类目:Computation and Language (cs.CL)

关键词:Financial reports repeat, metrics and accounting, accounting lines, allowing an LLM-generated, reports repeat

备注: counterfactual citation perturbation, evidence attribution verification, financial document question answering, Jev, LLM-as-a-judge, probabilistic source verification, tabular numerical reasoning

点击查看摘要

Abstract:Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.

13. 【2610.08670】Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment

链接:https://arxiv.org/abs/2610.08670

作者:Orion Reblitz-Richardson

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Language models increasingly, models increasingly act, Language models, increasingly act, model

备注: 33 pages

点击查看摘要

Abstract:Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.

14. 【2610.08660】Evidence-Bound Reasoning: Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics

链接:https://arxiv.org/abs/2610.08660

作者:Mariya Miteva,Maria Nisheva-Pavlova

类目:Computation and Language (cs.CL)

关键词:generate plausible explanations, generate plausible, reliably verifying, statement is supported, supported by patient-specific

备注: 15 pages, 4 figures, 4 tables. Preprint

点击查看摘要

Abstract:Background: Biomedical AI can generate plausible explanations without reliably verifying whether each statement is supported by patient-specific evidence. We developed a neuro-semantic verification framework that converts radiomic measurements into addressable evidence records and machine-checkable claims. Methods: UPenn-GBM radiomics were aligned with de novo CaPTk extraction from standardized MRI and expert-validated segmentations in an independent multicenter cohort. The shared space comprised 1,728 features from T1, T1GD, T2, and FLAIR MRI across three tumor regions. Reference-defined semantic states were derived from 611 UPenn cases. We evaluated cross-cohort transportability, model-linked provenance, deterministic verification, controlled predictive degradation, and an LLM claim-extraction pilot; MGMT prediction served only as a transport stress test. Results: Median semantic-state agreement was 0.786 (weighted kappa 0.709), ranging from 0.918 for morphologic to 0.252 for intensity features. The external evidence ledger contained 1,655 model-linked records for 331 patients. The verifier achieved 100% exact-set accuracy in a 6,620-claim corruption benchmark. In a 24-case pilot, GPT-5.6 Sol reproduced 72/72 prespecified atomic claims, and the frozen verifier recovered 24/24 expected conditions. During controlled degradation, ROC AUC declined from 0.899 to 0.500 while verification accuracy remained 1.000. External MGMT discrimination was weak (ROC AUC 0.543). Conclusions: Verifiability can be engineered and evaluated independently of predictive performance. LLMs may structure explanations, while final evidence-consistency checking remains deterministic.

15. 【2610.08647】SquidAgent: Parallelize Wisely, Coordinate Efficiently

链接:https://arxiv.org/abs/2610.08647

作者:Yexiong Lin,Shanshan Ye,Yu Yao,Zhen Fang,Bo Han,Tongliang Liu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:solve complex multi-step, LLM-based agents solve, agents solve complex, incurs substantial latency, complex multi-step tasks

备注: Accepted at NeurIPS 2026. 37 pages, including appendices

点击查看摘要

Abstract:LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2$\times$ mean throughput improvement and a 2.6$\times$ mean wall-time speedup over Claude Code, and a 2.0$\times$ throughput improvement over the strongest multi-agent baseline.

16. 【2610.08630】owards In-Parameter Memory Augmentation for Large Language Models

链接:https://arxiv.org/abs/2610.08630

作者:Haoyu Huang,Zhongwei Xie,Jiaxin Bai,Yisen Gao,Hong Ting Tsang,Wuganjing Song,Huihao Jing,Yufei Li,Yangqiu Song

类目:Computation and Language (cs.CL)

关键词:Recently Large Language, Large Language Models, Recently Large, Large Language, LLM-based agents increasingly

备注:

点击查看摘要

Abstract:Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbf{In-parameter memory} offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbf{Parameter Placement}, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbf{Parameter Acquisition Time}, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.

17. 【2610.08604】InterCorrect: Intersection-Aware Correction of Demographic Model Merging for Fair ASR

链接:https://arxiv.org/abs/2610.08604

作者:Ashley E. Bravo-Bravo,Yuchen Zhang,Haralambos Mouratidis,Ravi Shekhar,Monorama Swain

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:Automatic Speech Recognition, Automatic Speech, Speech Recognition, multiple demographic groups, demographic groups

备注: Under Review

点击查看摘要

Abstract:Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address for speakers belonging to multiple demographic groups. This work studies demographic-aware model merging for fair Speech-LLM-based ASR. Starting from a SLAM-ASR-based model, we fine-tune only the connector on demographic-specific subsets and merge the resulting subgroup-adapted connectors into a global model. We then identify critical cross-axis demographic pairs using subgroup WER and task-vector conflict, and apply intersection-specific correction vectors to the global merged model. Experiments on Fair-Speech show that global demographic merging improves overall WER over the base model, while intersection correction provides additional gains for several merging strategies. In particular, TIES with WER-based correction achieves the best overall WER, reducing it from 7.38\% to 5.13\%. Subgroup and disparity analyses further show that the proposed approach improves performance across demographic axes, while highlighting that lower average WER does not always imply reduced subgroup disparity.

18. 【2610.08601】Generative AI translations in high-stakes emergency messaging

链接:https://arxiv.org/abs/2610.08601

作者:Nune Ayvazan,Anthony Pym,Yu Hao

类目:Computation and Language (cs.CL)

关键词:involve high stakes, high stakes, tragic consequences, Emergency messaging, extreme-weather reports

备注:

点击查看摘要

Abstract:Emergency messaging such as extreme-weather reports and earthquake instructions can involve high stakes, to the extent that translation errors can lead to tragic consequences. The use of machine translation or generative artificial intelligence might therefore not be recommended. On the other hand, time savings in the initial translation can allow greater investments of resources in revision and authorization processes, as well as a wider range of target languages. An experiment with generative AI translations of an earthquake instruction text from English into Chinese and Spanish shows that use of discourse-specific prompts can considerably improve understandability and actionability, although the translations may still not be trusted by translators. Human revision is still required, not only to detect errors but also because of the ethical need for someone to take responsibility for any errors or delays in such messaging.

19. 【2610.08585】Incidental information contaminates patient notes and disrupts clinical reasoning in large language models

链接:https://arxiv.org/abs/2610.08585

作者:Krithik Vishwanath,Brandon Ye,Anton Alyakin,John E. Markert,Aaron Hsieh,Michał Mańkowski,Eric K. Oermann

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, support ambient documentation, increasingly relied, ambient documentation

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.

20. 【2610.08560】Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

链接:https://arxiv.org/abs/2610.08560

作者:Dan Ben-Ami,Kobi Cohen,Chaim Baskin

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Streaming video-language models, evidence needed, Streaming video-language, video-language models, current question

备注:

点击查看摘要

Abstract:Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.

21. 【2610.08559】Latent space bias directions in LLMs capture confidence, not fairness

链接:https://arxiv.org/abs/2610.08559

作者:Stephanie Buttigieg,Maeve Madigan,Parameswaran Kamalaruban,Stuart Burrell

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:lightweight inference-time debiasing, inference-time debiasing technique, large language models, gained popularity, lightweight inference-time

备注:

点击查看摘要

Abstract:Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model performance and limited transfer to new datasets. Our work analyses what the debiasing direction used for activation steering actually encodes, in order to shed light on its inconsistent performance. We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks. We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias. Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics. Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care. In short, steering appears to reduce bias, not by correcting the model's underlying preferences, but by making it less confident, even on tasks unrelated to bias.

22. 【2610.08553】DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory

链接:https://arxiv.org/abs/2610.08553

作者:Yining Li,Dongchen Han,Jie Fu,Gao Huang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Sequential test-time training, test-time training adapts, Sequential test-time, inner-loop gradient based, network previous state

备注:

点击查看摘要

Abstract:Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.

23. 【2610.08544】How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

链接:https://arxiv.org/abs/2610.08544

作者:Pranjal Garg

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:workhorse of interpretability, probe score, model, probe, input

备注:

点击查看摘要

Abstract:Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises $12$--$15\times$, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.

24. 【2610.08540】oward Alignment Scaling Laws: A Framework and First Preregistered Measurements

链接:https://arxiv.org/abs/2610.08540

作者:Jeremy Canale

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:isolated findings, easier or harder, argued from isolated, alpha, alignment

备注: 34 pages, 24 figures, 8 tables. Games: [this https URL](https://www.aisafety.fun) . Preregistrations: [this https URL](https://osf.io/wda8q) , [this https URL](https://osf.io/q2j3y) , [this https URL](https://osf.io/8kreb)

点击查看摘要

Abstract:Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run regime; that above 1 any policy holding headroom above a floor must grow super-exponentially; that, for burdens that are positive mixtures of power laws, fits on small models underestimate large-scale exponents; and that an audit that uncovers hidden failures without false positives never underestimates true alignment. We propose a pre-registrable protocol and apply reduced versions of it twice. A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60. A preregistered pilot on Qwen2.5 0.5B-72B finds exponents of -0.05 for truthfulness and 0.48 for stated dispositions (both scaling helps under its reduced rule, though local slopes approach 1 at the top; replicated on Qwen3 0.6B-14B), while sycophancy (0.89, or 0.83 with two seeds added at 72B) and a planted backdoor are undetermined: the backdoor is removed quickly when its trigger is known but survives blind safety training at four of five sizes. We release four browser games that play these laws (this http URL). We make no claim about which regime holds for current frontier models.

25. 【2610.08513】Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions

链接:https://arxiv.org/abs/2610.08513

作者:Dennis Fucci,Andrea Bacciu,Dong Liu,Weronika Łajewska,Saab Mansour

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:faithfully simulate human, LLMs are increasingly, social environments, making it critical, increasingly deployed

备注:

点击查看摘要

Abstract:LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human interactions. Central to this is grounding agents in realistic user personas, yet existing datasets rely on fictional personas and are limited to a handful of languages, lacking the empirical grounding necessary to evaluate behavioral fidelity across diverse populations. We introduce Wiki-Talkie, a multilingual dataset of real-world conversations from Wikipedia Talk pages across five languages spanning two language families: Germanic (German, English) and Romance (Spanish, French, Italian), paired with personas derived from real user communities and encompassing sociodemographic attributes, self-descriptions, and behaviorally grounded interaction traits. Using Wiki-Talkie, we evaluate agent interactional behavior on a next-turn generation task across various persona conditioning strategies. Our evaluation assesses whether agents collectively reproduce the distributional behavioral patterns observed in human discussions. Results show that user's comment history exemplifying interaction behavior consistently outperforms explicit persona information. In addition, models systematically underproduce negative or extreme sentiments, while over producing references and suggestions, revealing biases toward agreeableness and positivity. Crucially, these patterns hold robustly across languages, with small cross-lingual differences.

26. 【2610.08501】Language-model ratings of depression reflect the rater more than the patient

链接:https://arxiv.org/abs/2610.08501

作者:Baihan Lin

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)

关键词:diagnostic blood test, Patient Health Questionnaire, blood test, diagnostic blood, Depression

备注:

点击查看摘要

Abstract:Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) = 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.

27. 【2610.08463】UNREAL: Unifying Retrieval and Long-Context with a Single Model

链接:https://arxiv.org/abs/2610.08463

作者:Edan Kinderman,Elad Hoffer,Yochai Blau,Brian Chmiel,Ron Banner,Daniel Soudry,Boris Ginsburg

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:single long prompt, handle evidence selection, vastly different scales, long prompt, single long

备注:

点击查看摘要

Abstract:Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.

28. 【2610.08452】Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

链接:https://arxiv.org/abs/2610.08452

作者:Lasse B. Strand,Robert Jakob,Kevin O'Sullivan,Markus Kreft

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:grounding large language, large language models, Retrieval-augmented generation, widely used approach, approach for grounding

备注: Accepted at the Second Workshop for REsearch on Agent Language Models (REALM) at EMNLP 2026 and at the Machine Learning for Systems Workshop at NeurIPS 2026. 9 pages plus references and appendix (16 pages total), 4 figures, 6 tables. Code: [this https URL](https://github.com/Agentic-Systems-Lab/Agentic-AutoRAG)

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.

29. 【2610.08448】Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

链接:https://arxiv.org/abs/2610.08448

作者:Bingxi Hou,Guochao Jiang,Guofeng Quan,Weiqing Li,Wenfeng Feng,Guohua Liu,Yuewei Zhang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:On-Policy Distillation, teacher feedback, teacher, strict, student

备注:

点击查看摘要

Abstract:On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.

30. 【2610.08413】Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

链接:https://arxiv.org/abs/2610.08413

作者:Jerzy Kamiński,Ilya Galyukshev,Artem Kuznetsov,Danil Fedorov,Kirill Redko,Sergey Chuprin,Aidar Shumbalov,Stanislav Chumakov,Anna Kalyuzhnaya

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language models, Large language, language models routinely, routinely answer questions, models routinely answer

备注: 15 pages, 3 figures, 10 tables. Under review

点击查看摘要

Abstract:Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0-MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.

31. 【2610.08388】Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering

链接:https://arxiv.org/abs/2610.08388

作者:Yang Hong,Yajun Yang,Xin Wang,Liping Jing,Qinghua Hu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:demonstrated strong capabilities, Large language models, language models, knowledge-intensive tasks, demonstrated strong

备注: 25 pages, 10 figures. Accepted at NeurIPS 2026

点击查看摘要

Abstract:Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge source for reliable reasoning. However, existing LLM-guided graph reasoning methods typically rely on hop-wise greedy or beam-style pruning during evidence retrieval. Such local decision processes are inherently myopic: evidence that appears weak near the source may become crucial only after deeper graph context is explored, causing answer-critical branches to be discarded prematurely and making the reasoning chain difficult to recover. To address this limitation, we propose Foresight-over-Graph (FoG), a foresight-aware evidence retrieval framework for knowledge base question answering (KBQA). FoG iteratively constructs a question-relevant evidence subgraph and uses far-to-near feedback to guide path exploration, and maintains a compact memory subgraph to support continued exploration. Extensive experiments on widely used KBQA benchmarks demonstrate that FoG achieves state-of-the-art performance, with a particularly large improvement of 16.58% in Hit on CWQ, while also reducing LLM calls and token usage. Our code is available at this https URL .

32. 【2610.08312】CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLMs via Knowledge Consolidation and Decoupling

链接:https://arxiv.org/abs/2610.08312

作者:Maoqi Liu,Quan Fang,Yufei He

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large Language Models, Language Models, Large Language, essential for Large, Continual learning

备注: Accepted to EMNLP 2026 (Main Conference)

点击查看摘要

Abstract:Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-LoRA) to isolate task parameters. However, we reveal that such strict geometric constraints trigger an "Orthogonality Dilemma": rigid parameter isolation impedes the transfer and accumulation of shared representations across semantically related tasks. In this work, we propose a new replay-free method, called Consolidation and Decoupling LoRA (CoDe-LoRA), for CL of LLMs. CoDe-LoRA disentangles the learning process into Consolidating Universal Knowledge and Decoupling Task-Specific Knowledge. To achieve this, CoDe-LoRA leverages an adaptive null space projection mechanism and semantic routing to balance knowledge accumulation with task-specific adaptation. Experimental results across four backbones and three CL benchmarks show that CoDe-LoRA achieves the best average accuracy. Our code is available at this https URL.

33. 【2610.08303】Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation

链接:https://arxiv.org/abs/2610.08303

作者:Shu-Kai Hsieh,Da-Chen Lian

类目:Computation and Language (cs.CL)

关键词:Large Language Models, implicit Translation-Isomorphism Assumption, Current evaluation, multilingual Large Language, loss of information

备注: Position paper. 32 pages (10 pages main text), 6 figures, 12 tables

点击查看摘要

Abstract:Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in principle for a typologically identifiable class of concepts, including pragmatic markers, honorifics, and diachronically stratified terms. We formalize this failure using a usage-cloud framework, representing concepts as point sets of contextualized embeddings. We define $\alpha$-unalignability as the impossibility of any mapping that simultaneously preserves lexical faithfulness (centroid correspondence) and structural faithfulness (local neighborhood topology). We provide three layers of evidence. Behaviorally, we show that FLORES-200 translation failures are predicted by language family and resource class but not by script, and that LOBSTER reasoning scores vary by family. Mechanistically, we report a Representation-Intervention Gap (RIG) in a nine-model case study on Yami: the models' activations encode a regularity along which Yami groups with other low-resource and Austronesian languages, yet interventions on language-specific neurons show no demonstrated advantage over random masks: the regularity is visible but not usable by this intervention. Finally, we operationalize these findings into a multidimensional diagnostic profile: Cycle-Consistency, Pragmatic-Load Disagreement, Manifold-Curvature Mismatch, and RIG. We argue that collapsing cultural competence into a single scalar incentivizes "probabilistic flattening," and that recognizing the unalignable class is a precondition for AI that respects, rather than erases, cultural divergence. This suggests that multilingual alignment is not a single well-defined objective, but a set of mutually incompatible projections.

34. 【2610.08300】Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval

链接:https://arxiv.org/abs/2610.08300

作者:Michael Andreev

类目:Computation and Language (cs.CL)

关键词:modern LLM systems, Long-term conversational memory, LLM systems, Long-term conversational, modern LLM

备注: 4 pages, 1 figure. Accepted at the PALM Workshop at NeurIPS 2026

点击查看摘要

Abstract:Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural depth and the width of context supplied to the answer model. Using EverMemBench, we evaluate depths D1-D4, core budgets of 1,024/2,048/4,096 tokens, and additional Production and Oracle conditions up to the full archive. Increasing width from 1K to 4K improves Accuracy by 10.11-17.98 percentage points, whereas increasing depth provides no monotonic gain. Beyond 8-16K, Production performance reaches a plateau while tokens per correct answer continue to increase; Oracle preserves quality on full archives of 68-71K tokens. These results motivate further investigation of large, coherent context blocks instead of progressively deeper memory structures.

35. 【2610.08208】STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty

链接:https://arxiv.org/abs/2610.08208

作者:Nina Nusbaumer,Iria de-Dios-Flores,Corentin Bel,Christophe Pallier,Guillaume Wisniewski,Benoît Crabbé

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:subject-verb dependency resolution, self-paced reading dataset, long-distance subject-verb dependency, observations isolating, long-distance subject-verb

备注: Will be published at EMNLP 2026

点击查看摘要

Abstract:We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.

36. 【2610.08164】Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models

链接:https://arxiv.org/abs/2610.08164

作者:Seobin Song,Geonho Lee,Janghwan Lee,Jungwook Choi

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Low-rank quantization error, aggressive weight quantization, quantization error compensation, frozen quantized weight, Low-rank quantization

备注: 17 pages, 5 figures

点击查看摘要

Abstract:Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simplifications. They calibrate symmetrically, evaluating the full-precision and compensated weights on the same activation, which yields a compensation target that is inherently high-rank -- so a fixed rank budget captures only a small fraction of it. And they minimize only the second-order term of the loss, although the compensated model is not stationary: a first-order descent direction larger than the applied compensation itself remains in every layer, and no reconstruction objective can absorb it. We propose a two-stage closed-form framework that removes both simplifications. Stage 1 aligns each layer's output with the full-precision model under a Fisher-weighted asymmetric objective, concentrating the rank budget on a rank-compressible target. Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal. Every adapter is the result of a single truncated SVD; backward passes serve only to collect statistics. At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B. On the held-out C4 corpus, it recovers 51% and 84% of the gap to FP16, versus 31% and 63% for the strongest baseline, with consistent gains in the seven-task zero-shot average, at higher bit-widths, and under a distinct quantizer.

37. 【2610.08162】he Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

链接:https://arxiv.org/abs/2610.08162

作者:Tobias Hallmen,Fabian Deuser,Robin-Nico Kampa,Norbert Oswald,Elisabeth André

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Fine-grained emotion recognition, recognition supports therapy, supports therapy tools, emotion recognition supports, Fine-grained emotion

备注: Preprint. 19 pages, 6 figures

点击查看摘要

Abstract:Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at $\kappa_w = 0.468$ on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ($\kappa_w=0.268$-$0.486$). Under verification all eleven clear it, each of them significantly better at $\kappa_w=0.507$-$0.586$. Three also significantly beat EIF sitting at $\kappa_w = 0.551$ (Small; $0.534$ Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $\kappa_w=0.254$-$0.423$. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.

38. 【2610.08161】Symphony for Text Generation: Benchmarking Clinical Note Generation

链接:https://arxiv.org/abs/2610.08161

作者:Daniel Varab,Victor Petrén Bach Hansen,Asbjørn W. Helge,Kevin Pelgrims,Mathias Baltzersen,Adrian Young-San Roessler,Vanessa Klungtvedt,Maximilian Brand,Lasse Krogsbøll,Henrik Cullen,Lars Maaløe

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:rapidly gaining adoption, remains poorly characterized, note quality remains, quality remains poorly, Ambient Clinical Intelligence

备注:

点击查看摘要

Abstract:Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.

39. 【2610.08159】Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation

链接:https://arxiv.org/abs/2610.08159

作者:G. L. John Salvin(1),Swapnil Hingmire(1) ((1) Indian Institute of Technology Palakkad)

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:COMET reports translation, target languages written, assumes Script Invariance, routinely compared, COMET reports

备注: 18 pages, 2 figures. Camera-ready version, accepted at WMT 2026. Code and data: [this https URL](https://github.com/John-salvin/script-bias-comet-normalisation)

点击查看摘要

Abstract:COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.

40. 【2610.08153】Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices

链接:https://arxiv.org/abs/2610.08153

作者:Jinhyeok Kim,Hye-Young Jung

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Multiple-choice question answering, evaluate large language, Multiple-choice question, large language models, question answering

备注: Accepted to AACL-IJCNLP 2026 Main Conference (Short Paper)

点击查看摘要

Abstract:Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly. Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances. These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.

41. 【2610.08095】Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing

链接:https://arxiv.org/abs/2610.08095

作者:Remo Grillo,Lukas Klic,Giovanni Colavizza

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Semantic Web ambition, core Semantic Web, RDF knowledge graphs, Semantic Web, Web ambition

备注:

点击查看摘要

Abstract:Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 $\pm$ 0.026 with GPT-4.1 mini and 0.652 $\pm$ 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.

42. 【2610.08093】SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis

链接:https://arxiv.org/abs/2610.08093

作者:Chuan Li,Chengyu Wang,Cen Chen,Ye Lyu,Mingyuan Fan,Ming Gao

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Medical Question Answering, Question Answering, Developing reliable models, Developing reliable, Medical Question

备注: EMNLP 2026

点击查看摘要

Abstract:Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at this https URL.

43. 【2610.08085】DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs

链接:https://arxiv.org/abs/2610.08085

作者:Hemant Yadav,Sunayana Sitaram,Roger Zimmermann,Rajiv Ratn Shah

类目:Computation and Language (cs.CL)

关键词:exhibit prompt overfitting, Speech-LLMs often exhibit, prompt overfitting, automatic speech recognition, exhibit prompt

备注:

点击查看摘要

Abstract:Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.

44. 【2610.08082】POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

链接:https://arxiv.org/abs/2610.08082

作者:Yunju Kang,Seonghyeon Cho,Irene Li,Yeo-Chan Yoon,Chanjun Park

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:carry operational risk, LLM tool-use agents, LLM tool-use, actions carry operational, tool-use agents operate

备注: Accepted Findings of AACL-IJCNLP 2026

点击查看摘要

Abstract:LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on $\tau^2$-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.

45. 【2610.08077】Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

链接:https://arxiv.org/abs/2610.08077

作者:Haoxiang Zhang,Qinglin Chen,Hiroaki Hayashi,Zhuofeng Li,Siming Zhang,Jiaxin Zhang,Jixuan Chen,Fang Wu,Pan Lu,Silvio Savarese,Julian McAuley,Chien-Sheng Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:scalar outcome rewards, Reinforcement learning, turns agent experience, primarily through scalar, scalar outcome

备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to $24.2$ pp. Its advantage is especially pronounced when reward contrast is scarce: when $37$--$98\%$ of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where $98\%$ of groups are all-failure, the RLVR training ends up at $0.0\%$ success, while adding SRD reaches $60.6\%$ under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

46. 【2610.08055】Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training

链接:https://arxiv.org/abs/2610.08055

作者:Tobias Hallmen,Elisabeth André

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:dyadic counseling conversations, Automatic assessment, expensive to grow, quality in dyadic, conversations is bottlenecked

备注: Preprint. 25 pages, 2 figures

点击查看摘要

Abstract:Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; $n=195$ expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman $\rho = 0.54$ against $\le 0.48$ within the target domain, a paired session-level gap of $+0.15$ that holds at $+0.12$ when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from $0.32$ to $0.41$ over generic dialogue qualities, judges from three model families ensemble to $0.51$ language-only, and a nonverbal-dyadic block adds $+0.03$ more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth $+0.07$ there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.

47. 【2610.08048】DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

链接:https://arxiv.org/abs/2610.08048

作者:Antoine Edy,Max Conti,Victor Xing,Marc-Antoine Allard,Nawfal Benhamdane,Gautier Viaud

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:discover specific tool, specific tool behaviors, LLM agents, lack the operational, act reliably

备注: 9 pages (31 including Appendix), 8 figures (11 including Appendix). We release the code and artifacts, including generation and inference traces, at [this https URL](https://github.com/illuin-tech/daedalus)

点击查看摘要

Abstract:LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, $\tau^2$-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: this http URL.

48. 【2610.08037】Are Language Models Script-Aware?

链接:https://arxiv.org/abs/2610.08037

作者:David Kletz,Sandra Mitrović,Ljiljana Dolamić,Fabio Rinaldi

类目:Computation and Language (cs.CL)

关键词:Language models frequently, Large Language Models, off-target generation, models frequently generate, Language models

备注: Accepted to AACL-IJCNLP 2026

点击查看摘要

Abstract:Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.

49. 【2610.08026】he Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

链接:https://arxiv.org/abs/2610.08026

作者:Jorma Valjakka,Juhani Kivimäki,Juha Mylläri,Jukka K. Nurminen

类目:Computation and Language (cs.CL)

关键词:large language models, recent years, detecting when large, large language, reference answers

备注: 27 pages. Accepted at the NeurIPS 2026 Evaluations Datasets Track. Data: [this https URL](https://doi.org/10.7910/DVN/PCHISZ) . Code: [this https URL](https://github.com/jova486/LPHB)

点击查看摘要

Abstract:In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.

50. 【2610.08018】Structured but Silent: Probing Capability Requirements in LLM Hidden States

链接:https://arxiv.org/abs/2610.08018

作者:Kyojun Choo,Minsoo Song,Yunju Kang,Chanjun Park

类目:Computation and Language (cs.CL)

关键词:API description, triggering a mechanism, mechanism or matching, LLM hidden representations, API

备注: Accepted to AACL-IJCNLP 2026 Findings

点击查看摘要

Abstract:Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper, we investigate whether these query-side capability requirements are linearly decodable from LLM hidden representations prior to generation, and how this hidden-state accessibility compares with explicit verbal classification. We introduce TACIT, a framework that decomposes external requirements along three fundamental axes: Source, Transformation, and World Effect, defining eight structurally distinct capability classes. Using 1,600 balanced training queries from benchmarks, synthetic examples, and new domain scenarios, we train linear probes on pre-generation hidden states from four open-weight LLM families. Our empirical results demonstrate that fine-grained capability structures are linearly decodable with high accuracy across all models. Crucially, however, we expose a representation-to-verbalization gap: these same models are significantly less reliable when asked to explicitly classify the same queries in natural language. This disconnect indicates that information about required external capabilities is linearly accessible in LLM hidden representations but not reliably expressed, a phenomenon we define as "structured but silent."

51. 【2610.07990】A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

链接:https://arxiv.org/abs/2610.07990

作者:Sin-Han Yang,Shih-Cheng Huang,Chieh-Yen Lin,Yun-Nung Chen,Shao-Hua Sun,Hung-yi Lee

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:task-specific weight updates, individual task-specific models, Model merging aims, aims to build, cheaply by combining

备注: Preprint

点击查看摘要

Abstract:Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.

52. 【2610.07987】VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

链接:https://arxiv.org/abs/2610.07987

作者:Yuan Feng,Qize Yang,Ruizhe Chen,Sibo Song,Haolin He,Muzhi Zhu,Zihan Liu,Yunfei Chu,Xize Cheng,Yuxuan Wang,Jin Xu,Xike Xie

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:incur substantial costs, fixed-size patch tokens, Multimodal large language, large language models, inputs into dense

备注:

点击查看摘要

Abstract:Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.

53. 【2610.07948】Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents

链接:https://arxiv.org/abs/2610.07948

作者:Brendan King,Farima Fatahi Bayat,Jean-Flavien Bussotti,Pouya Pezeshkpour,Estevam Hruschka

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:intervene requires calibrated, consequential domain, trust its output, output or intervene, intervene requires

备注: 34 pages, 6 figures, 11 tables

点击查看摘要

Abstract:When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of an agent's trajectory. Practical agentic deployments introduce further challenges: frontier LLMs often provide limited access to internal signals, agent roll-outs are costly, and training data may be unavailable or quickly become outdated. To address these challenges, we introduce Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an agent accomplished its task from a single trajectory, without privileged model access or training data. Rather than compressing an execution into a single holistic judgment, a CRG begins with the claim that the agent accomplished its task, decomposes it into contextualized sub-claims grounded in trajectory evidence, estimates confidence for each terminal claim, and finally aggregates these into an overall confidence estimate. Across three agentic benchmarks, three backbone models, and three agent frameworks, CRGs yield better-calibrated confidence and stronger risk-aware decision making than verbalized, sampling-based, and white-box surrogate baselines. We further find that calibration error alone can be misleading: a white-box surrogate baseline appears well calibrated while providing near-chance discrimination. Ablations attribute CRG's improvements to claim-level confidence estimation and aggregation rather than graph construction alone. Finally, a CRG exposes the claims and trajectory evidence underlying each confidence estimate, enabling it to be audited at decision time.

54. 【2610.07940】Hybrid Latent Attention for Looped Language Models

链接:https://arxiv.org/abs/2610.07940

作者:Yuhan Chen,Siyuan Zhang,Nan Wang,Feiyang Kang,Ruoxi Jia

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:larger cache limits, Hybrid Latent Attention, language models apply, Looped language models, propose Hybrid Latent

备注:

点击查看摘要

Abstract:Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.

55. 【2610.07937】Leveraging a four-quadrant approach for evaluating Redpine Science

链接:https://arxiv.org/abs/2610.07937

作者:Filip Dorm,Leonora Vesterbacka

类目:Computation and Language (cs.CL)

关键词:Model Context Protocol, Context Protocol, Redpine Science, single access point, Model Context

备注:

点击查看摘要

Abstract:Redpine Science gives models and agents a single access point to a wide range of peer-reviewed literature, queried directly through the Model Context Protocol (MCP) and an API. This report evaluates Redpine Science on two levels: the relevance of the retrieved chunks, and a model's answer when it has access to Redpine Science compared to web search. Both public and expert-validated benchmarks are used. Public benchmarks are a widely accepted way to test model development and are comparable across labs, but risk saturation and memorization. To address this, we complement them with an expert-validated question set. In total, this report presents four evaluations. On ScholarQABench SciFact, the public answer-quality benchmark reported here, an agent with Redpine Science answers 94.4% of claims correctly against 87.6% with no retrieval. On the expert-validated question set, an agent with Redpine Science states 80.1% of the required claims against 70.2% for an agent restricted to web search. On the 668 queries of a public retrieval benchmark whose gold paper Redpine holds, stripped of any model reasoning, Redpine Science places the correct source paper in its top ten results for 83.1% of queries (Recall@10), against 79.3% for the benchmark's creator. A blinded expert relevance panel places Redpine Science's Precision@5 at 75.2% against 39.8% for the PubMed search tool. We release the expert-validated question set and instructions to reproduce every headline result above, at this https URL.

56. 【2610.07936】Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing

链接:https://arxiv.org/abs/2610.07936

作者:Jing Chen,Giulia Loca,Simona Amenta,Marco Marelli

类目:Computation and Language (cs.CL)

关键词:form to meaning, permeates language, probabilistic mapping, mapping of form, Systematicity

备注:

点击查看摘要

Abstract:Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseudoword processing. Yet whether LLMs exhibit comparable sensitivity to these cues remains unclear. We tested five LLMs on two Italian two-alternative forced-choice pseudoword experiments and compared their responses with a human behavioural baseline. LLMs aligned more reliably with humans when real-word options provided a lexical familiarity cue than in the pseudoword-only condition, where they fell substantially below fastText, a character-n-gram model. In addition, the sublexical cosine-similarity cue that reliably drove human--fastText agreement did not consistently transfer to human--LLM alignment, and reasoning-token expenditure bore no consistent relation to human processing difficulty. These findings suggest that LLMs do not necessarily share the sublexical cues that govern human pseudoword processing; we discuss tokenization and training-data coverage as candidate explanations.

57. 【2610.07906】Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations

链接:https://arxiv.org/abs/2610.07906

作者:K. P. Santoso,N. Z. Fadil,F. P. Harsanti,R. V. H. Ginardi,G. N. Iyer

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:sufficiency by investigating, representation retains, study sequential content, sequential content sufficiency, separates input ambiguity

备注:

点击查看摘要

Abstract:We study sequential content sufficiency by investigating whether a representation retains the ordered target information available in its input. An information-theoretic decomposition separates input ambiguity, representation loss, and readout mismatch. We construct recoverable views where perfect agreement and joint isotropic Gaussianity coexist with zero target information, and establish limits imposed by deterministic canonical anchors. Token log-loss provides a one-sided information-loss bound; a fixed-penalty ridge analysis shows why rank alone cannot determine prediction risk. These results motivate CANOPE, a nonautoregressive framework with ordered latent canvases, canonical-token supervision, and geometric regularization. On 40,000 validation sequences, latent-agreement (PL0) and token-grounded (PL2) have nearly identical pooled ranks but reach 13.5% and 98.8% positional Recall@1, respectively, under strong natural corruption when the correct target length is provided. On 3,930 LJSpeech validation utterances, frozen PL2 with a trained MatchaTTS readout yields 21.54% word error rate (WER) on corrupted text, versus 99.22% for frozen PL0, while end-to-end MatchaTTS reaches 10.93%. These results show that geometric regularity alone does not guarantee recoverable sequential content or effective downstream access in the text settings studied here.

58. 【2610.07902】ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring

链接:https://arxiv.org/abs/2610.07902

作者:Shengyu Li,Jinting Wang,Li Liu

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:writing requires close, requires close alignment, lyric writing requires, Cantonese lyric writing, writing requires

备注: Accepted for publication in Findings of EMNLP 2026. 24 pages, including references and appendices. Author-prepared version

点击查看摘要

Abstract:Cantonese lyric writing requires close alignment between lexical tones and melodic pitch. Existing melody-guided lyric generation methods typically rely on symbolic melody to generate lyrics. However, in real songwriting scenarios, melodies are often expressed as raw singing audio or hummed recordings, where pitch is implicit, noisy, and unstructured, making these methods difficult to apply directly. To address this limitation, we propose ARIA, a two-stage audio-driven melody-tone relation modeling framework for Cantonese lyric authoring that generates Cantonese lyrics from singing recordings with provided character-level timestamps. Specifically, we first design a Tri-Stream Relation-Aware Tone Estimator (TRATE) to predict 0243 sequences from timestamped singing audio by modeling multi-stream acoustic cues and relational tonal structure. We then propose a Decoupled Retrieval-Augmented Tone-Conditioned Lyric Generator (DRA-TCLG) to generate fluent lyrics conditioned on predicted tonal plans with retrieval-enhanced lexical guidance. Moreover, we construct a large-scale aligned audio-Jyutping-0243 dataset from real Cantonese singing recordings to support this new task. Experimental results demonstrate that ARIA achieves strong performance in both 0243 prediction and tone-consistent lyric generation, validating the effectiveness of the proposed framework.

59. 【2610.07894】Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective

链接:https://arxiv.org/abs/2610.07894

作者:Zizhuo Zhang,Xiong Peng,Jingwei Sun,Rong Yao,Borui Jiang,Bo Han

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large language models, Large language, questions faithfully based, faithfully based, answer questions faithfully

备注: 22 pages

点击查看摘要

Abstract:Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts. Our evaluations across thirty-nine models with seven model families demonstrate that faithfulness fundamentally involves a trade-off between answering and abstaining, and that most current models exhibit a strong bias toward answering, with most faithfulness errors arising from over-answering, i.e., models tend to fabricate a response even when the provided context is insufficient. We further conduct a series of studies on faithfulness training under different data constructions. Our results show that training outcomes are highly sensitive to the specific composition of answering and abstaining data. Constructing answering and abstaining data from mismatched sources can cause models to rely on dataset-specific shortcuts rather than actual context sufficiency. Moreover, increasing answer-supervised data improves answering performance but exacerbates over-answering, while increasing abstaining data reduces hallucination but leads to over-abstention. The code and data are released at this https URL.

60. 【2610.07887】Visual Abstention in Unified Multimodal Models

链接:https://arxiv.org/abs/2610.07887

作者:Chufan Shi,Cheng Yang,Tiannuo Yang,Isadora White,Yiwei Chen,Taylor Berg-Kirkpatrick,Xuezhe Ma

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Unified multimodal models, Unified multimodal, integrate understanding, understanding and generation, generative behavior

备注: 25 pages, 6 figures, 13 tables. Project page: [this https URL](https://visual-abstention.github.io)

点击查看摘要

Abstract:Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.

61. 【2610.07863】ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

链接:https://arxiv.org/abs/2610.07863

作者:Yupeng Su,Jiayi Tian,Zheng Zhang,Souvik Kundu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:append-only interaction history, LLM agents act, context window, context, additional model calls

备注: 27 pages, 6 figures, 14 tables

点击查看摘要

Abstract:Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.

62. 【2610.07853】Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

链接:https://arxiv.org/abs/2610.07853

作者:Avichal Sahai(Ofbusiness),Nishant Raj(Ofbusiness),Animesh Srivastava(Ofbusiness)

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:ternary codes produced, Ternary language models, higher-precision latent weights, Ternary language, labs' documented pipelines

备注: 14 pages, 4 figures, 17 tables

点击查看摘要

Abstract:Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16. We audit those pipelines across three labs. In released checkpoints, fp32 quantization of the shipped latents disagrees with the deployed codes on 0.83-1.77% of codes in Falcon-E and BitCPM and on 1.530% in BitNet 2B-4T; for Falcon-E and BitCPM most disagreements are products that bf16 rounding lands exactly on the threshold, which ties-to-even maps to zero, and the unmodified onebitllms exporter reproduces all four Falcon-E releases byte for byte. At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T's strict accuracy by 27.54 points while its last-number accuracy rises. Two compatibility remedies, writing the training quantizer's codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models. In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.

63. 【2610.07848】Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models

链接:https://arxiv.org/abs/2610.07848

作者:Dayan Pan,Jingyuan Wang,Xie Yu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:adapting large language, Parameter-efficient fine-tuning, large language models, downstream tasks, standard approach

备注: Accepted by KDD 2026

点击查看摘要

Abstract:Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.

64. 【2610.07847】OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement

链接:https://arxiv.org/abs/2610.07847

作者:Sihyeon Lee,Jihun Song,Chanwoo Kim,Jiwoo Kum,Chanjun Park

类目:Computation and Language (cs.CL)

关键词:reverse substantive outcomes, equivalent framings reverse, framings reverse substantive, LLMs increasingly assist, substantive outcomes

备注: Accepted to AACL-IJCNLP 2026 Findings

点击查看摘要

Abstract:As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a significant risk of skewed decision-making. Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts. To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel (utilitarianism, deontology, virtue ethics, care ethics, and contractualism). Evaluating eight LLMs, we find that omission bias is pervasive but inversely correlates with model size within families. We further evaluate four inference-time interventions and find that interventions encouraging models to consider moral principles before committing to a yes/no answer reduce omission bias and increase frame-consistent responses, although lower omission bias rates can also coincide with shifts toward action-biased responses. Ultimately, this work contributes not only the OMIT benchmark, but also a methodology for using diverse philosophical disagreement signals to evaluate framing-sensitive inaction preferences and the distributional effects of mitigation attempts in LLMs under complex moral conflicts.

65. 【2610.07832】Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

链接:https://arxiv.org/abs/2610.07832

作者:Haibo Jin,Xinjie Li,Peng Kuang,Haohan Wang

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:Large language models, demonstrated strong capabilities, equipped with terminal, terminal access, access have demonstrated

备注: 30 pages

点击查看摘要

Abstract:Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.

66. 【2610.07822】Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution

链接:https://arxiv.org/abs/2610.07822

作者:Shuhao Li,Fanghua Ye,Wanyu Lin,Tianyu Yuan,Xiaoyu Shen

类目:Computation and Language (cs.CL)

关键词:Speculative decoding, accelerates autoregressive generation, Speculative decoding accelerates, decoding, lightweight draft model

备注:

点击查看摘要

Abstract:Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel. However, the standard acceptance rule focuses on exact distribution correction and rejects tokens that remain highly plausible under the target model when the draft model assigns excess probability. This conservative verification limits the number of draft tokens retained after each verification forward pass. We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding. NSD accepts a draft token if it satisfies the standard acceptance rule or belongs to the target model's nucleus. We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus. We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding. Experiments across multiple target models and proposal mechanisms demonstrate that NSD consistently improves speculative decoding efficiency while maintaining competitive task performance. Our method achieves throughput speedups of up to $5.16\times$ over autoregressive decoding and up to $3.15\times$ over standard speculative decoding. These improvements coincide with longer accepted lengths, allowing more output tokens to share the cost of each target verification pass. Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency. Our code is available at this https URL.

67. 【2610.07819】$α$Transfer: Coefficient Transfer for Efficient Model Merging

链接:https://arxiv.org/abs/2610.07819

作者:Shih-Cheng Huang,Zhi Rui Tam,Chieh-Yen Lin,Yun-Nung Chen,Hung-yi Lee,Shao-Hua Sun

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:combining multiple fine-tuned, multiple fine-tuned checkpoints, parameter arithmetic, offers a promising, promising solution

备注: Under review

点击查看摘要

Abstract:Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$\alpha$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $\alpha$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $\alpha$Transfer as an efficient and generalizable approach to scaling model merging.

68. 【2610.07817】One Step at a Time: Trading LLM Autonomy for Process Predictability

链接:https://arxiv.org/abs/2610.07817

作者:Hans Schabert,Christoph Peters

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Organizations automating operational, automating operational processes, Organizations automating, Model Context Protocol, step

备注: 14 pages, 12 tables

点击查看摘要

Abstract:Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.

69. 【2610.07803】hinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models

链接:https://arxiv.org/abs/2610.07803

作者:Myunghoon Kang,Jungseob Lee,Jaehyung Seo,Heuiseok Lim

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:shown strong performance, Small reasoning models, complex reasoning tasks, Small reasoning, generating extended

备注: Accepted to EMNLP 2026 Findings

点击查看摘要

Abstract:Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model's trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at this https URL.

70. 【2610.07782】Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

链接:https://arxiv.org/abs/2610.07782

作者:Hochan Son,Kyungdoe Han,Jaehan Koh,Xiaowu Dai,Wenlu Xu,Guang Cheng

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)

关键词:Decomposing long-context inference, KV-cache memory binds, cooperating agents bounds, Decomposing long-context, memory binds

备注: 13 pages, 1 figure. Accepted as a poster at the Machine Learning for Systems Workshop, NeurIPS 2026

点击查看摘要

Abstract:Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching it took four measurement corrections -- three inflating the apparent benefit, the fourth making an effect that size look resolvable -- none visible in the results table. We give the conditions an agent-memory ablation must satisfy and detection procedures that need no knowledge of the specific defect.

71. 【2610.07781】Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs

链接:https://arxiv.org/abs/2610.07781

作者:Yuhe Hu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Post-training quantization reduces, deploying language-model agents, temporary tool failures, Post-training quantization, Llama

备注: Accepted at the NeurIPS 2026 Workshop on Small Language Models for Agentic Systems (SLM-Agents). 7 pages, 2 figures, 2 tables, plus appendix

点击查看摘要

Abstract:Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets. On tasks that both variants complete without faults under the same prompt, the difference ranges from 0 to +20.2 percentage points for Llama and from -50.0 to +35.0 points for Qwen. Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts. The evaluation target can also reverse the result. For Llama under one prompt, scoring each variant only on its own clean-passing tasks favors 4-bit by 17.5 points; scoring the same tasks for both variants gives no difference, while scoring the full pipeline favors 8-bit by 28.3 points. Executor leniency is a third such choice. Rescoring the same logs with strict output parsing, which 8-bit Llama violates far more often than 4-bit Llama under that prompt, turns that +28.3 into -15.0 while leaving Qwen essentially unchanged. These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness. Evaluations should compare variants on matched tasks, report full-pipeline success for deployment decisions, state the scoring policy, and quantify uncertainty across tasks rather than injected fault sites.

72. 【2610.07780】APEX: Speculate smarter, not deeper

链接:https://arxiv.org/abs/2610.07780

作者:Manvi Jha,Zach Zhang,Zhichao Xu,Linbo Liu,Sai Muralidhar Jayanthi,Vinayak Arannil

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Speculative decoding reduces, reduces large language, decoding reduces large, large language model, language model inference

备注:

点击查看摘要

Abstract:Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback. APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens. This allows the controller to adapt speculation while retaining the target model's verification procedure. We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding. Across the aggregate evaluation, APEX-S achieves 4.27X speedup, while APEX-B achieves 3.27X speedup with a 41.0% relative reduction in wasted-token percentage compared with fixed n-gram speculation at k=16, providing distinct operating points for balancing acceleration and draft-token utilization.

73. 【2610.07774】Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

链接:https://arxiv.org/abs/2610.07774

作者:Ziyuan Yang,Wenxuan Ding,Shangbin Feng,Yulia Tsvetkov

类目:Computation and Language (cs.CL)

关键词:harmful intent emerges, face compositional safety, compositional safety risks, Vision-language models, face compositional

备注: 15 pages, 4 tables, 11 figures

点击查看摘要

Abstract:Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.

74. 【2610.07767】RACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

链接:https://arxiv.org/abs/2610.07767

作者:Xin Wang,Hao Yu,Zhengyang Zhuge,Bochao Mao,Zheng Li,Junda Feng,Yuyan Luo,Yi Zhang,Yizhong Cao,Mi Zhang,Dayiheng Liu,Jianwei Zhang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:incurs substantial computation, Reinforcement learning, post-training large language, motivates low-precision rollout, incurs substantial

备注:

点击查看摘要

Abstract:Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.

75. 【2610.07764】No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays

链接:https://arxiv.org/abs/2610.07764

作者:Daniel Kua,Emrul Hasan,John-Jose Nunez,Frances Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Natural language processing, written twelve years, twelve years earlier, Child Development Study, text written twelve

备注:

点击查看摘要

Abstract:Natural language processing (NLP) models can detect depression-related language in text written near the time symptoms are measured, but whether pretrained transformers can predict depressive symptoms from text written twelve years earlier is largely untested. In the National Child Development Study, a British birth cohort, we predict probable depressive symptoms at age 23 from essays the same people wrote at age 11. Our baseline, a logistic regression on six childhood covariates, outperforms every text model that sees only the essay: seven fine-tuned transformers, a bag-of-words model, frozen embeddings and four zero-shot large language models. Its area under the receiver operating characteristic curve (AUC-ROC) is 0.737 against 0.670 for the best transformer on the primary seed, and no added text score detectably raises the baseline's AUC-ROC. None of the five domain-pretrained transformers detectably beats its general-domain control after Bonferroni correction. For long-horizon prediction, the baseline remains the model to beat.

76. 【2610.07753】From Evidence to Action: How Tool-Using Agents Fail

链接:https://arxiv.org/abs/2610.07753

作者:Hongzhan Lin,Shidong Cao,Ziyang Luo,Wenhao Chai,Mong-Li Lee,Wynne Hsu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Tool-using agents make, agents make consequential, Tool-using agents, external state, make consequential

备注: 36 pages. Project page: [this https URL](https://safeact.github.io)

点击查看摘要

Abstract:Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.

77. 【2610.07731】Learning to Retrieve via Reinforcement Learning in Embedding Space

链接:https://arxiv.org/abs/2610.07731

作者:Qi Liu,Fengming Liang,Yiqun Chen,Erhan Zhang,Jiaxin Mao

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Dense retrieval models, learn effective representations, optimize retrieval metrics, Dense retrieval, directly optimize retrieval

备注:

点击查看摘要

Abstract:Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.

78. 【2610.07730】SanSi: A Looped Typed Decision Model for System 1.5 Thinking

链接:https://arxiv.org/abs/2610.07730

作者:Shuyu Gan,Young-Jun Lee,Dongyeop Kang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:decision head returns, single forward pass, declared question, head returns, returns a probability

备注: 43 pages, 15 figures, 42 tables. Project page: [this https URL](https://minnesotanlp.github.io/Sansi/)

点击查看摘要

Abstract:Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.

79. 【2610.07722】Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods

链接:https://arxiv.org/abs/2610.07722

作者:Haotian Yang,Huikang Jiang,Yucheng Wu,Wen-Jie Jiang,Chenpeng Wang,Yibin Lou,Liangming Pan

类目:Computation and Language (cs.CL)

关键词:side effects, control large language, steering, Activation steering, side

备注:

点击查看摘要

Abstract:Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization and data dependence through steering-specific metrics for sample efficiency and sample sensitivity. Rather than comparing methods at a single operating point, we characterize the trade-offs between efficacy and side effects. Under matched models, tasks, and evaluation protocols, we benchmark 23 methods spanning 4 families, including prompting, LoRA, and SFT as baseline methods, and release the suite as an extensible codebase. We find that current activation steering methods do not yet surpass the Prompt Steering baseline in their overall balance between steering efficacy and side effects: across both model scales, no evaluated activation steering method achieves higher efficacy without incurring greater composite side effects. We further uncover a consistent coupling between steering efficacy and side effects. Under OOD prompts, target efficacy is often preserved, whereas side effects tend to become more pronounced, particularly through declines in instruction relevance and fluency. Methods also exhibit sharply different sample-efficiency profiles.

80. 【2610.07716】Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation

链接:https://arxiv.org/abs/2610.07716

作者:Ran Li,Lei Chen

类目:Computation and Language (cs.CL)

关键词:Jev model score, Prefill-only decision models, Jev model, single forward pass, Prefill-only decision

备注:

点击查看摘要

Abstract:Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label this http URL and data are available at this https URL.

81. 【2610.07715】When Old Facts Return: Re-Reads, Reverts, and the Limits of Temporal Memory

链接:https://arxiv.org/abs/2610.07715

作者:Neeraj Yadav(Called It Inc.)

类目:oftware Engineering (cs.SE); Computation and Language (cs.CL)

关键词:system can retire, retire an obsolete, memory system, Abstract, statement reduces accuracy

备注: 12 pages, 1 figure. Ancillary files contain retained aggregate evidence, derived scenario and annotation exports, reference code, and an offline verifier

点击查看摘要

Abstract:A memory system can retire an obsolete value and later restore it merely because the same old statement appears again. A re-read of an old source and a genuine revert can produce the same observed sequence of values while requiring opposite current answers. We study this ambiguity on 130 extractor-selected atomic transitions derived from software fixes. In the ordinary transition condition, identity-based temporal memory reaches 98.5% model-judged accuracy with zero observed errors under a literal stale-value proxy. Appending a verbatim re-read of the old statement reduces accuracy to 10.8% and raises the stale-value rate to 88.5%. A guard that refuses to reactivate a previously retired value restores accuracy to 97.7% and reduces that rate to 0.8% in this constructed re-read condition. The guard cannot also recognize a legitimate revert without additional change provenance. Two supporting studies examine exposing retired history to the answer model and supplying current source for changed behavior. An exploratory extraction study over 707 software fixes provides scope context, not a universal coverage estimate. The design implication is to distinguish an observation of a value from evidence that the value changed. Selected inputs, aggregate-only answer records, related-family judges and a post-failure guard evaluation limit the conclusions to the retained experiments.

82. 【2610.07700】Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulation

链接:https://arxiv.org/abs/2610.07700

作者:Thanh Dong,An Ngo,Minh Dau,Rajesh Kumar

类目:Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:detecting large language, large language model, dynamics for detecting, detecting large, large language

备注: 9 pages, 2 figures. Thanh Dong and An Ngo contributted equally. Accepted at the 2026 IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026)

点击查看摘要

Abstract:We study the robustness of keystroke dynamics for detecting large language model (LLM)-assisted writing. We introduce a Vietnamese keystroke dataset capturing realistic writing modes, including bona fide composition, transcription, and paraphrasing. We also define a behaviorally grounded threat model in which users deliberately alter typing patterns. To implement the threat model, we create behaviorally manipulated variants of the data designed to evade keystroke-based detection. We evaluate four keystroke modeling approaches: temporal and rhythmic representations, and sequential representations modeled with a one-dimensional convolutional neural network (1D-CNN) and TypeNet, under user-independent and context-independent settings. The results show that sequential models outperform feature-based approaches in most cases and that keystroke signals encode discriminative information about the writing process. However, detection is not uniformly robust: transcription is reliably identified, while paraphrasing and adversarially manipulated samples are frequently misclassified as bona fide when not explicitly modeled. To address this, we incorporate adversarial training using behaviorally manipulated data, which substantially improves separability and robustness. These results suggest that keystroke-based detection depends critically on exposure to diverse writing behaviors, and that strong performance under limited conditions does not generalize to realistic or adversarial settings without targeted modeling.

83. 【2610.07699】Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning

链接:https://arxiv.org/abs/2610.07699

作者:Zhijun Zhang,Qianlong Wang,Keyang Ding,Genan Dai,Bowen Zhang,Bin Liang,Ruifeng Xu,Yongsheng Liang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:high-quality structure-annotated datasets, Argument Mining, fundamentally constrained, scarcity of high-quality, high-quality structure-annotated

备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we revisit synthetic data generation for AM from a new perspective and propose a novel adversarial reinforcement learning framework for data synthesis. The proposed framework jointly optimizes the generator and the discriminator in an adversarial loop, in which the generator produces structured AM instances, and the discriminator provides learning signals by distinguishing real data from synthetic candidates. This enables the generator to progressively improve both the structural accuracy of generated argument data while maintaining diversity through adversarial feedback. Extensive experiments demonstrate that the proposed framework consistently improves AM performance on three benchmark datasets in both full-data and low-resource settings, validating its effectiveness and scalability.

84. 【2610.07659】DLoop: Looped Speculative Decoding

链接:https://arxiv.org/abs/2610.07659

作者:Geonmo Gu,Byeongho Heo,HeeJae Jun,Yoohoon Kang,Sangmin Lee,Sangdoo Yun,Dongyoon Han

类目:Computation and Language (cs.CL)

关键词:draft, large language models, accelerates autoregressive generation, drafting, generation in large

备注: 22 pages

点击查看摘要

Abstract:Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL.

85. 【2610.07657】Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

链接:https://arxiv.org/abs/2610.07657

作者:Shaswata Mitra,Raj Patel,Subash Neupane,Sudip Mittal,Md Rayhanur Rahman,Shahram Rahimi

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)

关键词:encountering adversarial content, LLM-based multi-agent systems, engage tools, share memory, LLM-based multi-agent

备注: 26 pages, 20 figures, 24 tables

点击查看摘要

Abstract:LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.

86. 【2610.07647】Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

链接:https://arxiv.org/abs/2610.07647

作者:Seymanur Akti,Alexander Waibel

类目:ound (cs.SD); Computation and Language (cs.CL)

关键词:humans naturally adapt, noisy environments, voice to compensate, humans naturally, naturally adapt

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.

87. 【2610.07643】Monte Carlo Estimation for KV Cache Eviction

链接:https://arxiv.org/abs/2610.07643

作者:Ahsan Bilal,Muhammad Ahmed Mohsin,Muhammad Umer,Wajih Hassan Raza,Atta Ul Asad,Young D. Kwon,Michal Valko,Dean F. Hougen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:memory appeared important, reading the prompt, appeared important, important while reading, KV-cache eviction methods

备注:

点击查看摘要

Abstract:Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projected leave-one-out attention-output deletion cost and aggregated across sampled futures with optional trajectory weighting. The temporary continuations are discarded before final decoding, requiring no auxiliary model or training. Ablations isolate the mechanism: at B=128, a single response-side continuation recovers about 89% of the gain over the prompt-window control, while additional futures provide smaller improvements. At B=128, LORE-KV raises the LongBench average on Qwen2.5-14B from 45.49 to 48.24 (+2.75) and the 16K RULER average on Mistral-7B from 45.20 to 51.05 (+5.85). Gains diminish at larger cache budgets and coexist with task-level regressions. LORE-KV incurs 1.46-2.77x AnDPro's per-sample wall-clock time as a one-time compression overhead across six dense and hybrid-attention backbones.

88. 【2610.07625】Stateless Language Agents: Scaling Long-Horizon Automated Research

链接:https://arxiv.org/abs/2610.07625

作者:Qizheng Zhang,Changxiu Ji,Isaac Sun,Yuetai Li,Shubhangi Upasani,Sherry Ruan,Boyuan Ma,Fenglu Hong,Vamsidhar Kamanuru,Yoonho Lee,Yuzhen Mao,Genghan Zhang,Rulin Shao,Qiuyang Mang,Andy Dimnaku,Changran Hu,Radha Poovendran,Kunle Olukotun

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:replay growing histories, increasingly run LLM, Automated research systems, run LLM agents, agents replay growing

备注: 32 pages

点击查看摘要

Abstract:Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA's progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.

89. 【2610.07591】Recurrent Looped Transformer

链接:https://arxiv.org/abs/2610.07591

作者:Yifan Zhang,Jichen Feng,Shihan Qin

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Recurrent Looped Transformer, State tracking requires, Transformer, requires an update, Transformer applies

备注: Project Page: [this https URL](https://github.com/yifanzhang-pro/recurrent-looped-tranformer)

点击查看摘要

Abstract:State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based $S_5$ permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based $S_5$ to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based $S_5$ from 100% to 20%.

90. 【2610.07587】Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

链接:https://arxiv.org/abs/2610.07587

作者:Shuqing Shi,Ziyan Wang,Milind Tambe,Yali Du

类目:Computation and Language (cs.CL)

关键词:maximize collective welfare, achieve common goals, LLM orchestration investigates, collective welfare, coordinates a group

备注:

点击查看摘要

Abstract:LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent's preference and updating it from the agents' observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbf{HARP} (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes' rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same $\tilde O(\sqrt K)$ Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript{+} augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript{+} is the strongest non-oracle method across the class our theory identifies.

91. 【2610.07580】LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems

链接:https://arxiv.org/abs/2610.07580

作者:Muhammad Faraz Shoaib,Muhammad Qasim,Raisulhaq Mohammed Rizwan,Rahmatullah Safdar,Muzammil Adnan Shaik,Abdul Aleem Mohammed

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Aerospace electrical-design revisions, Aerospace electrical-design, electrical-design revisions, multiple genuine, Aerospace

备注: 14 pages, 4 figures, 4 tables

点击查看摘要

Abstract:Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected difference can therefore produce overly broad impact reports. We present LOGIC, a controlled benchmark and evaluation framework in which locally deployable language models ground a request in a deterministic candidate-change inventory before selected changes are propagated through a typed electrical traceability graph. This separation permits candidate-selection errors to be distinguished from downstream propagation errors. LOGIC contains 168 scenarios, including 144 selection and 24 abstention cases. We evaluate three 7--8B models against intent-agnostic, lexical, and structured-evidence methods, with an oracle-root upper bound. On 96 explicitly anchored selection cases, gate-only structured evidence achieves candidate F1 of 1.0000, compared with 0.9677 for token-lexical matching. On 12 relational-paraphrase cases, token-lexical F1 is 0.1772 and gate-only F1 is 0.0000, compared with 0.5000--0.6400 for the large language models. Model grounding degrades as candidate inventories grow from 4 to 64 changes, while affected-element and typed-path accuracy remain comparatively stable when frozen selections are replayed over graphs of approximately 1K to 100K nodes. Strict evidence gating suppresses false positives but can remove correct semantic selections. An exploratory evidence-empty abstention policy raises strict abstention accuracy to 0.6667 for all three models and reduces unsafe-report rates to 0.1667, while decreasing answerable-case coverage by 16.0--27.1 percentage points. Four of six conflicting requests remain unsafe for each model. These findings support combining literal evidence and language-model reasoning with engineering review when intent cannot be established reliably.

92. 【2610.07572】wo Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

链接:https://arxiv.org/abs/2610.07572

作者:Xi Ding,Naichen Shi,Jiawei Zhang

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:In-context learning, adapts frozen large, adapts frozen, hundreds of visual, demo image adds

备注: Technical report

点击查看摘要

Abstract:In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.

93. 【2610.07563】HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior

链接:https://arxiv.org/abs/2610.07563

作者:Jin Huang,Diego Ferreras Garrucho,Yutong Xie,Walter M. Yuan,Qiaozhu Mei,Chen Lian,Jonathon Hazell

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, household decision making, decision making, variety of settings

备注:

点击查看摘要

Abstract:Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6 U.S. household surveys and 32 prediction tasks spanning numeric, categorical and probabilistic outcomes, related to consumption, income, labor, expectations, and housing. Using past behavior, demographics and macroeconomic conditions, the tasks test whether LLMs predict behavior, including how households adjust to changes in various policies. We evaluate 13 proprietary and open-weight LLMs against a no-change baseline and a gradient-boosted tree model. Most LLMs outperform the no-change baseline, including for policy response tasks -- with the best model lowering error for numeric outcomes by 12.2%. Across most tasks, gradient-boosted trees rank first; leading proprietary LLMs approach their performance, but open-weight models lag. LLMs exhibit systematic over- and underprediction across different tasks. We identify methods that enable a 4 billion parameter open-weight model to match proprietary models' performance: fine-tuning and aggregating 16 predictions per observation. Improvements generalize to policy-response tasks, which are excluded from fine-tuning. We release our datasets, code, and leaderboard on our website: this https URL

94. 【2610.07545】Quality-Aware Self-Correcting Speech Translation on an Edge Device

链接:https://arxiv.org/abs/2610.07545

作者:Zubair Ajmal Farooq,Diptesh Kanojia

类目:Computation and Language (cs.CL)

关键词:Jetson Nano, Whisper-tiny ASR feeds, fully offline, present a fully, pipeline that runs

备注: 7 pages, 2 figures, 4 tables. Full paper submitted to the Convergence 2026 proceedings; poster presented at Convergence 2026, University of Surrey. Code: [this https URL](https://github.com/juebae/speech-translation_edge_device)

点击查看摘要

Abstract:We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retraining. A Whisper-tiny ASR feeds an Opus-MT translator; multilingual BERT cosine similarity acts as a Quality Estimation (QE) gate, triggering a secondary-pass correction when confidence falls below a pre-defined threshold $\tau$. We compare three correction methods: QE reranking (M1), Minimum Bayes-Risk decoding (M2), and constrained beam search (M3). On 1,012 FLORES-200 sentences (English-Spanish), M2 at $\tau=0.90$ produces statistically significant improvements over greedy decoding on BLEU (+0.67, p0.001), ChrF (+0.51, p0.001), and COMET (+0.0020 at N=3, p=0.002); M1 yields no significant gains, and M3 is significantly worse than baseline (p0.99). Our central finding is that QE functions effectively as a gate but poorly as a ranker: removing the QE model from candidate selection (M1$\to$M2) does not hurt quality and frees 680 MB from the critical path. Using a gain-to-edit ratio adapted from the post-editing-effort literature, we further show that smaller candidate pools (N=3) yield more surgical corrections with better semantic adequacy, while larger pools (N=10) maximise lexical reward. We release the system and demonstrate live translation across six language pairs.

95. 【2610.07535】Disentangling Models from Personas in Heterogeneous LLM Simulations

链接:https://arxiv.org/abs/2610.07535

作者:Dani Roytburg,Daphne Ippolito

类目:Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:large language models, simulations with large, large language, single base model, base model

备注: Presented as a Spotlight Paper at the Second Workshop on Social Simulation with LLMS, Third Conference on Language Modeling, 2026

点击查看摘要

Abstract:Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network powered by several different base models and show that the amount of engagement an agent receives depends more on its base model than on its assigned persona. The attraction or repulsion effects of a base model strengthen dramatically when more models are added in the mix, suggesting that networks dynamics may converge to base model effects at scale. To help explain this effect, we conduct a series of content-mediating analyses, showing the predictability of base models across contexts as well as the relationship between a model's lexical patterns and an engagement-maximizing style. In light of recent developments in mass multi-agent interaction, this work underscores the relevance of heterogeneous compositions in driving the outcomes of those networks

96. 【2610.07532】Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

链接:https://arxiv.org/abs/2610.07532

作者:Wonjun Lee,Kyungsik Yang,Gaeun Ji,Vaidehi Patil,Haon Park,Bumsub Ham,Mohit Bansal,Suhyun Kim

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:raising growing concerns, Latent Safety Signals, Latent Safety, advanced rapidly, raising growing

备注: project page: [this https URL](https://wonjuun.github.io/LADE/)

点击查看摘要

Abstract:LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We show that these signals consistently align across LLMs, forming a model-agnostic direction that emerges from safety alignment. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across diverse LLMs and benchmarks, LADE is robust against a wide range of jailbreak attacks and lowers attack success rates while maintaining a competitive safety-utility trade-off.

97. 【2610.07519】Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AI

链接:https://arxiv.org/abs/2610.07519

作者:Muhammad Rafiullah Memon,Viet Vo,Wanlun Ma,Yang Xiang

类目:Computation and Language (cs.CL); Cryptography and Security (cs.CR); Computers and Society (cs.CY)

关键词:Automatic sign language, American Sign Language, turning American Sign, sign language translation, sign language

备注: 5 pages, 1 figure. Accepted as a poster at the NeurIPS 2026 Workshop on Child Safety in AI (non-archival)

点击查看摘要

Abstract:Automatic sign language translation (SLT) has entered consumer products, turning American Sign Language into English text for dictation, messaging, and queries put to a conversational assistant. Child-facing AI and platform trust-and-safety tooling decide on text, using filters on minor accounts and grooming classifiers that score chat messages. A signing child who uses SLT therefore reaches these safeguards through a translation. We found no publicly documented system in which the two have been jointly evaluated, and the leading deployed SLT model was neither trained nor formally evaluated on signers under 18. Errors that alter negation, participant roles, secrecy, urgency or help-seeking could change a safety decision without disturbing fluency. This paper proposes a Deaf-informed pre-deployment audit of that boundary, with a failure taxonomy, a sanitised scenario schema, four comparison conditions, and four outcome measures. Auslan is the planned first case study.

98. 【2610.07509】On Open-Ended Information Seeking for Information Elicitation Agents

链接:https://arxiv.org/abs/2610.07509

作者:Victor De Lima,Grace Hui Yang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:potentially valuable directions, open-ended information-seeking problem, valuable directions, potentially valuable, continually determine

备注:

点击查看摘要

Abstract:Information elicitation is an open-ended information-seeking problem in which an interaction can unfold in many potentially valuable directions, requiring an elicitor to continually determine which information to pursue as new information emerges. In agentic elicitation, these decisions may be delegated to a foundation model, yet how model choice shapes the resulting information-seeking behavior remains understudied. We study how judgments about information value vary across LLMs and how these differences shape sequential information seeking. We first examine these judgments across 11 LLMs spanning multiple model families and parameter scales, using a shared set of information and elicitation objectives. We then develop a controlled elicitation simulation in which different models encounter the same information space and use the same selection rule, isolating these judgments from question generation and respondent behavior. Using this setting, we characterize the breadth-depth behavior that emerges from model-specific information-seeking preferences over the course of elicitation. We further examine how interaction history changes the evaluation and subsequent selection of prospective information. We test the robustness and boundaries of these findings through sensitivity analyses and ablations over the opportunities available to the elicitor, the response labels used to operationalize information-seeking preferences, the presence of interaction history, and whether redundancy is explicitly relevant to the assessment. The project code, data, and trajectory files are available at this https URL.

99. 【2610.07502】Closing Ambient Clinical Documentation Gaps with Automated Provider Queries

链接:https://arxiv.org/abs/2610.07502

作者:Joseph Paul Cohen,Raj Shah,Han-Chin Shing,Fang Wang,Susan Nguyen,Chaitanya Shivade,Jack Moriarty

类目:Computation and Language (cs.CL)

关键词:ensure accurate billing, Provider queries, accurate billing, clinical documentation specialists, queries are clarifying

备注:

点击查看摘要

Abstract:Provider queries are clarifying requests sent by clinical documentation specialists to physicians to close gaps in the clinical note and ensure accurate billing. Prior work automates note drafting, ICD-10 coding, and order extraction assuming a complete transcript, leaving these gaps unaddressed. We study whether an LLM can automate the query loop, termed DAU (Draft, Ask, Update), across those three tasks. An audit of 3,000 real visits identifies the sources of missing documentation, from which we build five transcript-degradation benchmarks on public data. Analyzing 21k clarification turns on real conversations, we find useful-question predictors are task-specific: oracle confidence dominates, but note completeness needs only simple recall questions while ICD-10 coding needs harder, multi-option ones. About 9% of turns hurt performance, driven by redundant questions and non-answers that still trigger a rewrite. Deployment depends on learning "when not" as much as "what to" ask.

100. 【2610.07480】In With the Old: Enhancing 'Classical' Document Automation with Generative AI

链接:https://arxiv.org/abs/2610.07480

作者:Marc Lauritsen,Hannes Westermann

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Software-based legal assistance, legal assistance systems, Software-based legal, representation and reasoning, legal assistance

备注: 10 pages. AI for Access to Justice Workshop at ICAIL 2025

点击查看摘要

Abstract:Software-based legal assistance systems have leveraged many different forms of knowledge representation and reasoning. This article explores how document automation services rooted in expert system style and other symbolic approaches can usefully enhance and be enhanced by current generative AI approaches. We discuss the possible benefits and challenges, and report on preliminary experiments in using large language models to identify and fix issues in texts written by laypeople.

101. 【2610.07459】Auditable Claims about AI Agents

链接:https://arxiv.org/abs/2610.07459

作者:Yue Zhao,Jiate Li,Li Li,Yi Nian,Jinbo Liu,Xiaolin Zhou,Xiyang Hu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Multiagent Systems (cs.MA)

关键词:Organizations make claims, Organizations make, external email, safe to deploy, person approves

备注:

点击查看摘要

Abstract:Organizations make claims about their AI agents: a person approves every external email, every action is logged, an evaluation shows the agent is safe to deploy. Article 12 of the EU AI Act requires high-risk systems to allow the automatic recording of events but does not say which records settle a given claim. The position is one sentence: to be checked, a claim about an agent must first name its policy, its scope, the records that would settle it, and who writes them. Adapting the preconditions of an assurance engagement, we call a claim auditable when these elements and a decision rule are fixed before any verdict and the records are obtainable. This extends the Policy Checkability dimension of our Auditable Agents framework from single actions to claims. Agents add three conditions: coverage by an independent record, authorization bound to each action's arguments, and completeness beyond integrity. Under an explicit model, we prove that support is impossible without each wherever its hypotheses hold. A claim-check table applies the method to six common claims, anchored in current NIST, IETF, and OWASP drafts. A worked case follows one claim through five evidence states. We close with a practice box and steps for operators, buyers, auditors, and standard setters.

102. 【2610.07457】AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

链接:https://arxiv.org/abs/2610.07457

作者:Hanzhi Zhang,Qiao Zhang,Qinglei Cao,Heng Fan,Yan Huang,Kewei Sha,Yunhe Feng

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Fine-grained mixed-precision quantization, promises efficient large, efficient large language, Fine-grained mixed-precision, mixed-precision quantization promises

备注:

点击查看摘要

Abstract:Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores prioritize higher precision for tiles important to either phase under a model-wide weight-storage budget. Each tile stores one selected representation, while phase-specialized kernels reuse the packed model and expand lower-bit weights for INT8 computation with 8-bit activations. Across four LLMs spanning 3B to 14B parameters, AlignQuant achieves up to $2.50\times$ generation speedup over BF16 while preserving model quality. Evaluations further cover three GPUs and contexts up to 64K tokens. These results show that local precision flexibility and regular GPU execution can coexist through a shared tile unit. The implementation is available at this https URL.

103. 【2610.07426】AccentCL: Robust Accent Classification with Incremental Expansion

链接:https://arxiv.org/abs/2610.07426

作者:Mu-Ruei Tseng,Waris Quamer,Ghady Nasrallah,Ricardo Gutierrez-Osuna

类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)

关键词:fixed label inventory, Accent, classifiers are typically, typically trained, Chinese-accented English

备注: Published in Proceedings of IEEE Spoken Language Technology Workshop (SLT) 2026

点击查看摘要

Abstract:Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.

104. 【2610.07365】Who Wrote It Is Not Enough: Detecting Who Contributed the Insight

链接:https://arxiv.org/abs/2610.07365

作者:Zhuoyang Zou,Abolfazl Ansari,Jiaxi Yang,Delvin Ce Zhang,Qian Chen,Dongwon Lee,Wenpeng Yin

类目:Computation and Language (cs.CL)

关键词:LLMs increasingly assist, increasingly assist scientific, assist scientific writing, detecting who wrote, longer sufficient

备注:

点击查看摘要

Abstract:As LLMs increasingly assist scientific writing and peer review, detecting who wrote the text is no longer sufficient: we need to determine who contributed the underlying insight. We introduce Insight Provenance, the task of identifying whether a review insight originates from a human, an LLM, or their hybrid contribution. We construct InsightProv-v0 from 4,057 scientific papers and 12,660 human reviews, simulating different levels of LLM involvement with GPT-4o, Gemini, and DeepSeek and annotating provenance at the sentence level. We show that strong performance on raw data can be misleading, as models exploit linguistic and textual-authorship shortcuts that degrade substantially under progressively debiased evaluation. We therefore propose a two-stage adversarial framework that suppresses shortcut signals while preserving provenance-relevant information. Beyond detection, extensive analyses reveal what makes intellectual authorship identifiable: paper grounding and neighboring review context provide complementary provenance signals, while human, hybrid, and AI insights systematically differ in their information sources and failure modes. Most strikingly, AI insights predominantly remain close to generic or paper-provided information, whereas human insights more often introduce external knowledge and independent judgment. These findings suggest that while wording can be rewritten by an LLM, the provenance of an idea leaves a deeper and more persistent signal.

105. 【2610.07355】racking Is Not Permanence: What Video World Models Keep of a Hidden Object

链接:https://arxiv.org/abs/2610.07355

作者:Peng Xie,Amr Alanwar

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:world models track, models track objects, Video world models, models track, world models

备注:

点击查看摘要

Abstract:Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.

106. 【2610.07348】Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

链接:https://arxiv.org/abs/2610.07348

作者:Arnav Kundu,Zhaoyang Xu,Bairu Hou,Chang Gao,Reed Li,Tao Lei

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Training large language, Training large, diverse deployment scenarios, constraints remains challenging, varying computational constraints

备注: Apple Foundation Models, 15 Pages, Edge LLMs

点击查看摘要

Abstract:Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5\% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.

107. 【2610.07339】A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care

链接:https://arxiv.org/abs/2610.07339

作者:Junseob Kim,Jade Chng,Ayman Ali,Victor Moas,Yichun Lee,Po-Chun Chin,Sunil Hwang,Rishikesan Kamaleswaran

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Tactical Combat Casualty, Combat Casualty Care, Tactical Combat, Casualty Care, Combat Casualty

备注: 20 pages, 6 figures, 5 tables. Dataset: [this https URL](https://doi.org/10.5281/zenodo.23170287;) code: [this https URL](https://github.com/Kamaleswaran-Lab/TC3-VQA)

点击查看摘要

Abstract:Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.

108. 【2610.07332】Structuring MoE Expert Selection for Agentic Reinforcement Learning

链接:https://arxiv.org/abs/2610.07332

作者:Bolian Li,Ting-Yao Hu,Cheng-Yu Hsieh,Sanjoy Chowdhury,Oncel Tuzel,Raviteja Vemulapalli

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Long-horizon LLM agents, Long-horizon LLM, structures remains underexplored, LLM agents, implemented using sparse

备注:

点击查看摘要

Abstract:Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.

109. 【2610.07327】SharedKV-BT: Node-Local Typed Decisions for Behavior-Tree Agents

链接:https://arxiv.org/abs/2610.07327

作者:Naoki Wake,Justin Wagle

类目:Robotics (cs.RO); Computation and Language (cs.CL)

关键词:Agent tasks require, tasks require sequences, Agent tasks, require sequences, sequences of interdependent

备注: 8 pages, 5 figures, 1 table. Last updated on October 5th, 2026

点击查看摘要

Abstract:Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix methods reduce this cost by reusing encoded context and scoring multiple decisions in parallel, but do not model decision dependencies or verify execution. We propose SharedKV-BT, where each active node of a behavior tree (BT) exposes stage-local fields and candidates, and Shared-KV scores the candidates in parallel and passes the selected decision to a separate execution system. We tested SharedKV-BT on robot manipulation, mobile navigation, and computer-use tasks. Across three tasks, SharedKV-BT made typed decisions 2.36-4.15 times faster than prompt-matched autoregressive decoding. On the manipulation task, node-local Shared-KV improved joint decision accuracy from 75% to 94% and closed-loop success from 0% to 60%. Fixed-score policy replay showed that stage gating prevented out-of-order actions and external postconditions prevented premature completion.

110. 【2610.07306】Kurate: Scalable Scientific Quality Analysis

链接:https://arxiv.org/abs/2610.07306

作者:Matthew J. Vowels,Jamie Cummins

类目:Computation and Language (cs.CL)

关键词:Scientific search systems, Scientific search, search systems, Scientific, Kurate

备注:

点击查看摘要

Abstract:Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study's trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper's overall grade breaks down into separate judgments, with each linked to specific evidence from the trial's registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.

111. 【2610.07258】Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

链接:https://arxiv.org/abs/2610.07258

作者:Venkata M Sangaraju,Sudhir Vissa

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:key performance indicator, same-named key performance, memory store face, legitimately computed results, unaddressed risks

备注:

点击查看摘要

Abstract:Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic. Existing agent-memory systems (e.g., MemGPT, Zep, A-MEM) gate retrieval by content, ownership, and role, not derivation, missing a cached insight that embeds a forbidden column. We introduce the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result, gated by a retrieval policy that serves a hit only when the requester is authorised for every column touched. Provided lineage recording is complete, we prove by construction that the policy blocks retrieval of results derived from a sensitive column outside the requester's permissions, at O(n) worst case -- a conditional design guarantee, not an empirical claim, that excludes derived features encoding sensitive information without naming their source. Eliminating measured leakage required 75-90% recorded lineage completeness, so we treat 90% as a conservative deployment target. Across six experiments, lineage-gated retrieval removes the 18.8-25.5% cross-department leakage naive content-gated memory suffers, keeping 81.5-82.6% of memory reuse at 13.8 microsecond worst-case overhead. A real-agent proof-of-concept with LLM-generated SQL is consistent with the guarantee: zero leaks over 9 round-trips, two conflicts caught automatically -- though a feasibility demonstration, not evidence of production viability. This offers a practical governance layer for shared agent memory, complementing source-layer access control and supporting EU AI Act compliance.

112. 【2610.07226】Minimal Witness Reinforcement Learning

链接:https://arxiv.org/abs/2610.07226

作者:T. Y. Tsui,Zihao Ye,Pengxiang Cai,Yanchao Li,Yuqiang Li,Zhehong Ai

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Logic in Computer Science (cs.LO)

关键词:produce an outcome, computation and science, irreducible conditions, common questions, questions that recur

备注:

点击查看摘要

Abstract:``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problems usually ask for multiple minimal witnesses, yet standard RL methods may reveal only one solution or redundant ones. We formalize this problem as minimal-witness identification and introduce Minimal-Witness Reinforcement Learning (MWRL). MWRL takes the union of the sets certified by successful proposals sampled from the policy and credits each proposal for the coverage the group union would lose without that proposal. This credit assignment, derived directly from the problem definition, unifies the demands for minimality and recovery of alternatives from a single black-box verifier bit. Under this principle, we derive a value iteration planner that recovers the entire family of witnesses and a policy gradient method that can scale to large language models. Across different experimental settings, MWRL recovers most minimal witnesses, while other methods return redundant supersets or a single witness. By making witness families learnable from verifier feedback, MWRL expands the scope of reinforcement learning beyond single-solution optimization. Our code is available at this https URL.

113. 【2610.07224】IDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes

链接:https://arxiv.org/abs/2610.07224

作者:Jose D. Posada,Somalee Datta,Priya Desai

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:protected health information, Clinical notes capture, health information, research until protected, protected health

备注:

点击查看摘要

Abstract:Clinical notes capture most of what is documented about a patient's care, but they cannot be used for research until protected health information (PHI) is removed. De-identification is often treated as a detection problem. Detection alone is not sufficient: redaction strips clinical content along with identifiers, date blanking destroys the temporal intervals needed for longitudinal analysis, and assigning a fresh random surrogate at each occurrence breaks links between a patient's notes. We present TIDE 2.0, an MIT-licensed engine with two separable stages: an interchangeable recognizer and a keyed anonymizer. Both run on hardware the institution owns. Surrogates are generated cryptographically with no stored linkage table. Dates shift by a per-patient, interval-preserving offset; each value receives the same surrogate across all occurrences under a given key; and a release produced under a new key cannot be linked to earlier releases. We also release TIDE2-Sentry, a recognizer distilled from a large language model. On two gold-annotated corpora from two institutions, the default configuration reached span-level recall of 0.88 in-domain and 0.77 on the second institution's corpus, at precision 0.88 and 0.87. We report recall and precision per category alongside these aggregates. The engine is open source, and the recognizer is available under a gated research-use agreement, so institutions can run, inspect and extend both within their own environments.

114. 【2610.07209】Forecasting the Growth of Social Media Information Cascades: Towards Human-in-the-Loop Misinformation Triage

链接:https://arxiv.org/abs/2610.07209

作者:Ansh Gupta,Abhiram Gorle,Aayush Rajesh,Tsachy Weissman

类目:ocial and Information Networks (cs.SI); Computation and Language (cs.CL)

关键词:Limited review teams, Limited review, review teams, teams must identify, identify which emerging

备注: Work done as part of the SHTEM Internship at Stanford University

点击查看摘要

Abstract:Limited review teams must identify which emerging claims are likely to keep growing before their eventual reach is known. We center early misinformation triage on this continuation-forecasting problem: predicting subsequent recorded propagation-tree growth from the first 30 minutes of activity. On FibVID, we compare early node count with structural depth entropy, temporal arrival entropy, and their pair while keeping all propagation trees from each original claim in one partition. Across 352 test trees from 59 claim groups separate from training, the combined model raises $R^2$ for log-transformed future growth from 0.307 to 0.323 and reduces log-MAE by 2.4% (95% claim-bootstrap CI, -0.4% to 5.2%). The gain is especially pronounced among 97 high-activity trees: $R^2$ rises from 0.248 to 0.395, Spearman's $\rho$ from 0.394 to 0.529, and log-MAE falls by 11.7% (95% CI, -1.9% to 23.9%). Complementing the 30-minute growth forecast, we analyze the first 15 replies in 563 PHEME threads. In this cohort, the 15th reply arrives after a median of 28.8 minutes; 52.0% reach the fixed reply prefix within 30 minutes and 71.6% within one hour. Even without the LLM-generated factual-accuracy dimension, the remaining stance, communicative, and affective state composition retains cross-event ranking signal (ROC-AUC 0.538); including that dimension increases ROC-AUC to 0.562. In a separate Check-COVID evaluation of 229 claims, reciprocal-rank fusion retrieves a gold evidence document within the top five for 74.2% of claims and within the top 20 for 94.3%; sentence reranking reaches Recall@20 of 58.1%. We propose an integrated human-review system that brings these early forecasts, response patterns, and retrieved evidence together for misinformation triage relying on the potential virality of claims.

115. 【2610.07205】Responsible Institutional Analytics: Interpreting Bias with AI Support

链接:https://arxiv.org/abs/2610.07205

作者:Francielle Marques,Ariel Ortiz-Beltrán,Ishari Amarasinghe,Davinia Hernández-Leo

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:dashboards inform decision-making, missing contextual information, dashboards inform, higher education, constraints in analytical

备注: Accepted for publication in the Journal of Universal Computer Science ( [this http URL](http://J.UCS) )

点击查看摘要

Abstract:Institutional Analytics (IA) dashboards inform decision-making in higher education, yet data limitations, constraints in analytical techniques, and missing contextual information often affect their interpretation. To support more responsible interpretation of IA, we introduce FACTRIA, a framework that organizes potential biasing factors across four areas: the analytics pipeline, institutional context, course-level characteristics, and demographics. We used the FACTRIA framework as input to a generative-AI chatbot designed to prompt users to reflect on these factors while analyzing IA. A qualitative study with stakeholders, drawing on four authentic IA cases, and a transition network analysis showed that the chatbot prompted participants to recognize how overlooked factors influenced their initial interpretation. Findings indicated that combining a structured framework with AI-based guidance can enhance context-aware, responsible interpretation of institutional data.

116. 【2610.07186】Identifying Introspection From the Inside

链接:https://arxiv.org/abs/2610.07186

作者:David I. Atkinson,Dillon Plunkett,David Bau

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, consequential and increasingly, increasingly difficult, difficult to verify

备注: Published at COLM 2026. Project page: [this https URL](https://iii.baulab.info)

点击查看摘要

Abstract:Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks -- a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.

117. 【2610.07184】Learning Scientific Exploration from Human Research Decision Trajectories

链接:https://arxiv.org/abs/2610.07184

作者:Xuchen Gong,Shane Gu,Haokun Liu,Dixi Yao,Chenhao Tan,Tian Li

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:investigating unknown phenomena, research, key challenge, challenge in building, investigating unknown

备注:

点击查看摘要

Abstract:A key challenge in building AI systems for scientific research is enabling $\textit{scientific exploration}$: the systematic process of investigating unknown phenomena or ideas to gain new knowledge through sequences of research decisions and actions. Yet this process is largely missing from existing scientific corpora; for example, research papers primarily record final outcomes rather than the trajectories that produced them. In this work, we introduce $\textbf{ResearchTrails}$, a dataset of $\textbf{human research trajectories constructed from Git repositories}$, where $\textbf{commit histories}$ serve as proxies for research exploration. We develop an automated and scalable pipeline that extracts structured research trajectories from repository commits, capturing successive changes to methods, experiments, and ablations. We characterize the resulting dataset and show that these trajectories contain meaningful signals about intermediate research decisions beyond what final papers reveal. We further demonstrate utilities of ResearchTrails in multiple use cases, including retrieving human research experience as external skills at test time and training models on research trajectories to improve generalization to new research decisions. Our results suggest a path toward AI systems that learn not only from the products of science, but from the evolving process of discovery itself.

118. 【2610.07177】CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks

链接:https://arxiv.org/abs/2610.07177

作者:Gowthamkumar Nandakishore

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:open contrastive decision, contrastive decision model, RM-Bench and JudgeBench, open contrastive, statistically indistinguishable

备注:

点击查看摘要

Abstract:An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model's own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.

119. 【2610.07168】A theory of platonic representations in language models

链接:https://arxiv.org/abs/2610.07168

作者:Darshil Doshi,Wenjie Zhou,Corinna Elena Wegner,Daniel J. Korchinski,Santiago Acevedo,Matthieu Wyart

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:platonic representation hypothesis, representation hypothesis, platonic representation, unexplained theoretically, observation connected

备注: 10+14 pages, 7+12 figures

点击查看摘要

Abstract:Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract levels are shared across languages while surface levels are modality- or language-specific. Concretely, we generate synthetic languages from probabilistic context-free grammars sharing upper-level but not lower-level production rules. In this setting the Bayes-optimal next-token predictor is belief propagation (BP); encoding its messages in successive layers yields analytical predictions that agree well with transformers trained on the same data. The framework explains why cross-lingual similarity peaks in middle layers, coexists with language-specific structure, and strengthens with language proximity, model quality and data exposure. It distinguishes similarity (shared neighborhood geometry) from alignment (shared coordinates), showing that the latter occurs when code-switched data, i.e. mixed-language sentences, are abundant enough. It further predicts that subtracting from each layer the component linearly predictable from the preceding one increases cross-lingual similarity, which we confirm in pretrained LLMs.

120. 【2610.07132】CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

链接:https://arxiv.org/abs/2610.07132

作者:Berke Arda,Ahmetcan Yavuz,Paul Gerry,Sebastian Lobentanzer,Nobin Sarwar,Joan Giner-Miguelez,Kongtao Chen,Luyao Zhang,Mrinmaya Sachan,Mubashara Akhtar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:accompanying dataset documentation, machine-readable dataset metadata, requires careful reading, dataset documentation, machine-readable dataset

备注: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: [this https URL](https://berkearda.github.io/croissantminer/)

点击查看摘要

Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

121. 【2610.07125】Jailbreaking Open-Weight LLMs via Random Embedding Perturbations

链接:https://arxiv.org/abs/2610.07125

作者:Abhinav Sudhakar Dubey(University of California Santa Cruz),Scott Sirri(University of California Santa Cruz),Vaggos Chatziafratis(University of California Santa Cruz),C. Seshadhri(University of California Santa Cruz)

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:enjoyed steady progress, multiple domains, enjoyed steady, steady progress, progress in capabilities

备注: 15 pages, 4 figures, Code: [this https URL](https://github.com/AbhinavDubey30/Perturbed-Embedding-Vectors)

点击查看摘要

Abstract:While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabilities across six common open-weight LLMs of various sizes that consistently lead to harmful or unsafe responses on the JailbreakBench benchmark dataset. Our proposed attack, Perturbed Embedding Vector (PEV), is a simple and fast "jailbreaking" technique that is cheaper than prior approaches, which typically require gradient computations, per-prompt optimizations, or altering internal weights of the models. PEV just adds independent Gaussian noise in the embedding vector representations of the prompt, with no need for further manipulations. To generate unsafe responses, we repeatedly sample additive noise from this distribution. In our experiments, we observe that the average compute cost to get the first successful attack is up to an order of magnitude less than previous attacks. The first successful jailbreak on a new prompt typically arrives within one minute on every tested model, and PEV generates unsafe responses across all models for all prompts in JailbreakBench. No other tested method achieves such results, despite them taking longer to run. More broadly, we believe that understanding the behavior of LLMs under perturbations in the embedding vectors is an important research direction: while perturbations constitute a major security risk, they can also serve as a valuable tool for exploring the dynamical behavior of such models.

122. 【2610.07109】JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

链接:https://arxiv.org/abs/2610.07109

作者:Yiqi Liu,Joseph James,Yang Wang,Kun Zhao,Chenghao Xiao,Chenghua Lin

类目:Computation and Language (cs.CL)

关键词:distribution retains uncertainty, LLM judge scores, score distribution retains, scalar compression, LLM judge

备注:

点击查看摘要

Abstract:When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank $p=0.0091$. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.

123. 【2610.07091】Smart Content Ingestion for Generative AI Workloads

链接:https://arxiv.org/abs/2610.07091

作者:Abbas Raza Ali,Muhammad Ajmal Siddiqui,Moona Zahid

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:conventional machine learning, machine learning, progressively changed, changed where intelligence, intelligence resides

备注:

点击查看摘要

Abstract:The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any single task: one foundation model serves open-ended downstream tasks, and the generality gained on the model side is matched by heterogeneity on the data side, because enterprise knowledge is authored in the formats people use (PDF, presentations, spreadsheets, scanned documents, forms, tables, diagrams and mixed-layout files) that carry textual, visual, geometric and structural information at once. A language model or retriever cannot reason reliably over information misrepresented at this interface, so content extraction becomes a lifecycle stage in its own right whose errors no downstream retriever or re-ranker can repair. This paper presents a production-ready content-extraction system that makes this stage explicit, configurable, and measurable. The system incorporates selective OCR routing, a scarcity-first curation engine with a reference-based extraction scorer that measures character, word, and table-structure accuracy, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator that generates grounded questions from every page and reports Hit@k, mean reciprocal rank, and latency. On a 180-document corpus the best extractor scores 97.4 of 100 (character error rate 0.13%, table similarity 0.995) and the chunker reaches hit@1 of 68.6%, hit@10 of 92.8% and MRR 0.77 over 25,050 generated questions. We distil three design principles (structure before semantics, never mutate what you measure, budget your labels) and position measured content extraction as the perception layer of enterprise agentic systems.

124. 【2610.07070】urnslide: Scalable Multi-Turn Data Synthesis by Walking a Finite-State Machine

链接:https://arxiv.org/abs/2610.07070

作者:Aaron Fainman,Gabriela Kadlecová,Maciej Gryka,Bartosz Kruszczyński,Usman Zafar,Cédric Archambeau,Aaron Klein,David Salinas,Selim Nowicki,Jacek Golebiowski

类目:Computation and Language (cs.CL)

关键词:Small language models, Small language, multi-turn tool calling, private infrastructure, rarely exists

备注: Accepted at the SLM-Agents Workshop, NeurIPS 2026 (non-archival)

点击查看摘要

Abstract:Small language models are inexpensive to serve and can run on private infrastructure, but base models are often not good enough at multi-turn tool calling, and fine-tuning them needs per-API data that rarely exists. Existing synthesis methods are too expensive for high-scale fine-tuning, as they often require mock operational environments for different domains and multiple LLM calls per generated conversation turn. We introduce a fully automated, lightweight synthesis framework that models each API as a finite-state machine, representing the system as abstract states that determine when each tool may be called, producing state-valid sequences of tools; sequences are translated into complete examples with a single LLM call. Rather than optimize diversity, we set a target distribution over the number of turns, the tool sequence and task complexity. We measure data quality by fine-tuning SLMs on generated trajectories, showing that our FSM-based generation significantly improves downstream accuracy over an unmutated baseline and, against existing works, reaches 70.7% full accuracy over 63.4% and 53.7% with 3.6-6.6$\times$ fewer tokens.

125. 【2610.07062】Learning to Simulate Individuals from Macro Social Signals

链接:https://arxiv.org/abs/2610.07062

作者:Yining Zhao,Bushi Liu,Haofei Yu,Zhengyang Qi,Shanyong Wang,Chuyue Li,Yuxiang Liu,Jiaxuan You

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:Large language models, offer limited behavioral, limited behavioral diversity, Large language, individual-level annotations

备注:

点击查看摘要

Abstract:Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these responses is either inherited from pretraining or learned from individual-level annotations, which offer limited behavioral diversity and little supervision of the reasoning itself. We propose to learn behavioral reasoning from prediction markets, whose price trajectories record how populations respond to real-world events at scale. We introduce macro2mind, which trains a language model with GRPO using market signals. A social behavioral decomposition makes behavioral reasoning an explicit step of forecasting: the model infers representative groups of market participants, predicts how each interprets the news and updates its beliefs, reasons about their interactions, and aggregates these responses into a price. A hindsight-regret curriculum with difficulty-aware sampling focuses training on transitions where hindsight-identified groups substantially improve the forecast while prioritizing examples that remain learnable for the current policy. The learned reasoning applies to user simulation without further training. On SWM-Bench, macro2mind achieves state-of-the-art directional accuracy and correlation on Polymarket. Trained on market data, it transfers zero-shot to four user-simulation benchmarks (Humanual, OvertonBench, PRISM, and CAD) and has competitive performance among zero-shot methods. Used as a data generator, macro2mind also raises a downstream simulator's accuracy on unseen users by 15.5 points, outperforming data generated by its backbone by 13.2 points.

126. 【2610.07032】Investigating Model Compression for Neural Machine Translation in the Biomedical Domain

链接:https://arxiv.org/abs/2610.07032

作者:Maria Zafar,Souhail Bakkali,Rejwanul Haque

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large-scale pretrained transformer, Large-scale pretrained, including multilingual settings, pretrained transformer models, machine translation tasks

备注: Accepted at AICS 2025

点击查看摘要

Abstract:Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit to 8-bit representations) is widely used to accelerate inference, enabling models to run several times faster during deployment. However, both techniques face limitations when applied to specialized domain data, particularly under low-resource conditions. In knowledge distillation, the effectiveness of transfer is often constrained by the scarcity of domain-specific parallel data, while quantization can lead to performance degradation as bit precision decreases. In this work, we investigate the combined application of knowledge distillation and quantization for French-to-English biomedical translation, a domain characterized by specialized terminology and limited parallel resources. We develop and compare multiple fine-tuning strategies to adapt compressed student models to this challenging setting. Our experiments demonstrate that a collaboratively distilled and quantized student model achieves a 69% reduction in size, a 98.21% increase in inference speed, and a 98.46% reduction in CO2 emissions compared to the original baseline all without sacrificing translation quality. These results indicate that jointly optimized compression techniques can yield efficient, high-performance models suitable for translation service providers operating under resource constraints.

127. 【2610.07023】Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

链接:https://arxiv.org/abs/2610.07023

作者:Jinghao Pang,Jitai Hao,Qiang Huang,Zhaochun Ren,Jun Yu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Large Language Models, achieved remarkable capabilities, Large Language, safety alignment, unsafe outputs

备注: 27 pages,7 figures, under review

点击查看摘要

Abstract:Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model's general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.

128. 【2610.07019】Calibrated Answers About Randomized Trials From a 4-Billion-Parameter Open Model: A Registered Test and a License-Clean Release

链接:https://arxiv.org/abs/2610.07019

作者:Johann Emmanuel Li

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:answers typed questions, typed questions, Open Science Framework, Evidence Inference, answers typed

备注: 20 pages, 2 figures, 13 tables. Code, results files and the paper's sources: [this https URL](https://github.com/johann-e-li/fiorillo) . Model: [this https URL](https://doi.org/10.57967/hf/10722) . Registration: [this https URL](https://osf.io/kaxmn/)

点击查看摘要

Abstract:Fiorillo v0.5 is an open model that answers typed questions with a probability for each answer. Its main specialist reads a randomized trial's article, cut to 6,144 tokens, and answers whether an intervention significantly increased, significantly decreased or did not significantly change an outcome against a comparator (Evidence Inference 2.0, EI). It is Qwen3-4B-Base with low-rank adapters and a decision head, fine-tuned for EI only on the 1,431 of 2,657 training articles whose own license allows reuse. Four criteria registered on the Open Science Framework before this version's test predictions decided its release, the second bar judged on EI's test split, whose labels are public. On that split (1,218 prompts in 333 articles), the expected calibration error was 0.0168 against a limit of 0.05; log loss was below the prior's by 0.8603 (95 percent interval 0.8104 to 0.9078) and below that of Gemma 4 31B-it, reading the same input, by 0.1829 (0.1164 to 0.2598); and macro-F1 was 0.9248 against 0.8668, so all four criteria passed. Training the same recipe on clean articles alone cost 0.0123 in accuracy (0.0034 to 0.0207; descriptive). With no article, macro-F1 fell to 0.4384; the title alone raised it by 0.0939 (0.0655 to 0.1234), which a title stating the result or recall of the trial could explain; exchanging intervention and comparator reversed 0.6652 of its direction answers. Run as released, the files matched the evaluated predictions within limits set in advance. The release is under the Apache License 2.0 (digital object identifier https://doi.org/10.57967/hf/10722).

129. 【2610.06996】Mask-Guided KV Cache Eviction in Block Diffusion Language Models

链接:https://arxiv.org/abs/2610.06996

作者:Gleb Molodtsov,Ekaterina Alimaskina,Evgeny Uskov,Artur Zagitov,Aleksandr Beznosikov

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:diffusion language models, Block diffusion language, generation speed, large key-value, diffusion language

备注:

点击查看摘要

Abstract:Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for denoising the current block (selection) and which to keep in memory for future blocks (eviction). We propose MaskAhead, a training-free method that solves both tasks with a single mask-query-based ranking mechanism. Current-block masks guide selection, while probes of upcoming masked blocks guide eviction. Both rank KV entries by their estimated contribution to the attention output. Our quantized variant, Q-MaskAhead, computes selection and attention directly from low-bit KV, largely preserving the selected entries. Experiments on Fast-dLLM-v2, DreamReasoner, and LLaDA2.0-mini cover long-generation reasoning, long-prompt question answering, and needle-in-a-haystack retrieval. On long-prompt QA, MaskAhead reduces KV memory by $9.5\times$ on average with a 1.2-point mean F1 loss relative to dense inference. Q-MaskAhead increases the reduction to $20.1\times$ with a 2.3-point mean F1 loss. In a batch-32 systems profile, MaskAhead achieves $1.23\times$ end-to-end and $1.68\times$ decode-stage speedups over dense inference.

130. 【2610.06971】AegisFlow: A Multi-Agent Agentic AI Framework for Autonomous Remediation and Self-Healing in Fragile Data Ecosystems

链接:https://arxiv.org/abs/2610.06971

作者:Muhammad Bilal Awan,Zubair Hussain,Abdul Shahid

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:website DOM modifications, API contract, Traditional data pipelines, notoriously brittle, Traditional data

备注: 43 pages

点击查看摘要

Abstract:Traditional data pipelines are notoriously brittle, often failing due to upstream schema drift, API contract changes, or website DOM modifications. Present observability tools only raise alerts but for human engineers, resulting in a high Mean Time to Repair (MTTR) and operational fatigue. In this paper we propose AegisFlow (Agentic Engine for Intelligent Self-healing and Graph-driven Operations for Workload remediation), a novel agentic framework that closes the loop between detection and resolution. AegisFlow uses a Watchdog agent to collect runtime telemetry and has a Repair agent to automatically create, test and deploy code patches based on Large Language Models (LLMs). The framework presents the non-intrusive execution model called Parallel Shadow Patching, a non-intrusive execution model based on the Monitor, Analyze, Plan, Execute, Knowledge (MAPE-K) loop to generate and verify patches in digital twin environments. Through experimental testing, we have evaluated AegisFlow across five common failure scenarios, and see 98.1 percent improvement in MTTR (from an average of 170 minutes per patch to 3.2 minutes) and a patch success rate of 92 percent . In particular, the system is successful in dealing with changes in the JSON schema (96 percent ) and punctuation drift (98 percent ), and is least successful in Shadow DOM cases (85 percent ). AegisFlow frees up about 98 percent of data engineering on-call time from firefighting and reallocates it towards innovation. The framework is deployment agnostic consisting of a system that can be deployed in a plugin fashion into an existing pipeline orchestration system with minimal uplift to the existing system.

131. 【2610.06963】WavePrune: One period is often enough for RoPE

链接:https://arxiv.org/abs/2610.06963

作者:Guancheng Du,Luotian Huang,Shaowen Wang,Si Li,Kaifeng Lyu

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:Rotary Position Embedding, encodes token positions, Rotary Position, Position Embedding, attention logits invariant

备注:

点击查看摘要

Abstract:Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative positions separated by a full rotation period become hard to tell apart. To address this, we propose WavePrune, which restricts each channel to its first rotation period. We show that it removes the distractions in attention maps created by position aliasing and improves overall long-context performance. Specifically, WavePrune raises the HELMET score on four of five models we test without any extra tuning (e.g., 35.7 - 40.0 on Qwen3-8B). When pretraining models from scratch, WavePrune also achieves lower validation loss at extrapolated lengths than pretraining without it. Because WavePrune restricts each channel to a sliding window, it induces a fine-grained sparsity that our hardware-aligned CUDA kernels exploit for 1.15x prefill and 1.24x decoding speedups over FlashAttention-2 at 32K context. Together, these results show that RoPE's periodic structure, widely regarded as essential, is largely redundant beyond the first rotation period.

132. 【2610.06962】Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?

链接:https://arxiv.org/abs/2610.06962

作者:Nishanth Nayakanti,Prasang Gupta,Ashutosh Bilthare,Kevin Paul

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:review workflows, thing retained, human evidence spans, human evidence, evidence

备注: 16 pages, 5 figures

点击查看摘要

Abstract:In many review workflows the verdict is the only thing retained. The passages behind it are not marked, because that annotation costs far more than recording the decision. We measure how much of that evidence a small language model can recover when it is post-trained on the verdicts alone, with no human evidence labels at any stage. On ContractNLI the human evidence spans are held out until evaluation. Matching the recorded verdict and agreeing with those spans are not the same thing: across six systems the two scores are only weakly related and rank the systems differently, so accuracy is a poor guide when the citations have to be reviewable. Label-only training on the bare verdict reaches accuracy 0.896 and span F1 0.564. Rejection sampling, which keeps a generated trace only when its verdict matches the record and then picks one by an automatic source-grounding score, reaches 0.797 and 0.556, against 0.747 and 0.493 before training. Verbatim citation rises from 0.597 to 0.729 under label-only training and to 0.701 under rejection sampling. One seed on one corpus cannot say which method is better, but both improve the evidence without anyone annotating it.

133. 【2610.06956】EMODE: Dynamic Para-Semantic Experts for Emotion-Aware Speech Language Modeling

链接:https://arxiv.org/abs/2610.06956

作者:Jianan Pan,Yiwen Gu,Xinze Li,Rui Wang,Kejie Huang

类目:Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)

关键词:demonstrated strong capabilities, Large speech language, unified cross-modal understanding, Large speech, language model

备注:

点击查看摘要

Abstract:Large speech language models have demonstrated strong capabilities in unified cross-modal understanding and generation, yet paralinguistic cues, especially emotion, remain difficult to preserve. Existing systems typically rely on entangled acoustic representations, which allow the underlying language model to depend excessively on recovered lexical content instead of grounding its behavior in acoustic-prosodic evidence. We address this limitation with EMODE, an emotion-aware speech language model built around \textbf{Dynamic Para-Semantic Experts (DPSE)}. DPSE decomposes continuous speech features into semantic and paralinguistic pathways, routes them dynamically, and fuses them before integration into the language model. To turn this structural decomposition into functional specialization, EMODE is trained with a three-stage curriculum consisting of semantic warm-up, paralinguistic activation, and joint refinement, guided by Orthogonal Expert Guidance (OEG), Semantic-to-Acoustic Alignment (SAA), and Gating Diversity Regularization (GDR). Experiments on SER test, empathetic response evaluation, and the newly constructed bilingual MEPA benchmark show that EMODE improves the balance between lexical fidelity and emotional sensitivity, strengthens affect-grounded response generation, and exposes the value of explicit para-semantic factorization for robust cross-corpus emotion understanding.

134. 【2610.06940】Stabilizing language models under continual learning via condition-anchored distillation

链接:https://arxiv.org/abs/2610.06940

作者:Huan Li,Zhe Cao,Qinlei Xie,Fushun Cui,Xuechen Liang

类目:Computation and Language (cs.CL)

关键词:prompts learned earlier, learned earlier, undesirable or impossible, prompt-answer pair, output distribution

备注:

点击查看摘要

Abstract:Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillation (CAGD): retain a small set of old prompts, use a frozen previous model to reconstruct completions and generation states, and match its predictive distributions while learning the next task. The formulation separates three roles that ordinary replay conflates: conditions select the behavior to protect, teacher generations locate relevant states, and soft targets specify how predictions may change. For autoregressive language generation, teacher-rollout distillation admits an exact chain-rule decomposition of sequence divergence. For masked-diffusion language modeling, our implementation directly controls local denoising drift on teacher-generated completions. In continual adaptation of a 219M masked diffusion language model, CAGD reduces four-task final held-out loss from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse. The same soft targets lower final average loss by 0.055 over hard replay when teacher-generated support is held identical. The direction persists on fresh facts and natural instructions across SMDM and Qwen3. On GSM8K, Qwen adaptation preserves answer-format compliance, but exact-match retention is seed-mixed at 0.6B and worsens at 1.7B. These results support condition-anchored functional preservation as a common design principle across the tested language-generation objectives.

135. 【2610.06903】Component and Dimension Sparsity in Transformer Refusal Mechanisms

链接:https://arxiv.org/abs/2610.06903

作者:Vincent Siu,Glenn Grant-Richards,Vlad Pavlovich,Yizhou Sun,Dawn Song,Chenguang Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:remains poorly understood, manipulates large language, Activation steering manipulates, large language model, interventions remains poorly

备注: Accepted to COLM 2026

点击查看摘要

Abstract:Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48\% of upstream components, retaining 88--101\% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50\% of residual stream dimensions, retaining 85--98\% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in this https URL.

136. 【2610.06902】ree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA

链接:https://arxiv.org/abs/2610.06902

作者:Priyank Jayraj,Poonam Goyal,Navneet Goyal

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Retrieval-augmented generation grounds, miss complementary evidence, generation grounds language, grounds language models, Retrieval-augmented generation

备注:

点击查看摘要

Abstract:Retrieval-augmented generation grounds language models in external context, but for long documents flat top-$k$ retrieval can cluster on a single region and miss complementary evidence. RAPTOR-style summary trees address this by recursively clustering chunks and using a language model to summarize each cluster at indexing time, then ranking summary nodes alongside raw chunks at query time. We show the main benefit of summary trees in long-document QA can come from navigation rather than the generated summary content. We introduce NavTree, a leaves-only retriever that builds a deterministic balanced segment tree over chunks (zero language-model calls at indexing) and uses the tree purely as a navigation scaffold: a hybrid lexical-and-dense frontier walk, anchored on top retrieved leaves, descends from the root and emits only leaf chunks to the reader. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree is the strongest matched-cost hierarchical retriever in our evaluated grid and ties the strongest flat baseline. On long-document multi-hop QA, it is the only hierarchical method that significantly beats BM25 on a class-vs-class basis, corroborated by a reader-free retrieval-recall check. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still loses to NavTree at every multi-chunk budget, at zero indexing cost. The ranking carries across stronger and open-weight readers, a stronger encoder, and a full factorial that isolates leaves-only emission as the structural lever.

137. 【2610.06897】Capacity, Responsiveness and Alignment: What Makes a Latent Structure Actionable

链接:https://arxiv.org/abs/2610.06897

作者:Or Shafran,Mor Geva

类目:Computation and Language (cs.CL)

关键词:Localizing latent structures, Localizing latent, controlling their behavior, activation space, space of language

备注:

点击查看摘要

Abstract:Localizing latent structures in the activation space of language models (LMs) is central to understanding and controlling their behavior. Yet, localized structures can differ substantially in their causal influence, raising the question of what makes a structure actionable. We tackle this question by casting causal influence as a product of three factors and showing empirically that they act as interpretable, distinct constraints: capacity, measuring the sensitivity of the model's output to movement along the structure, responsiveness, capturing how promotable the concept is given the current context, and alignment, reflecting how well the structure aligns with the context-specific representation of the concept. Across 4 LM families and 50 concepts, we observe that causal effectiveness requires all factors to be high; low capacity and responsiveness reduce it by 84% and 95%, respectively, while low alignment can reverse it, suppressing concept expression. Moreover, we find that causality is context-dependent rather than an intrinsic property of the structure, with causally effective directions forming a low-dimensional subspace that varies across contexts. By restricting the training of linear probes to this subspace, we introduce causal probes that achieve 17%-118% improvement in steering across models, with only 3% reduction in concept detection.

138. 【2610.06889】Zero-Shot Visualization: Exploring Text Corpora with User-Prompted Axes

链接:https://arxiv.org/abs/2610.06889

作者:Arnau Bueno Tricas,Jose A. Rodríguez-Serrano

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:large language models, study the application, application of large, visual exploration, exploration of textual

备注:

点击查看摘要

Abstract:We study the application of large language models (LLMs) to the visual exploration of textual corpora. We introduce zero-shot visualization (ZSV), a task in which users specify concepts in natural language and documents are mapped onto the corresponding concept axes for visualization. Building a ZSV system of practical value is non-trivial, as it requires choices at the intersection of feature functions, efficient implementation tradeoffs, and pre/post-processing decisions affecting visualization quality. To that end, we establish a benchmark that compares methods spanning embedding similarity, direct semantic judgments, and conditional likelihood estimation in this setting. Across multiple datasets and use cases we evaluate the properties of different scoring methods and design choices in terms of semantic faithfulness, score fidelity, and computational cost. Our results identify that scoring based on next-token probabilities offers the strongest practical trade-off among the evaluated methods. We further apply this approach to unlabeled corpora to examine its behavior in realistic exploratory settings. These experiments highlight additional design considerations, including the use of graded axes together with binary relevance filtering, and reveal a compositional sentiment bias in off-topic documents. Based on these findings, we provide practical guidelines for constructing end-to-end ZSV baselines.

139. 【2610.06861】When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO

链接:https://arxiv.org/abs/2610.06861

作者:Sofia Torres,Gabriel Almeida,Carter Adams,Camila Rocha

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:large language models, retrieved thought patterns, eliciting multi-step reasoning, Reinforcement learning, expert traces

备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emph{external guidance} - expert traces, self-explanations, or retrieved thought patterns. Although each method reports empirical gains, none provides convergence rates, bias bounds, or an optimal weighting rule for the guidance signal. We close this gap with \emph{Guidance-Augmented GRPO} (GA-GRPO), a unified theoretical framework that casts external guidance as a stochastic guidance operator G re-writing the question distribution, and analyses the resulting policy-gradient estimator as a biased on-policy estimator whose bias is bounded by the total-variation guidance divergence delta\_G between the guidance-augmented sampling distribution and the policy's own distribution. The framework subsumes vanilla GRPO, LUFFY, ExPO, PAPO, and TAPO as special cases obtained by particular choices of G. Under smoothness and bounded-divergence assumptions we prove that GA-GRPO converges at rate O(1/sqrt(T)) to an O(delta sqrt(T))-neighbourhood of the GRPO stationary point, derive the closed-form MSE-optimal guidance weight lambda-star(T, delta, sigma\_0 squared) = sigma\_0 squared / (sigma\_0 squared + R\_max squared delta squared T), and prove a matching minimax lower bound showing the Omega(delta squared T) bias term is unavoidable. Experiments on Qwen2.5-Math-7B-Base across nine math and OOD benchmarks confirm that optimal-weight GA-GRPO matches or surpasses TAPO, LUFFY, ExPO, and vanilla GRPO while requiring 31\% fewer GPU-hours, and eight analysis experiments validate each theoretical prediction.

140. 【2610.05094】How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

链接:https://arxiv.org/abs/2610.05094

作者:Laurène Vaugrante,Thilo Hagendorff

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Large Language, increasingly use Large, Language Models, Researchers increasingly

备注:

点击查看摘要

Abstract:Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.

141. 【2610.03829】OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes

链接:https://arxiv.org/abs/2610.03829

作者:Wuraola Oyewusi,Eliana Vasquez Osorio,Goran Nenadic,Gareth Price

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)

关键词:tumour staging expressions, Real-world outpatient oncology, institution-specific de-identification markers, Real-world outpatient, outpatient oncology notes

备注: Accepted at the AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models Workshop at NeurIPS 2026

点击查看摘要

Abstract:Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.

142. 【2610.08125】Conversation Is a Two-Body Problem: Dyadic Evaluation of Full-Duplex Dialogue Models

链接:https://arxiv.org/abs/2610.08125

作者:Sungnyun Kim,Sungwoo Cho,Jihwan Oh,Se-Young Yun

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:enabling voice agents, Full-duplex spoken dialogue, spoken dialogue models, dialogue models listen, enabling voice

备注: Project page: [this https URL](https://dyafdb.github.io/)

点击查看摘要

Abstract:Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot react, or an automated examiner that reacts in real time but only administers a fixed sequence of tests and is never graded. These single-sided frameworks evaluate only half of a two-body problem, where turn-taking, overlap, and interruption are joint products of two coupled speakers. We propose DyaFDB, a framework that evaluates full-duplex models in a dyadic setup: two models converse directly under assigned roles with cooperative or conflicting goals, and both sides are scored offline with an external judge. DyaFDB probes how the two models behave toward each other, such as how they take turns or carry an assigned role under different interests. We instantiate four tasks as 140 scenarios and record 7,560 conversations, covering six self- and cross-play pairings. Throughout the experiments, we observe that how a model behaves continually reshapes its partner. We thus demonstrate that each model must be both the examiner and examinee of the other, and no single fixed interlocutor can play both parts. We will release the scenarios, role prompts, and recording protocols between two full-duplex models, without any pre-recorded audio.

143. 【2610.08063】HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR

链接:https://arxiv.org/abs/2610.08063

作者:Takanori Ashihara,Kohei Matsuura,Masato Mimura

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:Conversational Speech Language, Challenge and Workshop, Multilingual Conversational Speech, HINTT system submitted, Speech Language Model

备注:

点击查看摘要

Abstract:This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.

144. 【2610.07626】A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization

链接:https://arxiv.org/abs/2610.07626

作者:Tien-Hong Lo,Fong-Chun Tsai,Ting-An Hung,Yu-Hsuan Hsieh,Yao-Ting Sung,Berlin Chen

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:automatic pronunciation assessment, sentence stress detection, word stress detection, stress detection, pronunciation assessment

备注: Interspeech 2026

点击查看摘要

Abstract:Prosodic stress is a crucial aspect of automatic pronunciation assessment (APA), encompassing both sentence stress detection (SSD) and word stress detection (WSD). SSD highlights semantically salient words that shape discourse meaning, while WSD identifies the primary stressed syllable within each word to ensure lexical clarity. However, most prior work treats SSD and WSD as independent tasks, overlooking their shared reliance on prosodic cues such as pitch, duration, and intensity. To address this gap, we propose an effective SSD approach combining SSD with auxiliary WSD via a novel modeling paradigm. In addition, we introduce a word-span stress regularizer (WSR) that concentrates token-level SSD probabilities within each stressed word span. Experiments on the TinyStress-15K benchmark show that the proposed method outperforms strong baselines, with the complete configuration achieving the best SSD result.

145. 【2610.07338】Logbook: Extremely Long-form Audio Event Understanding

链接:https://arxiv.org/abs/2610.07338

作者:Kwanghee Choi,Suwon Shon,Dmitriy Serdyuk,Guitang Lan,Chao-Wei Huang,Mohammad Sadegh Rasooli,Sangeeta Srivastava,Zhaojiang Lin,Saurabh Adya,Ming Sun

类目:Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)

关键词:limiting model design, pre-segmented clips, built around short, limiting model, fixed vocabularies

备注: Submitted to ICASSP 2027. Source code available at [this https URL](https://github.com/facebookresearch/logbook)

点击查看摘要

Abstract:Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.

146. 【2610.07047】SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation

链接:https://arxiv.org/abs/2610.07047

作者:Shao-Chun Hu,Zi-Xiang Lin,Jeih-Weih Hung,Hung-Shin Lee

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

关键词:Compact time-frequency separators, Compact time-frequency, shared cell face, face two limits, shared cell

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Compact time-frequency separators that mask the mixture and refine through a shared cell face two limits. First, a bounded multiplicative mask only scales a mixture bin, so where overlapping components cancel, the estimate stays small. Second, a shared cell applies the same weights to every time-frequency token at every step, so enlarging it adds compute everywhere. We present SEAL (Sparse Expert routing with Additive Latent reconstruction) to address both. For reconstruction, a zero-sum additive residual bounded by the local mixture amplitude lets estimates be nonzero where components cancel yet still sum to the mixture. For routing, a query built from acoustic and inter-step evidence sends each token to one of six residual experts, and a norm cap keeps the step cue from overriding clear acoustic evidence. On EchoSet, SEAL (small) surpasses TIGER (small) by 0.31 dB SI-SDRi with 28% fewer parameters and 2.9 times fewer MACs, and SEAL (large) is within 0.07 dB SI-SDRi of TIGER (large) at 3.1 times fewer MACs.

147. 【2610.07046】GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting

链接:https://arxiv.org/abs/2610.07046

作者:Ming-Hsiang Hu,Kuan-Tang Huang,Hung-Shin Lee,Berlin Chen

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)

关键词:speech promises noise-robust, noise-robust keyword spotting, promises noise-robust keyword, Visual speech promises, keyword spotting

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder's lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bearing encoder, injection yields an effective SNR gain of 4.0-9.3 dB over masking at -10 dB, whereas under a phoneme-poor one it nearly vanishes. Relative to the benchmark system, GIVE-KWS reduces unseen-keyword EER by 72.9% at -10 dB and 62.8% on average.

148. 【2509.17988】Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech

链接:https://arxiv.org/abs/2509.17988

作者:Zirui Li,Jens Edlund,Yicheng Gu,Nhan Phan,Lauri Juvela,Mikko Kurimo

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:scarcity of high-quality, limited by scarcity, Finnish and Swedish, TTS, Finnish

备注: Accepted by ICASSP 2026. 5 pages, 2 figures

点击查看摘要

Abstract:Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Parl-TTS, an open TTS dataset for Finnish and Swedish based on speech found in the wild. Using recordings of Nordic parliamentary proceedings, we extract 900 hours of Finnish and 5090 hours of Swedish speech suitable for TTS training. The dataset is built using an adapted version of the Emilia data processing pipeline and includes unified evaluation sets to support model development and benchmarking. By offering open, large-scale data for Finnish and Swedish, Nord-Parl-TTS narrows the resource gap in TTS between high- and lower-resourced languages.

149. 【2507.02115】Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams

链接:https://arxiv.org/abs/2507.02115

作者:Zirui Li,Lauri Juvela,Mikko Kurimo

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:potentially highly valued, Synthesizing second-language, language learning experience, experience and feedback, potentially highly

备注: Accepted by Proceeding of 13th edition of the Speech Synthesis Workshop; 5 pages, 1 figure

点击查看摘要

Abstract:Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback. However, due to the lack of L2 speech synthesis datasets, it is difficult to synthesize L2 speech for low-resourced languages. In this paper, we provide a practical solution for editing native speech to approximate L2 speech and present PPG2Speech, a diffusion-based multispeaker Phonetic-Posteriorgrams-to-Speech model that is capable of editing a single phoneme without text alignment. We use Matcha-TTS's flow-matching decoder as the backbone, transforming Phonetic Posteriorgrams (PPGs) to mel-spectrograms conditioned on external speaker embeddings and pitch. PPG2Speech strengthens the Matcha-TTS's flow-matching decoder with Classifier-free Guidance (CFG) and Sway Sampling. We also propose a new task-specific objective evaluation metric, the Phonetic Aligned Consistency (PAC), between the edited PPGs and the PPGs extracted from the synthetic speech for editing effects. We validate the effectiveness of our method on Finnish, a low-resourced, nearly phonetic language, using approximately 60 hours of data. We conduct objective and subjective evaluations of our approach to compare its naturalness, speaker similarity, and editing effectiveness with TTS-based editing. Our source code is published at this https URL.

信息检索

1. 【2610.08732】A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

链接:https://arxiv.org/abs/2610.08732

作者:Alexia Allal,Hicham Randrianarivo,Sylvain Lamprier

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Generative Information Retrieval, predicts document identifiers, Generative Information, directly predicts document, shifting document retrieval

备注: 8 pages, 3 figures, 1 table

点击查看摘要

Abstract:Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs). While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID? Current approaches rely heavily on computationally expensive downstream evaluations, hindering systematic analysis and rapid iteration. In this work, we address this challenge by presenting a comprehensive study on the properties, metrics, and trade-offs that define effective numerical DocIDs. Specifically, our contributions are threefold: First, we propose a unified framework that unifies Product Quantization (PQ) and Residual Quantization (RQ), and their hybrid variants within a single design space. This enables us to systematically study key DocID properties, such as hierarchy versus parallelism, as well as the impact of hyperparameters like DocID length and codebook size. Second, we define a suite of training-free, intrinsic metrics, to quantify DocID quality and evaluate structural fidelity without the overhead of full model training. Through extensive experiments on MS MARCO 300K and NQ320K, we analyze how these structural properties influence retrieval effectiveness.

2. 【2610.08716】Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval

链接:https://arxiv.org/abs/2610.08716

作者:Hicham Randrianarivo,Logan Renaud,Alexia Allal

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Generative retrieval trains, Generative retrieval, Generative, diffusion, identifier

备注: 13 pages, 7 figures, 11 tables

点击查看摘要

Abstract:Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers. With identifier length and training budget fixed, we decode each model in several ways. Decoding alone moves a diffusion model's Hit@1 by 6.6 to 13.7 points. Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers. The generated identifier is right for 14-21% of NQ320K queries. We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes' probabilities. It matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead in Hit@1; on NQ320K, the lead comes from the model, not beam search. Starting from one sampled identifier, one-pass scoring removes 46-83% of masked diffusion's deficit to beam search; from generate-and-match, at most a quarter. On NQ320K, every paradigm largely memorises which identifier answers which query: random identifiers keep 83-90% of the Hit@1 of residual-quantised ones. There, product-quantised identifiers lead residual-quantised ones by 3.4 points in the autoregressive model and by -0.7 to +3.6 in diffusion models; across decodings, AR's gap exceeds diffusion's by 1.5-2.3 points, around our 2-point threshold. Paradigm comparisons must report each paradigm at its own recipe and best decoding.

3. 【2610.08463】UNREAL: Unifying Retrieval and Long-Context with a Single Model

链接:https://arxiv.org/abs/2610.08463

作者:Edan Kinderman,Elad Hoffer,Yochai Blau,Brian Chmiel,Ron Banner,Daniel Soudry,Boris Ginsburg

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:single long prompt, handle evidence selection, vastly different scales, long prompt, single long

备注:

点击查看摘要

Abstract:Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.

4. 【2610.08452】Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

链接:https://arxiv.org/abs/2610.08452

作者:Lasse B. Strand,Robert Jakob,Kevin O'Sullivan,Markus Kreft

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:grounding large language, large language models, Retrieval-augmented generation, widely used approach, approach for grounding

备注: Accepted at the Second Workshop for REsearch on Agent Language Models (REALM) at EMNLP 2026 and at the Machine Learning for Systems Workshop at NeurIPS 2026. 9 pages plus references and appendix (16 pages total), 4 figures, 6 tables. Code: [this https URL](https://github.com/Agentic-Systems-Lab/Agentic-AutoRAG)

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.

5. 【2610.08407】Seeing the Context: Enhancing Recommender Systems with Image-Derived Contextual Signals

链接:https://arxiv.org/abs/2610.08407

作者:Tal Cordova,Tomer Geva,Moshe Unger

类目:Information Retrieval (cs.IR)

关键词:Contextual information, capturing the circumstances, Review-aware Graph Contrastive, Graph Contrastive Learning, recommender systems

备注: Accepted at the CARS workshop, RecSys 2026. 8 pages, 2 figures

点击查看摘要

Abstract:Contextual information, capturing the circumstances of a user-item interaction, is central to recommender systems. Prior work draws context from location, time, or reviews, but not images; multimodal recommender systems mainly use images to enrich item or user representations, not identify situational context. We propose a new representation of context derived from images, spanning physical, social, and modal categories learned via a vision-language model. We introduce ICE-Fuse, a pipeline for evaluating this representation that fuses these categories and integrates them into a context-aware recommender system, using TripAdvisor data and Review-aware Graph Contrastive Learning as the recommendation algorithm. Image context does not outperform established signals standalone, but improves them combined, indicating complementary information. Semantic analysis shows image- and review-derived context capture distinct aspects of the interaction, positioning images as complementary context.

6. 【2610.08245】Aligning Performance with Contribution: Towards Contribution-Aware Fair Recommendation

链接:https://arxiv.org/abs/2610.08245

作者:Shuai Zhang,Hui Fang,Zun Sun

类目:Information Retrieval (cs.IR)

关键词:developed diverse objectives, diverse objectives, developed diverse, Existing research, recommendation

备注:

点击查看摘要

Abstract:Existing research on user fairness in recommender systems has developed diverse objectives. However, it has paid limited attention to a distinct distributive perspective: whether users' contributions to model learning should be reflected in the recommendation benefits they receive. We argue that, in addition to existing fairness protections, a fair system may account for the alignment between users' estimated contributions and the recommendation performance they receive. Such alignment can incentivize sustained and informative engagement, thereby supporting a sustainable recommendation ecosystem. To this end, we propose Contribution-Performance Fairness, a novel fairness perspective which requires recommendation performance to be aligned with estimated contribution across user groups and to remain equitable among users with comparable contributions within a same group. To instantiate this perspective, we introduce the Contribution-Performance Fair Recommender (CPFR), a framework applicable to different backbone recommenders. CPFR constructs ordered user groups from a training-dependent contribution considering interaction volume, loss alignment, and optimization intensity, and jointly optimizes recommendation accuracy with the two fairness requirements. A game-theoretic analysis shows that such alignment can strengthen contribution incentives and improve system-level recommendation accuracy under voluntary contribution. Experiments on three datasets and three backbone models demonstrate that CPFR achieves a strong accuracy--fairness trade-off under the proposed operational metric.

7. 【2610.08232】Behavior-Mining, Generative Conversations, and Collaborative Advisory: the Future of Travel and Tourism Recommender Systems

链接:https://arxiv.org/abs/2610.08232

作者:Alejandro Bellogín,Linus W. Dietz,Francesco Ricci,Pablo Sánchez

类目:Information Retrieval (cs.IR)

关键词:travelers choose destinations, adoption of e-commerce, choose destinations, tourism recommender systems, early adoption

备注:

点击查看摘要

Abstract:Since the early adoption of e-commerce, travel and tourism has been a lab for the design of recommender systems: tools that help travelers choose destinations, flights, accommodations, and combine them into itineraries. Data-driven recommendation techniques, ranging from case-based reasoning to reinforcement learning, have been adapted to travelers' needs. The research community has produced multifaceted prototypes of travel and tourism recommender systems (TTRSs), which are context-dependent, multistakeholder-oriented, and more recently, addressing sustainability issues, such as overtourism. Despite this enduring work, TTRSs are not widespread yet. We argue that three limitations can explain this: outdated and sparse data sets used to train and validate TTRSs, algorithms that prioritize prediction accuracy over domain-specific dimensions such as novelty and contextual relevance, and a failure to address the specific needs of travelers. Targeted incremental research could address these limitations, but a disruptive factor has meanwhile entered the ecosystem of tourism information and commercialization platforms: generative artificial intelligence. According to market research, GenAI applications are becoming the primary entry point for travelers planning their trips. This forces research to rethink how TTRSs should be designed and which core techniques should be integrated. We claim that future TTRSs, in addition to offering personalized information filtering, should become more flexible advisors that support decision making, integrating multiple data types and AI techniques, from data mining to natural language processing. Moreover, they must transparently balance the conflicting goals of travelers, service suppliers, platform owners, and local communities. We then outline research targets for building more effective TTRSs, fruitfully combining old and new recommendation techniques.

8. 【2610.08229】Confidence-Ordering Reversal under Contextual Priors in Neural Decoding

链接:https://arxiv.org/abs/2610.08229

作者:Xinyu Zhang,Sichao Liu

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Neurons and Cognition (q-bio.NC)

关键词:reshaping candidate scores, prior, scores, correct candidate, Contextual

备注: 28 pages, 4 figures, 18 tables

点击查看摘要

Abstract:Contextual priors improve neural-to-language decoding by reshaping candidate scores. However, confidence is read from the same reshaped scores, so the errors a prior leaves behind can become more confident with no change in accuracy to reveal it. We study how a prior shapes confidence in speech retrieval on MEG-MASC and MOUS using local decoding scores, a contextual prior combined by additive shallow fusion, and the fused top-two margin as confidence. Among initially incorrect predictions, we find a confidence-ordering reversal: a larger margin makes a repair more likely when the correct candidate starts near the top of the local ranking, but less likely when it starts lower. On MEG-MASC, pooled correctness AUROC is 0.87, yet AUROC separating repairs from residual errors falls from 0.70 at initial ranks 2-3 to 0.39 at ranks 21-50. Errors starting beyond rank 20, inside the reversed region, make up 46.6% of all post-fusion errors. We propose a score-level account: a repair must first close the correct candidate's initial deficit, limiting its final margin, whereas a residual error can build a large margin between two incorrect candidates. A causal intervention that changes only the fusion weight moves the reversal to deeper ranks as predicted. Under a word-level LM prior, it keeps moving after accuracy gain peaks, so a weight chosen for accuracy does not settle confidence. Reading local and prior scores separately improves selective decoding: the decoder answers on 74.5% of windows instead of 56.7%, while 92% of output sets still contain the correct candidate. Confidence after contextual fusion should retain the local and contextual evidence behind each prediction, not just the fused scores. Project website: this https URL Code: this https URL

9. 【2610.08136】Adapting Generative Recommenders for Multi-Turn Interaction

链接:https://arxiv.org/abs/2610.08136

作者:Yu-Chen Den,Zhi Rui Tam,Yung-Yu Shih,Shih-Hsin Wang,Yun-Nung Chen,Pu-Jen Cheng,Eugene Yang

类目:Information Retrieval (cs.IR)

关键词:Generative recommenders decode, recommenders decode items, misses their current, Amazon Beauty, decode items

备注:

点击查看摘要

Abstract:Generative recommenders decode items from a user's interaction history, but offer no way for users to correct a recommendation that misses their current intent. Adding conversation is natural since items and words share same output space, yet training the model to converse may overwrite the history-to-item mapping it relies on. We introduce INTEGER (**INTE**ractive **GE**nerative **R**ecommendation), which extends generative recommendation to multi-turn interaction with a learned routing token that lets the model decide when to recommend, history re-anchoring that conditions each item on both past behavior and the dialogue, and behavioral replay with instruction-data rehearsal that prevents forgetting during adaptation. Users can thus give feedback on recommendations within the dialogue, while recommendations stay grounded in behavioral history and accuracy is not traded for fluency. On Amazon Beauty and Toys, INTEGER matches or exceeds the strongest baselines in accuracy with competitive conversation quality, improving Hit@10 by 13.3% on Amazon Beauty, and significantly outperforms the generative recommender it starts from. Our analyses show that INTEGER learns behaviors that naive adaptation fails to acquire, recommending once the user's intent is clear and staying attentive to behavioral history at the moment of recommendation. INTEGER also learns an intent-agnostic replacement over the item space, which suppresses rejected items but points to attribute-aware feedback as the next step.

10. 【2610.08077】Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

链接:https://arxiv.org/abs/2610.08077

作者:Haoxiang Zhang,Qinglin Chen,Hiroaki Hayashi,Zhuofeng Li,Siming Zhang,Jiaxin Zhang,Jixuan Chen,Fang Wu,Pan Lu,Silvio Savarese,Julian McAuley,Chien-Sheng Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:scalar outcome rewards, Reinforcement learning, turns agent experience, primarily through scalar, scalar outcome

备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to $24.2$ pp. Its advantage is especially pronounced when reward contrast is scarce: when $37$--$98\%$ of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where $98\%$ of groups are all-failure, the RLVR training ends up at $0.0\%$ success, while adding SRD reaches $60.6\%$ under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

11. 【2610.07960】From Delivery to Stateful Exploration: Rethinking the Index for Agentic Search

链接:https://arxiv.org/abs/2610.07960

作者:Deogyong Kim,Sunghwan Kim,Sangam Lee,Wonjae Lee,Dongha Lee

类目:Information Retrieval (cs.IR)

关键词:large language model, Recent advances, agents finer control, language model, large language

备注: Work in Progress

点击查看摘要

Abstract:Recent advances in agentic search have given large language model (LLM) agents finer control over corpus exploration. However, search interfaces often return matching passages even when feedback about the candidate set would suffice for the next decision, coupling candidate refinement with source-text exposure. We propose IndexAct, an interface for Index-Native Corpus Interaction that separates candidate-set refinement from text inspection. Agents construct and manipulate persistent candidate sets through lexical conditions and set operations over an inverted index, receiving reusable state references and statistics such as candidate counts rather than matching passages. This feedback guides further refinement, while separately requested passages provide new clues or evidence that can inform subsequent operations on retained candidate sets. Experiments on five benchmarks spanning agentic search and multi-hop question answering show that IndexAct outperforms the evaluated baselines on each benchmark. On BrowseComp-Plus, it also achieves higher evidence coverage with a smaller average live context than terminal-based corpus interfaces, and maintains answer accuracy as the corpus expands. Further analyses suggest that informative refinement feedback and state reuse support continued evidence discovery, while shorter contexts or fewer search steps alone do not ensure better performance.

12. 【2610.07886】ShanLiangRen: A Nutrition Agent for Personalized Daily Meal Planning

链接:https://arxiv.org/abs/2610.07886

作者:Miao Xie,Xiao Zhang,Yuan Wang,Ruixin Zhu,Chunli Lv

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:chronic disease management, healthy body, plays an important, important role, role in chronic

备注:

点击查看摘要

Abstract:Dietary nutrition planning plays an important role in chronic disease management and maintaining a healthy body. In applications, it must simultaneously satisfy personalized constraints and reasonable multidimensional nutritional goals. These two aspects often conflict, and user constraints evolve with feedback, resulting in a substantial gap between generic guidelines and executable plans. To bridge this gap, we first propose the personalized fully quantified multiobjective dietary planning problem (MDP). To tackle MDP, we develop a nutrition agent, ShanLiangRen. The system first transforms dietary specifications, nutrient data, user attributes and natural language requirements into an individualized constrained planning instance. It then employs an exact retrieval-augmented generation method to shrink the feasible candidate set from a large scale ingredient and recipe space. Finally, it adopts a refinement guided by Pareto principles, where an LLM iteratively revises candidate plans under deterministic nutrition computation and feedback from constraint verification. The system outputs fully quantified meal plans with explicit ingredients and portion sizes, together with reports on nutrition compliance that show constraint satisfaction and nutrient interval attainment. We have released the system online as a WeChat Program, ShanLiangRen. A demo video is available at this https URL.

13. 【2610.07761】Contrastive Learning for Aspect Representation towards Explainable Recommendation

链接:https://arxiv.org/abs/2610.07761

作者:Emrul Hasan,Chen Ding

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:integrates aspect features, aspect features learned, integrates aspect, Contrastive Learning, Explainable Recommendation

备注: 8 pages. Published in WI-IAT 2025. Best Student Paper Award

点击查看摘要

Abstract:In this work, we propose a novel recommendation model, CLARER (Contrastive Learning for Aspect Representation towards Explainable Recommendation) that integrates aspect features learned from textual reviews with rating information to improve the accuracy and explainability of recommendations. Our proposed framework learns user and item representations by combining rating-based features and aspect-based features from reviews. Specifically, rating-based features are learned through a multi-layer perceptron (MLP) model, while aspect-specific review representations are learned using a transformer encoder to capture the semantic information and contrastive learning to better distinguish user preferences. To provide explanations, we train a transformer decoder, using the final representations of users and items from both rating and aspect-based features as context. Experimental results in three benchmark data sets demonstrate that our model achieves superior performance compared to baseline methods in both recommendation (accuracy) and explanation generation.

14. 【2610.07760】oken-Budgeted Escalation for Financial Document QA: Cost Is Predictable, Benefit Is the Bottleneck

链接:https://arxiv.org/abs/2610.07760

作者:Junru Zhu,Yixin Yang,Xiaoqing Ding,Ruoyu Qi

类目:Information Retrieval (cs.IR)

关键词:Retrieval-augmented generation systems, Retrieval-augmented generation, deeper context, vary by query, route difficult queries

备注: 6 pages, 3 figures, 6 tables. Code and aggregate artifacts: [this https URL](https://github.com/junru-zhu/token-budgeted-escalation-financial-qa)

点击查看摘要

Abstract:Retrieval-augmented generation systems can route difficult queries to deeper context, but batch deployments must allocate a shared token budget across calls whose costs vary by query. We formulate selective escalation as finite-batch allocation for financial document question answering. Each of 150 FinanceBench questions first receives a top-1 retrieval answer. Predictors estimate the adjudication-quality gain and token cost of an optional top-5 call, and the allocator prioritizes calls by predicted gain per token. At the nominal 10% budget, gain-per-token allocation improves adjudication quality over gain-only ranking by 0.034 (95% document-bootstrap CI [0.001, 0.072]) while using 46.6% fewer total tokens than one-pass top-5 retrieval. Additional-call cost is accurately predictable (R-squared 0.93), whereas beneficial escalation remains difficult to rank (AUROC 0.60). These results show that heterogeneous cost is actionable under tight constraints, while progress across the full budget frontier depends on stronger query-specific benefit estimates.

15. 【2610.07731】Learning to Retrieve via Reinforcement Learning in Embedding Space

链接:https://arxiv.org/abs/2610.07731

作者:Qi Liu,Fengming Liang,Yiqun Chen,Erhan Zhang,Jiaxin Mao

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Dense retrieval models, learn effective representations, optimize retrieval metrics, Dense retrieval, directly optimize retrieval

备注:

点击查看摘要

Abstract:Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.

16. 【2610.07622】DBRAG: Multi-Table Retrieval-Augmented Generation for Complex Database Queries

链接:https://arxiv.org/abs/2610.07622

作者:Prince Larbi Ampofo,Ryoji Kubo,Djellel Difallah

类目:Information Retrieval (cs.IR)

关键词:large language models, Recent advancements, advancements in large, large language, language models

备注:

点击查看摘要

Abstract:Recent advancements in large language models have introduced new capabilities for reasoning over structured data, particularly through program-aided tools that can analyze tables. However, many existing methods address single-table scenarios or assume that the relevant tables are already provided. In practice, users often issue complex data exploration queries over entire databases, where relevant information may be distributed across multiple relations. In this work, we introduce DBRAG, a retrieval-augmented generation framework tailored for multi-table question answering. DBRAG first retrieves candidate tables using an offline table index, enriches their summaries with query-relevant rows, and uses an LLM to rerank the candidates. A program-aided reasoner then selects the required tables and executes operations over their full contents, keeping the initial prompt context compact. Experiments on the Spider, GeoQuery, and ATIS datasets used in this study demonstrate improvements in table retrieval and multi-table question answering.

17. 【2610.07437】A Study of Prior Case Retrieval Using Lexical, Semantic, and Rhetorical Role Information in Indian Legal Documents

链接:https://arxiv.org/abs/2610.07437

作者:Sayed Ayaan Ahmed Sha,Sangeetha Sivanesan,Anand Kumar Madasamy,Navya Binu,Aniket Mani,Rishu Kumar

类目:Information Retrieval (cs.IR)

关键词:retrieving prior cases, involves distinguishing relevant, Indian legal judgments, Indian Legal Prior, problem involves distinguishing

备注:

点击查看摘要

Abstract:For retrieving prior cases in Indian legal judgments, the problem involves distinguishing relevant legal facts from mere lexical similarities because a prior case that shares a statute with the query judgment is not necessarily relevant. In this paper, we describe an empirical evaluation of three consecutive designs of retrieval systems for the IL PCR(Indian Legal Prior Case Retrieval) task. We demonstrate that the combination of the rhetorical roles (Fact, Ratio Of The Decision, Precedent, Argument, Statute) is better than either using individual roles or performing full text retrieval, that statute similarity is not discriminative, and that a legal entailment reranker with training data produced by an LLM is much worse in terms of official evaluation than its internal validation score. Two rankers, with excellent internal MRR up to 0.97, performed poorly in terms of official evaluation (MRR as low as 0.14). Thus, we propose a new design with query disjoint splitting and frozen validation fusion. The result is a four stage pipeline with Micro F1 0.2549, MRR 0.6081, and nDCG@10 0.4187.

18. 【2610.07402】Rethinking Semantic ID Construction for Generative Recommendation: SimHash with Parallel Decoding and Semantic Alignment

链接:https://arxiv.org/abs/2610.07402

作者:Yuqing Liu,Huiyuan Chen,Yibo Wang,Wooseong Yang,Philip S. Yu

类目:Information Retrieval (cs.IR)

关键词:enabling structured modeling, discrete tokens, enabling structured, sequence of discrete, structured modeling

备注: Accepted at NeurIPS 2026. Code: [this https URL](https://github.com/KevinC2015/Flash)

点击查看摘要

Abstract:Semantic ID-based generative recommendation represents each item as a sequence of discrete tokens, enabling structured modeling of item semantics. A critical challenge is constructing semantic IDs that are both semantically expressive and computationally efficient. While recent approaches favor complex learned quantization, simple hashing-based methods such as SimHash are widely regarded as fundamentally inferior. In this work, we challenge this consensus by showing that the apparent performance gap does not stem from inherent limitations of hashing, but rather from a structural mismatch with autoregressive decoding, coupled with the inevitable information loss during rigid discretization. Based on this insight, we propose FLASH, a two-stage framework that revitalizes training-free SimHash tokenization through parallel decoding and explicit semantic alignment. Despite its simplicity, FLASH achieves state-of-the-art performance across multiple datasets without requiring any tokenizer training, while exhibiting stronger generalization in cold-start scenarios. Notably, we demonstrate that semantic alignment acts as a universally effective mechanism across diverse paradigms. Our findings suggest that, with compatible decoding and semantic grounding, simple and efficient tokenizers can achieve performance comparable to complex learned counterparts in generative recommendation. Our code is available at this https URL.

19. 【2610.07384】WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

链接:https://arxiv.org/abs/2610.07384

作者:Turhan Can Kargin,Piotr Kubaty,Ekaterina Rostovskaya,Izabela Wierzbowska,Bartosz Zieliński,Marcin Przewięźlikowski

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:retrieval problem central, instance retrieval problem, instance retrieval, central to non-invasive, retrieve the correct

备注: 15 pages, 7 figures, 3 tables. Project page: [this https URL](https://wildmatch.gmum.net)

点击查看摘要

Abstract:Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin, or other visual markings. Current approaches either learn global embeddings as a classification problem, requiring many labeled images per individual while largely ignoring local evidence, or apply off-the-shelf, domain-agnostic image matchers. Although such matchers are pretrained on large and diverse image collections, adapting them to wildlife imagery is challenging because available datasets are small and lack correspondence-level annotations. We study weakly supervised adaptation of a pretrained keypoint matcher using only identity labels, without keypoint-level or geometric correspondence ground truth. We mine informative image pairs with the pretrained matcher, derive weak positive and negative supervision from identity agreement, and contrastively fine-tune the matching network to strengthen correspondences for same-identity pairs and suppress them for different identities. Across open-source wildlife re-identification datasets, our approach improves accuracy over off-the-shelf matchers and a state-of-the-art local--global fusion method. Under an open-world protocol with held-out individuals, it learns a transferable correspondence prior rather than memorizing training identities. To our knowledge, this is the first study of matcher-level, identity-supervised adaptation for animal re-identification. Our method enables data-efficient specialization of image matching models to wildlife domains using identity annotations already available in typical monitoring datasets.

20. 【2610.07309】he Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory

链接:https://arxiv.org/abs/2610.07309

作者:Zi Wang,Xingqiao Wang,Emmanuel Addai,Devika Ambekar,Xiaowei Xu

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multiagent Systems (cs.MA)

关键词:incompatible lifecycle state, retrieve relevant information, violates policy, agents can retrieve, lifecycle state

备注: 26 pages. Accepted at the NeurIPS 2026 Workshop "Who Verifies the Agents? Toward Reliable Agent Development". Code: [this https URL](https://github.com/ziwang11112/right-memory-wrong-context)

点击查看摘要

Abstract:Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.

21. 【2610.07266】RAGFlip: Measuring Query-Level Negative Flips in Retriever Upgrades

链接:https://arxiv.org/abs/2610.07266

作者:Elyas Irankhah,Muhammad Arif

类目:Information Retrieval (cs.IR)

关键词:served correctly, upgrades are typically, Natural Questions, hide regressions, regressions

备注: 19 pages, 4 figures. Code available at [this https URL](https://github.com/Elyasirankhah/RAGFlip)

点击查看摘要

Abstract:Retriever upgrades are typically evaluated using aggregate metrics, which can hide regressions on queries the previous retriever already served correctly. We study these regressions as negative flips: queries for which BM25 retrieves a judged relevant passage and the replacement does not. We evaluate BGE-large, E5-large-v2, and SPLADE on three BEIR collections: Natural Questions, HotpotQA, and FiQA, across five retrieval depths. All replacements improve overall retrieval coverage. Negative flips occur in every setting and vary substantially by corpus, retriever, and depth. At k=1, 8.6-37.5% of BM25 successes are lost across the evaluated settings. Negative-flip rates are lower in the larger-depth settings, where the BM25-supported cohort is defined separately at each depth. These rates use the any-relevant support label. On HotpotQA at k=10, requiring every positive qrel passage raises the negative-flip rate to 12.3-17.3%. Simple fixed-budget combinations with BM25 reduce these regressions, and a HotpotQA reader experiment provides a limited downstream check in which some retrieval flips are accompanied by answer regressions. These results motivate evaluating retriever updates using query-level compatibility alongside aggregate retrieval quality.

22. 【2610.07132】CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

链接:https://arxiv.org/abs/2610.07132

作者:Berke Arda,Ahmetcan Yavuz,Paul Gerry,Sebastian Lobentanzer,Nobin Sarwar,Joan Giner-Miguelez,Kongtao Chen,Luyao Zhang,Mrinmaya Sachan,Mubashara Akhtar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:accompanying dataset documentation, machine-readable dataset metadata, requires careful reading, dataset documentation, machine-readable dataset

备注: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: [this https URL](https://berkearda.github.io/croissantminer/)

点击查看摘要

Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

23. 【2610.07105】Beyond Successor Accuracy: State Retention for Recursive Self-Improvement in Recommendation

链接:https://arxiv.org/abs/2610.07105

作者:Jinfeng Xu,Zheyu Chen,Ziyue Peng,Zheng Lin,Wenhao Yuan,Jian Chen,Shujie Li,Edith Ngai

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:feeds recommender outputs, Recommendation recursive self-improvement, recursive self-improvement, feeds recommender, subsequent training

备注:

点击查看摘要

Abstract:Recommendation recursive self-improvement (Rec-RSI) feeds recommender outputs into subsequent training. Evaluating each round solely through its latest model assumes that the successor consolidates the update, although pre- and post-update models may retain complementary ranking decisions. We term this \emph{distributed progress} and quantify it using cross-generation advantage (CGA), a marginally matched contrast between cross- and within-generation model pairs. A rank-separation statistic, label-free at selection time, predicts which family to retain. Across four datasets and three sequential recommendation encoders, the preferred retention regime varies by architecture: cross-generation pairing benefits GRU4Rec and SASRec, whereas FMLP initially favors within-generation pairing and shifts toward cross-generation pairing after a second update. Rank separation selects the stronger family in 12/12 first-update and 5/6 second-update dataset-encoder settings; on held-out tests, the selected family outperforms the direct successor in 34/36 trajectories. Five transfer mechanisms do not consistently reproduce these gains in one model. These findings establish state retention as a distinct Rec-RSI problem: progress may reside in relations between generations as well as in the latest model. Code is available at \href{this https URL}{this https URL}.

24. 【2610.07091】Smart Content Ingestion for Generative AI Workloads

链接:https://arxiv.org/abs/2610.07091

作者:Abbas Raza Ali,Muhammad Ajmal Siddiqui,Moona Zahid

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:conventional machine learning, machine learning, progressively changed, changed where intelligence, intelligence resides

备注:

点击查看摘要

Abstract:The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any single task: one foundation model serves open-ended downstream tasks, and the generality gained on the model side is matched by heterogeneity on the data side, because enterprise knowledge is authored in the formats people use (PDF, presentations, spreadsheets, scanned documents, forms, tables, diagrams and mixed-layout files) that carry textual, visual, geometric and structural information at once. A language model or retriever cannot reason reliably over information misrepresented at this interface, so content extraction becomes a lifecycle stage in its own right whose errors no downstream retriever or re-ranker can repair. This paper presents a production-ready content-extraction system that makes this stage explicit, configurable, and measurable. The system incorporates selective OCR routing, a scarcity-first curation engine with a reference-based extraction scorer that measures character, word, and table-structure accuracy, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator that generates grounded questions from every page and reports Hit@k, mean reciprocal rank, and latency. On a 180-document corpus the best extractor scores 97.4 of 100 (character error rate 0.13%, table similarity 0.995) and the chunker reaches hit@1 of 68.6%, hit@10 of 92.8% and MRR 0.77 over 25,050 generated questions. We distil three design principles (structure before semantics, never mutate what you measure, budget your labels) and position measured content extraction as the perception layer of enterprise agentic systems.

25. 【2610.07023】Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

链接:https://arxiv.org/abs/2610.07023

作者:Jinghao Pang,Jitai Hao,Qiang Huang,Zhaochun Ren,Jun Yu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Large Language Models, achieved remarkable capabilities, Large Language, safety alignment, unsafe outputs

备注: 27 pages,7 figures, under review

点击查看摘要

Abstract:Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model's general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.

26. 【2610.06902】ree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA

链接:https://arxiv.org/abs/2610.06902

作者:Priyank Jayraj,Poonam Goyal,Navneet Goyal

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:Retrieval-augmented generation grounds, miss complementary evidence, generation grounds language, grounds language models, Retrieval-augmented generation

备注:

点击查看摘要

Abstract:Retrieval-augmented generation grounds language models in external context, but for long documents flat top-$k$ retrieval can cluster on a single region and miss complementary evidence. RAPTOR-style summary trees address this by recursively clustering chunks and using a language model to summarize each cluster at indexing time, then ranking summary nodes alongside raw chunks at query time. We show the main benefit of summary trees in long-document QA can come from navigation rather than the generated summary content. We introduce NavTree, a leaves-only retriever that builds a deterministic balanced segment tree over chunks (zero language-model calls at indexing) and uses the tree purely as a navigation scaffold: a hybrid lexical-and-dense frontier walk, anchored on top retrieved leaves, descends from the root and emits only leaf chunks to the reader. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree is the strongest matched-cost hierarchical retriever in our evaluated grid and ties the strongest flat baseline. On long-document multi-hop QA, it is the only hierarchical method that significantly beats BM25 on a class-vs-class basis, corroborated by a reader-free retrieval-recall check. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still loses to NavTree at every multi-chunk budget, at zero indexing cost. The ranking carries across stronger and open-weight readers, a stronger encoder, and a full factorial that isolates leaves-only emission as the structural lever.

27. 【2610.06857】Diff-SQL: SQL Efficiency Optimization via Patch Generation and Constraint Alignment

链接:https://arxiv.org/abs/2610.06857

作者:Shipei Lin,Duomin Zhang,Xiaolong Li,Bohan Hu,Bowen Qin,Jinyang Li,Chenhao Ma

类目:Databases (cs.DB); Information Retrieval (cs.IR)

关键词:transform slow queries, faster alternatives, SQL, aims to transform, transform slow

备注:

点击查看摘要

Abstract:SQL efficiency optimization aims to transform slow queries into semantically equivalent but faster alternatives. However, directly optimizing SQL with large language models in an end-to-end fashion often induces Objective Misalignment which creates a fundamental tension between optimization and correctness, making direct full SQL rewriting unreliable for execution-facing database applications. To address this problem, we propose Diff-SQL, a two-stage framework that decouples efficiency-oriented optimization from constraint-aware alignment. The first stage identifies optimization opportunities and proposes targeted edits in the form of a unified diff patch, while the second stage is trained with on-policy reinforcement learning to revise outputs under executability and semantic-equivalence constraints. To train and evaluate Diff-SQL, we construct an automated pipeline that mines optimization knowledge from StackOverflow and builds Slow-Fast SQL pairs through cascaded filtering. We further introduce Effi-SQL, a benchmark containing 1,100 human-verified Slow-Fast pairs across five SQL dialects. Experiments show that Objective Misalignment is widespread across existing LLM-based SQL optimization methods, where direct full SQL optimization causes an average 22.7% execution accuracy degradation across frontier models such as Claude-Opus-4.6, with the worst model dropping by 43.0%. Diff-SQL alleviates this trade-off. As an inference-only strategy, it improves R-VES by 10.0% on average while reducing execution accuracy degradation by 6.11% on average across three strong base models. With execution-grounded training, Diff-SQL further enables a 7B model to improve R-VES from 33.42% to 46.83%, demonstrating that the proposed two-stage optimization-and-alignment paradigm can deliver both stronger efficiency and better correctness in local, small model deployment settings.

计算机视觉

1. 【2610.08791】World Models' Last Exam in Physics

链接:https://arxiv.org/abs/2610.08791

作者:Mingju Gao,Qingle Liu,Yuzhao Peng,Xinjie Lin,Ziming Qin,Zheng Jiang,Wenyi Li,Calvin Xiao,Youjie Zheng,Kaisen Yang,Qinhuai Na

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:produce visually convincing, physically inconsistent sequences, Video world models, inconsistent sequences, raising concerns

备注:

点击查看摘要

Abstract:Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and planning in embodied AI systems. Existing evaluations often rely on model-based judgments or reference videos, while direct physical tests largely focus on mechanics. We introduce World Models' Last Exam in Physics, a measurement-based benchmark for evaluating physical consistency in video world models. The benchmark comprises 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism, and surface tension. Each task pairs an initial image and a generation prompt with predefined physical criteria, enabling interpretable tests of observable physical relationships without requiring reference videos. Its evaluator combines task-observability screening with task-specific quantitative physical measurements. Experiments on eight video generation models across 1,280 videos reveal persistent physical inconsistencies and substantial variation across tasks, with the best model achieving an overall score of 57.76 out of 100. Evaluation on synthetic videos with known physical relationships provides evidence for the validity of the measurement module under controlled conditions. The evaluator also achieves higher agreement with human judgments than a direct vision-language model baseline in both within-task rankings and pairwise comparisons. By combining coverage across physical domains with scores grounded in measurable evidence and explicit measurement limitations, the benchmark provides an interpretable basis for diagnosing physical inconsistencies and tracking progress toward physically consistent video world models.

2. 【2610.08790】Building Rome from a Single Image

链接:https://arxiv.org/abs/2610.08790

作者:Jiraphon Yenphraphai,Fang Li,Tianshuo Xu,Depu Meng,Quentin Herau,Yihan Hu,Raymond A. Yeh,Wei Zhan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Single-image scene generation, scene generation aims, Single-image scene, produce a complete, outdoor scenes

备注: Project page: [this https URL](https://build-rome.github.io/)

点击查看摘要

Abstract:Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b) making the generator capture explicit 2D-3D correspondence by lifting image features and making the model aware of the free space, observed surface, and unobserved region; (c) synthesizing around 4,000 outdoor scenes to broaden the training data, as existing scene datasets are largely indoor. Experiments on Tanks and Temples, ScanNet++, and in-the-wild images show that our method outperforms all baselines in geometric accuracy and perceptual quality across both indoor and outdoor scenes.

3. 【2610.08782】4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

链接:https://arxiv.org/abs/2610.08782

作者:Shiqi Li,Sean Cho,Yijie Li,Fengzhi Guo,Bowen Wen,Cheng Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:approaches typically synthesize, Existing methods, unstable interaction prediction, typically synthesize interactions, costly per-sequence optimization

备注: Project page: [this https URL](https://tamu-visual-ai.github.io/4D-HOF/)

点击查看摘要

Abstract:Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.

4. 【2610.08780】DepthWorld: 3D World Model for Robot Manipulation

链接:https://arxiv.org/abs/2610.08780

作者:Jai Bardhan,Josef Sivic,Vladimir Petrik

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:spanning policy evaluation, applications spanning policy, World models offer, simulators for robotics, policy evaluation

备注: Accepted at the Conference on Robot Learning (CoRL) 2026. Project page: [this https URL](https://www.jaibardhan.com/depthworld) . 32 pages including supplementary material, 15 figures, 7 tables

点击查看摘要

Abstract:World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving 0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.

5. 【2610.08779】ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

链接:https://arxiv.org/abs/2610.08779

作者:Zhenghong Zhou,Zhe Lin,Jiebo Luo,Yuqian Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Current video editors, Current video, Current, makes inserted objects, insert objects

备注: Project page: [this https URL](https://real-time-video-research.github.io/alive/)

点击查看摘要

Abstract:Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects "alive" through coherent interactions with the source video's contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object's presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess interaction fidelity, source preservation, and visual coherence using a unified VLM-based protocol, and evaluate on the general video object insertion benchmark. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% and 4.4% on the two benchmarks, respectively. VLM-predicted guidance further improves the ALIVE-interaction score by 0.95 points without additional user inputs.

6. 【2610.08777】CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

链接:https://arxiv.org/abs/2610.08777

作者:Shangye Song,Dong Gong,Hong Jia,Yun Sing Koh,Xinyu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Interactive video world, video chunk efficiently, video world models, video world, Interactive video

备注: 18 pages. Project page: [this https URL](https://wrecklong.github.io/CtrlCache/)

点击查看摘要

Abstract:Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.

7. 【2610.08772】Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation

链接:https://arxiv.org/abs/2610.08772

作者:Liao Ma,Jiayi Song,Yunfeng Wu,Songhua Liu,Peilin Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Diffusion Transformers, achieved strong performance, generation computationally expensive, full attention makes, attention makes high-resolution

备注:

点击查看摘要

Abstract:Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity. However, isolated windows block cross-window interaction, often introducing visible grid-like artifacts in the generated results. Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality. However, its irregular computation patterns create a substantial gap between theoretical and practical speedups and require specialized kernels tailored to each hardware backend. To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration. Specifically, BASA replaces visual self-attention with shifted local-window attention. By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts. Notably, our design introduces no additional irregular operators or customized kernels, making it readily deployable on existing attention backends and closing the gap between theoretical sparsity and practical acceleration. Experiments demonstrate that BASA achieves measured speedups exceeding 90\% of the theoretical estimates on FLUX and delivers a 4.52$\times$ attention speedup on Wan while maintaining competitive generation quality.

8. 【2610.08770】Data Leakage in Patch-Based Hyperspectral Image Classification: Quantifying the Impact of Spatial Overlap

链接:https://arxiv.org/abs/2610.08770

作者:Mohammed Q. Alkhatib

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:improves hyperspectral image, learning improves hyperspectral, local spectral-spatial information, exploiting local spectral-spatial, Patch-based learning improves

备注: paper accepted for presentation at IEEE-WHISPERS

点击查看摘要

Abstract:Patch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling from the same image can cause spatial patch overlap, leading to data leakage and optimistic performance estimates. This paper investigates same-class train-test spatial overlap in patch-based HSI classification using two measures: overlap percentage (OP), which quantifies the global amount of overlapped testing patch pixels, and average overlap ratio (AOR), which measures the local severity among affected testing patches. Experiments on the Pavia University dataset compare random and non-random spatial sampling using SVM, MLP, 2D-CNN, 3D-CNN, ViT, and MorpMamba. The results show that deep patch-based models achieve high accuracy under random sampling, with 3D-CNN reaching 96.17% Overall Accuracy (OA), but drop substantially under non-random spatial sampling, where 3D-CNN decreases to 55.20% and ViT and 2D-CNN drop by 40.71 and 38.81 percentage points (PP), respectively. Patch-size analysis further shows that increasing the patch size from 5x5 to 19x19 raises the random-sampling overlap percentage from 23.28% to 77.02%. These findings demonstrate that random patch-based evaluation can substantially inflate classification performance, especially for models that strongly exploit spatial context. The code associated with this paper is available at: this https URL.

9. 【2610.08760】WorldSonus: Bringing Sound to Worlds

链接:https://arxiv.org/abs/2610.08760

作者:Pengjun Fang,Jingyi Fa,Kam Man Wu,Jiaming Wang,Haoyuan Huang,Yaguang Wu,Xiangjun Huang,Ziyang Ma,Weijia Chen,Hongyu Liu,Zeyue Tian,Qifeng Chen

类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)

关键词:enabled increasingly realistic, increasingly realistic visual, Recent advances, realistic visual synthesis, world models

备注: 25 pages, 4 figures, 16 tables. Project page: [this https URL](https://noizai.github.io/WorldSonus/)

点击查看摘要

Abstract:Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: this https URL

10. 【2610.08756】Post-Training Semantic Lifting for 3D Gaussian Splatting: Separating Detector, Lifting and Representation Error

链接:https://arxiv.org/abs/2610.08756

作者:Iván Verdugo Guerra,Ezequiel López Rubio,Jorge García González

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting model, Gaussian Splatting, Splatting model, Gaussian, Gaussians

备注: 18 pages, 11 figures, 9 tables. Code: [this https URL](https://github.com/ivanver02/semantic-lifting-3dgs)

点击查看摘要

Abstract:The same Gaussian of a 3D Gaussian Splatting model is seen from many views, and these views do not always agree on the class it belongs to. The Gaussian may be occluded in some of them, and the confidence of the detector is not the same from one view to another. The ground truth, on the other hand, is given as an annotated mesh, because two training runs do not produce the same Gaussians. In this work, we propose a post-training lifting method that works with one target class at a time and combines the information coming from all the views. Target and non-target evidence are accumulated simultaneously, weighted by the visibility of each Gaussian in each view. After that, the Gaussians are filtered with two thresholds: a main threshold $\beta$ selects the high-confidence seeds, and a lower one $\gamma\beta$ adds the connected components around them. For the evaluation, the labels are transferred from the Gaussians to the mesh vertices that are both visible and annotated. With this design, we can separate three sources of error: the 2D detector, the lifting and the transfer between representations. The thresholds and the transfer operator are chosen on seven Replica validation scenes, and the method is evaluated on ten held-out ScanNet++ scenes with the same values for every scene and class. The mean mIoU on the validation scenes was 0.93 with masks from the dataset annotations and 0.65 with YOLO masks, and on the ScanNet++ test scenes it was 0.80 and 0.54. Compared with thresholding the evidence per view, as a previous version of the method did, the fraction improves the test mIoU by 0.24 and makes it possible to use a single threshold for all the classes and scenes of both datasets. Finally, the error analysis shows that most of the remaining error comes from the detector.

11. 【2610.08717】Co-Evolving Paths and Flows via Path-Flow Alignment

链接:https://arxiv.org/abs/2610.08717

作者:Zeyu Michael Li,William Xingxu Chen,Xiang Cheng

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:study path-flow alignment, alignment loss, path, flow matching, unified training objective

备注:

点击查看摘要

Abstract:We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We find that this failure is associated with low-entropy bottlenecks in the induced probability path, where the learned path routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularization gives an explicit entropy floor for the stochastic training-path marginals and empirically suppresses the bottleneck in the learned sampler, making joint path-flow training effective. On ImageNet-256x256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at this https URL

12. 【2610.08713】SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning

链接:https://arxiv.org/abs/2610.08713

作者:Hairong Yin,Huangying Zhan,Shin-Fang Chng,Yi Xu,Raymond A. Yeh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Embodied agents, answering questions, agents must reason, Embodied, Streaming VLMs process

备注:

点击查看摘要

Abstract:Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.

13. 【2610.08704】Local Content-Style Control for Diffusion-based Image Stylization

链接:https://arxiv.org/abs/2610.08704

作者:Amir Semmo

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:latent-diffusion models entangles, independently refined axes, Image stylization, refined axes, latent-diffusion models

备注: SIGGRAPH Asia 2026 Technical Communications. 4 pages, 4 figures, 1 table. Supplemental material included as an ancillary file

点击查看摘要

Abstract:Image stylization with latent-diffusion models entangles two independently refined axes: what a region depicts and how it is depicted. Such pipelines expose only global controls, yet professional retouching demands deliberate, region-specific control. We lift two conditioning weights already present in a ControlNet + IP-Adapter stylization pipeline from global scalars to per-location spatial maps, yielding local, per-axis control of content and style in a single generative pass. Because the two weights act on disjoint pathways, adjusting them independently spans a 2x2 retouching vocabulary, from free regeneration to identity preservation. We validate that edits stay confined to the retouched region and that each weight predominantly steers its own axis. Our approach requires no retraining and drops unchanged into any such pipeline.

14. 【2610.08684】RenderBench: Benchmarking Render-to-Real Video Transfer with Reconstructed Digital Twins

链接:https://arxiv.org/abs/2610.08684

作者:Dicong Qiu,Zhiyuan Xu,Yaosheng Liu,Feng Han,Bo Ye

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Modern video models, generate realistic videos, Modern video, generate realistic, real target video

备注: 10 pages, 4 figures, 2 tables

点击查看摘要

Abstract:Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video depicting the same scene evolution, paired with an editable, geometrically registered 3D replica. Such data has traditionally required substantial manual modeling, calibration, and animation effort. We introduce RenderBench, a benchmark of 12 reconstructed real-world scenes spanning large-scale indoor environments and egocentric viewpoints, with both static and dynamic settings. Our construction pipeline combines visual geometry, neural reconstruction, and assisted 3D authoring. Each scene is decomposed into static objects and dynamic actors, registered to the capture cameras, and accepted only after multi-view geometric and temporal validation. Each evaluation unit contains appearance reference images, a held-out real target video, an editable digital twin, a matched proxy render, and renderer-native scene annotations. We evaluate transfer models against paired real target videos, retain PAI-Bench-C-compatible structural projections, and use scene annotations to localize failures by object, visibility, articulation, and motion. The first release retains 12 of 14 registered samples (85.7%), comprising 1,496 paired real-proxy frames. All released scenes pass file-integrity and environment-edit audits, while proxy diagnostics yield a depth si-RMSE of 0.2170 and instance mIoU of 0.3673. RenderBench provides paired real observations and editable scene state for assessing both appearance fidelity and preservation of geometry and dynamics.

15. 【2610.08674】EC-RAG: Event Chain Retrieval-Augmented Generation for Long Video Understanding

链接:https://arxiv.org/abs/2610.08674

作者:Yuhao Qin,Junbo Wang,Yuke Li,Yining Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Current large video-language, Current large, large video-language models, processed independently, making it difficult

备注: 12 pages, 7 figures, 7 tables, including supplementary material

点击查看摘要

Abstract:Current large video-language models (LVLMs) still face challenges when dealing with long videos, mainly because frames are often processed independently, making it difficult to capture temporal dependencies across events. Although retrieval-augmented approaches have been introduced to provide additional context, most of them operate at the frame or snippet level, which limits their ability to model how events evolve over time and relate to each other. In this paper, we propose Event Chain Retrieval-Augmented Generation (EC-RAG), a training-free framework that organizes video content into an explicit event chain before question answering. Instead of retrieving isolated frames or text segments, EC-RAG first partitions the video into semantically coherent segments, represents each segment using multi-modal signals, and then links them into a structured chain that preserves temporal order and captures inter-event relationships. Given a query, the system identifies relevant events within this chain and gathers supporting evidence from the associated modalities. Our approach offers several practical advantages: (i) event-level abstraction that better reflects how video content is naturally structured, enabling more reliable localization compared to frame-level retrieval; (ii) structured multi-modal fusion that aggregates speech, text, and visual cues at the event level, allowing complementary information to be more effectively utilized during reasoning; and (iii) plug-and-play compatibility with existing LVLM backbones, requiring no additional training or reliance on proprietary models. Experiments on Video-MME, MLVU, and LongVideoBench show that this event-centric design consistently outperforms frame-level retrieval baselines, highlighting the importance of modeling temporal structure for long-video understanding.

16. 【2610.08672】PDB: Point-Based Deformation Blending for Facial Animation Retargeting

链接:https://arxiv.org/abs/2610.08672

作者:Sihun Cha,Hyeonseung Shin,Suah Yu,Junyong Noh

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Mesh-agnostic facial animation, facial animation retargeting, Mesh-agnostic facial, artifacts remains challenging, facial animation

备注:

点击查看摘要

Abstract:Mesh-agnostic facial animation retargeting transfers expressions across meshes with different structures, but preserving facial motion without surface artifacts remains challenging. To address this, we present PDB, Point-Based Deformation Blending for facial animation retargeting. PDB predicts a compact set of deformed control points from a source neutral-expression pair and blending weights from the target neutral mesh. The weights are computed once per target and reused across frames, while the control points vary with each source expression. ReLU enforces non-negative weights and permits exact zeros, followed by row-wise normalization. The target mesh is reconstructed directly by multiplying the weights and control points, without a predefined cage, precomputed coordinates, a learned per-element deformation decoder, or a global reconstruction solve. Trained only with self-retargeting reconstruction supervision, PDB supports cross-identity transfer without paired cross-identity training expressions. Experiments demonstrate accurate retargeting, fast inference, and localized support in the learned weights. Joint evaluation of expression accuracy and local surface preservation shows reduced surface artifacts relative to the evaluated dense displacement method while retaining the intended motion. Perceptual evaluations further support expression fidelity and visual quality in both self- and cross-retargeting.

17. 【2610.08663】Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction

链接:https://arxiv.org/abs/2610.08663

作者:Lichen Zhu,Yueqian Lin,Yiheng Wang,Hai "Helen" Li,Yiran Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video gaze prediction, priors carry signal, gaze-free priors carry, gaze-trained models, gaze prediction

备注:

点击查看摘要

Abstract:Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion's latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.

18. 【2610.08659】Selective Transfer of RL Updates for Visual Reasoning

链接:https://arxiv.org/abs/2610.08659

作者:Suxin Ji,Hungtao Wan,Mingjun Liu,An Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:transfer reasoning capabilities, conflate pre-existing model, pre-existing model differences, reasoning capabilities, conflate pre-existing

备注:

点击查看摘要

Abstract:Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at this https URL.

19. 【2610.08649】Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation

链接:https://arxiv.org/abs/2610.08649

作者:Lichen Zhu,Yiheng Wang,Yueqian Lin,Hai "Helen" Li,Yiran Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video-language models, models are ranked, ranked by multiple-choice, Video-language, phase

备注:

点击查看摘要

Abstract:Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchmarks report only the rate. The phase moves answers: two deployed samplers differing only by a half-step phase offset answer 23.6% of questions differently while scoring within a point, and across four releases from two families shifting only the phase changes roughly one answer in five after controlling option order. PHASEFUSION decodes three offset grids and averages the option posteriors. The grids are the polyphase components of the dense grid. Fusion matches a 32-frame single pass in accuracy within a prespecified margin (logit-scored) and cuts the answers a half-step shift of all three grids changes from 18.2% to 10.1%. Option order, which changes only the presentation, is flagged instead by a one-pass answer margin. Report the phase convention with the budget, or marginalize it.

20. 【2610.08639】Forensic Reserve: Eliciting Latent Knowledge for Image Forgery Detection

链接:https://arxiv.org/abs/2610.08639

作者:Jiahua Li,Zixu John,Tom Zhong,Fuping Wu,Tianhao Xu,Jianqing Zheng,Yuanhan Mo,Fei Shen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reliable forgery detection, increasingly realistic, reliable forgery, visual information, essential for maintaining

备注:

点击查看摘要

Abstract:As generated images become increasingly realistic, reliable forgery detection is essential for maintaining trust in visual information. However, existing methods primarily rely on task-specific supervision to adapt vision foundation model representations, without fully exploiting internal forensic knowledge to guide detection. To address this limitation, we propose Reserve-Guided Elicitation (RGE), a framework that treats sparse, origin-sensitive internal components in pretrained models as a forensic reserve and translates their localization into structural constraints for lightweight adaptation. Specifically, we first use the Forensic Lens (F-lens) to decompose activations across layers and token groups into independent components and globally screen them by their response differences between real and generated images, identifying reserve sites and directions. Next, we map the selected directions back to hidden-state space to construct fixed reserve subspaces and insert Forensic Reserve Adapters (FRA) only at the identified sites. Finally, with the backbone parameters, previously fitted reference classifier, and subspace bases fixed, we train only the FRA coefficient maps to generate input-dependent residual updates constrained to the corresponding subspaces, strengthening existing forensic responses. Using only 500 labeled training images and a trainable parameter budget below 0.2% of the backbone, RGE achieves competitive performance across three detection benchmarks without target-benchmark adaptation. Furthermore, RGE consistently improves over the corresponding frozen detectors across eight encoders spanning self-supervised and vision-language pretraining, eliciting a latent forensic capacity broadly shared across pretrained vision models.

21. 【2610.08620】LiDAR Resolution Recovery via Foundation-Model-Guided Diffusion

链接:https://arxiv.org/abs/2610.08620

作者:Samed Doğan,Nico Leuze,Alfred Schöttl

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dense angular sampling, perception pipelines require, pipelines require dense, require dense angular, pretrained Stable Diffusion

备注:

点击查看摘要

Abstract:High-beam-count LiDAR sensors are costly, yet many perception pipelines require dense angular sampling. Using a pretrained Stable Diffusion model as the backbone, we fine-tune a LiDAR-conditioned depth model with pseudo-depth targets from a 2D foundation model. During training, the LiDAR conditioning is randomly decimated at different beam budgets. We then investigate how much of a LiDAR scan can be recovered from heavily decimated input and characterize performance across the input beam budget. We evaluate against physically held-out real beams on nuScenes and report recovery separately from fit accuracy. Our model yields its largest advantage in very sparse regimes, achieving a $\delta_{1.25}$ accuracy of $66.8$% from $4$-beam input where scattered interpolation reaches only $45.1$%. A class-stratified error breakdown further reveals that planar surfaces recover first while objects introducing depth discontinuities degrade earliest. Together, these results quantify the recovery/resolution trade-off for foundation-model-guided LiDAR enhancement.

22. 【2610.08574】FedDermaSeg: Federated Learning for Dermatological Image Segmentation

链接:https://arxiv.org/abs/2610.08574

作者:Anabik Pal,Ganesh Patidar,Bikash Santra

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:major global health, accurate lesion delineation, skin lesion segmentation, global health concern, skin lesion

备注:

点击查看摘要

Abstract:Skin cancer is a major global health concern, and early detection and accurate lesion delineation are important for effective diagnosis and treatment planning. Automated skin lesion analysis can assist dermatologists, with lesion segmentation serving as a fundamental step in computer-aided diagnostic systems. Conventional deep learning-based segmentation models typically rely on centralized training, where images and their corresponding segmentation masks are collected on a central server. Such data aggregation raises privacy concerns in medical applications and requires substantial centralized computational resources. To address these limitations, we investigate the feasibility of federated learning for privacy-preserving skin lesion segmentation. The training and validation sets of the ISIC 2018 Skin Lesion Segmentation Challenge dataset are used to simulate a distributed learning environment and develop a federated segmentation model. The resulting model is evaluated on the ISIC 2018 test set and the PH2 dataset to assess its performance and generalizability. Experimental results demonstrate that the federated model achieves performance comparable to centralized training while consistently improving upon the locally trained models. These findings demonstrate the potential of federated learning for collaborative skin lesion segmentation without requiring centralized aggregation of medical images.

23. 【2610.08573】Sparse2comm: Towards Robust Cooperative 3D Object Detection

链接:https://arxiv.org/abs/2610.08573

作者:Lei Yang,Boqi Li,Chunmian Lin,Li Wang,Ziying Song,Shaoqing Xu,Heye Huang,Haibao Yu,Chen Lv

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:sharing complementary observations, perception improves autonomous, improves autonomous driving, Cooperative perception improves, autonomous driving

备注: 15 pages. Code: [this https URL](https://github.com/yanglei18/Sparse2comm)

点击查看摘要

Abstract:Cooperative perception improves autonomous driving by sharing complementary observations among vehicles and roadside infrastructure for 3D object detection. However, practical deployment is constrained by limited bandwidth and unreliable cooperation, where packet loss, transmission delay, and spatial misalignment jointly degrade the cooperative feature stream. Existing methods often reduce communication cost or compensate for one degradation type, leaving coupled disturbances insufficiently addressed. To address this problem, we propose Sparse2comm, a bandwidth-efficient and robust cooperative 3D object detection framework that treats unreliable cooperation as progressive restoration over degraded cooperative features. Sparse Feature Encoding first encodes communication as randomly mask-sampled foreground features transmitted by collaborating agents, from which the ego vehicle reconstructs dense semantic representations. This sparse-to-dense mechanism learns to infer missing object-centric content from sparse observations, enabling ultra-low-bandwidth communication and packet-loss recovery within the same representation. On the semantically restored features, Latency-Aware Alignment predicts motion flow to compensate delayed messages, and Self-Calibrating Fusion estimates residual spatial offsets in a self-supervised manner before adaptive cross-agent fusion. Sparse2comm therefore restores semantic completeness, temporal consistency, and spatial alignment in an ordered pipeline. Extensive experiments on DAIR-V2X, OpenV2V, and V2V4Real show that Sparse2comm maintains competitive clean accuracy and consistently improves robustness under individual and mixed real-world degradations. Compared with the selective feature communication baseline Where2comm, Sparse2comm improves mixed-setting AP@0.5/AP@0.7 by +20.15/+11.79, +12.66/+11.07, and +15.36/+12.61 on the three datasets, respectively.

24. 【2610.08570】Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26

链接:https://arxiv.org/abs/2610.08570

作者:Truong Viet Vu,Nguyen Chi Hai,Nguyen Phuc Nguyen,Ngo Hoang Tu,Vo Nguyen Quoc Bao,Nguyen Thai Anh

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:suppress imaging artifacts, enhance lesion visibility, automated dermoscopic analysis, Handcrafted preprocessing, widely employed

备注: 6 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheless, its actual contribution to modern real-time models remains unclear, particularly when evaluation protocols do not adequately control correlations among images of the same lesion. This study presents a leakage-controlled, lesion-disjoint evaluation of dermoscopic preprocessing and augmentation for joint multi-class lesion classification and instance segmentation using a fixed nano-scale YOLO26 segmentation model (YOLO26n-seg). From HAM10000 (10,015 images), quality control yields 10,013 valid image-mask pairs from 7,468 unique lesions, partitioned into mutually exclusive sets by lesion identity. With the architecture, resolution, training budget, and evaluation protocol held fixed, we compare minimally processed images plus online augmentation against offline class balancing, DullRazor-CLAHE preprocessing, and raw-processed hybrid views, over three random seeds. On the lesion-disjoint test set, the raw baseline achieves a mask mAP$_{50:95}$ of $0.5636 \pm 0.0234$, a Dice score of $0.9356 \pm 0.0024$, and a macro-F1 score of $0.6917 \pm 0.0202$. Offline augmentation does not improve the mean performance, while the combined and hybrid strategies reduce both class-aware segmentation and classification accuracy. At only 2.69 million parameters, the model runs at approximately 50 frames per second. Under a leakage-controlled, lesion-disjoint protocol with all non-input factors held fixed, minimally processed dermoscopic images combined with standard online augmentation deliver a better accuracy-efficiency trade-off than increasingly complex deterministic preprocessing, which yields no consistent joint benefit across three seeds on HAM10000.

25. 【2610.08560】Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

链接:https://arxiv.org/abs/2610.08560

作者:Dan Ben-Ami,Kobi Cohen,Chaim Baskin

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Streaming video-language models, evidence needed, Streaming video-language, video-language models, current question

备注:

点击查看摘要

Abstract:Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.

26. 【2610.08539】RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models

链接:https://arxiv.org/abs/2610.08539

作者:Dongchen Si,Di Wang,Mingzhen Xu,Jing Zhang,Bo Du,Liangpei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Earth observation, Remote sensing scene, sensing scene classification, geospatial analysis, Remote sensing

备注:

点击查看摘要

Abstract:Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradigms: task-specific visual classification, vision-language similarity matching, and autoregressive multimodal generation. However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories. To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification. Unlike conventional MLLMs that formulate classification as autoregressive text generation, RSJEV reformulates scene classification as a candidate-conditioned multimodal discriminative decision process, where visual representations, task instructions, and candidate category semantics are jointly modeled. Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions. Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods. Moreover, RSJEV significantly reduces inference costs and achieves a better accuracy-efficiency trade-off with only a compact 0.8B-parameter model. These results demonstrate the effectiveness of state-conditioned multimodal decision making for efficient remote sensing image understanding. The code will be available at this https URL.

27. 【2610.08533】Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations

链接:https://arxiv.org/abs/2610.08533

作者:Yongsheng Luo,Wengan He,Yu Li,Rouying Wu,Wei Lv

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD)

关键词:Gram determinants provide, model higher-order consistency, alignment scores based, Geometric alignment scores, based on Gram

备注: Submitted to IEEE Transactions on Multimedia (TMM). 12 pages, 6 figures, 3 tables

点击查看摘要

Abstract:Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.

28. 【2610.08528】MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis

链接:https://arxiv.org/abs/2610.08528

作者:Asim Khan,Samee Ullah Khan,Dwarikanath Mahapatra

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:clinicians systematically evaluate, existing deep learning, mapping image features, image features directly, deep learning models

备注: 16 pages, 4 figures, conference

点击查看摘要

Abstract:Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic transparency and may compromise safe clinical deployment. We present MedCORE (Medical Criteria-Oriented Reasoning and Evidence), a structured diagnostic framework that operationalizes clinical reasoning within a vision-language architecture. For each input image, MedCORE decomposes the diagnostic process into clinically defined criteria, spatially localizes each criterion to diagnostically relevant image regions, encodes evidence through multi-scale representations that capture macro-structural and micro-textural pathological characteristics, and refines criterion representations using a Graph Attention Network that explicitly models inter-criteria dependencies. Criterion representations are further aligned with clinical text descriptors, reinforced through class-wise visual prototypes, and aggregated using uncertainty-calibrated weighting that proportionally discounts low-confidence diagnostic evidence. MedCORE is validated across three clinically heterogeneous imaging modalities, including dermoscopic lesion classification on ISIC 2018, breast ultrasound lesion characterization on BUSI, and diabetic retinopathy grading on IDRiD. Quantitatively, MedCORE achieves 89.2% accuracy, 85.7% macro-F1, and 96.4% AUC on ISIC 2018; 96.1% accuracy, 95.2% macro-F1, and 98.4% AUC on BUSI; and 84.3% accuracy, 80.2% macro-F1, and 92.8% AUC on IDRiD. These results demonstrate consistent improvements over strong CNN, transformer, biomedical vision-language, concept-based, and prototype-based baselines.

29. 【2610.08526】WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses

链接:https://arxiv.org/abs/2610.08526

作者:Thinh D. Le,Son T. Nguyen,Duong Q. Nguyen,Dung D. Le,Ngo Anh Vien,H. Nguyen-Xuan

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:remains largely unexplored, realistic industrial environments, fine-grained natural-language target, warehouses remains largely, photorealistic UAV VLA

备注: 41 pages, 35 figures, 11 tables

点击查看摘要

Abstract:Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.

30. 【2610.08518】2D Spatial Reasoning with Adaptive Neural Cellular Automata

链接:https://arxiv.org/abs/2610.08518

作者:Martin Spitznagel,Janis Keuper

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:modern learning approaches, utilize geometric information, spatial reasoning tasks, Neural Cellular Automata, Adaptive Neural Cellular

备注:

点击查看摘要

Abstract:Many modern learning approaches are still struggling with spatial reasoning tasks, i.e. they lack the ability to utilize geometric information of perceived entities and their spatial relation to each other to solve problems. We introduce a novel Adaptive Neural Cellular Automata (aNCA) architecture which uses deformable convolutions to dynamically adapt the perceptive field and iteratively reason over 2D spatial relations on grid-like data structures (e.g. images). Empirical results on public benchmarks show state of the art comprehensible results with high generalization abilities for solving image based puzzles like Sudoku or finding the shortest path in a maze.

31. 【2610.08482】Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment

链接:https://arxiv.org/abs/2610.08482

作者:Maryam Baizhigitova,Andrew Seohwan Yu,Po-Hao Chen,Naveen Subhas,Sixu Chen,Xinxin Wang,Kunio Nakamura,Richard Lartey,Xiaojuan Li,Mingrui Yang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:three-dimensional medical imaging, MRI remains limited, knee MRI remains, MRI Osteoarthritis Knee, Osteoarthritis Knee Score

备注: 11 pages, 2 figures, 5 tables

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.

32. 【2610.08433】HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

链接:https://arxiv.org/abs/2610.08433

作者:Ricardo Pizarro,Roberto Valle,José M. Buenaposada,Luis M. Bergasa,Luis Baumela

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Modern action recognition, Modern action, massive collections, collections of web-crawled, NTU RGB

备注:

点击查看摘要

Abstract:Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the use of such data raises ethical concerns, as subjects' consent is typically not obtained. Recent high-quality synthetic video datasets generated from motion-capture data, such as BEDLAM2.0, offer a promising ethical alternative. In this work, we investigate self-supervised pretraining of video transformers on synthetic human-motion datasets. We first show that directly applying the standard VideoMAE masking strategy leads to substantially worse performance than pretraining on Kinetics. To address this limitation, we propose a human-centric masking scheme that leverages body keypoints and person bounding box regions. Our approach encourages the model to focus on the structure and dynamics of human motion during pretraining. Experiments on NTU RGB+D and Toyota-Smarthome demonstrate that our method significantly outperforms standard VideoMAE pretraining on synthetic data, closing 49% of the gap to Kinetics pretraining on NTU RGB+D cross-view-subject without using a single real frame during pretraining. To promote the use of ethical action recognition models, we will publicly release our pretrained models.

33. 【2610.08419】Deformable CT-US Registration via Anatomy-Aware Implicit Neural Representations

链接:https://arxiv.org/abs/2610.08419

作者:Agnieszka Lach,Magdalena Wysocki,Feng Li,Mohammad Farid Azampour,Benjamin D. Killeen,Felix Ginzinger,Mathias Braun,Philipp Steininger,Heinz Deutschmann,Nassir Navab

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:preoperative computed tomography, minimally invasive interventions, computed tomography, imaging would enhance, invasive interventions

备注: 10 pages, 3 figures. Accepted at the 7th International Workshop on Advances in Simplifying Medical UltraSound (ASMUS 2026), held with MICCAI 2026; to appear in Springer LNCS 17276 (MICCAI 2026 Workshops and Challenges). Open-access camera-ready: [this https URL](https://papers.miccai.org/miccai-2026-sat/ASMUS_047.html)

点击查看摘要

Abstract:Slice-to-volume registration between ultrasound (US) and preoperative computed tomography (CT) imaging would enhance many minimally invasive interventions, for example by locating soft tissue structures intra-operatively that are discernible in CT. While optical tracking enables initial rigid registration, contact from the probe induces soft tissue deformations that inhibit accurate alignment. In this work, we introduce a deformable CT-ultrasound registration framework that incorporates anatomical priors derived from CT to improve registration under deformation. Rigid registration is first established using a robot-assisted optical tracking system, after which a deformable transformation is estimated using a sinusoidal implicit neural representation (SIREN) optimized per frame. Tissue stiffness is approximated from CT-based HU values and used as spatially varying regularization, suppressing deformation in rigid structures such as bone while allowing more flexibility in soft tissue. Two additional constraints capture the physics of probe contact: a contact-zone displacement prior that drives the displacement field to compress tissue below the probe face, and a fan-geometry regularization term based on beam direction and convex transducer field of view. Model parameters are optimized with a normalized gradient field (NGF). The proposed approach improves alignment over rigid initialisation by 17% and outperforms classical deformable baselines while maintaining near-zero topological folding.

34. 【2610.08418】From the Drosophila Visual Connectome to General-Purpose Computer Vision

链接:https://arxiv.org/abs/2610.08418

作者:Zongyu Li,Akito Yamauchi,Huaizhi Liu,Vishwanatha Rao,Jia Guo, for theFrontotemporal Lobar Degeneration Neuroimaging Initiative, for theAlzheimer's Disease Neuroimaging Initiative

类目:Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Neurons and Cognition (q-bio.NC)

关键词:Biological connectomes encode, connectomes encode structured, encode structured solutions, provide reusable inductive, reusable inductive biases

备注: 27 pages, 17 figures, 7 tables

点击查看摘要

Abstract:Biological connectomes encode structured solutions to visual computation that may provide reusable inductive biases for artificial vision. We develop ConnectomeX around FlyVision, a trainable architecture that preserves parallel ON/OFF processing, recurrent computation and population-level graph interaction while scaling model capacity across tasks. FlyVision reached 99.34% accuracy on MNIST with 80,608 parameters and 78.03% on CIFAR-10 with 81,408 parameters. On ImageNet-1K, FlyVision Base and Large reached 60.79% and 66.25% top-1 accuracy with 1.8 and 3.7 million parameters, while a Large local-k7 model with a learned low-frequency branch reached 66.53%, compared with 69.25% for ResNet18 with 11.7 million parameters. On a 22-class skin-disease benchmark, FlyVision Large achieved 63.78% accuracy and 95.28% macro-AUROC with 2.99 million parameters. In four-class chest radiography, ImageNet-pretrained FlyVision Base and Large reached 92.60% and 92.76% accuracy with 1.33 and 2.97 million parameters, compared with 91.56% for ImageNet-pretrained ResNet18 with 11.18 million. BrainAGE extends FlyVision to volumetric T1-weighted MRI by applying a shared ImageNet-pretrained FlyVision Large encoder to 24 sagittal, coronal and axial slices per scan and combining slice-level age estimates by confidence-modulated Gaussian voting. On 433 held-out scans, three-axis fusion achieved a mean absolute error of 5.98 years and R^2 = 0.868. Across the 224x224 classification tasks, the best FlyVision configuration remained within three percentage points of ResNet18 on ImageNet-1K and skin-disease classification and exceeded it on chest radiography with substantially fewer parameters. These results show that a conserved connectome-informed computation can scale from compact recognition to large-scale natural and biomedical vision.

35. 【2610.08417】Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

链接:https://arxiv.org/abs/2610.08417

作者:Tianyi She,Jiawei Liu,Weifeng Liu,Hanqing Zhao,Weiming Zhang,Kejiang Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)

关键词:posing severe societal, severe societal risks, Recent advances, highly realistic videos, LipSync generation technology

备注: 24 pages, Accepted at ICML 2026

点击查看摘要

Abstract:Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97\% AUC in detection and 97.5\% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at this https URL.

36. 【2610.08414】Image Bitstream Fine-grained Understanding for Privacy-Friendly AIoT

链接:https://arxiv.org/abs/2610.08414

作者:Zhen Yu,Wenyang Liu,Kejun Wu,Chengwang Xiao,Renjie Qiao,Chengtao Cai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image byte sequences, perform fine-grained classification, byte sequences, Fine-grained Understanding, Fine-grained

备注:

点击查看摘要

Abstract:Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences. In contrast to conventional pixel-domain visual understanding, IBFU conducts semantic analysis without fully decoding images into the pixel domain. Since pixel-level visual content is not explicitly reconstructed during inference, this paradigm reduces visual exposure within the processing pipeline and suits privacy-friendly Artificial Intelligence of Things (AIoT) applications. In this paper, we propose Bitstream Fine-grained Generator (BFG), a novel foundation model tailored for IBFU. BFG consists of two main components: a Bitstream Semantic Encoder (BSeE) and a Fine-grained Semantic Generator (FSeG). BSeE directly models semantic representations from encoded image bitstreams without explicit pixel reconstruction, while FSeG transforms the extracted bitstream semantics into detailed natural-language descriptions through autoregressive generation. To train BFG and comprehensively evaluate IBFU in practical AIoT scenarios, where image bitstreams may suffer corruption during transmission and storage, we construct a large-scale Corrupted-bitstream Fine-grained Understanding dataset (CFU-D), containing both intact bitstreams and corrupted variants across multiple corruption types and severity levels. Experiments show that BFG maintains stable fine-grained caption generation under bitstream corruption. For example, the performance only has slight change from 0.6339 to 0.6077 in terms of average CIDEr score on Stanford Dogs Caption dataset, while vision-language models, such as Qwen-VL-Chat, BLIP-2, GLM, Gemini, and GPT suffer severe performance decrease. This paper provides a practical paradigm for privacy-friendly fine-grained understanding in AIoT.

37. 【2610.08410】Decoy and disclosure radii of invariant shape descriptors

链接:https://arxiv.org/abs/2610.08410

作者:Tanush Shaska,Lubjana Beshaj

类目:Computer Vision and Pattern Recognition (cs.CV); Earth and Planetary Astrophysics (astro-ph.EP); Cryptography and Security (cs.CR); Algebraic Geometry (math.AG)

关键词:compares rotation-invariant descriptors, recognizer that compares, compares rotation-invariant, rotation-invariant descriptors, radius

备注:

点击查看摘要

Abstract:A recognizer that compares rotation-invariant descriptors sees a surface only up to the fiber of the descriptor. We measure this fiber by its radius in the orbit distance from the enrolled surface. A large radius admits decoys, that is, distant shapes that pass the matcher. A small radius discloses the enrolled shape to anyone who captures the stored value. For star-shaped surfaces truncated to spherical harmonics of degree at most $L$, with $n$ coefficients, a descriptor of generic rank $r$ has generic fibers of dimension $n-3-r$ modulo rotations. The standard pool of band powers, even bispectra, and three invariants of the degree-three band therefore admits decoy families of dimension $5$, $13$, $20$ at $L=4,6,8$. Its rank first reaches $n-3$ at $L=16$, and a mirror decoy remains at every $L$. The odd bispectra remove the mirror decoy generically for $L \geq 4$. Yet at fixed mean radius the same pool determines the enclosed volume exactly, and it does not determine whether a surface meets a clearance requirement. We certify two cases by exact and interval arithmetic. At $L=6$ a decoy matches all $32$ invariants to relative precision $2 \cdot 10^{-18}$ at orbit distance at least $0.87$ times the norm of the enrolled tuple. For the radar shape model of asteroid (101955) Bennu, the pool recovers the modeled volume, misses the handedness, and leaves the keep-out radius uncertain by more than $7 \, \mathrm{m}$.

38. 【2610.08401】GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

链接:https://arxiv.org/abs/2610.08401

作者:Seulgi Kim,Zhixiong Zhang,Xinwei Zhang,Jie Ling,Ronn Shaw

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:shown outstanding performance, recent vision-language models, diverse applications, recent vision-language, shown outstanding

备注: Under Review

点击查看摘要

Abstract:While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63\%.

39. 【2610.08398】UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation

链接:https://arxiv.org/abs/2610.08398

作者:Taojie Zhu,Jing Jin,Yuan Xia,Chenyang Ding,Qunshan He,Wanke Xia,Tao Sun,Yan Chen,Jian Wang,Jinjie Gu,Tao Feng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multiple teachers combines, teachers combines expertise, single student, hinder this integration, On-policy distillation

备注:

点击查看摘要

Abstract:On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval this http URL results support projecting optimizer updates to reduce interference between domains.

40. 【2610.08379】UniCounting: Instance-Aware Proposal Consolidation for Image-Query-Free Multi-Category Counting

链接:https://arxiv.org/abs/2610.08379

作者:Jinshi Liu,Pan Liu,Lei He,Weichao Luo,Rui Qian

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual counting, returning a single, text query, image-specific exemplar, commonly formulated

备注:

点击查看摘要

Abstract:Visual counting is commonly formulated as counting a single specified target, with a model receiving an image-specific exemplar, text query, or target category and returning a single count. We instead study fixed-vocabulary image-query-free multi-category counting. A global vocabulary is fixed for each run, and, given only an RGB image, the model predicts a complete category--count vector without being told which categories appear. We present UniCounting, which casts counting as instance-aware structural inference over an over-complete proposal set. Generic segmenters produce duplicate masks, partial views, and proposals from neighboring instances; semantic scores can name them but cannot determine which denote the same object. Frozen SAM~2.1 generates masks, while frozen DINOv2 and OpenCLIP provide relation and category features. A 3,267-parameter category-shared relation head predicts same-instance affinities from instance-mask-derived supervision. Sparse graph construction, representative selection, labeling, and background-margin admission then convert each admitted component into one count with replayable group evidence. Only the relation head is trained, without count or density-map targets. On COCO clean500, UniCounting obtains lower point-estimate vector $\ell_1$ error and absent-class false mass than calibrated OWLv2-All80, with comparable micro presence F1. Under a matched decoder, the learned relation reduces both errors relative to mask containment, mask IoU, CLIP, and DINO, while revealing a fragmentation--merge trade-off. We also report transfer diagnostics on OmniCount-sub, FSC-147, and CARPK.

41. 【2610.08365】A Stevens's Power Law Check-up of GPT-5.5's Image-Based Visualization Reading

链接:https://arxiv.org/abs/2610.08365

作者:Kaichun Yang,Jian Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:adapt Stevens power, Stevens power law, built-in perceptual mechanisms, adapt Stevens, Stevens power

备注: 9 pages, 6 figures, including supplementary material. Accepted by the VISxGenAI workshop at IEEE VIS 2026

点击查看摘要

Abstract:We adapt Stevens's power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, models see no legend. A model first views a reference visual representation and estimates its magnitude, then estimates the magnitude of each subsequent image of the same representation relative to that reference. Our evaluation of twelve visual variables makes how algorithmic models read visual encodings measurable, comparable with human perception, and more interpretable to humans.

42. 【2610.08358】st-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration

链接:https://arxiv.org/abs/2610.08358

作者:Hyeongheon Cha,Young D. Kwon,Sung-Ju Lee

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:fitting vision transformers, Post-training quantization, vision transformers, standard route, route to fitting

备注: 44 pages, 6 figures. Code at [this https URL](https://github.com/chahh9808/QuAR)

点击查看摘要

Abstract:Post-training quantization is a standard route to fitting vision transformers (ViTs) into edge compute and memory budgets, yet quantized models become especially brittle under distribution shift. Test-time adaptation (TTA) addresses such shifts without labels, but most existing approaches are poorly aligned with the constraints of quantized inference. Prevailing TTA methods recover accuracy through backpropagation, while backprop-free methods often still incur overhead from extra forward passes or parameter updates, and lightweight feature- or logit-level methods recover only part of the loss. Across these approaches, a quantization-specific failure mode that amplifies the drop is not directly targeted: under shift, activations occupy frozen quantizers' calibrated ranges differently, distorting their code distribution. We propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters. QuAR recalibrates activations at the input to a frozen quantizer, mapping the test stream's running per-channel statistics back toward the source calibration. On ImageNet-C with ViT-B, QuAR achieves the highest mean accuracy among state-of-the-art backprop-free TTA methods at 3-, 4-, 6- and 8-bit weight/activation precision, outperforming the strongest baseline by 2.28 points at 8 bits and 4.00 at 3 bits, with 46% lower latency and a memory overhead of only 0.17 MB (0.01% of peak inference memory). Analysis and diagnostics trace the gain to a reduced per-channel mismatch at these quantizers, which restores the code distribution the baselines leave unchanged or distort further. A single fixed configuration remains ahead across continual streams, non-i.i.d. label shift, seven out-of-distribution suites, and three other backbones.

43. 【2610.08346】PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

链接:https://arxiv.org/abs/2610.08346

作者:Beibei Lin,Tingting Chen,Xin Zhang,Wenhao Zhao,Dongjun Li,Zifeng Yuan

类目:Computer Vision and Pattern Recognition (cs.CV); Optics (physics.optics)

关键词:requires specialized hardware, typically requires specialized, normalized Stokes components, Polarization imaging, intensity imaging

备注: 22 pages, 17 figures, 8 tables. Accepted to NeurIPS 2026

点击查看摘要

Abstract:Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image $s_0$ (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.

44. 【2610.08341】DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

链接:https://arxiv.org/abs/2610.08341

作者:Shuo Yang,Changbai Li,Linlin Yang,Huobin Tan,Rongyu Chen,Tongfei Chen,Tian Wang,Sheng Xu,Baochang Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Large Language Models, Multimodal Large Language, effectively cut computational, Language Models, cut computational overhead

备注:

点击查看摘要

Abstract:Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.

45. 【2610.08339】Digital Twin-Driven Real2Sim2Real: Simulator-Conditioned Generation via Paired Driving-Scene Reconstruction

链接:https://arxiv.org/abs/2610.08339

作者:Hojun Lim,Hyeongseok Jeon,Donghyun Kim,Soonyoung Jung,Heecheol Yoo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large annotated datasets, region typically requires, typically requires data, autonomous driving relies, driving relies heavily

备注: 8 pages, 6 figures

点击查看摘要

Abstract:Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.

46. 【2610.08331】ransferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving

链接:https://arxiv.org/abs/2610.08331

作者:Heyam Bin Jahlan Areej Alhothali Abeer Alhothali

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Vision Language Models, critical safety vulnerabilities, integration of Vision, sensitive systems introduces, systems introduces critical

备注:

点击查看摘要

Abstract:The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in order to ensure that adversarial perturbation target the most semantically significant frames. this http URL spatial attack, we craft effective perturbation and preserve high similarity. Then the perturbed video generated fed into STCA stage that disrupt cross-frame temporal coherence using motion guided mask. Our method operate under black box threat model against victim target VLMs, relying solely on transferability from white-box surrogate this http URL conduct our experiments on the BDD100K and nuScenes autonomous driving datasets across three VLM models: Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. Experimental results demonstrate spatial attack achieves an ASR with high SSIM. Our finding reveal that existing video language model, remain highly susceptible to adversarial attack in autonomous driving scenarios, underscoring the urgent need for robust defense for VLM models.

47. 【2610.08315】Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance

链接:https://arxiv.org/abs/2610.08315

作者:Khac Duc Giang Nguyen,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Counter-UAV systems based, operational datasets evolve, remains insufficiently characterized, thermal infrared detection, Counter-UAV systems

备注:

点击查看摘要

Abstract:Counter-UAV systems based on thermal infrared detection must stay accurate as operational datasets evolve, yet sequential fine-tuning causes catastrophic forgetting of prior tasks, a problem that remains insufficiently characterized in this domain. This continual-learning study measures the stability-plasticity trade-off in YOLOMG, a YOLOv5-based detector run as a single thermal-infrared stream with the motion channel disabled, trained sequentially across three anti-UAV benchmarks of rising scale difficulty: Anti-UAV-RGBT, Anti-UAV410, and CST Anti-UAV. Naive fine-tuning on CST yields a Forgetting Measure of -0.605 against the Stage 1 ceiling, corresponding to a 90% capability loss, with -0.572 occurring in Stage 3 alone. In contrast, knowledge distillation from a frozen teacher is associated with FM = -0.033 +/- 0.004 across three seeds, corresponding to 95% retention. Because no Stage 2 no-KD control is included, this result establishes retention under KD training rather than a causal KD effect. Per-stratum analysis shows large-target detection collapsing to near zero within the first epoch, despite an inter-stage cosine similarity of 0.987 over the gradient-updated weights, pointing to scale-conditioned gradient imbalance, rather than weight drift, as a candidate mechanism. Scale-Stratified Herding (SSH), a 300-exemplar buffer balanced across four UAV size strata, roughly halves the forgetting (FM = -0.605 to -0.311) and keeps large-target detection non-zero. An ablation attributes the gain primarily to scale stratification rather than herding: random-stratified replay performs at least as well (FM = -0.221 versus -0.311 for SSH). These replay results are single-seed and should therefore be treated as preliminary.

48. 【2610.08286】Event Detection in Table Tennis Videos using 2D Keypoints

链接:https://arxiv.org/abs/2610.08286

作者:Rainer Lienhart,Daniel Kienzle,Shin'ichi Satoh,Anastasiia Bilinska

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:frame-accurate event detection, paper addresses, addresses the challenge, table tennis videos, frame-accurate event

备注: Accepted at the 9th International ACM Workshop on Multimedia Content Analysis in Sports

点击查看摘要

Abstract:This paper addresses the challenge of automatic, frame-accurate event detection in table tennis videos. Current methods for estimating 3d ball trajectories and ball spin typically require that key events, such as ball-racket contacts, have already been identified in advance. This requirement makes it difficult to apply these methods to longer, unedited video recordings. To overcome this limitation, we propose EventNet, a two-stage pipeline to detect key events: (1) 2d keypoints are extracted of the upper-body poses for both players, table corners and ball center. A small keypoint transformer combines them into a compact representation that is robust to changes in viewpoint, lighting, and background clutter. (2) The temporal sequences of these frame-based representations are processed by a transformer encoder that predicts two time-to-event values for each frame, indicating how close the current frame is to the next and previous ball-racket contact. One novelty is a new, temporal cosine-like target signal. Furthermore, we introduce viewpoint augmentation via 3D reprojection and frame-rate augmentation to improve robustness and generalization. Our extensive ablation study gives deeper insights into the importance of various architectural and training aspects. Experimental results show that the proposed approach achieves an F1 score of 91.16% and a mean frame deviation between ground truth and predicted frame of 0.42 on the Latte-MV dataset and 73.08% / 1.16 on the challenging TTHQ dataset. Overall, our work demonstrates that 2d keypoint-based temporal modeling with our EventNet architecture is a promising and practical approach for automatic event detection in table tennis videos.

49. 【2610.08279】Whose Face Is It Anyway? A Multi-Model Audit of Facial Affect Recognition on Children, and Why the Gap Is the Head, Not the Features

链接:https://arxiv.org/abs/2610.08279

作者:Tobias Hallmen,Robin-Nico Kampa,Elisabeth André

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Facial affect models, Facial affect, developmental research, increasingly applied, affect models

备注: Preprint. 10 pages, 5 figures

点击查看摘要

Abstract:Facial affect models are trained almost entirely on adults, yet are increasingly applied to children in education, health, and developmental research. We present a controlled, multi-model audit of five AffectNet-pretrained expression models (EmoNet, EmotiEffLib, DDAMFN++, OpenFace 3.0, LibreFace) on children, across four child image datasets, the AffectNet-8 validation set, and two spontaneous child video datasets, through one shared harness. Three findings emerge. First, the child gap is model-agnostic: every architecture degrades from posed to naturalistic faces and shares the fear$\rightarrow$surprise confusion. Second, it is concentrated and corroborated across all five models: open-mouth faces (read as surprise, correlating with the AU26 jaw drop) and South-Asian children degrade systematically, with a smaller averted-gaze penalty, while closed-mouth faces, White and Black children, and direct gaze do not; the bias tracks expression morphology and specific populations, not skin tone. Third, the gap is diagnosable: a linear probe on frozen features reaches 0.75-0.91 on unseen children versus 0.48-0.66 zero-shot, so it lies largely in the classifier head, not the representation, whereas dimensional valence/arousal regression degrades sharply under domain shift. Building on this, recalibrating only the head on a little target data recovers $+0.13$ to $+0.28$ on the two largest child sets across all five models at negligible adult cost, though the gain is in-distribution and does not transfer across child collections. We will release the harness, per-sample predictions, and analysis code; the child face data stays license-locked and is never redistributed.

50. 【2610.08230】SRN-RTVD: Real-Time Video Deblurring System

链接:https://arxiv.org/abs/2610.08230

作者:Nikita Alutis,Danila Evsyukov,Egor Chistov,Mikhail Voronin,Evgeney Bogatyrev,Dmitriy Vatolin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:lowers perceptual quality, downstream vision tasks, harms downstream vision, video capture moves, capture moves

备注:

点击查看摘要

Abstract:As video capture moves to handheld and edge devices, motion blur from camera shake has become a pervasive degradation that lowers perceptual quality and harms downstream vision tasks. The strongest deblurring networks recover impressive detail, yet they remain computationally heavy and overwhelmingly complex, so their quality comes at a cost that consumer hardware cannot pay in real time. This gap between restoration quality and on-device speed is exactly what makes real-time deblurring difficult. We developed and implemented TSRN-RTVD, an efficient video deblurring system that explicitly reconstructs the underlying camera trajectory during exposure and uses the recovered motion to guide restoration. This approach turns the physical cause of blur into a signal that drives sharpening. Our system runs on a single consumer GPU and restores the video at 30 FPS while reaching 30.08 dB PSNR on the GoPro dataset. We demonstrate TSRN-RTVD on consumer devices with interactive side-by-side visualization of the blurry input and the deblurred output, live throughput, and an on-screen view of the recovered camera trajectory. Demo video is available at this https URL.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.08230 [cs.CV]

(or
arXiv:2610.08230v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.08230

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
51. 【2610.08213】RACE-FPP: A Robust AI-assisted Characterisation Enhancement for Fringe Projection Profilometry

链接:https://arxiv.org/abs/2610.08213

作者:Osman Ali(1),Xiangjun Kong(1),Tibebe Yalew(1),Waiel Elmadih(2),Samanta Piano(1) ((1) Manufacturing Metrology Team, University of Nottingham, Nottingham, United Kingdom, (2) Taraz Metrology Ltd., Nottingham, United Kingdom)

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV); Optics (physics.optics)

关键词:Fringe Projection Profilometry, Fringe Projection, Projection Profilometry, requires precise system, robust checkerboard feature

备注: 19 pages, 9 figures, 7 tables

点击查看摘要

Abstract:Fringe Projection Profilometry (FPP) requires precise system characterisation to achieve reliable three-dimensional (3D) reconstructions; however, characterisation accuracy strongly depends on robust checkerboard feature localisation, which can deteriorate under challenging imaging conditions such as lens blur and characterisation target orientations. Existing deep learning-based corner detectors are typically assessed using detection metrics and camera reprojection error alone, without considering their wider impact on projector characterisation, camera-projector stereo characterisation consistency, or overall measurement accuracy. In this work, we introduce a complete FPP characterisation pipeline that incorporates deep learning-based corner detection into the standard camera characterisation workflow. We also characterise the projector by sampling phase values at the centres of the white squares in the characterisation target. Rather than treating corner detection as an isolated task, the proposed framework explicitly analyses how localisation errors propagate throughout the entire FPP characterisation chain. Performance is evaluated using detection metrics (e.g., precision and recall), camera and projector reprojection errors, and the camera and projector stereo characterisation. Across a mixed dataset of clean and degraded images, the camera reprojection error is reduced from 1.237 pixels to 0.259 pixels, while the projector reprojection error is reduced by roughly 50%. Dimensional evaluation of reconstructed artefacts shows improved geometric accuracy compared with those resulting from the conventional pipeline. Overall, the findings indicate increased robustness of system-level characterisation under challenging imaging conditions, thereby enabling more reliable industrial FPP measurements.

52. 【2610.08192】MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos

链接:https://arxiv.org/abs/2610.08192

作者:Souptik Sen,Zahra Ahmadi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:exploiting complementary cues, complementary cues, exploiting complementary, typically assume, improve egocentric action

备注:

点击查看摘要

Abstract:Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbf{MacJEPA}, a missing-modality-robust \textbf{Ma}sked-\textbf{c}ontext query \textbf{JEPA} that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.

53. 【2610.08188】PIE-PS: Photometric Stereo from Physical Irradiance Event Streams

链接:https://arxiv.org/abs/2610.08188

作者:Xiangze Meng,Guangyu Li,Jing Li,Di Mei,Songchen Ma,Mingkun Xu,Rui Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cameras record asynchronous, high dynamic range, Event cameras record, record asynchronous, dynamic range

备注: 9 pages, 7 figures. Accepted to SIGGRAPH Asia 2026 Conference Papers

点击查看摘要

Abstract:Event cameras record asynchronous log-image-irradiance changes with microsecond latency and high dynamic range. These properties are useful for photometric stereo under moving illumination, but raw events are sparse and depend on an unknown contrast threshold. We start from the event trigger model and derive a physical relation between adjacent events, light motion, and surface normals. This relation gives a direct physics-only solver, but the solver needs the threshold, enough events at each pixel, and independent per-pixel optimization. To address these limits, we introduce PIE-PS, a learning-based framework for dense surface normal reconstruction from raw event streams and known lighting. We form Physical Irradiance Events (PIEs) by pairing two adjacent events at the same pixel with their corresponding light directions. Each PIE provides a Physical Irradiance Event Feature (PIEF), defined as the signed event rate. PIEF does not require the unknown contrast threshold. To share spatial and temporal context across nearby PIEs, we introduce PIE-GNN, which treats each PIE as a graph node and encodes it with its light-pair geometry. Since the reliability of PIE observations can vary with local appearance, illumination geometry, and sensor noise, Reliability-Grading Attention (RGA) predicts reliability weights to down-weight unreliable PIEs. Pixel aggregation then produces dense normals. Experiments on synthetic and real data show that PIE-PS outperforms prior event-based photometric stereo methods and the direct solver baseline.

54. 【2610.08179】View Matters: Keyframe-Guided Text-Driven 3D Gaussian Editing

链接:https://arxiv.org/abs/2610.08179

作者:Kaizhe Zhang,Yijie Zhou,Weizhan Zhang,Xuanyu Wang,Feng Lei,Sha Gong

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:Gaussian editing commonly, viewpoints provide supervision, Gaussian editing, substantially different quality, supervision of substantially

备注: 14 pages, 11 figures, including appendices

点击查看摘要

Abstract:Text-driven 3D Gaussian editing commonly does not distinguish the editing reliability of rendered views, although different viewpoints provide supervision of substantially different quality. Views that clearly show the scene and match the edit instruction provide reliable guidance, while less informative views may weaken the edit when all views are treated equally. We present View Matters, a view-importance-aware framework that conducts editing around reliable keyframes. Keyframe Importance Estimation (KIE) identifies reliable views using geometric visibility, semantic distinctiveness, and edit relevance. Keyframe-Guided Editing (KGE) then propagates their editing signals asymmetrically to non-keyframes without noisy reverse influence, while Importance-Aware Optimization (IAO) preserves this reliability preference during 3DGS optimization. Across 23 scene-prompt pairs, View Matters achieves the highest average CLIP text-image similarity of 0.2822 and directional similarity of 0.2564 among the evaluated methods, with a four-minute editing time. Additional adjacent-view analysis indicates that the fidelity-oriented editing process maintains cross-view coherence.

55. 【2610.08170】Visual Orchestration Tax in Agentic VLM Pipelines: Auditing and Certifying Visual Evidence Reuse

链接:https://arxiv.org/abs/2610.08170

作者:Lingteng Zeng

类目:Multiagent Systems (cs.MA); Computer Vision and Pattern Recognition (cs.CV)

关键词:Agentic VLM pipelines, VLM pipelines increasingly, pipelines increasingly pass, VLM API boundary, Agentic VLM

备注: 9 pages, 2 figures, 5 tables

点击查看摘要

Abstract:Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools. This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary. We call this phenomenon visual orchestration tax and develop a measurement-to-certification framework for visual evidence reuse in agentic VLM pipelines. The audit side defines $\mathrm{M1}_{\mathrm{trace}}$ to count raw visual-evidence touches and M2 to measure structural touch redundancy, with query-level distributions, bootstrap confidence intervals, and paired quality tests. Across SeeingEye and MAMMQA on chart, document, general-VQA, and multi-modal-QA tasks, audits reveal 66.8-75.6% visual-evidence touch redundancy, and every audited query exceeds the predefined gate. The certification side introduces SharedVisCache, a contract-aware evidence reuse hook keyed by image content, preprocessing fingerprint, and encoder assumptions. On SeeingEye, contract validation certifies 75.0-75.5% repeated touches as reusable while preserving 350/350 output strings and $\Delta\mathrm{M5}{=}0$. At the physical layer, certified hits reduce $F_{\mathrm{vision}}$ from 800 to 200 in ChartQA-200 trace replay and from 200 to 50 inside live SeeingEye translator-stage physical integration, preserving 800/800 replay strings and 200/200 integrated call outputs. The results position visual reuse as a measurable, behavior-preserving property of agent orchestration and define an agent-layer contract that makes backend prefix or token reuse semantically interpretable.

56. 【2610.08162】he Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

链接:https://arxiv.org/abs/2610.08162

作者:Tobias Hallmen,Fabian Deuser,Robin-Nico Kampa,Norbert Oswald,Elisabeth André

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:Fine-grained emotion recognition, recognition supports therapy, supports therapy tools, emotion recognition supports, Fine-grained emotion

备注: Preprint. 19 pages, 6 figures

点击查看摘要

Abstract:Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at $\kappa_w = 0.468$ on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ($\kappa_w=0.268$-$0.486$). Under verification all eleven clear it, each of them significantly better at $\kappa_w=0.507$-$0.586$. Three also significantly beat EIF sitting at $\kappa_w = 0.551$ (Small; $0.534$ Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $\kappa_w=0.254$-$0.423$. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.

57. 【2610.08137】Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering

链接:https://arxiv.org/abs/2610.08137

作者:Zheng Gao,Xiaoyu Li,Zhicheng Bao,Yang Song,Jiaojiao Jiang

类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)

关键词:video generation models, generation models, writing code, code and graphics, graphics descriptions

备注: 46 pages, 6 figures, 4 tables. Conceptual research agenda; no experiments. Video: [this https URL](https://youtu.be/14SMl0d_e48) . Project page: [this https URL](https://zhenggao-30.github.io/Rethinking-Visual-Provenance/)

点击查看摘要

Abstract:AI systems create images and videos with image/video generation models or by writing code and graphics descriptions that are then rendered. These routes can produce similar visible artifacts but expose different representations, intervention points, and provenance evidence. We develop a production-centered framework that compares detection and watermarking across both routes. An explicit verification specification distinguishes passive inference, message recovery, and authenticated provenance. We organize image, video, source-code, and rendering-aware watermarks by production stage. We examine the different requirements of generated images and video, plots and SVG, programmable video, and agent-composed workflows. Documented Claude, OpenAI, and rendering-tool interfaces connect the framework to concrete systems. We pose ten scoped research questions on identifiability, observability, fair comparison across stages, recoverable payload, reconstruction, synchronization, composition, hybrid local contribution, and private production-event authentication. The result is a conceptual research agenda grounded in published methods, inspected interfaces, and elementary boundary examples. It reports no experiments and claims no new theorems; its appendix results are elementary calculations, and documentation and source inspection establish interfaces, not empirical robustness.

58. 【2610.08133】VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

链接:https://arxiv.org/abs/2610.08133

作者:Owen Du,Yang Yue,Jie Zhang,Jiaqi Pi,Chi Bene Chen,Gao Huang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:limiting real-time deployment, incur high computational, high computational costs, processing long token, base VLA model

备注:

点击查看摘要

Abstract:Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at this https URL.

59. 【2610.08131】Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors

链接:https://arxiv.org/abs/2610.08131

作者:Mina Abbaszadeh,Matilda Karabina Moore,Raem Haq,Martha Lewis,Mehrnoosh Sadrzadeh

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Achieving compositional concept, compositional concept generalization, recombining learned primitives, Compositional Distributional Semantics, Achieving compositional

备注:

点击查看摘要

Abstract:Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.

60. 【2610.08126】Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery

链接:https://arxiv.org/abs/2610.08126

作者:Mayank Sah,Jimson Mathew

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:retail industry, Identification, Object, object detection models, challenging problems

备注: 10 Pages, 7 Figures, 5 Tables

点击查看摘要

Abstract:Product Identification has sprung up to become one of the most challenging problems in the automation of the retail industry. With the new industry 5.0 standards, automated inventory management, and catalog creation tasks are vitally important. Object identification models have emerged as a viable answer with their unprecedented identification and localization accuracy. However, the close-knit rack design of supermarkets generates the problem of angle variation in capturing images. The angle-variant densely packed images(a single image contains many objects) become overwhelming for these models alone. In this paper, we try to supplement object detection models with traditional Hough transform (HT) and homogeneous estimation concepts. We study the effect of rectified images using homography estimation and hough transform and their limitations on the problem of grocery identification. We make a case for creating a new dataset to test the effects of such rectification and produce analytical results on different scenarios of angle variation and object densities per image. Extensive experiments on different object detection models suggest that image rectification of angled images improves the detection accuracy of grocery products in images. The results also highlight the limitation of rectification on the angle of image capture and the object density of the image.

61. 【2610.08109】Beyond Training from Scratch: Foundation Models for Data-Efficient and Generalizable Cardiac MRI Reconstruction

链接:https://arxiv.org/abs/2610.08109

作者:Anam Hashmi,Mayug Maniparambil,Julia Dietlmeier,Kathleen M. Curran,Noel E. O'Connor

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:preserving diagnostic fidelity, enabling faster scans, magnetic resonance imaging, recover high-quality images, Cardiac magnetic resonance

备注: Accepted at ECCVW 2026

点击查看摘要

Abstract:Cardiac magnetic resonance imaging reconstruction aims to recover high-quality images from undersampled acquisitions, enabling faster scans while preserving diagnostic fidelity. Recent reconstruction methods are typically trained from scratch and often require large amounts of task-specific data, limiting their robustness under data scarcity and distribution shifts. In this work, we investigate whether pretrained vision foundation models can serve as effective priors for accelerated cardiac MRI reconstruction. We propose a reconstruction framework that integrates frozen and parameter-efficiently adapted visual encoders, including CLIP, BiomedCLIP, and DINOv2, within a transformer-based reconstruction architecture. Extensive experiments on the CMRxRecon2023 and CMRxRecon2024 benchmarks demonstrate that pretrained representations consistently outperform a transformer trained from scratch across multiple acceleration factors. We further evaluate performance under limited supervision and cross-dataset transfer, showing that foundation models provide superior data efficiency and generalization. While frozen representations are particularly effective in extreme low-data regimes, Low-Rank Adaptation (LoRA) yields additional gains when moderate amounts of training data are available. Among the evaluated backbones, DINOv2 achieves the strongest overall performance. These findings highlight the potential of vision foundation models as robust and transferable priors for cardiac MRI reconstruction.

62. 【2610.08086】Multi-Dataset Diagnostic Utility of Clinical Visual Concepts in AI Systems for Dermatology

链接:https://arxiv.org/abs/2610.08086

作者:Linda Wermelinger,Simone Lionetti,Fabian Gröger,Nipun Ranasekara,Philippe Gottfrois,Ludovic Amruthalingam,Labelling Consortium,Marc Pouly,Alexander A. Navarini

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:digital dermatology relies, dermatology relies heavily, systems in digital, digital dermatology, dermatology relies

备注: Accepted at the MICCAI ISIC Workshop 2026. 11 pages, 3 figures, 3 tables. Code and dataset: [this https URL](https://github.com/Digital-Dermatology/SkinLex)

点击查看摘要

Abstract:The clinical integration of AI systems in digital dermatology relies heavily on human trust. Clinically interpretable visual concepts can act as intermediate representations enhancing trust and reliability. However, research in this domain is currently limited by scattered, heterogeneous dataset annotations. In this work, we introduce SkinLex, a harmonized dataset of 48 clinical morphological attributes across four public datasets (SkinCon, DermaCon-IN, MM-Skin, and PASSION) for a total of 20,411 records. Supervised nine-partition classification of skin conditions shows that limiting features to specific visual groups, like shapes or colors alone, reduces diagnostic accuracy. Bootstrapped backward elimination reveals that the set of 48 visual concepts has some degree of redundancy for algorithmic nine-partition diagnosis on the examined dataset. This demonstrates that coarse diagnosis on the selected dataset requires a relatively small but varied combination of clinical concepts, and motivates further research to improve concept taxonomy. Results can be translated into clinical benefits by reducing inputs for concept-based models, improving efficiency for annotation and modeling, and further enhancing interpretability. Code and prompt templates are available at this https URL.

63. 【2610.08075】Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields

链接:https://arxiv.org/abs/2610.08075

作者:Rudolf L.M. van Herten,Soufiane Ben Haddou,Rachit Saluja,Johannes C. Paetzold

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:represent signals continuously, fields represent signals, Conditional neural fields, conditional latent representations, signals continuously

备注:

点击查看摘要

Abstract:Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We formalize this connection by interpreting latent optimization as an optimization encoder, unifying the roles of second-order differentiation, latent parameterization, and task supervision. This concept enables second-order meta-learning for end-to-end training of the encoding procedure alongside the decoder, and clarifies which learning pathway first-order approximations discard. Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention. These interactions shape both field predictions and the updates that construct their representation, allowing local observations to inform coherent non-local structure. Disentangling the inner encoding objective from outer task supervision unifies reconstruction, classification, and segmentation within an end-to-end meta-learning framework, using reconstruction-only latent adaptation at test time. Controlled experiments on polynomial fields link latent coordination to lower effective rank and stronger alignment with the underlying function space. Across image and 3D shape reconstruction, MetaLF improves fidelity within three to five gradient updates, while supporting semantic prediction across images, shapes, and volumes. Together, these findings position the optimization encoder perspective as a unified basis for designing neural fields around how representations are constructed, coordinated, and used.

64. 【2610.08070】wo Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation

链接:https://arxiv.org/abs/2610.08070

作者:Zhen Guo,Rongyuan Wu,Qiaosi Yi,Chenxi Xie,Xinyu Wei,Lei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent diffusion-based image, cost increase rapidly, network inference cost, inference cost increase, Recent diffusion-based

备注:

点击查看摘要

Abstract:Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly. While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging. Existing one-step methods typically allocate this budget to a single evaluation of a monolithic student. However, approximating the heterogeneous coarse-to-fine transport with a single monolithic mapping is difficult and often leads to over-smoothed outputs. To address this issue, we propose Phase-wise Velocity Distillation (PVD), which partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative computation equivalent to one full-backbone forward pass. We show that the use of two half-sized phase-specific experts outperforms a single full-size monolithic student. On class-conditional image generation, PVD achieves an FID of 1.48 on ImageNet 256 x 256. On more complex text-to-image (T2I) tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev, Qwen-Image) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Moreover, across the evaluated T2I backbones, PVD reduces active parameters by 49.10-50.89% and peak VRAM by 45.76-48.36% compared to the corresponding teachers. Source code and distilled models are available at this https URL.

65. 【2610.08068】PhysTacGen: Physics-Aware Visual-Tactile Sensor Image Generation

链接:https://arxiv.org/abs/2610.08068

作者:Guo Tang,Yongtao Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Realistic physical interaction, data remains costly, Realistic physical, collecting paired visual, embodied intelligence

备注:

点击查看摘要

Abstract:Realistic physical interaction is a cornerstone of embodied intelligence, yet collecting paired visual--tactile data remains costly. Visual-to-tactile synthesis offers a promising approach to augmenting such data, but learning this mapping is complicated by the gap between visual appearance and contact-related material properties, as well as spatial misalignment in paired observations. To address these challenges, we present \textbf{PhysTacGen}, a visual-to-optical-tactile image generation framework that integrates material-aware descriptions with geometric conditioning. First, we introduce Group Tactile Policy Optimization (GTPO), a reinforcement learning strategy that refines a vision--language model to generate structured material descriptions using task-specific rewards. Second, we combine DINOv2-based pair curation with monocular relative-depth estimation to select training pairs and provide geometric priors. Finally, an SDXL ControlNet synthesizes optical tactile images conditioned on RGB, relative depth, and GTPO-generated text. Experiments on curated SSVTP data demonstrate improved structural similarity over the compared baselines, while a blinded user study shows a preference for GTPO-generated descriptions. Generated tactile inputs also improve performance on an attribute-derived force-coefficient prediction proxy. Together, these results demonstrate the effectiveness of PhysTacGen for optical tactile image synthesis and its utility in the evaluated downstream this http URL code will be available at this https URL.

66. 【2610.07990】A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

链接:https://arxiv.org/abs/2610.07990

作者:Sin-Han Yang,Shih-Cheng Huang,Chieh-Yen Lin,Yun-Nung Chen,Shao-Hua Sun,Hung-yi Lee

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:task-specific weight updates, individual task-specific models, Model merging aims, aims to build, cheaply by combining

备注: Preprint

点击查看摘要

Abstract:Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.

67. 【2610.07987】VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

链接:https://arxiv.org/abs/2610.07987

作者:Yuan Feng,Qize Yang,Ruizhe Chen,Sibo Song,Haolin He,Muzhi Zhu,Zihan Liu,Yunfei Chu,Xize Cheng,Yuxuan Wang,Jin Xu,Xike Xie

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:incur substantial costs, fixed-size patch tokens, Multimodal large language, large language models, inputs into dense

备注:

点击查看摘要

Abstract:Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.

68. 【2610.07984】Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

链接:https://arxiv.org/abs/2610.07984

作者:Youxing LI

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Multimodal assistants answer, Multimodal assistants, answering model, assistants answer questions, model

备注:

点击查看摘要

Abstract:Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of M$^3$Exam, DMV and MemEye and uses 11--23\% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.

69. 【2610.07982】M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding

链接:https://arxiv.org/abs/2610.07982

作者:Jinsong Zhang,Kejun Wu,Ming Zhu,Renjie Qiao,Chengtao Cai,Zhengguo Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:spatial information required, metric depth estimation, depth, Monocular, spatial

备注: 13 pages. 7 figures, submitted to IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

点击查看摘要

Abstract:Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ($\delta 0.25$). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.

70. 【2610.07969】EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation

链接:https://arxiv.org/abs/2610.07969

作者:Yikai Qin,Yifei Deng,Mingjian Liang,Wenxuan Song,Zepeng Lin,Zhiyi Jiang,Jiajun Fu,Qiao Sun,Huashuo Lei,Xicheng Gong,Jiayi Chen,Han Zhao,Shuanghao Bai,Pengxiang Ding,Pengwei Wang,Haoang Li

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Scaling robotic foundation, robotic foundation models, foundation models requires, models requires diverse, requires diverse training

备注:

点击查看摘要

Abstract:Scaling robotic foundation models requires diverse training data and reliable evaluation environments. Simulation offers a scalable solution, yet existing generation pipelines remain constrained by predefined assets and skills, a disconnect between scene generation and task generation, and limited support for complex embodiments and physics. We introduce EmbodiedSmith, a framework for scalable embodied data generation through recursive self-improvement (RSI). EmbodiedSmith unifies asset, scene, and task generation in a pipeline that supports autonomous creation and language-driven customization. Its core is an agentic refinement loop: scene generation anticipates downstream task requirements, while task generation guides targeted scene edits, allowing scenes and tasks to iteratively improve one another. This joint refinement improves task generation success, including for long-horizon tasks. The framework further supports mobile manipulators, humanoids, and dexterous hands, as well as interactions involving deformable objects and fluids, broadening the range of behaviors and physical phenomena represented in generated data. Together, these capabilities provide a flexible simulation engine for both robot pretraining and evaluation. Extensive experiments validate the quality, diversity, and generation efficiency of the resulting data, while downstream policy experiments demonstrate that increased data diversity improves generalization.

71. 【2610.07958】DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given

链接:https://arxiv.org/abs/2610.07958

作者:Minhyeok Lee,Jungho Lee,Minseok Kang,Heeseung Choi,Ig-Jae Kim,Sangyoun Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, replacing per-scene optimization, single forward pass, forward pass, replacing per-scene

备注:

点击查看摘要

Abstract:Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.

72. 【2610.07954】Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction

链接:https://arxiv.org/abs/2610.07954

作者:JunGyu Lee,Inhwan Bae,Hae-Gon Jeon

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:learn auxiliary tasks, predictors represent coordinates, group reasoning, discrete tokens, tokens and learn

备注: 35 pages, 15 figures. Project page: [this https URL](https://jungyu0413.github.io/MoRE/)

点击查看摘要

Abstract:Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning. This formulation enables the model to capture behavioral intent and social context beyond coordinate dynamics alone. However, token-level objectives provide only indirect guidance for continuous coordinate-space dynamics. To address this limitation, we introduce MoRE (Mixture of Reward Experts), a refinement framework that transfers numerical forecasting priors into a pretrained language-based predictor through reinforcement learning. Five frozen numerical predictors provide complementary coordinate-level knowledge of motion and interactions. Their predictions are converted into expert rewards and combined through an uncertainty-weighted consensus that penalizes disagreement. A ground-truth reward anchors the prediction to the target trajectory. To focus refinement on difficult cases, MoRE refines the policy using the top 1% of training samples ranked by predictive entropy. Expert predictions are computed once and cached before PPO training, so the experts are not run during policy updates or inference. In this way, MoRE combines the contextual modeling of the language-based predictor with coordinate-level feedback from numerical experts. On ETH-UCY, MoRE reduces ADE from 0.22 to 0.20 m and FDE from 0.32 to 0.29 m. Relative to the base policy, ADE decreases by 17.9% on SDD and 12.7% on NBA. On ETH-UCY, MoRE also reduces collision rates and better matches ground-truth pedestrian spacing, without increasing measured inference memory or latency. The project page is available at this https URL.

73. 【2610.07941】UltraDiff: Differentiable Ray Tracing in Ultrasound for Shape Optimization

链接:https://arxiv.org/abs/2610.07941

作者:Felix Duelmer,Magdalena Wysocki,Nassir Navab,Mohammad Farid Azampour

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:enables gradient-based optimization, Physically-based differentiable rendering, rendering enables gradient-based, matching rendered images, differentiable rendering enables

备注: 4 pages, 4 figures, 1 table. Accepted at SIGGRAPH Asia 2026 Technical Communications

点击查看摘要

Abstract:Physically-based differentiable rendering enables gradient-based optimization of scene parameters by matching rendered images to measurements, but has so far mainly focused on light transport. We extend this paradigm to medical ultrasound, where image formation resembles transient rendering: echoes are binned by time-of-flight rather than projected onto an image plane. We present UltraDiff, a modular framework for differentiable ultrasound ray tracing. UltraDiff formulates ultrasound image formation as a path-space integral, gated by travel time between the transducer and tissue interfaces, and derives a Monte Carlo estimator of both the forward model and its gradients with respect to scene parameters. We demonstrate this on an inverse geometry estimation: starting from a sphere, an SDF is optimized until simulated echoes match measured ones, recovering vertebral surfaces from simulated B-mode sweeps and from a real robotic acquisition of a spine phantom. Unlike state-of-the-art ultrasound shape reconstruction methods, which rely on pre-segmented images, our approach operates unsupervised on B-mode images through analysis-by-synthesis, while achieving competitive geometric accuracy. Implemented on top of Mitsuba 3, UltraDiff brings differentiable path tracing to a new sensing modality and provides a foundation for inverse problems in acoustic imaging.

74. 【2610.07939】CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage

链接:https://arxiv.org/abs/2610.07939

作者:Baptiste Chopin,Thomas Swearingen,Arun Ross,Antitza Dantcheva,Christian Rathgeb

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:produce fabricated surveillance, Due to rapid, fabricated surveillance footage, produce fabricated, fabricated surveillance

备注:

点击查看摘要

Abstract:Due to rapid advances in Generative AI, commercial video generation tools can be used to produce fabricated surveillance footage that can fool both human viewers and automated synthetic video detectors. Since these tools are so widely accessible, a malicious user can create a harmful video clip at minimal cost. The production and dissemination of such videos in high-stakes settings, such as crime reporting and elections, can misdirect emergency response efforts or distort political discourse. Existing deepfake video datasets, used by the research community to develop deepfake detection algorithms, exhibit two limitations: (1) they emphasize benign web content rather than footage of possibly malicious activity, and (2) they rely on older or open-source generators that do not represent recent advances in generative systems. We assemble CCtv DeepFakes (CCDF), a video deepfake dataset, to address both gaps. CCDF contains 1840 videos (460 real and 1380 generated) spanning 16 crime and accident categories, with generated content produced using three leading commercial systems: Grok Imagine, Google VEO 3.1, and OpenAI Sora 2. CCDF is a highly realistic, small-scale, manually annotated dataset targeting evaluation of detection models. We release three versions of the dataset: the raw generated data, a cleaned version in which video metadata are standardized between real and synthetic samples to prevent detectors from exploiting trivial cues, and an altered version simulating low-effort post-processing attacks. We evaluate CCDF with ten recent state-of-the-art detectors covering different detection approaches. Our results suggest that these approaches do not reliably distinguish CCDF's generated videos from real ones, despite their strong reported performance on existing datasets. These results further confirm that existing datasets are not well-suited to evaluating certain threats.

75. 【2610.07928】Dynamic Alignment and Calibration for Multimodal Learning

链接:https://arxiv.org/abs/2610.07928

作者:Jinghao Xu,Zhenhua Guo,Xiaofeng Zhu,Xiaoshuang Shi

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:modeling information discrepancies, learn robust representations, adaptively modeling information, discrepancies across modalities, multimodal learning aims

备注: 17 pages

点击查看摘要

Abstract:Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.

76. 【2610.07925】F-PRVR: Training-Free Partially Relevant Video Retrieval

链接:https://arxiv.org/abs/2610.07925

作者:Giyeol Kim,Chanho Eom

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Partially Relevant Video, Relevant Video Retrieval, Partially Relevant, retrieve untrimmed videos, Video Retrieval

备注:

点击查看摘要

Abstract:Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and source-domain overfitting induced by task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments demonstrate consistent performance across datasets with diverse visual and temporal characteristics, suggesting a practical direction for training-free PRVR.

77. 【2610.07922】OpenWAM: An Open Framework for Composable World-Action Models

链接:https://arxiv.org/abs/2610.07922

作者:Heng Yu,David D. Yuan,Juze Zhang,Changan Chen,Yao Feng,Michelle Baldonado,Steve Cousins,Li Fei-Fei,Jiajun Wu,Ehsan Adeli

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:inference procedure simultaneously, couple future prediction, design choices difficult, couple future, robot control

备注: 18 pages, 5 figures, 14 tables. Project page: [this https URL](https://openwam.stanford.edu) ; Code: [this https URL](https://github.com/OpenWAM/OpenWAM) ; Code and project page released June 4, 2026. Equal contribution: Heng Yu, David D. Yuan, Juze Zhang

点击查看摘要

Abstract:World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.

78. 【2610.07916】Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization

链接:https://arxiv.org/abs/2610.07916

作者:Xuekang Zhu,Kaiwen Feng,Ruifeng Wang,Xiwen Wang,Xiaochen Ma,Bo Du,Changjiang Jiang,Chenfan Qu,Songyu Ye,Xia Du,Wentao Feng,Jian Liu,Ji-Zhe Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Image Manipulation Localization, optimal manipulation mask, Image Manipulation, manipulation mask, optimal manipulation

备注: NeurIPS 2026 (Oral)

点击查看摘要

Abstract:Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y|x)=\int P(y|z)\,P(z|x)\,dz$, where $z$ denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling $z$ in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate $P(z|x)$ and $P(y|z)$ via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at this https URL

79. 【2610.07913】Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images

链接:https://arxiv.org/abs/2610.07913

作者:Shrihari Dumbre,Bikash Santra

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:cancer-related mortality worldwide, effective treatment planning, Gastric adenocarcinoma, mortality worldwide, treatment planning

备注:

点击查看摘要

Abstract:Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at this https URL.

80. 【2610.07911】Diverse Motion Customization via Control-based Dynamic Optimization

链接:https://arxiv.org/abs/2610.07911

作者:Youngyoon Choi,Kihyun Kim,Jeongwoo Shin,Joonseok Lee

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:video unintentionally propagate, reference video unintentionally, reference video, remains challenging due, customization remains challenging

备注: Preprint

点击查看摘要

Abstract:Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video, which arises from formulating the learning objective as a direct regression on the reference. To address this, we propose Control-based Motion Customization (CMC), a principled training framework that is structurally robust to content leakage. Our key idea is to steer generative dynamics toward desired motion while avoiding collapse toward the reference video, which we formalize using Stochastic Optimal Control (SOC). Under this formulation, customized videos acquire the target motion yet remain within the pre-trained model's prompt-conditional distribution, where appearance is determined by the text prompt rather than the reference video. Furthermore, to improve efficiency, we tailor the SOC formulation to motion customization by eliminating the need for an explicit reward and introducing a timestep-adaptive motion cost that focuses only on early generative stages, accelerating training by 2.5 times. Extensive experiments demonstrate that CMC effectively mitigates content leakage and achieves competitive motion fidelity while preserving the diversity of the base model across diverse scenarios.

81. 【2610.07903】Unsupervised Long-Tailed Adaptation of Vision-Language Models

链接:https://arxiv.org/abs/2610.07903

作者:Keliang Chen,Yaxin Hou,Hui Liu,Yuheng Jia

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieved remarkable success, Adapting vision-language models, Adapting vision-language, leveraging pseudo-labels generated, downstream tasks

备注: 18 pages, 6 figures

点击查看摘要

Abstract:Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.

82. 【2610.07887】Visual Abstention in Unified Multimodal Models

链接:https://arxiv.org/abs/2610.07887

作者:Chufan Shi,Cheng Yang,Tiannuo Yang,Isadora White,Yiwei Chen,Taylor Berg-Kirkpatrick,Xuezhe Ma

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Unified multimodal models, Unified multimodal, integrate understanding, understanding and generation, generative behavior

备注: 25 pages, 6 figures, 13 tables. Project page: [this https URL](https://visual-abstention.github.io)

点击查看摘要

Abstract:Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.

83. 【2610.07885】Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools

链接:https://arxiv.org/abs/2610.07885

作者:Jeonghwa Lim,Minje Park,Yeongyeon Na,Yujin Eom,Soyeon Lim,Young Ho Lee,Yu Jeong Kim,Sunghoon Joo,Ki Hong Lee

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)

关键词:clinically interpretable measurements, translates raw ECG, raw ECG signals, waveform boundaries, interpretable measurements

备注: 20 pages, 5 figures. First two authors contributed equally

点击查看摘要

Abstract:Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised pretraining and semi-supervised learning are expected to ease this burden, yet it remains unclear whether they yield reliable delineation and whether the deep models they produce outperform the delineation tools used in practice. We address this in two stages. First, comparing self-supervised objectives with supervised or semi-supervised fine-tuning across one internal and four external datasets, we find that pretraining helps but the objective matters, and that the value of semi-supervised fine-tuning depends on the pretraining objective. Second, we benchmark the selected deep learning model against widely used open-source (NeuroKit2, Prominence, ECGdeli) and commercial (CalECG) tools using three complementary metrics. The model ranks best on every metric and dataset, outperforming the strongest tool by a clear margin on the rhythm-diverse set (mIoU 71.3 vs. 54.8%; averaged point-wise sensitivity 92.6 vs. 76.4%), and degrades the least from sinus to arrhythmia. A rhythm-stratified and point-wise analysis further characterizes the distinctive behavior of each tool, yielding practical guidance for tool selection. These results provide systematic, multi-dataset evidence that self-supervised pretraining is effective for ECG delineation and enables a label-efficiently trained deep learning model to outperform widely used delineation tools by leveraging abundant unlabeled data. This supports adopting such models in diverse, real-world clinical settings.

84. 【2610.07883】Revar3r: gauge-aware perturbation uncertainty for feed-forward 3d reconstruction

链接:https://arxiv.org/abs/2610.07883

作者:Sammam Mahdi,Fariha Binta Salim,Rakin Bin Rabbani,Aniqua Nusrat Zereen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:correctly reconstructed distant, processes equivalent inputs, frame rotates fractionally, reconstructed distant point, model processes equivalent

备注:

点击查看摘要

Abstract:A correctly reconstructed distant point appears uncertain even when a frozen 3D model processes equivalent inputs because its output frame rotates fractionally. This exposes a weakness of trainingfree perturbation uncertainty: when outputs contain an unobserved symmetry, run-to-run variation potentially reflects symmetry rather than error. Existing alternatives have trade-offs: built-in confidence is outperformed in most evaluated conditions, while trained evidential heads require modelspecific supervision. For point maps, this research derives a closed-form, error-independent variance term that grows with scene extent and potentially overwhelms the desired signal. Simulation reproduces the effect; all 30 real VGGT view-sets tested exhibit its predicted $\|x_p\|^2$ signature. ReVar3R robustly registers predictions to a common similarity frame before computing per-point variance, without retraining or modifying the frozen model. Optional calibration and fusion use a held-out split. Across VGGT, {\pi}3, and MASt3R on six datasets, the same estimator on every backbone lowers AUSE below built-in confidence in 15 of 18 conditions. The staged evaluation yields 11 of 18 wins for the label-free core, 12/18 for label-free equal-weight fusion, 14/18 with held-out weights, and 15/18 when the built-in signal is included. Against a trained evidential head, the result is a trade-off: the head calibrates magnitude better and leads in its training domain, whereas ReVar3R transfers across backbones without adaptation. Its ranking improves point filtering, but it does not detect stable systematic bias, aid novel-view synthesis, or transfer calibration across domains.

85. 【2610.07868】CueRator: Agentic Search for Symbolic Rules to Adapt Frozen Multimodal Encoders

链接:https://arxiv.org/abs/2610.07868

作者:Sunchan Park,Beomkwon Cho,Kyeongbo Kong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large language model, language model agents, Large language, programs and equations, language model

备注: 40 pages, 18 figures

点击查看摘要

Abstract:Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework for policy-aware decision-rule discovery, which adapts frozen contrastive multimodal encoders by searching for the decision rule that converts their cross-modal similarities into predictions. We validate it on open-vocabulary audio-visual event perception, where existing methods involve a trade-off between adaptivity and generalization to unseen categories: trained modules adapt at the cost of generalization, and fixed rules the reverse. The framework pairs a symbolic formulation for generalization with a lightweight policy that predicts its parameters per video for adaptivity. A report-guided multi-agent loop discovers the formulation offline, evaluating each candidate on its expressive ceiling and on whether a trained policy can realize it. On OV-AVEBench, CueRator raises the total average from 57.8 to 60.2 and unseen-category performance from 55.8 to 59.9 over the best existing method, reducing the seen-unseen gap from 7.1 to 1.2. Ablations attribute the gains to both the formulation and the policy and show that both feedback signals are necessary for effective search. CueRator also improves over the respective baselines on two further audio-visual event perception tasks, and the discovered rule remains competitive across encoders with only the policy retrained. Code is available at this https URL.

86. 【2610.07843】CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology

链接:https://arxiv.org/abs/2610.07843

作者:Hyun Do Jung,Jungwon Choi,Soojung Choi,Yujin Oh,Hwiyoung Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:instance learning models, whole-slide image multiple, image multiple instance, multiple instance learning, original full-bag prediction

备注:

点击查看摘要

Abstract:In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models. In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction. If the evaluation reference changes while the intended target remains the original full-bag prediction, however, not only can the measured fidelity of the same compact evidence change, but comparisons between competing candidate strategies can also change. To make this dependence explicit, we introduce CHARTER, a reference-aware evaluation charter that asks researchers to DECLARE the intended target and reference, QUANTIFY candidate-induced prediction shift, and AUDIT the stability of comparative conclusions. Across the 15 comparisons in our main five-seed Random-K audit, 4 showed determinate reversals; in a matched native-ranking stress test, the ACMIL comparison changed from REVERSED to PRESERVED. CHARTER turns otherwise implicit candidate-filtering and reference choices into an auditable evaluation specification, helping distinguish genuine preservation of the intended prediction from apparent gains induced by changing the prediction being explained.

87. 【2610.07819】$α$Transfer: Coefficient Transfer for Efficient Model Merging

链接:https://arxiv.org/abs/2610.07819

作者:Shih-Cheng Huang,Zhi Rui Tam,Chieh-Yen Lin,Yun-Nung Chen,Hung-yi Lee,Shao-Hua Sun

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:combining multiple fine-tuned, multiple fine-tuned checkpoints, parameter arithmetic, offers a promising, promising solution

备注: Under review

点击查看摘要

Abstract:Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$\alpha$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $\alpha$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $\alpha$Transfer as an efficient and generalizable approach to scaling model merging.

88. 【2610.07802】owards benchmarking Western Bluebird detection in the wild

链接:https://arxiv.org/abs/2610.07802

作者:Estela Monserrat Arriaga Santana,Julian Rosas Scull,Ibeth P. Alarcón,Bibiana Montoya,Aylin Sosa Mejía,Hugo Jair Escalante

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:variability in illumination, observers' viewpoint, monitoring in natural, natural environments, environments is challenging

备注: 15 pages, 4 figures, 10 tables

点击查看摘要

Abstract:Bird monitoring in natural environments is challenging due to the small size of some species of birds relative to the scene, background clutter, variability in illumination, and the observers' viewpoint. Progress is further limited by the scarcity of large-scale, realistic datasets, which are essential for understanding behavioral patterns. To address this gap, we introduce a new benchmark dataset for the detection and segmentation of Western bluebirds (Sialia Mexicana), comprising over 6,000 labeled images from 41 recording sessions. The dataset features high-resolution (4K) in-the-wild images in which birds occupy only a small fraction of the image. We evaluated supervised detectors, open-vocabulary models under zero-shot and fine-tuned settings, and segmentation approaches. Supervised detectors remain the most reliable overall, with Faster R-CNN achieving the highest detection mAP and RT-DETR offering the best precision-recall trade-off. Open-vocabulary models perform poorly in zero-shot settings; however, fine-tuning substantially improves their performance, with YOLO-World becoming competitive with supervised methods and achieving the highest precision, F1-score, and mAP@0.5. For segmentation, supervised methods significantly outperform Grounded-SAM and SAM 3: Mask R-CNN achieves the highest mask mAP, while YOLOv8-Seg provides the best precision and fastest inference. A diagnostic analysis further shows that failures are not explained by object size alone, but by a combination of apparent scale, brightness, contrast, clutter, blur, crowding, and recording-session variation. Overall, our findings highlight the difficulty of zero-shot bird detection in cluttered ecological scenes and underscore the importance of domain adaptation in small-object settings.

89. 【2610.07795】Efficient Gaussian Splatting Sequence Compression with Standard Video Codecs

链接:https://arxiv.org/abs/2610.07795

作者:Qi Yang,Shuting Xia,Le Yang,Geert Van Der Auwera,Zhu Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:effective Gaussian Splatting, Gaussian Splatting, Linear Assignment Sorting, Parallel Linear Assignment, effective Gaussian

备注: Accepted by MM Asia 2026

点击查看摘要

Abstract:This paper presents a novel effective Gaussian Splatting (GS) sequence Compression method that utilizes the Video codec (GSCV). Existing video-based GS sequence compression relies on the Parallel Linear Assignment Sorting (PLAS) and tracked primitive information to convert GS into smooth 2D videos. However, tracked information is not available for most practical applications, and without it, using the vanilla PLAS can generate images exhibiting weak inter-frame correlation, due to its stochastic nature. GSCV incorporates a simple yet efficient Inter-PLAS method to produce close images between the I- and P-frames of GS, enhancing the inter-frame performance of video codec greatly. GSCV also realizes a new pipeline based on the state-of-the-art video codecs with high bit-depth GS images, achieving higher compressibility while simultaneously providing a higher quality upper bound. Experimental results show that the proposed GSCV exhibits obviously improved performance over MPEG video and point cloud-based anchors in GS sequence compression. The code is available at this https URL.

90. 【2610.07793】Geometry-Constrained Bidirectional Point Cloud Registration for Thin, Sheet-Like Heritage Artifacts

链接:https://arxiv.org/abs/2610.07793

作者:Yuezhe Zhang,Lei Wei,Jingnan Du,Shuai Wan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:artifacts poses significant, Non-contact three-dimensional reconstruction, poses significant geometric, sheet-like heritage artifacts, heritage artifacts poses

备注: 26 pages, 8 figures. Accepted for publication in ACM Journal on Computing and Cultural Heritage

点击查看摘要

Abstract:Non-contact three-dimensional reconstruction of thin, sheet-like heritage artifacts poses significant geometric and registration challenges. Due to their fragility, these artifacts cannot be suspended or equipped with artificial markers, necessitating independent acquisition of their front and back surfaces. Subsequent registration proves difficult due to the limited number of shared geometric features and the scarcity of explicit physical constraints, which may result in rotational ambiguity, instability, and structural collapse during iterative optimization. To address these challenges, we propose a geometry-constrained bidirectional point cloud registration method specifically tailored for thin, sheet-like heritage artifacts. The method integrates semantic-guided preprocessing, Principal Component Analysis (PCA)-based geometric normalization, and a thickness-aware registration strategy. The estimated physical thickness is incorporated as a geometric constraint to preserve structural integrity during registration. Rotational ambiguity is resolved by evaluating a finite set of global rotation hypotheses, each refined using the point-to-plane Iterative Closest Point (ICP) algorithm, with the optimal transformation selected via a geometry-aware fitness criterion consistent with the thickness scale. Experimental results show that the proposed method achieves competitive or improved performance in most cases, particularly in projected area consistency and physically plausible front-back alignment. In addition, the thickness-aware constraint and rotation hypothesis evaluation reduce the risk of degenerate configurations in which the two surfaces are incorrectly flipped while still yielding deceptively acceptable numerical scores, supporting reliable non-contact digitization of delicate and thin heritage artifacts. Implementation details are available at this https URL.

91. 【2610.07788】Image-Space Refraction Correction for Underwater 3D Reconstruction: Warping Flat-Port Views into Pinhole Perspective

链接:https://arxiv.org/abs/2610.07788

作者:Chelim Lim,Tobias Fischer,Emilio Olivastri,Beverley Gorry,Michael Milford,Alejandro Fontan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:seafloor habitats due, Consumer-grade cameras, cost and accessibility, housings are widely, coral reefs

备注:

点击查看摘要

Abstract:Consumer-grade cameras in flat-port housings are widely used for underwater exploration and mapping of coral reefs and seafloor habitats due to their low cost and accessibility. However, refraction at flat-port interfaces causes bowl-shaped deformation in reconstructed scenes and camera trajectories, compromising the metric accuracy required for mapping and navigation. To remove the dominant refractive distortion before reconstruction, we introduce a physics-based refraction correction in image space. Our method is downstream-agnostic: the refraction-corrected images can be directly used as input to existing reconstruction and SLAM algorithms. We characterize the refractive distortion through ray-tracing simulations and validate our correction on two real underwater datasets with differing scene structures. Compared with conventional and refractive Structure-from-Motion (SfM), our approach removes reconstruction deformation while registering more frames and maintaining low reprojection error. The correction further generalizes across diverse reconstruction and VSLAM backends, demonstrating its broad applicability to downstream vision pipelines.

92. 【2610.07783】From Laboratory to Road: Evaluating Wearable Gaze Accuracy for Driving

链接:https://arxiv.org/abs/2610.07783

作者:William Engel,Fabian Flohr

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:widely used interface, interface between perception, perception and planning, planning in autonomous, behaviorally relevant

备注: Peer-reviewed and accepted as an Extended Abstract at the German Conference on Pattern Recognition (GCPR 2026). Presented as a poster at GCPR 2026

点击查看摘要

Abstract:Bird's-eye-view (BEV) representations have become a widely used interface between perception and planning in autonomous driving, but they encode what is in a scene, not what is behaviorally relevant to a human driver. Gaze offers a compelling behavioral signal for this gap, yet wearable eye trackers are routinely deployed as if their spatial output were ground truth, despite known sensitivity to head motion, illumination, and calibration drift. We present, to our knowledge, the first unified framework for quantifying wearable gaze accuracy under real driving conditions. Our on-road study contains 41 validated scenes in which one driver fixated a vehicle's license plate. Gaze error is measured as the angular difference between the plate center and the gaze direction estimated by the glasses. Separate indoor studies with the same driver and device systematically analyze how distance, illumination, head motion, target motion, and gaze eccentricity affect both systematic bias and gaze precision. The mean on-road error was 4.58 degrees. Applying an offset estimated from the indoor recordings reduced it to 1.10 degrees and improved all 41 scenes. Because this offset varied between sessions, reliable BEV supervision may require online recalibration and condition-dependent estimates of gaze uncertainty.

93. 【2610.07758】Later Is Better: Token Reduction for ViTs Under Distribution Shift

链接:https://arxiv.org/abs/2610.07758

作者:Hyeongheon Cha,Hyungjun Yoon,Sung-Ju Lee

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:accelerates vision transformers, Training-free token reduction, removing redundant tokens, reduction accelerates vision, Training-free token

备注: 35 pages. Code: [this https URL](https://github.com/chahh9808/LaterIsBetter)

点击查看摘要

Abstract:Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the uncompressed model widens with the removal rate. We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail. Concretely, we introduce a one-parameter late-concentrated power-law schedule that consistently improves out-of-distribution accuracy over flat at no extra inference cost. On ImageNet-C with DeiT-S, the late schedule closes 83% of that gap at a 26% compute reduction (+1.17pp), and 99% of it at a lighter 7% reduction (+0.26pp). The gain cannot be attributed to retaining more tokens or using extra compute: held to flat's compute, the late schedule removes more tokens in total and leaves fewer tokens at the end, yet still wins. Single-layer probes point to a mechanism: earlier reductions perturb features that pass through more remaining layers, front-loading reduction error in depth. The effect is broad, holding across five token-reduction methods (ToMe, EViT, ATS, ATC, PiToMe), nine backbones, all ImageNet-C corruption types, eight further shift suites, and two further modalities, video and vision-language QA. It is also specific to shift, still positive on clean and rising monotonically to ~4x that at the highest severity 5. The schedule keeps its gain under six test-time adaptation methods, and needs no per-input or per-domain tuning.

94. 【2610.07754】Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

链接:https://arxiv.org/abs/2610.07754

作者:Soichiro Kumano

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:high computational cost, reliable defenses, high computational, computational cost, cost must generally

备注:

点击查看摘要

Abstract:Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for each task. Robust foundation models offer a promising alternative: adversarially pretrain a model once and then transfer its robustness to downstream tasks through lightweight adaptation. However, a fundamental question remains open: can robustness acquired during pretraining transfer to unseen tasks without further adversarial training? In this study, we answer this question affirmatively. A single model adversarially pretrained at scale can achieve optimal robustness on new tasks without additional task-specific training. Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations. By contrast, a standardly trained model cannot. We further analyze convergence under gradient flow, an accuracy--robustness trade-off, and demonstration complexity.

95. 【2610.07729】Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs

链接:https://arxiv.org/abs/2610.07729

作者:Donghyun Han,Jangho Park,Yuseok Bae

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual tokens, vision-language models, strong compression baseline, major source, source of inference

备注:

点击查看摘要

Abstract:Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.

96. 【2610.07726】Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study

链接:https://arxiv.org/abs/2610.07726

作者:Kai Zhou,Chuanshen Chen,Runhao Zeng,Meng Dai,Yifan Yang,Jinwu Hu,Daiyuan Li,Mingkui Tan,Fei Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Videofluoroscopic Swallowing Study, providing dynamic X-ray, diagnosing swallowing disorders, dynamic X-ray imaging, Videofluoroscopic Swallowing

备注: Accepted by ICME 2026 Oral

点击查看摘要

Abstract:Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S$^3$KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: this https URL.

97. 【2610.07720】RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing

链接:https://arxiv.org/abs/2610.07720

作者:Wanning He,Yuyao Zhang,Yu-Wing Tai

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Multi-reference image generation, generation requires preserving, Multi-reference image, requires preserving, multiple subjects

备注: 19 pages. Wanning He and Yuyao Zhang contributed equally and share first authorship

点击查看摘要

Abstract:Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve $18.3\times$ and $14.2\times$ speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.

98. 【2610.07711】Comprehensive Evaluation and Fine-Tuning of Foundational Cell Nuclei Segmentation Models in Renal Pathology

链接:https://arxiv.org/abs/2610.07711

作者:Ruijie Wu,Junlin Guo,Ruining Deng,Yu Wang,Shilin Zhao,Haichun Yang,Yuankai Huo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate nuclei instance, quantitative renal pathology, strong background staining, Accurate nuclei, nuclei instance segmentation

备注: 11 pages, 5 figures, 4 tables. Submitted to SPIE Medical Imaging 2027

点击查看摘要

Abstract:Accurate nuclei instance segmentation is essential for quantitative renal pathology, yet general-purpose models often struggle with low contrast, dense nuclei, complex morphology, and strong background staining. In this work, we extended a human-in-the-loop framework by combining 5,901 foundation-model-generated pseudo-labels from well-segmented cases (Easy), 860 newly expert-annotated unresolved challenging cases (Medium), and 198 expert-annotated consensus failure cases (Hard). These annotations, spanning different levels of segmentation difficulty, enabled the systematic evaluation of seven single-source and mixed-source fine-tuning strategies across nine cell segmentation model configurations. Fine-tuning improved all models, with Medium data included in seven of the nine best-performing strategies. LSP-DETR achieved the highest F1 score of 0.8725 with Hard-only fine-tuning, while StarDist showed the largest improvement, increasing from 0.7380 to 0.8332 with Medium-only fine-tuning. These findings show that annotations spanning multiple difficulty levels support effective model adaptation, although the optimal annotation composition remains model dependent.

99. 【2610.07705】What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video

链接:https://arxiv.org/abs/2610.07705

作者:Wonbin Son,Gyumum Choi,Junil Seo,Hyungjoon Kim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:unmanned aerial vehicles, image-based UAV detection, aerial vehicles, unmanned aerial, increased the importance

备注:

点击查看摘要

Abstract:The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.

100. 【2610.07698】RBMatch: Dual-Level Class Rebalancing for Semi-Supervised Building Footprint Extraction

链接:https://arxiv.org/abs/2610.07698

作者:Akil Ahmad Taki,Shaikh Anowarul Fattah

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate building footprint, high-resolution remote sensing, remote sensing imagery, disaster response, Accurate building

备注:

点击查看摘要

Abstract:Accurate building footprint extraction from high-resolution remote sensing imagery is essential for urban planning, disaster response, and environmental monitoring. However, obtaining dense pixel-level annotations is costly, motivating the use of semi-supervised learning (SSL) to leverage unlabeled imagery. In remote sensing, severe foreground--background imbalance poses a particular challenge for self-training, as it can bias pseudo-label generation and the resulting unsupervised optimization toward the majority background class. We show that addressing this imbalance at only one stage is insufficient: balancing pseudo-label selection alone does not prevent background bias from re-emerging during unsupervised loss optimization, a failure mode we term \emph{imbalance leak}. To address this issue, we propose \textbf{RBMatch}, a dual-level class-rebalancing framework that jointly regulates pseudo-label generation and unsupervised optimization. RBMatch combines a supervised learning pathway with a self-training module comprising three components: adaptive class-specific thresholding (ACT) for balanced pseudo-label selection, confidence-aware class-balanced reweighting (CACBR) for mitigating class bias in the unsupervised loss, and distribution alignment (DAL) for matching the predicted unlabeled-data distribution to the labeled-data prior. Experiments on the WHU, INRIA, and Massachusetts building footprint datasets across labeled ratios of 1%--10% show that RBMatch consistently achieves the best building IoU and F1-score among the evaluated methods. The improvement is most pronounced on the highly imbalanced Massachusetts dataset, where RBMatch improves IoU by 1.37 points over the strongest baseline at a 1% labeling ratio and is the only method to outperform the fully supervised baseline across all twelve dataset--ratio settings.

101. 【2610.07694】Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction

链接:https://arxiv.org/abs/2610.07694

作者:Tao Zhou,Ying Hu,Huazhu Fu,Yi Zhou,Xiao-Jun Wu,Haibin Ling

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:histopathological Whole-Slide Images, profiles holds significant, holds significant promise, Whole-Slide Images, transcriptomic profiles holds

备注: 15 pages, 6 figures, 7 tables

点击查看摘要

Abstract:The integrative analysis of histopathological Whole-Slide Images (WSIs) and transcriptomic profiles holds significant promise for cancer survival prediction. However, existing methods typically project multi-modal features directly into a shared latent space without explicit alignment, leading to the entanglement of mismatched morphological cues and molecular signals. Furthermore, current fusion strategies often treat the extreme spatial heterogeneity of WSIs uniformly, lacking mechanisms to adaptively prioritize clinically relevant tissue scales for individual patients. To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM$^2$ES) framework for survival prediction. Specifically, we present an Anchor-driven Multi-modal Fusion (AMF) module, which introduces learnable semantic anchors as cross-modal mediators to bridge the semantic gap by enforcing a structurally regularized alignment between transcriptomic features and multi-scale pathology representations. Built upon this aligned semantic space, we further design a Hierarchical Mixture-of-Experts (H-MoE) selection module to decouple the hierarchical prognostic selection process. Mimicking the pathologist's diagnostic workflow, H-MoE performs (i) Intra-scale Expert Filtering to discriminatively identify salient tumor regions within each magnification, and (ii) Inter-scale Hierarchy Routing to dynamically weight and select the most informative resolution levels. Extensive experiments on multiple TCGA cancer cohorts demonstrate that our AM$^2$ES achieves state-of-the-art performance while offering fine-grained interpretability by visualizing how specific molecular pathways drive the expert routing decisions across tissue scales. The code will be released at this https URL.

102. 【2610.07689】Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling

链接:https://arxiv.org/abs/2610.07689

作者:Juntong Li,Lingwei Dang,Haomin Wu,Ziyan Qiu,Qingxin Xiao,Qingyao Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Multimodal Large Language, Large Language Models, lack fine-grained perceptual, Vision-Language Models, Multimodal Large

备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at this https URL.

103. 【2610.07684】Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation

链接:https://arxiv.org/abs/2610.07684

作者:Haipeng Liu,Yang Wang,Meng Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:style transfer, customization style transfer, synthesize text-driven images, text-driven images conditioned, Personalized image generation

备注: 28 pages, 14 figures, to appear at NeurIPS 2026, Sydney, Australia

点击查看摘要

Abstract:Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized this http URL experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from this https URL.

104. 【2610.07609】PhysLDM: Latent Diffusion for High-Fidelity Deformable Simulation

链接:https://arxiv.org/abs/2610.07609

作者:Yu Zhang,Xudong Xu,Xingang Pan

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:computer graphics, graphics and physical, high-fidelity deformable bodies, foundational challenge, deformable bodies

备注:

点击查看摘要

Abstract:Neural simulation of high-fidelity deformable bodies is a foundational challenge in computer graphics and physical AI. Long-horizon prediction for high-resolution 3D volumetric meshes is hard: autoregressive methods are susceptible to error accumulation, while direct multi-frame prediction at native resolution is computationally prohibitive. This motivates a compact spatiotemporal latent representation, which is largely unexplored for mesh-based volumetric physics. Meanwhile, it remains unclear whether deterministic regression or generative diffusion is the more appropriate predictive paradigm. To address these coupled challenges, we introduce PhysLDM, a unified latent-diffusion paradigm for one-shot volumetric deformable simulation. Its core is a holistic spatiotemporal VAE that avoids the "staircase" artifacts of standard temporal compression (as in common video VAEs), achieving ~2.48 mm reconstruction precision on meter-scale scenes at up to 78x token compression. Based on this reliable latent space, we systematically compare regression and diffusion methods. Our experiments uncover a key modeling insight: complex deformable dynamics are often chaotic, and in this regime deterministic regression tends to produce non-physical averages, whereas diffusion better models their distribution. Accordingly, we employ a latent diffusion model that effectively learns from the chaotic data to generate physically plausible trajectories. Trained purely kinematically on an Objaverse-scale dataset, a single PhysLDM generalizes zero-shot to unseen OOD datasets (GSO and Toys4K). Its differentiability further enables efficient solution of inverse problems and higher-order design optimization. To our knowledge, PhysLDM is the first high-fidelity spatiotemporal autoencoder and latent-diffusion paradigm for volumetric deformable dynamics, offering a scalable and robust approach to neural simulation.

105. 【2610.07603】Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage

链接:https://arxiv.org/abs/2610.07603

作者:Yi Xia,Ibrahim Khan,Mury Fajar Dewantoro,Wenwen Ouyang,Ruck Thawonmas

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Video Vision Transformers, article proposes Selective, Vision Transformers, proposes Selective Affective, Selective Affective Layer

备注:

点击查看摘要

Abstract:This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters after brief adaptation, directly measuring representational shifts and providing a more stable basis than gradient-based alternatives. Evaluated via five-fold cross-validation on the Arousal Video Game AnnotatIoN dataset, SALFT achieves performance comparable to full fine-tuning across all games without statistically significant degradation ($p0.05$), while updating only $\approx$8% of parameters (over 92% reduction). Notably, in one game, SALFT consistently outperforms both full fine-tuning and the best baseline across all metrics and folds, reaching the theoretical minimum p-value (p=0.0625, exact two-sided Wilcoxon signed-rank test). In addition, we introduce an interpretability method to trace attention patterns, enhancing model transparency. These results establish SALFT as an effective and efficient approach for affective game computing.

106. 【2610.07585】REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

链接:https://arxiv.org/abs/2610.07585

作者:Sheir A. Zaheer,Jihwan Moon,Chan Y. Park

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:hierarchical feature architecture, windowed group-convolutional self-attention, vision transformer based, propose a scalable, feature architecture

备注: 7 pages, Accepted for presentation at NeurIPS NeurREPS workshop 2026

点击查看摘要

Abstract:We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at this https URL.

107. 【2610.07576】CETUS: How Far Do Representations Trained on Earth Transfer to Cassini SAR of Titan?

链接:https://arxiv.org/abs/2610.07576

作者:Kevin Lee

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)

关键词:Survey Cassini SAR, Cassini synthetic aperture, Geological Survey Cassini, Cassini SAR mosaic, learned from Earth

备注: Research work at NASA Jet Propulsion Laboratory. Available at: [this https URL](https://github.com/magnaprog/CETUS)

点击查看摘要

Abstract:Cassini synthetic aperture radar (SAR) images reveal the dunes, plains, and lake basins of Titan, providing an instance of representations learned from Earth imagery for planetary terrain classification. Cross-domain Evaluation of Earth-to-Titan Transfer Using SAR (CETUS) compares features from DINOv2, DOFA and CROMA with classical image measurements and features from an untrained vision transformer on the U.S. Geological Survey's Cassini SAR mosaic. The classifiers learn terrain labels from an expert geomorphological map and predict those labels in geographically separate Titan regions. Under logistic regression settings, pretrained encoders achieve higher mean macro F1 than the combined classical features. Encoder rankings change when feature scaling, optimization, and regularization change together. Further training on Titan improves DINOv2 performance, degrades DOFA performance, and leads to mixed results for CROMA under the tested settings. Architectural and input processing differences prevent these comparisons from isolating the effect of pretraining. Classifier fitting and performance on individual terrain classes matter when assessing representation transfer for planetary mapping. Since the map draws partly on the same radar observations, the scores measure agreement with expert interpretation.

108. 【2610.07572】wo Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

链接:https://arxiv.org/abs/2610.07572

作者:Xi Ding,Naichen Shi,Jiawei Zhang

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:In-context learning, adapts frozen large, adapts frozen, hundreds of visual, demo image adds

备注: Technical report

点击查看摘要

Abstract:In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.

109. 【2610.07569】OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception

链接:https://arxiv.org/abs/2610.07569

作者:Binh Long Nguyen,Kien Nguyen,Clinton Fookes,Peyman Moghadam

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting-based mapping, complex environments, Gaussian Splatting-based, Splatting-based mapping approaches, dense semantic

备注: Accepted to ACCV 2026

点击查看摘要

Abstract:Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping approaches enable high-fidelity geometry and efficient open-vocabulary perception, but typically represent semantics as unstructured feature fields that limit object-centric reasoning. In contrast, 3D scene graphs explicitly model objects and their relationships for structured reasoning, but are commonly constructed from sparse geometric representations that do not fully exploit dense semantic maps. In this work, we present OpenSplatGraph, a unified framework that constructs persistent 3D scene graphs directly from an online Gaussian-based open-vocabulary semantic map. The proposed framework augments the dense semantic map with a reliability-aware semantic field that maintains lightweight observation statistics for confidence-aware, query-conditioned object extraction. Extracted object instances are associated with persistent graph nodes, allowing object attributes and relationships to be incrementally updated across observations and queries. By tightly coupling dense semantic mapping with persistent object-centric representations, our framework supports both language-guided object grounding and structured relational reasoning while preserving the geometric fidelity of Gaussian-based mapping. Comprehensive evaluations on standard 3D scene understanding benchmarks and real-world robotic experiments demonstrate that OpenSplatGraph achieves competitive performance for online open-vocabulary perception and downstream robotic tasks. Project page: https://csiro-robotics.github.io/OpenSplatGraph.

110. 【2610.07566】AIMS: Anchor-Integrated Multi-View Synthesis for Scalable Novel View Rendering

链接:https://arxiv.org/abs/2610.07566

作者:JooHyun Park,HanYoung Jang,HyeongYeop Kang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:posed multi-view inputs, sets remains challenging, large input view, methods achieve strong, achieve strong generalization

备注:

点击查看摘要

Abstract:Feed-forward novel view synthesis methods achieve strong generalization from posed multi-view inputs, but scaling them to large input view sets remains challenging. Transformer-based approaches that jointly process all input-view tokens incur rapidly increasing computation and memory as the number of views grows, while simple view subsampling discards potentially useful observations. We introduce Anchor-Integrated Multi-View Synthesis (AIMS), a scalable framework that decouples the number of available observations from the number of views processed by the global synthesis model. AIMS selects a fixed set of spatially distributed anchor views using farthest point sampling, groups nearby observations around each anchor, and uses a lightweight learnable integrator to fuse their information into enriched anchor representations. This allows additional observations to contribute to synthesis while keeping the downstream global view budget fixed. Evaluations on RealEstate10K and ScanNet demonstrate a favorable quality--efficiency trade-off against transformer-based and Gaussian-based baselines. AIMS achieves 29.41 dB and 17.73 dB PSNR on the two datasets, respectively, with rendering averaging 7.24 ms per view.

111. 【2610.07561】LARK: A Low-Cost, Accurate, Occlusion-Resilient, Kalman Filter-Assisted Tracking System for Image-Guided Surgery

链接:https://arxiv.org/abs/2610.07561

作者:George Sideris,Justin Cree,Andrew Stirling,Mamadou Ly,Étienne Léger,D. Louis Collins

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:real-time navigation relative, provide real-time navigation, surgical instruments, real-time navigation, navigation relative

备注: 24 pages, 15 figures, including appendices. Supplementary document included as an ancillary file. Supplementary video: [this https URL](https://youtu.be/ApZ8q9DjB-4)

点击查看摘要

Abstract:Image-guided surgery (IGS) depends on accurate tracking of surgical instruments to provide real-time navigation relative to anatomical structures. Commercial stereo infrared trackers are accurate but prone to occlusion and cost-prohibitive for many settings. This work presents LARK, a multi-camera optical tracking system using commodity RGB hardware and multi-view redundancy and fusion. We develop and evaluate two complete tracking methods: multi-view monocular pose fusion and multi-view triangulation. Both methods are assessed under varying occlusion levels using a precision-machined grid and an anatomical head phantom, and compared against a gold-standard stereo infrared system. With five cameras and adaptive Kalman filtering, LARK achieves median target registration errors of 0.64 mm for point localization with triangulation and 0.73 mm for trajectory tracking with pose fusion on the machined grid. Camera-subset experiments show graceful degradation in adaptive pose-fusion accuracy as fewer views remain available. With tracking hardware costing under $1,000 USD, LARK provides a low-cost platform for image-guided surgery research. Hardware designs and software are publicly available at this https URL , and datasets at this https URL .

112. 【2610.07511】MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation

链接:https://arxiv.org/abs/2610.07511

作者:Suzannah Wistreich,Stephen Tian,Isabella Huang,Vitor Campagnolo Guizilini,Sergey Zakharov,Katherine Liu,Jiajun Wu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Mobile manipulators, perform dexterous manipulation, deployed in dynamic, unstructured environments, dexterous manipulation tasks

备注:

点击查看摘要

Abstract:Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot's kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA's benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: this https URL

113. 【2610.07464】Protective Perturbations Must Survive the Resize: Scale-Robust Image Immunization against Malicious Editing

链接:https://arxiv.org/abs/2610.07464

作者:Zhongliang Guo,Yan Lin,Yifei Qian,Daizong Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Protective perturbations aim, stop malicious instruction-guided, malicious instruction-guided editing, editor working resolution, Protective perturbations

备注: 16 pages, 9 figures

点击查看摘要

Abstract:Protective perturbations aim to stop malicious instruction-guided editing of personal photos, but they are optimized and evaluated at the editor's working resolution, whereas shared photos have 10 megapixels or more and editors first downscale them by an unknown factor. We model this resize as a frequency-selective channel. In this model, a perturbation computed at the native resolution decays with the downscaling factor and is weak even without a resize, and a perturbation computed at a fixed working resolution protects only a window of scales. The best worst-case protection over an unknown range of scales degrades only logarithmically with the width of the range, and averaging over scales does not reach it. Guided by this analysis, we propose SRIM, which samples a grid of anchor scales covering the whole range, with weights that favor the currently weakest scale, at the cost of standard expectation over transformation. On full-resolution photos of 9 to 30 megapixels and downscaling factors from 2 to 8, SRIM raises the worst-case disruption of FLUX.2-klein edits from 0.192 LPIPS, attained by the strongest published protection, to 0.463. At equal visibility, it roughly doubles the protection. The same protected photos also protect against the 9B model and against FLUX.2-dev, with worst cases of 0.450 and 0.386 against at most 0.184 for published protections, and SRIM leads on InstructPix2Pix as well.

114. 【2610.07460】ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation

链接:https://arxiv.org/abs/2610.07460

作者:Tzu-Hsin Hsieh,Ricardo Marroquim

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:preserving semantic intent, Toggle, Inserting objects, fit local geometry, Toggle Hugging Face

备注: Accepted at NeurIPS 2026. Project page: [this https URL](https://celine-hsieh.github.io/elasticfit/)

点击查看摘要

Abstract:Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce \textbf{ElasticFit}, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic fitting). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8\% to 69.7\% and support success from 48.3\% to 91.7\% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.

Comments:
Accepted at NeurIPS 2026. Project page: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.07460 [cs.CV]

(or
arXiv:2610.07460v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.07460

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Tzu Hsin Hsieh [view email] [v1]
Mon, 5 Oct 2026 22:12:18 UTC (47,540 KB)

Full-text links:
Access Paper:

View a PDF of the paper titled ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation, by Tzu-Hsin Hsieh and 1 other authorsView PDFHTML (experimental)TeX Source

view license

Additional Features

Audio Summary

Current browse context:
cs.CV

prev

|
next

new
|
recent
| 2026-10

Change to browse by:

cs
cs.AI

References Citations

NASA ADSGoogle Scholar
Semantic Scholar

export BibTeX citation
Loading…

BibTeX formatted citation

loading…

Data provided by:

Bookmark

checked="checked"class=“labs-tab-input”>
Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author
Venue
Institution
Topic

    About arXivLabs

arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.

Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)

mathjaxToggle();

    We gratefully acknowledge support from
    our major funders,
    member institutions, ,
    and all contributors.

About

Help

Contact

Subscribe

Copyright

Privacy

Accessibility

Operational Status (opens in new tab)

Major funding support from

115. 【2610.07444】Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?

链接:https://arxiv.org/abs/2610.07444

作者:Aadi Chauhan,Arthur Ilyasov

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:GUI agent decides, GUI agent, auxiliary loss, agent decides, hard-routed action word

备注: 18 pages, 4 figures, 12 tables. Code and per-example logs are available at [this https URL](https://github.com/aadcha/action-conditioned-gui-agent)

点击查看摘要

Abstract:A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard-routed action word, an additive learned embedding, a prepended learned token, and the type written into the prompt. With five seeds, an episode-clustered bootstrap, and seed-level paired tests, the ranking on a mixed stream is clear: the auxiliary loss, the additive embedding, and the prompt word each gain five to seven hit@0.10 points over the baseline, while hard routing and the prepended token are not distinguishable from it. Much of that gain is protection from a preprocessing choice of ours rather than a spatial prior. Our serializer clamps the off-screen touch point AITW records for type events to the origin; that class degrades the baseline's click grounding, and removing it lifts the baseline by nearly seven points, after which no mechanism's hit rate beats it and the intervals exclude a two-point effect, though the auxiliary loss still shortens the average miss; on a stream of taps and swipes none helps. Whether this generalizes beyond one serialization is open. For deployment, the pipeline's margin over the baseline with predicted rather than gold types is not established (+0.016, 95% interval [-0.017, +0.052]), and a wrong type collapses every model conditioned at inference. The prepended token does not help at the shared learning rate, where its rows barely move from initialization; trained ten times faster it reaches the level of the other three, with a margin three seeds do not establish. We also document a silent failure: injecting conditioning through inputs_embeds makes Qwen2-VL fall back to 1-D positions for image tokens, costing nine points.

116. 【2610.07419】Learnable Spectral Activations

链接:https://arxiv.org/abs/2610.07419

作者:Tamir Shor,Or Litany,Alex Bronstein

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Implicit neural representations, spectral structure induced, Implicit neural, input encodings, structure induced

备注:

点击查看摘要

Abstract:Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency access is not the only bottleneck: signals with localized or spatially varying structure require the network to efficiently compose frequencies into multi-harmonic internal responses. We introduce learnable spectral activations (LSA), which replace fixed neuron-level nonlinearities with a residual truncated Fourier series whose harmonic amplitudes are learned during training. LSA does not expand the asymptotic function class. Instead, it changes the factorization of the representation: linear weights select features while activation coefficients control spectral shaping, and the two are updated by separate gradients. Because the activation output is affine in the coefficients given fixed pre-activations, spectral tuning becomes a more direct subproblem compared to architectures where it is entangled with feature selection. Empirically, this factorization concentrates more target-signal energy in the leading eigenmodes of the neural tangent kernel, consistent with improved optimization behavior. Across audio, image, neural radiance field, and neural acoustic field tasks, LSA also improves reconstruction quality.

117. 【2610.07385】Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning

链接:https://arxiv.org/abs/2610.07385

作者:Suguru Onda,Matthew Bailey,Ryan Farrell

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Fine-tuning vision-language models, single downstream checkpoint, vision-language models, typically evaluated, single downstream

备注: 67 pages, including supplementary material

点击查看摘要

Abstract:Fine-tuning vision-language models (VLMs) is typically evaluated at a single downstream checkpoint, obscuring whether a semantic capability was never acquired or emerged earlier and later declined during specialization. We ask how semantic capabilities are acquired, when they peak, how well they transfer, and what remains at deployment. We study these dynamics as a semantic capability trajectory, tracking identity- and attribute-based capabilities over training. We formulate a trajectory-based framework that separates capability acquisition, capability-specific optima, and later specialization, and introduce Structured Semantic Routing (SSR) to study how the representation of supervision shapes what is acquired. Across six pretrained backbones spanning DFN, MetaCLIP, and OpenAI CLIP, we show that fine-tuning can acquire semantic capability beyond the pretrained state, including gains observed on held-out evaluations. Unstructured name-and-attribute supervision produces strong name-and-attribute retrieval with comparatively weak name-free attribute-profile retrieval, whereas SSR yields substantially stronger name-free attribute-profile retrieval and is further strengthened by stochastic name-branch dropout. Different capabilities can peak at different stages, so a checkpoint selected by target class-name retrieval need not coincide with a transferable semantic optimum. Continued optimization can therefore preserve strong target class-name retrieval while reducing previously acquired transferable semantic capability. In a representative diagnostic study, this late specialization is consistent with reduced cross-modal semantic accessibility while substantial image-only class structure remains available.

Comments:
67 pages, including supplementary material

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.07385 [cs.CV]

(or
arXiv:2610.07385v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.07385

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
118. 【2610.07384】WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

链接:https://arxiv.org/abs/2610.07384

作者:Turhan Can Kargin,Piotr Kubaty,Ekaterina Rostovskaya,Izabela Wierzbowska,Bartosz Zieliński,Marcin Przewięźlikowski

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:retrieval problem central, instance retrieval problem, instance retrieval, central to non-invasive, retrieve the correct

备注: 15 pages, 7 figures, 3 tables. Project page: [this https URL](https://wildmatch.gmum.net)

点击查看摘要

Abstract:Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin, or other visual markings. Current approaches either learn global embeddings as a classification problem, requiring many labeled images per individual while largely ignoring local evidence, or apply off-the-shelf, domain-agnostic image matchers. Although such matchers are pretrained on large and diverse image collections, adapting them to wildlife imagery is challenging because available datasets are small and lack correspondence-level annotations. We study weakly supervised adaptation of a pretrained keypoint matcher using only identity labels, without keypoint-level or geometric correspondence ground truth. We mine informative image pairs with the pretrained matcher, derive weak positive and negative supervision from identity agreement, and contrastively fine-tune the matching network to strengthen correspondences for same-identity pairs and suppress them for different identities. Across open-source wildlife re-identification datasets, our approach improves accuracy over off-the-shelf matchers and a state-of-the-art local--global fusion method. Under an open-world protocol with held-out individuals, it learns a transferable correspondence prior rather than memorizing training identities. To our knowledge, this is the first study of matcher-level, identity-supervised adaptation for animal re-identification. Our method enables data-efficient specialization of image matching models to wildlife domains using identity annotations already available in typical monitoring datasets.

119. 【2610.07381】GeoWM: Efficient Direct World Modeling in Explicit Geometry

链接:https://arxiv.org/abs/2610.07381

作者:Mehrdad Noori,Guile Wu,Sam Hosseini,Dongfeng Bai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:essential for autonomous, geometry, Modeling, scene geometry, future

备注:

点击查看摘要

Abstract:Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to predict future images or latent representations of the environment and subsequently recover geometry from these predictions. However, this paradigm does not explicitly model geometric structure and typically relies on recursive rollouts to reach longer prediction horizons, leading to error accumulation and increasing computational cost. To address these limitations, we present GeoWM, a geometry world model that directly forecasts future scene geometry at specified future horizons without recursive rollout. The key idea is to leverage a geometry foundation model to transform observed RGB frames into a geometric history, which conditions a flow-matching transformer to predict the scene geometry at a specified future horizon. We further show that a lightweight camera-motion predictor can accurately estimate the future viewpoint, and that projecting the observed geometry into the predicted viewpoint provides an effective geometric prior for future geometry forecasting. Extensive experiments on four datasets spanning urban driving, aerial flight, and dynamic manipulation demonstrate that GeoWM outperforms the evaluated world models in forecasting depth, camera pose, and 3D scene geometry, while substantially reducing inference time at longer horizons.

120. 【2610.07378】SimCortex v2: Joint Cortical Surface Reconstruction with Near-Zero Collisions and Self-Intersections

链接:https://arxiv.org/abs/2610.07378

作者:Kaveh Moradkhani,Sylvain Bouix

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:magnetic resonance imaging, surface-based neuroanatomical analysis, structural magnetic resonance, Reconstructing cortical, resonance imaging

备注: Submitted to Medical Image Analysis

点击查看摘要

Abstract:Reconstructing cortical WM and pial surfaces from structural magnetic resonance imaging (MRI) is a prerequisite for surface-based neuroanatomical analysis, yet remains challenging because the cortex is thin and tightly folded. Reconstruction methods can produce geometric artifacts such as mesh self-intersections and collisions between cortical surfaces, and although recent deep learning methods have reduced reconstruction time from hours to minutes, these artifacts persist. We propose SimCortex v2, a deep learning framework for simultaneous reconstruction of the left and right WM and pial surfaces from T1-weighted MRI. SimCortex v2 estimates topologically correct initial surfaces from a volumetric segmentation and refines all four jointly using multi-scale stationary velocity fields predicted by a ribbon-conditioned, U-Net-like network. We evaluated SimCortex v2 on 560 cases from 14 cohorts, thirteen of them unseen during training, spanning ages 6-89, healthy and clinical populations, and scanners from three vendors. SimCortex v2 matched the surface-distance accuracy of the strongest baseline (average symmetric surface distance 0.253 mm) while showing no detected inter-surface collision in 92.14% of cases and the lowest self-intersection fraction (0.044%) among learning-based methods, whereas every baseline produced at least one collision in every case. Source code, configuration files, pretrained weights, preprocessed data, and the exact evaluation splits are publicly released.

121. 【2610.07366】Identity-Conditioned Score Fusion for Open-Set Person Re-Identification

链接:https://arxiv.org/abs/2610.07366

作者:Manyi Yao,Jurijs Nazarovs,Eunji Chong,Abhishek Sharma,Rohan Sarkar,Yue Guo,Christian R. Shelton,Amit K. Roy-Chowdhury,Debashish Pal

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:combines complementary cues, Robust person re-identification, body shape, Robust person, combines complementary

备注:

点击查看摘要

Abstract:Robust person re-identification often combines complementary cues such as face, gait, and body shape. While adaptive fusion typically targets query quality, model strength also varies across identities. We introduce identity-conditioned score fusion, a framework that tailors weights to each gallery identity without training. By contrasting intra-identity consistency against cross-identity impostors, it extracts identity-specific profiles that couple with query-conditioned adaptation via a parameter-free rule. This widens the separation between true and false matches while preserving score calibration. Evaluations on three clothes-changing person re-identification benchmarks show that our method consistently outperforms statistical, rank-based, and learned baselines, achieving up to an 8.8% absolute reduction in the false non-identification rate and demonstrating the value of identity-conditioned fusion in open-set person re-identification.

122. 【2610.07355】racking Is Not Permanence: What Video World Models Keep of a Hidden Object

链接:https://arxiv.org/abs/2610.07355

作者:Peng Xie,Amr Alanwar

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:world models track, models track objects, Video world models, models track, world models

备注:

点击查看摘要

Abstract:Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.

123. 【2610.07339】A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care

链接:https://arxiv.org/abs/2610.07339

作者:Junseob Kim,Jade Chng,Ayman Ali,Victor Moas,Yichun Lee,Po-Chun Chin,Sunil Hwang,Rishikesan Kamaleswaran

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Tactical Combat Casualty, Combat Casualty Care, Tactical Combat, Casualty Care, Combat Casualty

备注: 20 pages, 6 figures, 5 tables. Dataset: [this https URL](https://doi.org/10.5281/zenodo.23170287;) code: [this https URL](https://github.com/Kamaleswaran-Lab/TC3-VQA)

点击查看摘要

Abstract:Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.

124. 【2610.07337】Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery

链接:https://arxiv.org/abs/2610.07337

作者:Chen Dai,Ganyu Zou,Nathan Self,Kevin Piper,Ramachandra Rao Seethiraju,Karthik Shyamsunder,Chang-Tien Lu,Naren Ramakrishnan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Removing copyrighted, practical requirement, diffusion model, Grounded Semantic Surgery, Hierarchically Grounded Semantic

备注: Accepted at BMVC 2026

点击查看摘要

Abstract:Removing copyrighted, unsafe, or user-specified concepts from a deployed text-to-image diffusion model is now a practical requirement. Weight-editing methods can suppress fixed targets, but they require per-target retraining and modify the model checkpoint. Training-free methods, on the other hand, are deployment-friendly, but they suffer from text-side routing failures on compositional prompts. In such prompts, the erase target may be invoked through a related class rather than its lexical name, and its modifiers may migrate onto preserved objects. This paper proposes Hierarchically Grounded Semantic Surgery (HGSS), a training-free framework for compositional concept erasure. The framework lifts both the routing signal and the edit operator used by text-side erasure. First, hierarchical span grounding resolves erase-target spans through lexical, taxonomic, and semantic evidence, while guarding against broad-hypernym and compound-head false positives. Second, dynamic attribute binding refines the text conditioning during early denoising via a counterfactual reference and a preserve-aware cross-attention objective, keeping surviving attribute-noun bindings intact. HGSS selectively removes the erase target without updating model weights or adding learned parameters. On SEE, HGSS cuts hierarchical evasion from 29.54 to 10.02 and roughly halves pairwise attribute leakage, achieving the best Neighbor E and AttrP scores among the reported erasure methods. On UnlearnCanvas, HGSS slightly improves the six-metric average over the matched Semantic Surgery baseline, reaching state-of-the-art.

125. 【2610.07326】Localize Any Object in X-Ray Security Scans without Human Annotation

链接:https://arxiv.org/abs/2610.07326

作者:Yaqi Cai,Mingxuan Liu,Lorenzo Vaquero,Ning Wang,Nan Pu,Feng Xue,Elisa Ricci,Nicu Sebe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:automated threat detection, X-ray, Universal object localization, X-ray security inspection, Universal object

备注:

点击查看摘要

Abstract:Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder the direct zero-shot transfer of dense perception foundation models trained on web-scale RGB data. Moreover, annotated X-ray data is scarce and requires expert labeling, limiting both the training of generalizable X-ray native models and the adaptation of RGB foundation models for X-ray data via fine-tuning. Given these challenges, the bright promise of highly generalizable perception models, enabled by data scaling laws in the RGB domain, remains largely out of reach for X-ray inspection. To this end, we introduce LAO-X, a self-supervised adaptation framework that Locates Any Object in X-ray scans using diverse synthesized image--annotation pairs with granularity-aware supervision. LAO-X first designs a saliency-guided X-ray object mining module to separate diverse object instances, which are then used for physics-guided synthesis in the absorbance domain. LAO-X further incorporates an occlusion-controlled curriculum strategy to fine-tune a Segment Anything Model 2 (SAM2) localizer, progressively adapting it to X-ray scans with increasing object counts and overlap levels. Experiments on six X-ray benchmarks show that LAO-X substantially improves category-agnostic localization, achieving 2\% to 23\% mAP gains over SAM2 and X-ray specific baselines in heavily cluttered scenarios, entirely without human-annotated labels.

126. 【2610.07323】ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning

链接:https://arxiv.org/abs/2610.07323

作者:Marsalis Gibson,Claire Tomlin,Shankar Sastry

类目:Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)

关键词:Security evaluation, Latent Adversarial Searches, fixed collection, ATLAS, Security

备注: 9 pages main body, plus 10 additional pages for references and appendix

点击查看摘要

Abstract:Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textit{sets} of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), which is a query-based framework that discovers adversarial input sets for black-box learning systems. ATLAS casts attack generation as an active learning level set estimation problem then combines calibrated approximations with a local-global sampling architecture to find regions of the input space that contain adversarial examples. Once discovered, ATLAS is designed to sample points within these adversarial regions to build adversarial sets that accurately represent the state of robustness of the target model. When applied on toy experiments, we find that ATLAS is able to recover more of the adversarial region under a limited query budget than does previous work. When applied to standard and adversarially trained MNIST, CIFAR, and ImageNet model targets, ATLAS produces better representative attacks than other query-based black-box attacks (NES, SignHunter, BayesOpt). ATLAS represents an automated red-teaming framework that can be used for both analyzing the robustness of learning-based systems under development and continuous auditing to see how the robustness of a system changes over time.

127. 【2610.07269】What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization

链接:https://arxiv.org/abs/2610.07269

作者:Ayesh Abu Lehyeh,Jay Hwasung Jung,Safwan Wshah

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Cross-view geo-localization, image retrieval problem, retrieval problem, matching a ground-level, geo-localization is commonly

备注: Accepted at NeurIPS 2026 Workshop Physical World AI: Geometry, Characteristics, and Multimodal Sensing

点击查看摘要

Abstract:Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree closely across the two views, yet ranking the full pool by description similarity almost never returns the correct tile (0.39% Recall@1). Second, we narrow the pool to ten neighboring tiles, as a coarse prior would do. The same descriptions now become useful: an MLLM judge that scores structural consistency doubles random ranking and matches a strong lexical baseline. It also states which fields of the two descriptions agree and which conflict, which an embedding distance cannot do, and which we see as a step toward interpretable localization. Third, we place the judge on a trained visual retriever. On the queries it ranks wrongly, reranking from images works, while reranking from our descriptions does not (23.5% against 10.7% Recall@1). Scene structure survives the conversion into language, while the fine appearance detail needed to separate nearby places does not. Code and prompts are publicly available at this https URL.

128. 【2610.07243】Hybrid Cross-Modal Attention Network for Early Breast Cancer Detection in Low-Resource Clinical Settings

链接:https://arxiv.org/abs/2610.07243

作者:Simon Hadush Nrea(1),Filimon Gidey Gebremichael(1),Gebrekirstos Hagos Gebrekirstos(2),Yaecob Girmay Gezahegn(1) ((1) Mekelle University, Mekelle, Ethiopia (2) Clinical Oncologist London School of Hygiene and Tropical Medicine London, UK)

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:delayed diagnosis results, limited radiology expertise, cancer-related mortality, mortality among women, women in Sub-Saharan

备注: 5 double pages numbers, conference paper presented at AI4SD 2026 ( [this https URL](https://www.mu.edu.et/index.php/component/content/article/artificial-intelligence-for-sustainable-development-conference-participants?catid=69&Itemid=101) )

点击查看摘要

Abstract:Breast cancer is the leading cause of cancer-related mortality among women in Sub-Saharan Africa, where delayed diagnosis results from limited radiology expertise and fragmented clinical data systems. Although deep learning models have demonstrated strong performance in mammographic analysis, most rely solely on imaging data and are trained on Western populations, limiting their applicability in African healthcare settings. This paper presents a Hybrid Cross-Modal Attention Network (HCMAN) that integrates mammogram images with structured clinical data using transformer-based cross-modal attention mechanisms. The model was developed and validated using a locally collected dataset of 2,560 mammogram images from 1,024 patients across four Ethiopian referral hospitals, with biopsy-confirmed ground truth labels. The proposed framework achieves 97.8% accuracy, 97.2% sensitivity, 98.3% specificity, and an AUC of 0.987, significantly outperforming image-only baselines. The system demonstrates robustness to low-quality images typical of resource-limited settings, with only 3.2% performance degradation compared to 8.7% for image-only models. Cross-modal attention analysis reveals clinically appropriate behavior: higher reliance on clinical features for ambiguous cases such as dense breasts and young patients. The model's lightweight architecture enables deployment on standard hospital workstations (2 seconds inference on CPU). This work advances sustainable, context-aware AI solutions for equitable breast cancer diagnostics in Africa.

129. 【2610.07231】Monocular Navigation Relative to Unknown Spacecraft Using a Transformer-Aided Kalman Filter

链接:https://arxiv.org/abs/2610.07231

作者:Pol Francesch Huc,Simone D'Amico

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Constraint Kalman Filter, Multi-State Constraint Kalman, single monocular camera, work presents, target

备注:

点击查看摘要

Abstract:This work presents a novel learning-based pipeline for pose estimation of unknown spacecraft using only monocular images from a single servicer. The approach combines a transformer-based neural network with a Multi-State Constraint Kalman Filter (MSCKF) to estimate the pose (i.e., position and orientation of the target spacecraft relative to the camera) throughout rendezvous and proximity operations. Unlike existing vision-based methods that require prior knowledge of the target shape or inertia properties, rely on additional sensing modalities such as depth, lidar, or stereo, or only recover translation up to scale, the proposed pipeline generalizes to previously unseen spacecraft using a single monocular camera. The transformer network estimates the odometry, the change in pose between images up to scale, from SuperPoint features matched by LightGlue. The MSCKF uses these pseudo-measurements along with an orbit and attitude kinematics model to estimate the pose of the target. In particular, the relative orbit elements, the target's attitude with respect to the servicer's camera, and the associated angular velocity are estimated directly by the filter. Given the monocular approach and short distance to the target, the full observability of the range to the target is recovered via attitude maneuvers by the servicer. The method is trained and evaluated on a re-rendered high-resolution version of the SPE3R dataset, which includes synthetic images of 103 spacecraft. Eleven of these spacecraft are held out during training to evaluate the generalization to unseen targets. Monte Carlo simulations are then used to evaluate the navigation pipeline on rendered trajectories of the held out spacecraft. The results demonstrate that learned vision pipelines as a front-end for Kalman filters provide median errors of 3.7° in attitude and 2.2% of range in ROE when navigating about unknown targets.

130. 【2610.07223】Deep Learning Based Illegal Bowling Action Detection

链接:https://arxiv.org/abs/2610.07223

作者:Debopom Sutradhar,Niful Islam,Sudipto Mondal,Tasmima Hossain Jamim,Jubaer Muhammad Shufol,Swakkhar Shatabda

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:strict rule set, illegal bowling actions, illegal bowling, bowling actions, gentleman game

备注:

点击查看摘要

Abstract:Cricket, often referred to as the "gentleman's game," adheres to a strict rule set for both batsmen and bowlers, where each delivery can significantly impact the match outcome. Detecting illegal bowling actions is crucial for maintaining fair play, yet it remains challenging for umpires to monitor in real time. Existing sensor-based solutions have limitations in live match scenarios, making real-time assessment difficult. This paper proposes a computer vision-based deep learning solution to detect illegal bowling actions in live cricket matches. To develop and evaluate our approach, we compiled a dataset of 62 videos featuring 11 male bowlers, capturing both legal and illegal bowling actions from multiple angles-front, back, and side. However, the dataset predominantly comprises right-handed bowlers with conventional actions. The proposed system identifies two key frames, the shoulder frame and the release frame from video footage of a bowler's delivery and analyzes the change in the bowling arm's angle between these frames. If the angle difference exceeds a predefined threshold (e.g., 15 degrees), the delivery is flagged as potentially illegal. We evaluated the system on a custom dataset and achieved a high true positive rate, suggesting the system's potential effectiveness in real-time match settings. However, further research is required to validate the system across diverse environmental conditions and larger datasets to ensure generalizability and robustness in various live match scenarios. To the best of our knowledge, this is the first AI-based computer vision method for detecting illegal bowling actions in cricket.

131. 【2610.07217】RoboCap: A New Platform for Egocentric Robot Learning

链接:https://arxiv.org/abs/2610.07217

作者:Grounded Superintelligence,BitRobot

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:scaling robot learning, robot learning, scarce today, promise for scaling, scaling robot

备注:

点击查看摘要

Abstract:Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrating ergonomic hardware with centimeter-precise 3D algorithms, at a precision that has not been publicly demonstrated. To address this gap, we introduce RoboCap, a 250\,g six-camera dual-IMU hat designed for in-the-wild egocentric data capture, and the Grounded API, a suite of device-agnostic 3D algorithms tuned for RoboCap. In this report, we demonstrate how hardware, calibration, and 3D algorithms interact to achieve state-of-the-art performance on the public benchmarks: our SLAM across diverse settings and rigs, our depth estimation on egocentric settings, and our hand tracking when adapted to third-party devices.

132. 【2610.07206】Energy-Conditioned Noise Schedule and Whitening for Spectral Diffusion

链接:https://arxiv.org/abs/2610.07206

作者:Bata Vasic,Bane Vasic

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET)

关键词:transform-domain diffusion models, energy-adaptive noise scheduling, paper introduces, introduces an energy-adaptive, strategy for transform-domain

备注: 7 pages, 5 figures, 2 tables, Manuscript submitted for publication in Elsevier Pattern Recognition Letters

点击查看摘要

Abstract:This paper introduces an energy-adaptive noise scheduling and whitening strategy for transform-domain diffusion models. Existing spectral diffusion methods account for the non-uniform statistics of transform coefficients through coefficient scaling, normalization, or frequency prioritization, while the forward diffusion noise schedule remains largely independent of the underlying spectral-energy distribution. We investigate whether the temporal evolution of the forward diffusion process should also follow the spectral organization of natural images. The proposed formulation combines global spectral whitening with energy-conditioned noise allocation that jointly modulates the injected noise according to the energy of individual transform coefficients and an image-dependent energy path over diffusion time. The resulting forward process preserves Gaussian transitions with closed-form marginals and remains compatible with standard DDPM and DDIM procedures without modifying the diffusion architecture. Experiments on CIFAR-10 demonstrate the contribution of the proposed energy-conditioned noise schedule and spectral whitening, reducing Fréchet Inception Distance from 142.48 for a compact DCTdiff U-Net variant to 100.45.

133. 【2610.07175】CALR: Continuous Anchored Latent Reasoning via Render-of-Thought Compression

链接:https://arxiv.org/abs/2610.07175

作者:Zhaoyang Wei,Bowen Jiang,Yanchao Hao,Wenchao Ding,Zheng Wei,Shaocheng Wu,Zhenjun Han,Jianbin Jiao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual latent reasoning, textual reasoning overhead, reducing textual reasoning, Visual latent, reasoning compresses rendered

备注:

点击查看摘要

Abstract:Visual latent reasoning compresses rendered derivations into compact intermediate states, reducing textual reasoning overhead. Existing approaches differ in how they represent these states: continuous methods avoid vocabulary constraints, whereas discrete methods improve accuracy through quantization into a finite codebook. Our analysis of representative continuous and discrete systems identifies two functional requirements: answers must rely on latent states, and those states must carry valid, problem-specific reasoning. Continuous latents influence answers despite collapsed reasoning content, whereas discrete latents retain recoverable intermediate reasoning that answer prediction largely bypasses. To address these challenges, we propose Continuous Anchored Latent Reasoning (CALR), which connects latent formation with answer use through functional anchoring. With reference latents from information-balanced compression, CALR couples latent-mediated answer supervision with derivation-level semantic anchoring: the former routes answer supervision through intermediate states, while the latter grounds their decoded content in problem-specific derivations. A parallel-to-autoregressive curriculum develops sequential reasoning by conditioning subsequent latent blocks on generated prefixes. Evaluations on five mathematical reasoning benchmarks across model families show substantial accuracy gains. Under matched budgets, CALR gains 26.0 percentage points over a comparable continuous latent reasoning method. Further analyses show that its latents support answer prediction and carry problem-specific intermediate reasoning.

134. 【2610.07127】PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

链接:https://arxiv.org/abs/2610.07127

作者:Dheeraj Varghese,Anna Vettoruzzo,Walter Simoncini,Michelle Lorena Acevedo Callejas,Mohammad Mahdi Derakhshani,Kristof Meding,Joaquin Vanschoren,Cees G. M. Snoek

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:extended time horizons, acting competently, time horizons, evaluations largely overlook, advances in multimodal

备注:

点击查看摘要

Abstract:Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and this http URL. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.

135. 【2610.07117】R2RI: A Multi-View Event and RGB Dataset for Robot-to-Robot Interaction

链接:https://arxiv.org/abs/2610.07117

作者:Gabriele Magrini,Riccardo Catalini,Federico Becattini,Guido Borghi,Pietro Pala,Roberto Vezzani,Lorenzo Seidenari

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:collaborative systems, human-robot coexistence, autonomous agents, fundamental challenge, broad implications

备注:

点击查看摘要

Abstract:Understanding and modeling interactions between autonomous agents is a fundamental challenge in robotics, with broad implications for collaborative systems, social robotics, and human-robot coexistence. Although the study of robot interactions has emerged as a compelling research direction, progress has been severely hampered by the absence of large-scale benchmarks. In this paper, we introduce Robot-to-Robot Interaction (R2RI), the first dataset specifically designed to address the Robot-Robot Interaction (RRI) task. R2RI consists of different humanoid robots and realistic interactions modeled on real human social behaviors. Complementary viewpoints are available, \textit{i.e.}, an egocentric perspective from each robot's onboard sensors, and an exocentric perspective from external fixed cameras, thus enabling rich spatial and contextual understanding of the interaction dynamics. The dataset comprises more than $6.5$M frames and $\approx5000$ videos at $120$ fps, including Event and RGB domains. We investigate pros and cons of each domain, comparing state-of-the-art approaches for a number of key sensing and interaction based tasks. We publicly release the dataset and its annotations for all tasks and modalities at this https URL.

136. 【2610.07114】Sample-Optimal Estimation of the Fréchet Inception Distance

链接:https://arxiv.org/abs/2610.07114

作者:Ziyun Chen,Jerry Li,Kevin Tian,Yusong Zhu

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST); Machine Learning (stat.ML)

关键词:Fréchet Inception Distance, evaluate generative models, Fréchet Inception, Theta, frac

备注: Our code is available at [this https URL](https://github.com/zys996/sample-optimal-fid-estimation)

点击查看摘要

Abstract:The Fréchet Inception Distance (FID) is widely used to evaluate generative models, but its empirical plug-in estimator suffers from finite-sample bias [BSAG18, CF20]. We study the sample complexity $n$ of estimating FID to error $\epsilon$ between $d$-dimensional Gaussians with bounded mean distance and covariances, when one distribution is known. Our contributions are threefold. (1) We establish tight finite-sample $\Theta(\frac{d^2}{n})$ bias and $\Theta(\frac{d}{n} + \frac {d^2} {n^2})$ variance bounds for the empirical plug-in estimator, establishing a $\gtrsim d^2$ sample complexity. (2) To debias the empirical plug-in estimator, we generalize the ${\rm FID}_\infty$ estimator of [CF20] to extrapolation methods of arbitrary order $k$. We further prove tight bias and variance bounds of $\Theta(\frac{d^{k + 2}}{n^{k + 1}})$ and $\Theta(\frac d n + \frac{d^2}{n^2})$ for any order-$k$ extrapolation under our framework. (3) We introduce Relative Taylor Debiasing (RTD), a new, computationally efficient FID estimation algorithm using debiasing techniques inspired by U-statistics. We show that RTD achieves an $O(\frac d {\epsilon^2})$ sample complexity, and prove that this is optimal. We provide a complementary empirical evaluation of our new estimators. Our experiments on synthetic Gaussians validate the predicted residual bias and support the tightness of our bounds. On ImageNet with Inception embeddings, RTD achieves the lowest mean estimation error at the standard 50K sample budget, while our second-order variance-aware extrapolation estimator (VALE$_2$) uses only 10K samples to achieve accuracy comparable to FID$_\infty$ at 50K samples.

137. 【2610.07110】MoonGS: High-quality Representation of the Lunar Surface via Gaussian Splatting Using Robust Depth Features from Image Pairs

链接:https://arxiv.org/abs/2610.07110

作者:Yun Jiang,Bo Zheng,Yingying Zhang,Xueming Xiao,Tao Hu,Hutao Cui,Zhiguo Meng,Ke Gao,Yang Gao,Meibao Yao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous lunar exploration, overlap is insufficient, surface textures, textures are weak, volume is limited

备注:

点击查看摘要

Abstract:High-quality 3D reconstruction of lunar terrain from sparse rover images is indispensable for autonomous lunar exploration, but remains challenging because viewpoint overlap is insufficient, surface textures are weak, and data volume is limited. We propose MoonGS, the first feed-forward 3D Gaussian Splatting framework tailored to lunar scenes. Given only two input images, MoonGS predicts pixel-aligned Gaussian primitives in a single forward pass and renders photorealistic novel views without any per-scene optimization. MoonGS (i) adopts an adaptable backbone design that seamlessly integrates advanced vision foundation models to extract robust depth features; (ii) integrates semantic priors in two manners: merging semantic cues with visual features to refine Gaussian parameter estimation, and adopting a semantic ranking loss that regularizes background depth; and (iii) employs an entropy-guided heuristic resampling strategy to augment sparse observations by selecting the most informative distant viewpoints with negligible overhead. Experiments on the LuSNAR benchmark and our synthetic weak-texture MoonBlender dataset show that MoonGS surpasses state-of-the-art feed-forward NeRF/3DGS baselines by +4.9 dB PSNR, +0.29 SSIM, and 40\% lower LPIPS while maintaining sub-second inference. Furthermore, we validate the broad applicability of our framework by demonstrating that it effectively leverages state-of-the-art backbones, including VGGT, to significantly boost performance. Qualitative evaluations on Chang'e mission imagery also show the best visual quality among compared methods, indicating robustness on real lunar data. The source code and dataset are publicly available at this https URL.

138. 【2610.07091】Smart Content Ingestion for Generative AI Workloads

链接:https://arxiv.org/abs/2610.07091

作者:Abbas Raza Ali,Muhammad Ajmal Siddiqui,Moona Zahid

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:conventional machine learning, machine learning, progressively changed, changed where intelligence, intelligence resides

备注:

点击查看摘要

Abstract:The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any single task: one foundation model serves open-ended downstream tasks, and the generality gained on the model side is matched by heterogeneity on the data side, because enterprise knowledge is authored in the formats people use (PDF, presentations, spreadsheets, scanned documents, forms, tables, diagrams and mixed-layout files) that carry textual, visual, geometric and structural information at once. A language model or retriever cannot reason reliably over information misrepresented at this interface, so content extraction becomes a lifecycle stage in its own right whose errors no downstream retriever or re-ranker can repair. This paper presents a production-ready content-extraction system that makes this stage explicit, configurable, and measurable. The system incorporates selective OCR routing, a scarcity-first curation engine with a reference-based extraction scorer that measures character, word, and table-structure accuracy, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator that generates grounded questions from every page and reports Hit@k, mean reciprocal rank, and latency. On a 180-document corpus the best extractor scores 97.4 of 100 (character error rate 0.13%, table similarity 0.995) and the chunker reaches hit@1 of 68.6%, hit@10 of 92.8% and MRR 0.77 over 25,050 generated questions. We distil three design principles (structure before semantics, never mutate what you measure, budget your labels) and position measured content extraction as the perception layer of enterprise agentic systems.

139. 【2610.07087】A Data-Centric Review of Plant Disease Datasets: Taxonomy, Critical Analysis, Environmental Variability, and Implications for Precision Agriculture

链接:https://arxiv.org/abs/2610.07087

作者:Aamir Hilal,Shabir Ahmad Sofi,Neeraj Goel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:artificial intelligence, rapid advances, advances in artificial, reliable real-world plant, plant disease detection

备注:

点击查看摘要

Abstract:Despite rapid advances in artificial intelligence, reliable real-world plant disease detection remains a persistent challenge. Visual and deep learning approaches have shown promising results, but their deployment under field conditions remains limited. A key bottleneck is the reliance on laboratory-generated datasets that lack environmental diversity, realistic backgrounds, and balanced class distributions, resulting in poor generalization. In contrast, datasets collected directly from agricultural environments capture natural variability and better reflect challenges faced by farmers across regions. This review presents a critical analysis of visual and deep learning approaches for plant disease detection, with emphasis on plant disease datasets. It establishes a taxonomy based on acquisition setting, accessibility, plant diversity, disease composition, class structure, and imbalance severity, and examines their implications for model generalization and real-world deployment. A comparative analysis of laboratory and real-field datasets identifies critical gaps that hinder disease detection. The review further analyzes how multi-level dataset imbalance, including intra-class, inter-crop, and cross-dataset imbalance, and limited environmental variability affect model performance and robustness, an area insufficiently examined in existing surveys. Beyond image-based approaches, it highlights the importance of integrating environmental parameters such as temperature, humidity, and leaf wetness with image data to improve prediction under dynamic field conditions. Finally, the review identifies key challenges, research gaps, and future directions concerning dataset construction, environmental variability, structural imbalance, standardization, and multimodal disease monitoring. It provides a foundation for developing next-generation multimodal frameworks for precision agriculture.

140. 【2610.07083】Graph-Based Recognition of Simulated Train-Driver States From Facial and Upper-Body Keypoints

链接:https://arxiv.org/abs/2610.07083

作者:Olivia Nocentini,Marta Lagomarsino,Gokhan Solak,Younggeol Cho,Qiyi Tong,Sara Zeynalpour,Marta Lorenzini,Alessandro Ledda,Arash Ajoudani

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Driver fatigue poses, basic alertness checks, dead-man switch offering, switch offering limited, Driver fatigue

备注: 11 pages,5 figurees

点击查看摘要

Abstract:Driver fatigue poses a significant challenge to railway safety, with traditional systems like the dead-man switch offering limited and basic alertness checks. This study presents a vision-based monitoring system that relies solely on a single front-facing RGB camera and a graph neural network to classify simulated train-driver states into alert, not-alert, and an emergency class comprising acted emergency-like behaviours. To optimize input representations for the model, an ablation study was performed, comparing three feature configurations: skeletal-only, facial-only, and a combination of both. Experimental results show that combining facial and skeletal features yields the highest accuracy (81%) for the three-class model under the light condition, outperforming models that use only facial or skeletal features. Furthermore, the combination of facial and skeletal features achieves 99% accuracy in the alert/not alert classification in light condition. Additionally, we introduced a controlled RGB video dataset containing alert, not alert, and acted emergency-like behaviours recorded under three illumination conditions. These contributions represent a step toward passive and non-contact train-driver state recognition based on facial and upper-body dynamics.

141. 【2610.07072】On Color Alignment in VAE Latent Spaces and Its Applications

链接:https://arxiv.org/abs/2610.07072

作者:Julian D. Santamaria,Kai Wang,Jesús Malo,Javier Vazquez-Corral,Alexandra Gómez-Villa

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Variational autoencoders, part of modern, key part, VAE latent space, generate images

备注:

点击查看摘要

Abstract:Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance axis and two opponent-color axes. Color should therefore be expected to emerge as a distinct factor in the VAE latent space. Yet how these latent spaces represent color remains largely unexplored. In this work, we show that the VAEs of text-to-image models share a color subspace aligned with brightness and opponent-colors. Through a linear approximation of the encoder and targeted latent steering, we find this subspace consistently across a broad range of VAEs, from SD1.5 to FLUX.2 and Z-Image. Building on this characterization, we propose three applications: ColorTuning, which achieves state-of-the-art in precise numerical color generation on the fine-grained CSS3/X11 system of GenColorBench, saturation control, to adjust the global chromatic intensity, and color transfer, to change the palette to match a reference. The code and models are publicly available at this https URL

142. 【2610.07071】A BEMD-Based Quaternion Filtering Approach Sharp-to-Soft Kernel CT Image Conversion

链接:https://arxiv.org/abs/2610.07071

作者:Mahmoud Nasr,Jan K. Argasinski,Krzysztof Brzostowski,Adam Piorkowski

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Quaternion Bilateral Filtering, improve spatial resolution, Empirical Mode Decomposition, Bidimensional Empirical Mode, sharp kernels improve

备注:

点击查看摘要

Abstract:The quality of computed tomography (CT) images is significantly affected by the selection of reconstruction kernels: sharp kernels improve spatial resolution but increase noise, whereas soft kernels diminish noise at the expense of edge clarity. This study presents an innovative enhancement framework utilising Bidimensional Empirical Mode Decomposition in conjunction with Quaternion Bilateral Filtering (BEMD--QBF) to convert sharp-kernel CT images into representations resembling soft-kernels, while maintaining critical anatomical structures. The technique disaggregates each image into intrinsic mode functions via BEMD and analyzes them inside a cohesive quaternion framework to attain efficient noise reduction and structural integrity. The proposed methodology is evaluated using several reconstruction kernels (B50, B46, B41, B36, B35, B31) and compared with recognised filtering strategies, including Non-Local Means, Anisotropic Diffusion, Bilateral Filtering, and Quaternion Bilateral Filtering. Quantitative evaluations of the Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR) indicate that BEMD-QBF consistently attains superior structural fidelity and competitive noise reduction across all evaluated kernels. The results underscore the efficacy of the proposed strategy as a viable approach to enhancing post-reconstruction CT images, yielding superior image quality without requiring access to raw projection data.

143. 【2610.07067】Decomposition-Guided Curvelet Thresholding for Sharp-to-Soft CT Kernel Conversion

链接:https://arxiv.org/abs/2610.07067

作者:Mahmoud Nasr,Jan K. Argasinski,Krzysztof Brzostowski,Adam Piorkowski

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:essential structural elements, maintaining essential structural, focused on improving, crucial task, maintaining essential

备注:

点击查看摘要

Abstract:Image denoising is a crucial task in image processing, focused on improving image quality by minimizing noise while maintaining essential structural elements. This study presents a hybrid denoising framework that combines several decomposition techniques, including empirical mode decomposition (EMD), variational mode decomposition (VMD), multichannel EMD (MEMD), and bidimensional EMD (BEMD), with curvelet transform thresholding. Each decomposition mode undergoes processing through both soft and hard thresholding, and the denoised modes are combined to rebuild the final image. Comprehensive evaluations of standard CT image datasets reconstructed with various kernels (B50, B46, B41, B36) reveal substantial enhancements in denoising efficacy. VMD consistently achieves the highest peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), signifying exceptional noise reduction and feature preservation. The study analyses the trade-offs between soft and hard thresholding: soft thresholding maintains intricate visual details, whilst harsh thresholding provides enhanced noise reduction. The suggested method surpasses traditional techniques in both reference and non-reference quality criteria, indicating its potential for broader application in medical imaging and future incorporation with adaptive thresholding algorithms.

144. 【2610.07059】Crop Yield Prediction for Punjab, Pakistan: A Tree-Ensemble and Leaf-Health Prototype, and What Random Validation Hides

链接:https://arxiv.org/abs/2610.07059

作者:Amina Asif,Qurat ul ain Asif,Noor Bakhat Asif

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:small agricultural tables, reported accuracy fail, make reported accuracy, storage and imports, combines Random Forest

备注:

点击查看摘要

Abstract:Yield forecasts help planners and farmers decide on inputs, storage and imports, but small agricultural tables can make reported accuracy fail on a new season. We built a crop-yield prototype for Punjab, Pakistan that combines Random Forest, XGBoost, support vector regression and a Ridge-stacked ensemble with a MobileNetV2 leaf-health classifier, and deployed it as a web application with SHAP explanations. A random 80/20 split of a merged Kaggle-derived table (414 rows) gives the ensemble an R2 of 0.991. An audit showed that the table contains only 46 independent observations: a join with nine temperature records per year copied every crop-year nine times. Holding out whole years takes XGBoost on the same rows from R2 = 0.994 to -0.20. On deduplicated data, and on a longer FAOSTAT table (1990-2024, 70 observations), a per-crop linear trend (leave-one-year-out RMSE 0.29 Ton/Ha) beats every model not given the year (0.84-1.06). Neither pesticide use nor national temperature change explains the trend residuals. An apparent pesticide gain in the Kaggle table disappears on FAOSTAT, where the pesticide series is mostly imputed and the two releases disagree. An independent district-level wheat panel (36 districts, 13 seasons) shows that most variation in Punjab is spatial and that a district mean with a common trend matches the learned models. The leaf classifier reaches 99.87% accuracy on held-out PlantVillage images and recalls 96.4% of 336 unseen diseased leaves after near-duplicates were removed. However, no leaf image is paired with a yield record, so the health score used in the yield models had to be constructed and adds nothing. We report these negative findings together with the prototype.

145. 【2610.07031】Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update

链接:https://arxiv.org/abs/2610.07031

作者:Sitian Shen,Jiuming Liu,Mengmeng Liu,Yian Wang,Michael Ying Yang,Francesco Nex,Hao Cheng,Daniele De Martini,Ayush Tewari,Per Ola Kristensson

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent video world, Recent video, witnessed the paradigm, paradigm shift, shift from single-agent

备注: The first three authors contributed equally, and their order was determined by drawing lots. Project Lead: Jiuming Liu. Corresponding Author: Ayush Tewari. Project page: [this https URL](https://liujiuming123.github.io/Artemis/)

点击查看摘要

Abstract:Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics. To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues. Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents and multi-camera setting.

146. 【2610.07025】WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI

链接:https://arxiv.org/abs/2610.07025

作者:Gabriel Lee Jun Rong,Shanhong Liu,Pai Chet Ng,Konstantinos N. Plataniotis,Jamal Seyedmohammadi,S. Mohammad Sheikholeslami

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:channel state information, WiFi channel state, directly identifying individual, identifying individual joints, state information

备注:

点击查看摘要

Abstract:Multi-person 3D pose estimation with WiFi channel state information (CSI) is challenging because reflections from different people overlap without directly identifying individual joints. Existing masked embedding objectives capture wireless relationships without explicit pose supervision, while structured decoders can retain coordinate errors. We propose WiSPER, a two-stage framework combining pose-aware predictive pretraining with conditional residual flow refinement. Pose-Aware Masked Embedding Learning (PAMEL) couples masked latent prediction with auxiliary pose-set supervision on the same CSI context, guiding the encoder toward joint localization from partial observations. Residual Flow refinement with Transformer (ReFT) generates a set of pose candidates to accommodate a variable number of people and refines each candidate through a conditional flow guided by its coarse coordinates and per-joint decoder features. Both stages use paired CSI and pose annotations during training, while inference requires only CSI. Experiments on the PiW3D dataset show that WiSPER achieves an overall mean per-joint position error of 63.72 mm, a 40.0% reduction relative to WiFi-JEPA. For experiments with two and three people, WiSPER reduces MPJPE by 42.1% and 38.1%, respectively. Pose-supervised pretraining configurations obtain lower errors than CSI-only JEPA, and enabling the trained residual refiner reduces overall MPJPE by 13.8-15.6% across the evaluated configurations.

147. 【2610.07016】Anchor and Adapt: Asymmetric Prompt Adaptation for Few-Shot Industrial Anomaly Detection

链接:https://arxiv.org/abs/2610.07016

作者:Mengyang Zhao,Teng Fu,Haiyang Yu,Ke Niu,Bin Li,Xiangyang Xue

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:direct defect supervision, few-shot industrial anomaly, target images provide, defect supervision, few-shot industrial

备注:

点击查看摘要

Abstract:In few-shot industrial anomaly detection, the few normal target images provide no direct defect supervision, making anomaly prompts difficult to learn from these samples alone. Some vision-language methods therefore use manually specified descriptions to supply explicit anomaly semantics. However, constructing these descriptions requires product-specific effort, and their effectiveness depends on prompt selection. We propose Anchor and Adapt, a two-stage prompt learning framework that separates the acquisition of anomaly semantics from adaptation to target normal appearance. Stage I learns transferable normal and abnormal anchors from annotated auxiliary data. Stage II keeps these anchors fixed and adapts an additional normal branch using the few target normal samples. The inherited and adapted normal branches jointly characterize target normality, with text-anchor regularization encouraging consistency with the generic normal prior and separation from the abnormal anchors. This design retains learned anomaly knowledge while reducing dependence on category-specific anomaly templates, without requiring synthetic anomaly generation. Cross-dataset experiments between MVTec-AD and VisA under 1-, 2-, and 4-shot settings demonstrate competitive detection and localization performance. Controlled ablations assess the roles of transferred anchors, asymmetric adaptation, dual-normal representations, and anchor regularization.

148. 【2610.07014】DTFormer: Text-Guided Semantic Alignment for RGB-D Segmentation

链接:https://arxiv.org/abs/2610.07014

作者:Ziang Wei,Yinlong Liu,Yan Xia,Alois Knoll,Hu Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:RGB and Depth, lacking direct high-level, made notable progress, fusing RGB, high-level semantic constraints

备注: This is submitted to IEEE Journal

点击查看摘要

Abstract:RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel-level supervision, lacking direct high-level semantic constraints. This raises a central question-can external knowledge such as language priors inject stronger semantic discriminability into mainstream RGB-D segmentation models. We present DTFormer, a novel tri-modal (RGB-D-Text) semantic segmentation framework. At its core is Text-guided Semantic Alignment Module (TSAM) that first encodes textual cues into a set of semantic prototypes and then explicitly aligns multi-modal RGB-D features with these prototypes at multiple encoder and decoder layers. This design imposes strong semantic regularization on representation learning, guiding the network toward more discriminative features. Extensive experiments on multiple benchmarks show that DTFormer delivers consistent gains while remaining simple and efficient. Our results demonstrate that explicit semantic alignment offers an effective and practical route to improving RGB-D semantic segmentation. The code will be released upon acceptance.

149. 【2610.07008】Learning to Curate What You Generate for Generalizable Few-Shot Class-Incremental Learning

链接:https://arxiv.org/abs/2610.07008

作者:Junhui Yin,Yuchen Yang,Yilin Yin,Shuai Na,Haoran Xi,Jianhua Yang,Muyi Sun,Man Zhang,Shengfeng He

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Few-shot class-incremental learning, Few-shot class-incremental, preserving prior knowledge, class-incremental learning, limited annotations

备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Few-shot class-incremental learning (FSCIL) aims to learn novel classes from limited annotations while preserving prior knowledge. Existing methods typically assume a sufficiently large base session, but this assumption fails when both base and incremental data are scarce, leading to weak initial representations, semantic drift, and unstable boundaries. We study this underexplored yet realistic setting, termed Generalizable FSCIL (G-FSCIL), where the base session itself contains only a few classes. Although synthetic data can alleviate supervision scarcity, naively mixing generated samples often introduces semantic noise and exacerbates old-new boundary conflicts. To address this, we propose a framework that curates trustworthy synthetic knowledge for stable G-FSCIL. Specifically, we first construct class-specific synthetic candidate pools using a frozen latent diffusion model, where class inversion is performed at the first observation and the resulting condition embeddings are reused for on-demand generation. Building on these candidates, we learn a knowledge curation strategy that selects samples with both semantic consistency and visual diversity, and distill this process into a transferable selection policy during the base session, which is then reused without further optimization. Leveraging the curated synthetic data, we further design a boundary-stable incremental adaptation scheme, including synthetic-informed prototype initialization and bidirectional boundary calibration to mitigate old-new conflicts. Extensive experiments demonstrate that our method consistently outperforms existing FSCIL baselines, with reduced forgetting and improved balance between old and new classes. Code is available at this https URL.

150. 【2610.07002】Should We Skip Diffusion?

链接:https://arxiv.org/abs/2610.07002

作者:Yiping Ji,James Martens,Simon Lucey

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:Decoupled Diffusion Transformer, Diffusion models learn, Diffusion Transformer, Diffusion models, Decoupled Diffusion

备注:

点击查看摘要

Abstract:Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides features that guide a velocity decoder in denoising. To enable effective denoising at all noise levels, these features must capture both high-level abstract structures and low-level details. However, skip/residual connections in the encoder allow shallow features to bypass successive transformations, which may limit progressive abstraction, or at least make it difficult to disentangle different levels of abstraction. We propose DDT-RFE, which removes the residual connections around the Self-Attention and MLP operations in each encoder block while maintaining stable training. To retain the information that abstraction discards but that the decoder still needs, we fuse the input patch embedding with intermediate and final encoder features to form the encoder output. The decoder thus has access to information from multiple encoder depths, while each encoder block is able to learn more abstract representations. DDT-RFE achieves overall improvements over DDT across visual understanding tasks, including image classification, semantic segmentation, object discovery, and semantic correspondence, while using fewer encoder blocks. It also achieves a lower FID for image generation on ImageNet.

151. 【2610.06991】State-Aware Interaction MIL for Rare Joint Molecular Phenotype Prediction in Colorectal Cancer and Lung Adenocarcinoma

链接:https://arxiv.org/abs/2610.06991

作者:Dasari Naga Raju,Tripti Bameta

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:small joint-positive populations, overlapping histological features, alternative molecular states, State-Aware Interaction MIL, complicated by small

备注: Accepted at the NeurIPS 2026 Workshop on AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models

点击查看摘要

Abstract:Joint molecular phenotype prediction is complicated by small joint-positive populations and overlapping histological features across alternative molecular states. Existing computational pathology approaches typically predict biomarkers independently or formulate the joint-positive phenotype as a binary endpoint. Independent prediction does not model interactions between biomarker-specific histological representations, whereas binary joint prediction collapses the double-negative and two single-positive configurations into a single negative class. We propose State-Aware Interaction MIL, a weakly supervised method that preserves biomarker-specific histological representations, models their interaction, and supervises the complete four-state molecular configuration. We evaluate the proposed approach for joint BRAF+/MSI+ prediction in colorectal cancer and EGFR+/TP53+ prediction in lung adenocarcinoma using frozen UNI2-h and CONCH pathology foundation-model representations. With UNI2-h, State-Aware Interaction MIL achieved an average precision of 0.5566 in colorectal cancer (joint-positive prevalence 6.8%) compared with 0.5161 for NaiveMTL, and 0.2784 in lung adenocarcinoma (joint-positive prevalence 8.6%) compared with 0.2525 for IndependentPair. With CONCH, State-Aware achieved an average precision of 0.4410 compared with 0.3932 for DirectJoint in colorectal cancer and 0.1659 compared with 0.1226 for DirectJoint in lung adenocarcinoma. These results indicate that pathology foundation-model representations contain predictive information for rare joint molecular phenotypes and that preserving biomarker-specific representations within a structured molecular-state formulation can improve prediction of these phenotypes from histopathology.

152. 【2610.06978】Hierarchy-GBP: Accelerating Factor Graph Inference via Abstraction and Recovery

链接:https://arxiv.org/abs/2610.06978

作者:Yuzhou Cheng,Tom Yates,Ignacio Alzugaray,Danyal Akarca,Pedro A. M. Mediano,Andrew J. Davison

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)

关键词:Gaussian Belief Propagation, Gaussian Belief, distributed inference algorithm, Belief Propagation, scalable spatial intelligence

备注: 33 pages, 10 figures, including appendices

点击查看摘要

Abstract:Gaussian Belief Propagation (GBP) is a distributed inference algorithm that passes messages in graphical models, making it attractive for scalable spatial intelligence. However, we find GBP most effective locally: it rapidly smooths message errors that vary sharply between neighbor variables, but corrects global errors across distant graph regions incrementally through long-range message propagations. We propose Hierarchy-GBP (H-GBP), an iterative, two-stage framework that accelerates GBP by first solving these global errors with a coarse graph approximation (abstraction) and projecting the results back to the original graph (recovery), then refining the remaining local errors with GBP. We prove H-GBP convergence to the optimum by deriving the combined matrix operator of our abstraction and recovery steps and analyzing its spectral radius. Experiments on linear sparse graphs show that H-GBP converges fundamentally faster than standard GBP. Moreover, we validate H-GBP on two important spatial problems: Pose Graph Optimization (PGO) and Bundle Adjustment (BA). H-GBP markedly accelerates large-scale PGO and achieves state-of-the-art runtime across all tested BA scales.

153. 【2610.06977】Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

链接:https://arxiv.org/abs/2610.06977

作者:Xiaojun Jia,Simeng Qin,Yiming Li,Jie Liao,Sensen Gao,Ke Ma,Yang Liu,Xiaochun Cao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Multimodal large language, large language models, Multimodal large, open-source surrogate models, language models

备注:

点击查看摘要

Abstract:Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS] embeddings. However, such coarse alignment insufficiently exploits patch-level visual structures, limiting transferability across heterogeneous closed-source MLLMs. We propose IAU-FOA, a visual-invariance-augmented feature optimal alignment attack with adaptive unbalanced transport, to improve targeted transferability against closed-source MLLMs. IAU-FOA aligns adversarial and target samples at both global and local levels: a cosine-based objective narrows their global semantic gap, while patch tokens are clustered into compact local patterns and matched through optimal transport for fine-grained feature alignment. Balanced optimal transport enforces fixed marginal masses even for local clusters without reliable counterparts, potentially introducing misleading alignment gradients. We therefore introduce confidence-adaptive unbalanced transport to relax these constraints for weakly matched clusters, aiming to reduce unreliable local alignment and improve adversarial transferability. We further study the effect of input transformations and propose visual-invariance augmentation, which applies bidirectional pixel-intensity rescaling and per-channel white-balance adjustment to simulate exposure, contrast, illumination, and color-temperature variations. This strategy encourages adversarial perturbations to generalize across different visual encoders. Extensive experiments on open-source and closed-source MLLMs show that IAU-FOA consistently outperforms state-of-the-art transferable attack methods. Code is available at this https URL.

154. 【2610.06973】Event Cameras for Melt-Pool Monitoring in Additive Manufacturing: A Benchmark and a Cross-Machine Transfer Analysis

链接:https://arxiv.org/abs/2610.06973

作者:Mohamad Yazan Sadoun,Sarah Sharif,Yingtao Liu,Zahed Siddique,Yaser Mike Banad

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:qualifying metal additive, metal additive manufacturing, event-camera benchmark exists, central to qualifying, qualifying metal

备注:

点击查看摘要

Abstract:Melt-pool monitoring is central to qualifying metal additive manufacturing (AM), yet no public event-camera benchmark exists for this domain. Event cameras report per-pixel brightness changes with microsecond timing instead of reading full frames, giving the temporal resolution AM transients demand at a fraction of the data rate. We present SynAM-E (Synthetic AM Events), the first public multi-source simulated event-camera benchmark for metal-AM melt-pool monitoring: 85 physics-calibrated event shards from 15 sources across 8 institutions, with public baselines and fixed cross-machine evaluation splits. On a single-machine case study, event-spatial monitoring matches dense-frame accuracy (0.874 versus 0.863 macro-F1), and the absolute intensity that events discard adds only +0.006 under fusion. On the NIST Additive Manufacturing Metrology Testbed (AMMT) build, a near-sensor event-rate counter recovers a raw-frame-confirmed 528.7 Hz intensity oscillation at ~380 times less sensor readout than the frame stream requires. A compact 93 k-parameter spiking model runs at 15 times lower modeled inference energy for a 0.073 macro-F1 cost. Every cross-source task includes a built-in trust test against camera identity shortcuts: process-type classification passes while material classification remains confounded by camera band, a corpus-structural limitation the release documents and the trust test exposes.

155. 【2610.06972】BoT-Feedback: Grounding Multimodal Reasoning in Biomechanical Evidence for Explainable Human Action Feedback

链接:https://arxiv.org/abs/2610.06972

作者:Xu Dong,Wanqing Li,Anthony Adeyemi-Ejeye,Andrew Gilbert

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Multimodal Large Language, Language Models, Large Language, Multimodal Large

备注:

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual understanding and multimodal reasoning, yet they remain fundamentally limited in Human Action Feedback Generation. Existing methods infer coaching feedback directly from visual observations, producing generic advice, limited interpretability, and physically implausible hallucinations. In contrast, expert human coaches diagnose performance through explicit biomechanical reasoning over joint kinematics, posture, and body dynamics. We introduce BoT-Feedback, a framework that grounds MLLM reasoning in structured biomechanical evidence. Our key contribution is Biomechanics of Thought (BoT), a four-stage reasoning framework that progressively identifies the action, localises the critical body regions, analyses quantitative biomechanical differences between expert and student performances, and synthesises interpretable coaching feedback. To support this reasoning process, we develop a plug-and-play Biomechanical Data Parser (BDP) that converts videos into structured biomechanical descriptors and an alignment strategy that temporally matches expert and student motions. We further introduce BiomAF, a benchmark containing paired teacher-student videos, 3D skeletons, biomechanical attributes, and expert-coaching annotations. Experiments across twelve open- and closed-source MLLMs demonstrate that grounding reasoning in biomechanical evidence consistently improves feedback quality, interpretability, and robustness while substantially reducing biomechanical hallucinations. BoT-Feedback improves the average expert evaluation score from 2.07 to 2.95 (+40%), enabling compact open-source MLLMs to approach the performance of substantially larger proprietary systems for explainable action feedback generation.

156. 【2610.06960】DistScene: Object-to-Scene Distillation for 3D Scene Generation

链接:https://arxiv.org/abs/2610.06960

作者:Kunming Luo,Hongyu Yan,Ken Deng,Chengcheng Zhou,Tianyu Liu,Haipeng Li,Haibin Huang,Xuelong Li,Ping Tan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:present DistScene, single-image compositional, framework for single-image, jointly modeling, environment

备注: Project page: [this https URL](https://coolbeam.github.io/DistScene/)

点击查看摘要

Abstract:We present DistScene, a framework for single-image compositional 3D scene generation by jointly modeling the environment and individual objects. Unlike existing methods that represent scenes primarily as collections of objects, we model the environment as an explicit scene component to provide geometric context for object placement. Specifically, we introduce Scene-Frame Generation, which jointly generates separate environment and object components in a shared coordinate frame, allowing their geometry and relative placement to be learned together. Then we introduce Object-Centric Refinement to refine each object in a local frame with scene context. Finally, we develop Object-to-Scene Distillation to transfer pretrained object-generation priors to scene generation through automatically composed and rendered synthetic scenes. Evaluations on indoor and outdoor benchmarks demonstrate improved scene-level spatial coherence over the evaluated baselines. Project page: this https URL

157. 【2610.06955】ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

链接:https://arxiv.org/abs/2610.06955

作者:Ruoxuan Feng,Yutong Chen,Ruihua Song,Huan Yang,Zhongyuan Wang,Guocai Yao,Di Hu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Humans inherently understand, Humans inherently, inherently understand, Humans, evidence

备注:

点击查看摘要

Abstract:Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.

158. 【2610.06945】Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models

链接:https://arxiv.org/abs/2610.06945

作者:Muhammad Atif Butt,Paweł Skierś,Joost Van De Weijer,Kamil Deja

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Linear Representation Hypothesis, Mechanistic interpretability, Representation Hypothesis, interpretability often relies, assumes that high-level

备注:

点击查看摘要

Abstract:Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear directions in activation space. Yet a natural visual concept does not necessarily require a linear visual transition: between sunny and stormy lies an intermediate weather state such as a sky with a few white clouds, not simply a weaker storm; between a caterpillar and a butterfly, the progression is not a caterpillar with continuously growing wings. This raises the question of whether such true intermediate states are also represented nonlinearly by the model. Indeed, when we prompt text-to-image models directly for intermediate attributes, their activations rarely fall along the straight direction connecting the endpoints. Therefore, we propose KANSteer, which models concept traversal as a curve passing through its intermediate states. Seeking a representation that is both simple and interpretable, we propose to use Kolmogorov-Arnold Networks (KANs), which provide a one-dimensional coordinate whose learned functions define the trajectory. This allows the steering direction to vary along the concept while preserving an interpretable representation. Across several concepts and text-to-image diffusion transformers, we find that their activation trajectories substantially deviate from straight lines, and that KANSteer provide a closer fit and smoother traversal of intermediate attributes than linear steering.

159. 【2610.06938】UniPro: Unified Multi-Mode Medical Image Segmentation from 2D Images to 3D Volumes via Propagation

链接:https://arxiv.org/abs/2610.06938

作者:Bangwei Guo,Yunhe Gao,Meng Ye,Yang Zhou,Difei Gu,Guoning Zhang,Leon Axel,Dimitris Metaxas

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Medical image segmentation, Medical image, segmentation remains fragmented, image segmentation remains, remains fragmented

备注:

点击查看摘要

Abstract:Medical image segmentation remains fragmented along two axes: segmentation paradigms and data dimensionality. Existing methods are typically developed separately for semantic, in-context, and interactive segmentation, and are further specialized to either native 2D images or 3D volumetric data. In clinical practice, however, segmentation workflows take many forms: a case may be initialized by semantic prediction, reference-guided segmentation, or user interaction. Regardless of how it begins, fine-grained refinement is naturally performed on 2D views; for volumetric scans, such 2D edits must propagate coherently to the rest of the volume. We present UniPro, a unified model that bridges segmentation paradigms and data dimensionality, using propagation to extend 2D segmentation to 3D volumes. Our key insight is that volumetric propagation and in-context segmentation share the same reference-conditioned prediction mechanism, differing only in whether the reference image-mask pairs come from other cases or from previously segmented neighboring slices. Building on this view, UniPro supports semantic, in-context, interactive, and propagation-based segmentation within a single slice-based framework, using class priors, reference exemplars, user clicks, and neighboring-slice predictions as mode-specific conditioning inputs. To improve propagation reliability, UniPro further incorporates bidirectional and 3D supervision to regularize slice-wise propagation beyond per-slice losses. Extensive experiments across diverse modalities and anatomies show that UniPro achieves strong performance across all segmentation settings, enabling annotation-efficient 3D segmentation from sparse 2D initialization and reducing slice-by-slice correction effort.

160. 【2610.06932】RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation

链接:https://arxiv.org/abs/2610.06932

作者:Siyu Huang,Yueyong Chen,Xuejiao Li,Jun Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Cache-based test-time adaptation, Cache-based test-time, unreliable entropy-based cache, test-time adaptation, unreliable entropy-based

备注:

点击查看摘要

Abstract:Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy-based cache admission under representation variations. To address these limitations, we propose RADC, which enhances prototype learning through reliable dual caching. RADC introduces a Semantic Foreground Cache that aggregates category-consistent spatial evidence from CLIP representations, yielding foreground prototypes that complement the global cache while mitigating background interference. To reliably manage both caches, Gaussian Risk Admission models multi-view representations as diagonal Gaussian distributions and jointly considers class separation and feature uncertainty to prioritize reliable cache candidates. RADC integrates zero-shot logits with complementary global- and foreground-cache predictions for robust inference. Extensive experiments on cross-domain and out-of-distribution benchmarks demonstrate consistent state-of-the-art performance.

161. 【2610.06919】Anchor Divergence for Semantic Geometry in Contrastive Learning

链接:https://arxiv.org/abs/2610.06919

作者:Akash Kannan,Kiho Park,Victor Veitch

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)

关键词:learned vector representations, paper concerns, learned vector, semantic context determines, context determines geometry

备注: Code is available at [this https URL](https://github.com/sky1712/Anchor-Divergence)

点击查看摘要

Abstract:This paper concerns how semantic context determines geometry in learned vector representations. Similarity is typically measured using cosine similarity, which provides a single fixed geometry. Semantic similarity, however, is inherently context dependent: two images may be similar because they depict the same object, share a visual style, or are relevant to the same clinical finding. We show that contrastive representations naturally encompass a family of geometries that can be specialized to particular semantic structure. The key idea is to use an interplay between contrastive learning, exponential families, and information geometry to establish a correspondence between probability distributions over "anchors" and Bregman geometries on the representation space. We use this correspondence to define "Anchor Divergences", a method for specifying context-specific semantic geometries on fixed representations. Under this correspondence, modeling the anchor distribution models the geometry itself. Experiments on retrieval show that anchor divergences provide an effective and efficient way to specify context-specific semantic similarity.

162. 【2610.06896】Medical Image Alignment Assessment as a Test of Generalist Visual Reasoning in Frontier Multimodal Models

链接:https://arxiv.org/abs/2610.06896

作者:Ross Callaghan,Niannu Gao,Hojjat Azadbakht,Hui Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:artificial general intelligence, medical image alignment, image alignment, medical image, increasingly positioned

备注:

点击查看摘要

Abstract:Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial general intelligence. A key test of this generality is whether they can perform novel visual judgments that humans can make reliably from visual evidence and task instructions, without task-specific parameter optimisation. We investigate this question through the task of medical image alignment assessment, where the goal is to establish whether there is anatomical correspondence between two images. Human visual assessment of image alignment is still the gold standard and most common approach; however, it requires trained operators and is impractical to scale for large datasets. We evaluate recent generations of MLLMs on two exemplar medical image alignment tasks, varying both prompting strategies and image-presentation methods. We compare against a locally fine-tuned MLLM and a task-specific CNN to examine the trade-off between frontier general purpose models and smaller models that require specific task optimisation but can be used locally. We show that are reaching an inflection point, where frontier MLLMs can now perform effective visual assessment of medical image alignment. Models released only a few months ago generalise poorly and, in some settings, perform barely above chance, whereas GPT-6 achieves over 85% across almost all scenarios tested. Fine-tuned local models can match or exceed frontier-model performance on the tasks on which they are trained, but transfer substantially less effectively to unseen settings. These findings identify medical image alignment as a useful test bed for generalist visual reasoning and suggest that frontier multimodal models are beginning to acquire capabilities that could support a common quality-control mechanism across heterogeneous medical-imaging pipelines.

163. 【2610.06855】EMPEST: Temporal Embeddings for Scalable Driver Identification via Angular Margin Learning

链接:https://arxiv.org/abs/2610.06855

作者:Kyle Musgrove,Dylan B. Lewis,Sarah Powers,Emma J. Reid,Hector Santos-Villalobos

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:fleet size grows, maintain discriminative performance, Temporal Convolutional Network, Convolutional Network embedding, existing triplet-loss formulations

备注:

点击查看摘要

Abstract:Scalable driver identification requires embedding models that maintain discriminative performance as fleet size grows, yet existing triplet-loss formulations degrade rapidly with driver pool size and overfit to session-specific patterns under rigorous temporal evaluation. We introduce TEMPEST, a Temporal Convolutional Network embedding model trained with an additive angular margin (ArcFace) loss that enforces global class-level separation in a normalized angular space. TEMPEST maps 60-second multimodal driving windows to compact 96-dimensional embeddings, supporting truly dynamic enrollment without any retraining or classifier refitting. Under rigorous temporal evaluation on a 45-driver dataset, TEMPEST achieves 91.71% Rank-1 accuracy, outperforming the best classical model by 17.9 pp and the strongest triplet-loss baseline by 58.4 pp. TEMPEST degrades by only 4.3 pp when growing the subject pool from 10 to 45 drivers, compared to 22 pp and 32.5 pp for supervised and unsupervised triplet-loss baselines, and its cross-session advantage is corroborated on the public KIA Soul dataset, where it outperforms the best classical model by 7.3 pp within-session and 14.3 pp cross-session. With 720K parameters, a 2.80 MB footprint, and 50-epoch convergence, TEMPEST establishes a rigorous, reproducible baseline for scalable behavioral driver biometric identification.

164. 【2610.08227】How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data

链接:https://arxiv.org/abs/2610.08227

作者:Robin Young

类目:Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Machine learning classifiers, remote sensing imagery, Machine learning, remote sensing, sensing imagery

备注:

点击查看摘要

Abstract:Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actually contain? For an $n \times n$ image whose spatial correlation persists over a range of $r$ pixels, the effective sample size is $\Theta(n^2/r^2)$, not $n^2$. We prove this as a finite-sample upper bound for classifiers on spatially correlated data, and show via a matching lower bound that the rate is tight, and no algorithm can do better. We extend the results to images with directional correlation and spatially varying correlation structure. Our result justifies spatial cross-validation since block holdout with separation proportional to the correlation range achieves optimal generalization guarantees, while random holdout can underestimate confidence interval widths by a factor proportional to $r$. We validate the theory on synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1).